AI model accuracy on the study's accounting tasks compared to accountants working without AI assistance, broken down by model release date. Image: Mercor
Source: the-decoder.com
What the benchmark measures
The full APEX Accounting benchmark contains 160 tasks across 10 simulated companies. More than 40 specialists prepared it, with an average of 11 years of experience.
Mercor says the tasks focus on areas where AI is particularly strong: finding details and following instructions precisely. The benchmark leaves out client communication, coordination with colleagues and context accountants accumulate over years.
That boundary matters. A model can perform well on bounded tasks without taking responsibility for an entire accounting workflow. The results do not show that AI can independently close a company’s books.
The gap between task performance and replacement
Mercor’s conclusion is that accountants cannot yet be replaced, though AI could substantially raise productivity across the industry. I think the more revealing result is not the narrow gap between the top two models, but how much of the job the benchmark does not test. Its scores measure performance on prepared tasks; they do not establish how well a model handles the human coordination and accumulated context that real accounting work depends on.
The benchmark makes a case for AI as a tool inside accounting teams, not as a substitute for them. The unanswered question is whether future systems can bridge that gap without human oversight.
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X