i
News
News · 2026-10-02

Mercor’s accounting benchmark puts Claude Opus 5.5 ahead

@neuronium_ai @neuronium_ai

Mercor’s APEX Accounting benchmark puts Claude Opus 5.5 ahead of licensed accountants on speed and accuracy, but the results point to a narrower claim than the headline suggests: the leading model met 61.8% of the benchmark’s criteria, and no model fully solved nearly 60% of the tasks. The test measures selected accounting skills, not the full work of closing books.

Cover: Mercor’s accounting benchmark puts Claude Opus 5.5 ahead
AI model accuracy on the study's accounting tasks compared to accountants working without AI assistance, broken down by model release date. Image: Mercor

AI model accuracy on the study's accounting tasks compared to accountants working without AI assistance, broken down by model release date. Image: Mercor

Source: the-decoder.com

What the benchmark measures

The full APEX Accounting benchmark contains 160 tasks across 10 simulated companies. More than 40 specialists prepared it, with an average of 11 years of experience.

Claude Opus 5.5: 61.8% of evaluation criteria
Fable 5.1: 61.0%
GPT-6 Astra: 57.9%

Mercor says the tasks focus on areas where AI is particularly strong: finding details and following instructions precisely. The benchmark leaves out client communication, coordination with colleagues and context accountants accumulate over years.

That boundary matters. A model can perform well on bounded tasks without taking responsibility for an entire accounting workflow. The results do not show that AI can independently close a company’s books.

The gap between task performance and replacement

Mercor’s conclusion is that accountants cannot yet be replaced, though AI could substantially raise productivity across the industry. I think the more revealing result is not the narrow gap between the top two models, but how much of the job the benchmark does not test. Its scores measure performance on prepared tasks; they do not establish how well a model handles the human coordination and accumulated context that real accounting work depends on.

The benchmark makes a case for AI as a tool inside accounting teams, not as a substitute for them. The unanswered question is whether future systems can bridge that gap without human oversight.

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X