AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision
Mercor finds leading models now beat licensed CPAs on structured APEX accounting tasks but still need supervision.
Mercor compared AI models with 12 licensed CPAs, averaging 5.5 years of experience, on simplified tasks from the APEX Accounting Benchmark. Eighteen months ago the best models scored below the accountants’ roughly 37 percent average; Mercor says current models now handle those same tasks almost flawlessly. On the full benchmark of 160 tasks across 10 simulated companies, Claude Opus 5.5 led at 61.8 percent of grading criteria, followed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent. Mercor said no model fully solved nearly 60 percent of tasks, and the study left out client interaction and long-built professional context.
- Twelve licensed CPAs averaged about 37 percent on simplified APEX tasks.
- Claude Opus 5.5 leads the full benchmark at 61.8 percent of criteria met.
- Fable 5.1 scored 61.0 percent and GPT-6 Astra 57.9 percent.
- No model fully solved nearly 60 percent of the 160 tasks.
- The test omitted client contact and years of professional context.
Full article274 words · extracted from the-decoder.com · click to collapse
AI models are faster, more accurate, and far cheaper than accountants at structured bookkeeping tasks. That's the finding of a study by Mercor in which 12 licensed CPAs with an average of five and a half years of experience worked through simplified tasks from the APEX Accounting Benchmark. Eighteen months ago, the best models still scored below the accountants' average of about 37 percent. Today, they solve the same tasks almost flawlessly.

The full APEX Accounting benchmark is much bigger, with 160 tasks across 10 simulated companies, built by more than 40 professionals who average 11 years of experience. Claude Opus 5.5 currently leads with 61.8 percent of grading criteria met, followed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent. Still, Mercor says no model fully solved almost 60 percent of the tasks. AI models can't close the books without oversight yet.
Mercor also admits the study's tasks test exactly what AI does best, which is hunting down details and following instructions precisely. The study left out key parts of the job, such as talking with clients, checking in with colleagues, and drawing on context built up over years. Mercor says that's why accountants can't be replaced, though it expects major productivity gains across the industry.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/ai-beats-licensed-accountants-on-speed-and-accuracy-but-still-cant-close-the-books-without-supervision/