AI beats licensed accountants on speed and accuracy, but still can't close the books without supervision
Mercor finds leading models now beat licensed CPAs on structured APEX accounting tasks but still need supervision.
Mercor compared AI models with 12 licensed CPAs, averaging 5.5 years of experience, on simplified tasks from the APEX Accounting Benchmark. Eighteen months ago the best models scored below the accountants’ roughly 37 percent average; Mercor says current models now handle those same tasks almost flawlessly. On the full benchmark of 160 tasks across 10 simulated companies, Claude Opus 5.5 led at 61.8 percent of grading criteria, followed by Fable 5.1 at 61.0 percent and GPT-6 Astra at 57.9 percent. Mercor said no model fully solved nearly 60 percent of tasks, and the study left out client interaction and long-built professional context.