Artificial intelligence models are faster, more accurate, and much cheaper than human accountants at structured bookkeeping tasks. A study by Mercor tested 12 licensed certified public accountants who had an average of five and a half years of experience. These professionals worked through simplified tasks drawn from the APEX Accounting Benchmark.
Eighteen months ago, the top models scored below the accountant average of about 37 percent. Today, those same models solve the identical tasks almost flawlessly.
Benchmark results show the leading models
The full APEX Accounting benchmark is much larger. It features 160 tasks across 10 simulated companies, built by more than 40 professionals averaging 11 years of experience.
Claude Opus 5.5 currently leads the benchmark with 61.8 percent of grading criteria met. Fable 5.1 follows closely behind at 61.0 percent. GPT-6 Astra secures the third spot with 57.9 percent.
Despite these high scores, Mercor notes that no model fully solved almost 60 percent of the tasks. Artificial intelligence models still cannot close the books without human oversight.

Some accounting tasks remain outside reach
Mercor acknowledges that the study tasks test the exact areas where artificial intelligence excels. These strengths include hunting down details and following instructions precisely.
The study omitted key parts of the accounting profession. Left out were tasks like talking with clients, checking in with colleagues, and drawing on context built up over years.
Mercor states that accountants cannot be replaced for these reasons. However, the company expects major productivity gains across the industry.



