A new benchmark called APEX-Accounting has been developed by Mercor in collaboration with Ramp to evaluate the capabilities of frontier AI models in performing accounting tasks. The benchmark includes 160 expert-authored tasks covering account reconciliation, expense accrual, transaction posting, and report generation. In evaluations across nine models, Claude Fable-5 achieved the highest score with 56.4% Mean Criteria@3, followed by Muse Spark 1.1 at 52.6%. The study also observed a Simpson's paradox related to token budgets, where higher budgets generally improved scores, but within a constrained harness, increased token usage on specific tasks led to lower performance. AI
IMPACT This benchmark could drive improvements in AI's ability to handle complex, domain-specific tasks like accounting, potentially leading to new tools for financial professionals.
RANK_REASON The cluster describes a new benchmark and research paper evaluating AI models on specific tasks.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →