Databricks has released OfficeQA Pro V2, a new benchmark designed to evaluate the grounded reasoning capabilities of AI agents on enterprise-style tasks. This benchmark utilizes a new corpus of approximately 120,000 pages from the U.S. Treasury's Accounts of Receipts and Expenditures, spanning from 1793 to 2024. Initial evaluations of frontier models showed an average accuracy of 26.0%, with notable performance variations between models and persistent challenges in parsing, temporal reconciliation, and entity scope interpretation. AI
IMPACT This benchmark will help drive progress in AI agent capabilities for enterprise document analysis and reasoning.
RANK_REASON The item describes the release of a new benchmark for evaluating AI capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- AI agents
- Anthropic
- Claude Code
- Codex
- Databricks
- Google DeepMind
- GPT-5.6 Sol
- OfficeQA Pro V2
- OpenAI
- Sonnet 5
- U.S. Treasury
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →