PulseAugur
EN
LIVE 15:50:36

Databricks launches OfficeQA Pro V2 benchmark for enterprise AI reasoning

Databricks has released OfficeQA Pro V2, a new benchmark designed to evaluate the grounded reasoning capabilities of AI agents on enterprise-style tasks. This benchmark utilizes a new corpus of approximately 120,000 pages from the U.S. Treasury's Accounts of Receipts and Expenditures, spanning from 1793 to 2024. Initial evaluations of frontier models showed an average accuracy of 26.0%, with notable performance variations between models and persistent challenges in parsing, temporal reconciliation, and entity scope interpretation. AI

IMPACT This benchmark will help drive progress in AI agent capabilities for enterprise document analysis and reasoning.

RANK_REASON The item describes the release of a new benchmark for evaluating AI capabilities. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Databricks Blog →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Databricks launches OfficeQA Pro V2 benchmark for enterprise AI reasoning

COVERAGE [1]

  1. Databricks Blog TIER_1 English(EN) ·

    Introducing OfficeQA Pro V2: A New Benchmark for Enterprise Grounded-Reasoning

    Today, we are releasing OfficeQA Pro V2, a new benchmark designed to evaluate whether...