A new benchmark called AcoustiClaim has been developed to evaluate the accuracy of audio language models in stating numerical claims about acoustic quantities. The benchmark extracts numeric claims from text and scores them against instrument references, classifying quantities by their reference source. Experiments with five models, including one closed model, revealed significant error rates, with many models failing to exceed a constant-predictor floor. Even with a calibrated threshold, error reduction was limited, and some metrics remained consistently high. AI
IMPACT Highlights limitations in current audio language models' ability to accurately quantify acoustic properties, suggesting a need for improved numerical reasoning and grounding.
RANK_REASON The item is a research paper detailing a new benchmark for evaluating audio language models. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →