Researchers have introduced GAMUT, a new benchmark designed to evaluate the factual completeness of long-form AI-generated text. Unlike previous methods that focused on the accuracy of individual claims, GAMUT assesses whether a response includes all necessary information, addressing the 'missing half' of factuality. The benchmark utilizes a two-level meta-rubric system that can be mechanically compiled into a machine-gradable checklist, proving effective even with LLM judges. In evaluations, GAMUT demonstrated its challenging nature, with the best-performing model, Gemini 3.1 Pro, achieving only 58.7% accuracy. AI
IMPACT This benchmark could drive improvements in AI models' ability to generate comprehensive and factually complete long-form content.
RANK_REASON The cluster describes a new academic paper introducing a novel benchmark for evaluating AI model outputs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →