Researchers have developed a new benchmark called Probity to measure the instability of language models when answering questions about documents. The benchmark, comprising 60 tasks and 470 items derived from venture-financing filings, revealed that models sometimes provide different answers to the same question about the same document. An audit of the corpus identified a defect where evidence crucial for answering questions was missing from the text window provided to the model, leading to increased instability. The study found that even after attempting to correct for this missing evidence, model instability persisted, suggesting limitations in what document-grounded benchmarks can reveal about model stability. AI
IMPACT Highlights potential issues with current LLM evaluation methods and the need for more robust benchmarking for document-grounded tasks.
RANK_REASON The cluster contains an academic paper detailing a new benchmark for evaluating language model behavior. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →