PulseAugur
EN
LIVE 18:37:46

Google Research: LLMs struggle with recall, not encoding, for factual errors

Google Research has introduced a new framework called knowledge profiling to better understand why Large Language Models (LLMs) make factual errors. This framework distinguishes between facts that a model fails to encode (an "empty shelf" problem) and facts that are encoded but inaccessible (a "lost keys" problem). Their research, using the WikiProfile benchmark, indicates that frontier LLMs like Gemini3 and GPT-5 primarily suffer from recall failures rather than encoding failures, suggesting that improving retrieval mechanisms could significantly enhance LLM factuality. AI

IMPACT Suggests that improving LLM recall mechanisms could be a more effective path to enhanced factuality than simply scaling model size.

RANK_REASON Research paper introducing a new framework and benchmark for evaluating LLM factuality. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Google AI / Research →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Google Research: LLMs struggle with recall, not encoding, for factual errors

COVERAGE [1]

  1. Google AI / Research TIER_1 English(EN) ·

    Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

    Generative AI