A new research paper introduces LLMPEDIA, a system designed to measure and browse the encyclopedic knowledge embedded within large language models. LLMPEDIA recursively extracts approximately 1.3 million articles from the parametric memory of models like GPT-5 mini, DeepSeek V3.2, and Llama 3.3 70B. The system then verifies a sample of these claims against Wikipedia and other web sources, categorizing them as supported, refuted, or insufficient. Findings indicate that LLMs have a true factual rate of 68.4% on this broader knowledge set, significantly lower than benchmark scores, with a substantial portion of claims being unresolvable. AI
IMPACT This research highlights a significant gap between benchmark performance and real-world factual accuracy in LLMs, suggesting a need for more robust evaluation methods.
RANK_REASON The cluster is a research paper detailing a new methodology and system for evaluating LLM knowledge. [lever_c_demoted from research: ic=1 ai=1.0]
- DeepSeek V3.2
- GPT-5 mini
- GPTKB
- Hendrycks et al.
- Hu et al.
- Llama 3.3 70B Instruct
- LLMPEDIA
- Massive Multitask Language Understanding
- Saeed and Razniewski
- Wikipedia
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →