A new benchmark called ConsistencyAI has been developed to evaluate how factually consistent large language models (LLMs) are when responding to users from different demographic groups. The benchmark tests whether LLMs provide the same factual information regardless of the persona asking the question. In experiments with 19 LLMs, scores for factual consistency ranged from 0.7896 to 0.9065, with a mean of 0.8656. xAI's Grok-3 performed most consistently, while smaller models were less consistent. The study also found that consistency varies by topic, with the job market being the least consistent and world leaders being the most consistent. AI
IMPACT Highlights potential biases in LLMs and the need for persona-invariant prompting strategies.
RANK_REASON This is a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- ConsistencyAI
- DagsHub
- Gotit.pub
- Grok-3
- Hugging Face
- LLMs
- Peter Banyasz
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →