A new benchmark called Polistemics has been developed to evaluate how well Large Language Models (LLMs) mediate political information during elections. This benchmark, grounded in the theory of Epistemic Modesty, assesses LLMs' performance across various informational conditions like clarity, noise, and consistency. When applied to three state-of-the-art LLMs using data from the 2025 German and Dutch elections, the study found that while models performed well with clear information, they struggled with absent, vague, or contradictory data, and tended to reduce the intensity of political language. These issues appear to be influenced by party priors and output language, indicating that reliable mediation is possible but not yet consistently achieved by current models. AI
IMPACT Highlights potential risks of LLMs in political discourse and the need for robust evaluation frameworks.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- DagsHub
- German Elections of 2002: Aftermath and Implications for the United States
- Hugging Face
- LLMs
- Polistemics
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →