The Estonian Language Institute has developed a new benchmark to evaluate how well large language models resist Russian propaganda. The test ranks dozens of LLMs on their ability to avoid taking positions on topics frequently used in Russian strategic narratives. Anthropic's Claude models, particularly Opus 4.7, performed best among proprietary frontier models, achieving a high score by consistently pushing back against misinformation. AI
IMPACT Establishes a new evaluation standard for LLM safety and resistance to state-sponsored disinformation campaigns.
RANK_REASON The cluster describes a new benchmark developed by a government-sponsored institute to evaluate LLM performance on a specific safety/policy-related task.
Read on Mastodon — fosstodon.org →
- Anthropic
- Claude Opus 4.7
- Claude Sonnet
- Estonian Language Institute
- Propastop
- Claude
- Claude models
- Propaganda Resistance benchmark
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →