A new benchmark called GulliBench aims to measure the "gullibility" of AI models. The top performers in this benchmark, indicating the most gullible models, include Opus 5 and Fable 5. Other models like Muse Spark, Gemini 3.1 Pro, and Kimi K3 also showed varying degrees of gullibility. AI
IMPACT This benchmark may highlight potential vulnerabilities in AI models, prompting further research into their robustness and safety.
RANK_REASON The cluster describes a new benchmark for AI models. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
- DeepSeek V4 Flash
- DeepSeek V4 Pro
- Gemini 3.1 Pro
- GLM 5.2
- Grok 4.6
- GulliBench
- Kimi K3
- Muse Spark
- Opus 4.8
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →