PulseAugur
EN
LIVE 04:13:07

New GulliBench benchmark ranks AI models by "gullibility"

A new benchmark called GulliBench aims to measure the "gullibility" of AI models. The top performers in this benchmark, indicating the most gullible models, include Opus 5 and Fable 5. Other models like Muse Spark, Gemini 3.1 Pro, and Kimi K3 also showed varying degrees of gullibility. AI

IMPACT This benchmark may highlight potential vulnerabilities in AI models, prompting further research into their robustness and safety.

RANK_REASON The cluster describes a new benchmark for AI models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New GulliBench benchmark ranks AI models by "gullibility"

COVERAGE [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    GulliBench is a benchmark that purports to measure the "gullibility" of AI models. Here's there most "gullible" top 10: 1. Opus 5 - 49 2. Fable 5 - 48 3. Muse S

    GulliBench is a benchmark that purports to measure the "gullibility" of AI models. Here's there most "gullible" top 10: 1. Opus 5 - 49 2. Fable 5 - 48 3. Muse Spark 1.2 - 42 4. Gemini 3.1 Pro - 23 5. Kimi K3 - 18 6. Grok 4.6 - 16 7. Opus 4.8 - 16 8. DeepSeek V4 Flash - 16 9. GLM …