PulseAugur
实时 03:52:39
English(EN) GulliBench is a benchmark that purports to measure the "gullibility" of AI models. Here's there most "gullible" top 10: 1. Opus 5 - 49 2. Fable 5 - 48 3. Muse S

新的 GulliBench 基准测试按“易受骗性”对 AI 模型进行排名

一个名为 GulliBench 的新基准测试旨在衡量 AI 模型的“易受骗性”。在此基准测试中表现最佳(表明最易受骗的模型)的模型包括 Opus 5Fable 5Muse SparkGemini 3.1 ProKimi K3 等其他模型也表现出不同程度的易受骗性。 AI

影响 该基准测试可能会凸显 AI 模型潜在的漏洞,促使对其鲁棒性和安全性进行进一步研究。

排序理由 该集群描述了一个新的 AI 模型基准测试。[lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 GulliBench 基准测试按“易受骗性”对 AI 模型进行排名

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    GulliBench 是一个旨在衡量 AI 模型“易受骗性”的基准测试。以下是其“最易受骗”的前 10 名:1. Opus 5 - 49 2. Fable 5 - 48 3. Muse S

    GulliBench is a benchmark that purports to measure the "gullibility" of AI models. Here's there most "gullible" top 10: 1. Opus 5 - 49 2. Fable 5 - 48 3. Muse Spark 1.2 - 42 4. Gemini 3.1 Pro - 23 5. Kimi K3 - 18 6. Grok 4.6 - 16 7. Opus 4.8 - 16 8. DeepSeek V4 Flash - 16 9. GLM …