PulseAugur
实时 21:54:45
English(EN) Capabilities of Claude Fable 5 on Biomedical Challenge Problems

Claude Fable 5 准确率高但拒绝回答大多数生物医学问题

一项评估 Anthropic 的 Claude Fable 5 模型在生物医学挑战方面的新研究论文揭示了该模型在回答问题意愿方面存在一个显著问题。尽管 Claude Fable 5 在 MedQA 和 RareBench 等基准测试中表现出高准确率(当它确实提供答案时),但它拒绝回答 8.0% 到 99.4% 的问题,具体比例取决于特定基准。这种拒绝模式与其前代模型和 GPT-5 不同,表明可能存在限制其在生物医学领域实际效用的安全或对齐机制。 AI

影响 这项研究强调了模型安全/对齐与效用之间可能存在的权衡,表明未来的 LLM 可能需要在专业领域内平衡拒绝机制与任务完成度。

排序理由 该集群包含一篇评估 LLM 在特定基准上能力的学术论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Claude Fable 5 准确率高但拒绝回答大多数生物医学问题

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Dominic Okonkwo, Magnus Hodgson, Temitope I. David, Susan Adanna Ihejirika ·

    Claude Fable 5 在生物医学挑战问题上的能力

    arXiv:2607.10849v1 Announce Type: new Abstract: Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language models.…

  2. arXiv cs.CL TIER_1 English(EN) · Susan Adanna Ihejirika ·

    Claude Fable 5 在生物医学挑战问题上的能力

    Frontier language models are increasingly evaluated on biomedical benchmarks, but two problems undermine most published evaluations: legacy benchmarks are near-saturated, and open-ended responses are graded by other language models. We evaluate Claude Fable 5, Anthropic's most ca…