PulseAugur
中
实时 09:58:13
English(EN) A new medical AI study found the same flaw in OpenEvidence, OpenAI, Anthropic, and Doximity

一项新研究显示,包括OpenAI和Anthropic模型在内的医疗AI工具存在高遗漏错误率

斯坦福大学、哈佛大学和ARISE网络的研究人员最近进行的一项名为NOHARM的研究,评估了几种AI工具在临床环境中的表现。该研究使用了1100个真实的临床案例和医生注释,测试了OpenEvidence、Doximity的Ask、OpenAI的GPT-5.6 Sol以及Anthropic的Claude Fable 5。尽管Doximity的Ask表现最佳,但所有接受测试的AI系统都存在一个显著缺陷:76.6%的有害错误是遗漏,意味着AI未能包含关键信息,而不是陈述错误事实。这凸显了在医疗保健领域确保AI可靠性所面临的持续挑战,即使监管机构和法律框架仍在努力解决AI问责问题。 AI

影响 强调了医疗AI中关键的遗漏错误,突出了人工监督和健全监管框架的必要性。

排序理由 该集群报告了一项新的独立基准研究,该研究评估了临床环境中的AI工具,包括特定的AI模型及其性能指标。[lever_c_demoted from research: ic=1 ai=1.0]

在 Fortune 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

一项新研究显示,包括OpenAI和Anthropic模型在内的医疗AI工具存在高遗漏错误率

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群报告了一项新的独立基准研究,该研究评估了临床环境中的AI工具,包括特定的AI模型及其性能指标。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, safety, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
70 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Fortune TIER_1 English(EN) · Lily Mae Lazarus ·

    一项新的医疗AI研究发现OpenEvidence、OpenAI、Anthropic和Doximity存在相同缺陷

    The findings suggest healthcare AI's next challenge isn't adoption or funding—it's proving the technology can avoid dangerous omissions at the point of care.