PulseAugur
中
实时 15:56:03

DR. INFO 临床 AI 在 HealthBench 上击败 GPT-5、Gemini · arXiv 论文

一项新的研究论文介绍 DR. INFO,一个基于代理 RAG 的临床助手,在 HealthBench 基准测试中表现显著优于领先的 LLM。DR. INFO 在具有挑战性的 HealthBench Hard 子集上取得了 0.68 的分数,超过了 GPT-5 (0.46)、Grok 3 (0.23)、Gemini 2.5 Pro (0.19) 和 Claude 3.7 Sonnet (0.02) 等模型。评估突出了 DR. INFO 在准确性和指令遵循方面的优势,同时也指出了在上下文感知和响应完整性方面的改进空间,强调了在构建值得信赖的 AI 医疗助手时需要基于评分标准的评估。 AI

影响 为 AI 临床助手设定了新的基准,强调了超越多项选择题的先进评估方法的必要性。

排序理由 该集群是一篇研究论文,详细介绍了一个新的人工智能模型及其在基准测试上的表现。 [lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

DR. INFO 临床 AI 在 HealthBench 上击败 GPT-5、Gemini · arXiv 论文

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群是一篇研究论文,详细介绍了一个新的人工智能模型及其在基准测试上的表现。 [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
72 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Sandhanakrishnan Ravichandran, Shivesh Kumar, Rogerio Corga Da Silva, Miguel Romano, Reinhard Berkels, Michiel van der Heijden, Olivier Fail, Valentine Emmanuel Gnanapragasam ·

    OpenAIs HealthBench 实践:在真实的临床查询中评估基于 LLM 的医疗助手

    arXiv:2509.02594v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) on their ability to generate high-quality, accurate, situationally aware answers to clinical questions requires going beyond conventional benchmarks to assess how these systems behav…