PulseAugur
实时 20:44:36
English(EN) Clinical Reasoning Graphs: Structured Evaluation of LLM Diagnostic Reasoning Reveals Competence Without Consistency

新框架揭示大语言模型临床推理缺陷,尽管诊断准确性尚可

两篇新研究论文介绍了评估大语言模型(LLMs)临床推理能力的框架。第一篇,CLExEval,采用一种人工干预的循环方法,通过渐进式信息屏蔽来揭示诸如冗余偏见和推理到输出不匹配等失败模式,涉及GPT-4o-mini等模型。第二篇,临床推理图谱(Clinical Reasoning Graphs),采用结构化图表示来分析大语言模型的诊断轨迹,揭示模型虽然表现出诊断能力,但在相似病例中缺乏一致的推理。两项研究都强调,除了简单的准确性指标外,还需要进行过程级别的评估,以确保大语言模型在临床上的可靠应用。 AI

影响 强调了大语言模型临床推理的关键局限性,表明当前的评估方法可能高估了其可靠性,并警示不要在未经核实的情况下将其部署到医疗保健领域。

排序理由 两篇介绍大语言模型新评估框架的学术论文。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 10 个来源。 我们如何撰写摘要 →

新框架揭示大语言模型临床推理缺陷,尽管诊断准确性尚可

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
两篇介绍大语言模型新评估框架的学术论文。
Source corroboration
10 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
66 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+6 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [10]

  1. arXiv cs.CL TIER_1 English(EN) · Zhiyun Zhang, Liwen Sun, Xiang Qian, Chenyan Xiong ·

    FaithMed:训练LLM以实现忠实于循证医学推理

    arXiv:2607.01440v1 Announce Type: new Abstract: Faithful reasoning is essential in medicine, where clinical decisions require transparent justification grounded in reliable evidence. Current medical LLMs either lack active access to evidence or use retrieved evidence without supe…

  2. arXiv cs.AI TIER_1 English(EN) · Samiha A. Ismail, Fan X. Chen, Ali Merali ·

    基于评分标准的受控比较:前沿语言模型在专家撰写的临床推理任务上的表现

    arXiv:2607.02175v1 Announce Type: new Abstract: Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its "Hard" subset top score remains 32%. We …

  3. arXiv cs.AI TIER_1 English(EN) · Ali Merali ·

    基于评分标准的受控比较:前沿语言模型在专家撰写的临床推理任务上的表现

    Multiple-choice medical benchmarks are increasingly saturated, and recent rubric-based evaluations such as HealthBench have shown that open-ended clinical performance is far from solved - its "Hard" subset top score remains 32%. We present a small, deliberately difficult evaluati…

  4. arXiv cs.CL TIER_1 English(EN) · William Philipp, Finn Fassbender, Thorsten Langer, Martje Pauly, Rebecca Herzog, Alexander Baumann, Markus Hobert, Theresa Paulus, Ip Chi Wang, Lukas Goede, Johanna Reimer, Sebastian L\"ons, Ronald B\"ock, Sebastian Fudickar ·

    无需临床谨慎即可达到临床医生级别的一致性:医疗AI基准测试中的LLM评估器局限性

    arXiv:2607.01103v1 Announce Type: new Abstract: Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whether such evaluators replicate clinical calibration …

  5. arXiv cs.CL TIER_1 English(EN) · Chenyan Xiong ·

    FaithMed:为忠实于循证医学推理训练大型语言模型

    Faithful reasoning is essential in medicine, where clinical decisions require transparent justification grounded in reliable evidence. Current medical LLMs either lack active access to evidence or use retrieved evidence without supervising how it should be appraised and applied d…

  6. arXiv cs.CL TIER_1 English(EN) · Sebastian Fudickar ·

    无需临床谨慎即可达到临床医生级别的一致性:医疗AI基准测试中的LLM评估器局限性

    Open-response evaluation provides stronger clinical validity than multiple-choice benchmarks but creates a scoring bottleneck that motivates automated LLM-asa-Judge approaches. Whether such evaluators replicate clinical calibration and caution, however, remains untested. We intro…

  7. arXiv cs.CL TIER_1 English(EN) · Ajmal M., Abin Roy, Afthab Salam Kanniyan, Jawadh Abdul Kabeer, Jerin James, Preslav Nakov, Zhuohan Xie ·

    CLExEval:用于 LLM 临床推理定性评估的“人在回路”框架

    arXiv:2606.31608v1 Announce Type: new Abstract: Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations c…

  8. arXiv cs.CL TIER_1 English(EN) · Zhuohan Xie ·

    CLExEval:LLM临床推理定性评估的人工干预框架

    Large Language Models (LLMs) achieve strong results on many medical benchmarks, but their clinical reasoning remains difficult to evaluate reliably. A central risk is an evaluation illusion: fluent and well-structured explanations can appear clinically convincing even when the fi…

  9. arXiv cs.AI TIER_1 English(EN) · Nisarg A. Patel (University of California, San Francisco) ·

    临床推理图:LLM诊断推理的结构化评估显示出能力但缺乏一致性

    arXiv:2606.29876v1 Announce Type: cross Abstract: Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot distinguish stable clinically-grounded reasoning from pattern matching. We introduce clinical reas…

  10. arXiv cs.CL TIER_1 English(EN) · Nisarg A. Patel ·

    临床推理图:对LLM诊断推理的结构化评估揭示了缺乏一致性的能力

    Modern large language models (LLMs) reach 60-70% diagnostic accuracy on complex clinical case benchmarks, but accuracy alone cannot distinguish stable clinically-grounded reasoning from pattern matching. We introduce clinical reasoning graphs, structured graph representations ext…