PulseAugur
实时 06:19:05

New benchmark ClinMM-Bench evaluates LLMs on complex clinical diagnostic reasoning

研究人员开发了 ClinMM-Bench,这是一个旨在评估大型语言模型 (LLM) 在复杂临床场景中进行多轮多模态诊断推理能力的新基准。该基准包含 1,089 个真实临床病例和 3,760 张医学影像,涵盖八个专科。对 15 个代表性 LLM 的初步评估显示,尽管专有模型表现出更高的诊断准确性,但没有一个模型能做出完美诊断,所有模型在生成可靠的诊断推理方面都存在局限性,常见的失败模式包括信息综合问题和视觉幻觉。 AI

影响 该基准通过突出特定的推理失败,有望推动 LLM 在医疗应用中的改进。

排序理由 该集群包含一篇介绍用于评估 AI 模型的新基准的研究论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

New benchmark ClinMM-Bench evaluates LLMs on complex clinical diagnostic reasoning

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Rui Yang, Weihao Xuan, Yi Lin, Zhuhan Bao, Jonathan Chong Kai Liew, Matthew Yu Heng Wong, Nicol\'as Lescano, Nikita R. Paripati, Emily Ling-Lin Pai, Jiarui Liu, Heli Qi, Heng-Jui Chang, Benny Kai Guo Loo, Huitao Li, Kunyu Yu, Yufan Wang, Chuan Hong, Shij… ·

    对具有挑战性的真实世界临床病例进行多轮多模态诊断推理评估

    arXiv:2607.25933v1 Announce Type: cross Abstract: Clinical diagnostic evaluation should not only assess whether models can provide correct diagnoses, but also reflect the realities of clinical practice, including progressive disclosure of multimodal information, dynamic updating …