PulseAugur
中
实时 06:23:51
English(EN) A rubric landscape for evaluating clinical reasoning in large language models: what exists, what is missing, and what needs to be combined

新研究探讨大型语言模型在临床医学中的评估与应用 · 追踪 8 个来源

多篇研究论文正在探索大型语言模型(LLMs)在临床环境中的评估与应用。其中一篇论文介绍了 STEP-CTS,一种从临床时间序列中为 LLM 预测选择可追溯证据的方法,其表现优于现有的基于文本的基线。另一篇论文回顾了用于评估 LLM 临床推理的现有评分标准,指出了在时间综合和忠实度等方面的差距。此外,研究还在调查临床分诊中开源 LLM 的偏见,开发用于医学图像推理的多模态 LLM,并创建 KlinikeBench 等基准来评估 LLM 的能力,而不仅仅是简单的诊断准确性。LLM 在临床医学中的更广泛应用,包括其可靠性和安全性,也是一个日益关注的领域。 AI

影响 这些研究强调了在医疗保健领域安全有效地部署 LLM,需要更强大的评估框架和专用模型的日益增长的需求。

排序理由 多篇 arXiv 论文介绍了临床环境中 LLM 的新基准、方法和分析。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 8 个来源。 我们如何撰写摘要 →

新研究探讨大型语言模型在临床医学中的评估与应用 · 追踪 8 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇 arXiv 论文介绍了临床环境中 LLM 的新基准、方法和分析。
Source corroboration
8 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, product, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
8 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [8]

  1. arXiv cs.LG TIER_1 English(EN) · Kwanhyung Lee, Juhwan Choi, Jongheon Kim, Joohyung Lee, Hyeongwon Jang, Jeonguk Lee, Jisoo Jung, Eunho Yang ·

    从不规则临床时间序列中学习选择可追溯源证据以进行语言模型预测

    arXiv:2605.20292v2 Announce Type: replace Abstract: Numerical time-series models effectively process irregular electronic health record (EHR) trajectories, but do not expose which temporal patterns support each prediction as readable evidence. Existing text-based interfaces eithe…

  2. arXiv cs.AI TIER_1 English(EN) · Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo ·

    用于评估大型语言模型临床推理的评分标准概览:现有内容、缺失部分及需要整合之处

    arXiv:2610.01938v1 Announce Type: cross Abstract: Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a…

  3. arXiv cs.AI TIER_1 English(EN) · Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang ·

    面向临床分诊的开源大型语言模型中的偏见反事实审计

    arXiv:2610.01963v1 Announce Type: new Abstract: Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) …

  4. arXiv cs.AI TIER_1 English(EN) · Zhen Chen, Yihang Fu, Rong Zhou, Serina Applebaum, Min Kyu Kim, Aidan Gilson, Morten Lee, Salahudeen Mirza, Gabriel Madera, Mauro Giuffre, Yuanting Pan, Roy Jiang, Hyunjae Kim, Hua Xu, Qingyu Chen ·

    从复合图表到医学多图推理:利用生物医学文献扩展多模态大语言模型

    arXiv:2511.22232v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly capable in medical imaging, yet most focus on single-image settings. Clinical interpretation often requires integrating evidence across multiple images, such as dif…

  5. arXiv cs.AI TIER_1 English(EN) · Xueting Fang, Zehui Li, Yang Yang, Camilla Giovino, Shubh K. Patel, Shailly Prajapati, Vallijah Subasri, Caihua Shan ·

    KlinikeBench: 超越诊断准确性评估语言模型

    arXiv:2609.38480v1 Announce Type: cross Abstract: Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and …

  6. arXiv cs.AI TIER_1 English(EN) · Naoto Iwase, Hiroki Okuyama, Junichiro Iwasawa ·

    MedRECT:用于临床文本错误纠正的双语医学推理基准

    arXiv:2511.00421v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show promise in medical applications, but their ability to detect and correct errors in clinical texts remains under-evaluated, particularly beyond English. We introduce MedRECT, a bilingual be…

  7. arXiv cs.AI TIER_1 English(EN) · Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo ·

    大型语言模型响应中表达的临床推理评估拟议标准

    arXiv:2609.37788v1 Announce Type: cross Abstract: Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, K…

  8. arXiv cs.AI TIER_1 English(EN) · Erik Aerts ·

    在临床医学中应用语言模型:近期趋势与展望

    arXiv:2609.34780v2 Announce Type: replace Abstract: The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expande…