PulseAugur
中
实时 08:45:13
English(EN) OpenMTB-Audit: Exposing Over-Refusal and Clinical Expert Perspectives in LLM-Based Molecular Tumor Board Safety Evaluation

新基准揭示LLM在癌症治疗建议中存在过度拒绝问题

一项名为OpenMTB-Audit的新开源基准已被开发出来,用于评估大型语言模型(LLM)在分子肿瘤委员会(MTB)工作流程中的安全性。该基准包含500个合成的非小细胞肺癌病例,揭示出当前的LLM存在显著的过度拒绝问题,未能正确识别部分支持的治疗建议。为解决此问题,创建了一个名为MTB-AuditAgent的框架,该框架将过度拒绝率降低至6.7%,并在安全性分类方面达到91.2%的准确率。 AI

影响 凸显了LLM在医疗应用中的关键安全局限性,需要专门的框架来实现可靠的临床决策支持。

排序理由 该集群包含一篇学术论文,详细介绍了用于评估特定领域LLM安全性的新基准和框架。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准揭示LLM在癌症治疗建议中存在过度拒绝问题

本文如何被排名

Signal score
16 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇学术论文,详细介绍了用于评估特定领域LLM安全性的新基准和框架。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Negin Ashrafi, Jia Luo, Stacey M. Frumm, Roxana Daneshjou ·

    OpenMTB-Audit:揭示基于LLM的分子肿瘤委员会安全评估中的过度拒绝和临床专家视角

    arXiv:2610.01497v1 Announce Type: new Abstract: Molecular tumor boards integrate genomic findings, clinical context, and therapeutic evidence to support precision oncology. As AI enters this workflow, a key safety challenge is distinguishing truly unsupported recommendations from…