PulseAugur
实时 10:14:37

Vision-language models struggle with institutional agreement on chest X-rays

研究人员评估了三种视觉语言模型(VLM)在不同机构准确识别胸部X光片(chest radiographs)病灶的能力。研究发现,现有的VLM缺乏置信度分数,导致接收机构难以衡量单个预测的可靠性。在跨三个语料库和六种病灶的超过345,000次预测中进行测试时,即使有少量本地标签用于估算,模型的性能也因地点和接口而异,这表明需要进行特定地点的重新评估,而不是依赖通用的默认估算器。 AI

影响 强调了在医疗AI中进行特定地点校准和置信度评分的必要性,以确保在不同医疗保健环境中的可靠部署。

排序理由 该集群包含一篇详细介绍AI模型性能研究结果的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.LG 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Vision-language models struggle with institutional agreement on chest X-rays

报道来源 [1]

  1. arXiv cs.LG TIER_1 English(EN) · Pengyang Yu, Yiou Wang, Zhongping Dong, Sahraoui Dhelim, Chun-Mei Feng, M. Tahar Kechadi ·

    审计胸部放射影像的医学视觉语言模型:估算跨机构的参考一致性

    arXiv:2608.07550v1 Announce Type: cross Abstract: Vision-language models return structured chest-radiograph findings through interfaces exposing no confidence score, so a receiving institution cannot read off how far to trust an individual judgment. Whether agreement with an inst…