研究人员开发了PatientAgentBench,一个用于评估旨在与医疗保健环境中患者互动的人工智能代理的新框架。该基准使用LLM-as-a-Jury系统,在六个维度上评估代理,包括分诊质量和临床安全性,该系统与持证临床医生高度一致。初步测试显示,十种不同AI模型的能力存在显著差距,这凸显了随着这些系统变得更加自主,需要超越静态基准的更强大的评估方法。同时,对医学领域代理式AI的审查强调了临床转化的挑战,呼吁更清晰的定义、可重复的评估以及在真实工作流程中的前瞻性验证。
AI
arXiv:2607.25485v1 Announce Type: new Abstract: Health AI is evolving from answering questions to agentic systems that converse with patients, reason about health records, and act on their behalf. Primary care guards against diagnostic errors and unsafe care; agents assisting in …
arXiv:2607.25489v1 Announce Type: new Abstract: Large language models and multimodal foundation models are enabling medical artificial intelligence (AI) systems to move beyond isolated prediction and undertake multistep clinical tasks that require planning, tool use, memory, iter…
<p>Modern AI systems are becoming increasingly capable of reasoning over complex clinical information. They can summarize medical literature, generate differential diagnoses, connect symptoms to diseases, and assist clinicians in navigating enormous amounts of evidence.</p> <p>Bu…