arXiv:2605.20292v2 Announce Type: replace Abstract: Numerical time-series models effectively process irregular electronic health record (EHR) trajectories, but do not expose which temporal patterns support each prediction as readable evidence. Existing text-based interfaces eithe…
arXiv cs.AI
TIER_1English(EN)·Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo·
arXiv:2610.01938v1 Announce Type: cross Abstract: Exam-style accuracy does not establish whether large language models (LLMs) reason well over clinical records. We define clinical reasoning as integrating and updating evidence across time and sources to form, revise and justify a…
arXiv cs.AI
TIER_1English(EN)·Manar Aljohani, Brandon Ho, Kenneth McKinley, Dennis Ren, Xuan Wang·
arXiv:2610.01963v1 Announce Type: new Abstract: Emergency department (ED) triage is a high-stakes prioritization task in which demographic, socioeconomic, and system-context information may improperly influence acuity assignment. Although open-source large language models (LLMs) …
arXiv cs.AI
TIER_1English(EN)·Zhen Chen, Yihang Fu, Rong Zhou, Serina Applebaum, Min Kyu Kim, Aidan Gilson, Morten Lee, Salahudeen Mirza, Gabriel Madera, Mauro Giuffre, Yuanting Pan, Roy Jiang, Hyunjae Kim, Hua Xu, Qingyu Chen·
arXiv:2511.22232v2 Announce Type: replace-cross Abstract: Multimodal large language models (MLLMs) are increasingly capable in medical imaging, yet most focus on single-image settings. Clinical interpretation often requires integrating evidence across multiple images, such as dif…
arXiv cs.AI
TIER_1English(EN)·Xueting Fang, Zehui Li, Yang Yang, Camilla Giovino, Shubh K. Patel, Shailly Prajapati, Vallijah Subasri, Caihua Shan·
arXiv:2609.38480v1 Announce Type: cross Abstract: Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and …
arXiv:2511.00421v2 Announce Type: replace-cross Abstract: Large language models (LLMs) show promise in medical applications, but their ability to detect and correct errors in clinical texts remains under-evaluated, particularly beyond English. We introduce MedRECT, a bilingual be…
arXiv cs.AI
TIER_1English(EN)·Zhangshu Joshua Jiang, Zina Ibrahim, James T. Teo·
arXiv:2609.37788v1 Announce Type: cross Abstract: Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, K…
arXiv:2609.34780v2 Announce Type: replace Abstract: The use and applicability of artificial intelligence (AI) in medical research and clinical practice has received increasing attention in the literature over recent years. The emergence of large language models (LLMs) has expande…