A benchmark test of three large language models (LLMs) on clinical care-management tasks revealed significant safety concerns. While Gemini 3.7 Flash, Gemma 4 31B, and GPT-5.4 Mini performed perfectly on medication safety review and triage prioritization, they all failed to include crucial elements in care-plan generation. Notably, all models missed the guideline-recommended depression screening for post-myocardial infarction patients, demonstrating a critical blind spot for psychosocial factors despite generating otherwise fluent and authoritative care plans. AI
IMPACT Highlights critical safety gaps in current LLMs for healthcare, suggesting that fluency does not equate to clinical safety and that specialized evaluation is necessary.
RANK_REASON The item details a new benchmark for evaluating LLMs in a specific domain (clinical care management) and reports findings on their performance and limitations. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →