A new benchmark study has evaluated the reliability of large language models (LLMs) in clinical triage scenarios, comparing their recommendations to those of practicing physicians. The research found that LLMs are more prone to suggesting unnecessary care, and this tendency worsens when clinical text is perturbed. Furthermore, LLM recommendations showed greater sensitivity to irrelevant textual changes, such as gender and tone variations, compared to human physicians. These findings underscore the need for LLM evaluations that are grounded in expert physician behavior and consider real-world variations before deployment in clinical settings. AI
IMPACT Highlights potential risks of LLM deployment in healthcare, emphasizing the need for robust, expert-validated evaluations.
RANK_REASON Academic paper presenting a new benchmark and findings on LLM performance in a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →