Researchers have introduced AcuityBench, a new benchmark designed to evaluate the ability of large language models to accurately identify the urgency of medical situations presented by users. Unlike existing health benchmarks that focus on question answering or triage, AcuityBench harmonizes five datasets into a unified four-level acuity framework, encompassing 914 cases. Initial testing across 12 frontier models revealed significant variations in performance, with models showing a trade-off between over-triage and under-triage depending on the response format. The benchmark also highlighted that current models do not closely match physician judgments on ambiguous cases and exhibit less nuanced uncertainty. AI
IMPACT Highlights a critical safety gap in current LLMs for healthcare applications, necessitating further research into uncertainty alignment.
RANK_REASON Academic paper introducing a new benchmark for LLM evaluation. [lever_c_demoted from research: ic=1 ai=1.0]
- AcuityBench
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Litmaps
- Robin Linzmayer
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →