A new benchmark called WHBench has been developed to evaluate large language models (LLMs) specifically on women's health topics. This benchmark, created by experts, includes 47 scenarios designed to identify critical failures such as outdated medical advice, unsafe omissions, and equity issues. When tested on 22 different LLMs, none achieved an average performance above 75 percent, with even the top-performing models exhibiting significant harm rates and low accuracy. AI
IMPACT Highlights critical safety and equity gaps in current LLMs for specialized medical domains, necessitating further development for safe clinical deployment.
RANK_REASON The cluster describes a new academic paper introducing a novel benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- DagsHub
- Gotit.pub
- Hugging Face
- ScienceCast
- Spandana Govindgari
- WHBench
- Women's Health
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →