Researchers have introduced ObGynLongBench, a new benchmark designed to evaluate the performance of large language models (LLMs) in making clinical decisions using longitudinal electronic health records (EHRs). The benchmark, which includes 1,500 cases derived from real pregnancy EHR histories, reveals a significant gap between LLM performance when evidence is directly provided and when it must be extracted from patient records. The study found that model accuracy decreases with longer EHR contexts and more complex evidence requirements, with active-search agents showing the best performance among EHR access strategies. AI
IMPACT Highlights the challenges LLMs face in real-world clinical decision-making with complex patient data, indicating a need for improved evidence extraction capabilities.
RANK_REASON The cluster describes a new benchmark and research paper published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →