PulseAugur
EN
LIVE 07:25:24

LLMs struggle with clinical registry abstraction due to ambiguity

A new study published on arXiv evaluates the performance of large language models (LLMs) in abstracting information from clinical registry data. Researchers found that LLMs achieved significantly lower accuracy compared to human abstractors when processing unprocessed electronic medical record data. The LLM's accuracy decreased notably as the ambiguity and clinical reasoning required for the questions increased, performing best on straightforward tasks like medication flagging and worst on complex event timing questions. AI

IMPACT Highlights limitations of current LLMs in complex, high-stakes domains like healthcare, indicating a need for improved reasoning and ambiguity handling.

RANK_REASON The cluster contains an academic paper detailing a study on LLM performance in a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

LLMs struggle with clinical registry abstraction due to ambiguity

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · James Matheson, Betsy Castillo, Andrew Y. Shin, David Scheinker ·

    An ambiguity taxonomy for evaluating large language model performance on clinical registry abstraction: a multi-site prospective study

    arXiv:2608.20373v1 Announce Type: new Abstract: Objective: To evaluate large language model (LLM) performance on unprocessed electronic medical record (EMR) data for clinical registry abstraction. Methods: We evaluated LLM performance answering registry questions for the American…