A new study published on arXiv evaluates the performance of large language models (LLMs) in abstracting information from clinical registry data. Researchers found that LLMs achieved significantly lower accuracy compared to human abstractors when processing unprocessed electronic medical record data. The LLM's accuracy decreased notably as the ambiguity and clinical reasoning required for the questions increased, performing best on straightforward tasks like medication flagging and worst on complex event timing questions. AI
IMPACT Highlights limitations of current LLMs in complex, high-stakes domains like healthcare, indicating a need for improved reasoning and ambiguity handling.
RANK_REASON The cluster contains an academic paper detailing a study on LLM performance in a specific domain. [lever_c_demoted from research: ic=1 ai=1.0]
- American College of Cardiology National Cardiovascular Data Registry
- Hugging Face
- large language model
- Medication/Event Flag
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →