A new study published on arXiv investigates how sensitive large language models are to prompt, model, and schema choices when extracting structured data from clinical notes. Researchers found that while prompt variations had a moderate impact on categorization, the choice of model size significantly influenced the reassignment of primary admission tags. The study also revealed that collapsing the extraction schema to a binary format resolved most disagreements, highlighting that the distinction between absence and silence was a key factor in the model's interpretation. AI
IMPACT Highlights the need for careful configuration and validation of LLMs in sensitive domains like healthcare to ensure reliable data extraction.
RANK_REASON Academic paper detailing methodology and findings on LLM performance.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →