A new paper introduces NarrativeShield SDoH MedQA, a dataset designed to evaluate bias in medical large language models. The study assesses how models respond to the same clinical case presented with different patient narrative styles, focusing on "SDoH aware narrative anchoring bias." Three models from the Qwen2.5 family were tested, with the 7B version showing the best accuracy and consistency, though significant narrative sensitivity errors persisted. AI
IMPACT Highlights the need for robust evaluation of medical LLMs beyond simple accuracy, focusing on fairness and reliability in clinical decision support.
RANK_REASON Research paper introducing a new dataset and evaluation methodology for LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →