A new study published on arXiv investigated the reliability of language models in assessing depression symptoms. Researchers found that the choice of language model and prompting strategy significantly influenced the depression scores, accounting for 30% of the variance in summed-symptom scores. Even among models with high accuracy, two randomly selected raters disagreed on screening decisions for approximately 40% of participants. While calibration improved accuracy and reduced disagreement, a substantial portion of individuals were still assessed differently by various raters. AI
IMPACT Highlights the need for robust calibration and standardization in LLM-driven diagnostic tools to ensure consistent and reliable patient assessments.
RANK_REASON The cluster contains a research paper published on arXiv detailing findings about language model capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Language Models
- Patient Health Questionnaire
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →