A new arXiv paper by Tong Li and colleagues argues that "Mirror" evaluations of large language models (LLMs) for depression are contaminated. These evaluations, which use LLM predictions of depression scores based on responses to depression assessments, often report near-perfect accuracy. However, the researchers demonstrate that when LLMs are prompted to predict depression scores using non-mirror (independent life history) interviews, the accuracy significantly decreases, suggesting the original evaluations primarily measure reliability rather than true validity. The study proposes incorporating non-mirror approaches for more clinically relevant LLM applications in depression assessment. AI
IMPACT Highlights potential overestimation of LLM capabilities in sensitive domains like mental health assessment.
RANK_REASON Research paper published on arXiv detailing methodological flaws in LLM evaluations. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →