PulseAugur
EN
LIVE 09:32:28

arXiv paper flags criterion contamination in LLM depression evaluations

A new arXiv paper by Tong Li and colleagues argues that "Mirror" evaluations of large language models (LLMs) for depression are contaminated. These evaluations, which use LLM predictions of depression scores based on responses to depression assessments, often report near-perfect accuracy. However, the researchers demonstrate that when LLMs are prompted to predict depression scores using non-mirror (independent life history) interviews, the accuracy significantly decreases, suggesting the original evaluations primarily measure reliability rather than true validity. The study proposes incorporating non-mirror approaches for more clinically relevant LLM applications in depression assessment. AI

IMPACT Highlights potential overestimation of LLM capabilities in sensitive domains like mental health assessment.

RANK_REASON Research paper published on arXiv detailing methodological flaws in LLM evaluations. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

arXiv paper flags criterion contamination in LLM depression evaluations

How we ranked this

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
Research paper published on arXiv detailing methodological flaws in LLM evaluations. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Tong Li, Rasiq Hussain, Mehak Gupta, Joshua R. Oltmanns ·

    "Mirror" Large Language Model Evaluations of Depression are Criterion Contaminated

    arXiv:2508.05830v3 Announce Type: replace Abstract: Large Language Model (LLM) studies that use language responses elicited from depression assessments to predict scores on those same assessments often report near-perfect prediction of depression. We refer to these as "Mirror" ev…