A systematic scoping review titled "Mental Health Benchmarks for Large Language Models" has identified and analyzed 173 benchmarks. The review found that the majority of these benchmarks utilize data from social media or are generated by LLMs themselves. Notably, only a small fraction, specifically 2%, of the reviewed benchmarks report the use of hidden test sets. AI
IMPACT Highlights limitations in current LLM mental health evaluation, suggesting a need for more robust and diverse datasets.
RANK_REASON The cluster contains an academic paper detailing a systematic review of benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →