PulseAugur
实时 07:43:19
English(EN) "Mental Health Benchmarks for Large Language Models: A Systematic Scoping Review" maps 173 benchmarks. Most rely on social media or LLM-generated data; only 2%

审查梳理173个LLM心理健康基准,发现数据局限性

一项题为《大型语言模型心理健康基准》的系统性范围审查,识别并分析了173个基准。审查发现,这些基准中的大多数利用社交媒体数据或由LLM本身生成。值得注意的是,在审查的基准中,仅有很小一部分,具体为2%,报告使用了隐藏测试集。 AI

影响 强调了当前LLM心理健康评估的局限性,表明需要更强大和多样化的数据集。

排序理由 该集群包含一篇详细介绍基准系统性审查的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

审查梳理173个LLM心理健康基准,发现数据局限性

本文如何被排名

Signal score
13 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍基准系统性审查的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    大型语言模型心理健康基准:系统性范围审查,梳理173个基准。多数依赖社交媒体或LLM生成数据;仅2%

    "Mental Health Benchmarks for Large Language Models: A Systematic Scoping Review" maps 173 benchmarks. Most rely on social media or LLM-generated data; only 2% report hidden test sets. # LLM # MentalHealth # AI https:// doi.org/10.17605/OSF.IO/CZB7V