PulseAugur
实时 07:24:22
English(EN) Whose Assessment of Distress? Community Perspectives and LLM Alignment on Well-Being Posts

谁来评估痛苦?社区视角与大型语言模型在福祉帖子上的对齐研究

一项新研究表明,像GPT-5、Gemini 2.5 Pro和Claude Opus 4这样的大型语言模型常常误解在线帖子中的心理痛苦,特别是来自基于身份的社区的帖子。研究人员发现,开放权重模型甚至前沿模型倾向于高估痛苦,尤其是在无痛苦到轻度痛苦的情况下,产生了大量假阳性。这与人类判断形成对比,人类的群体外评估更为平衡,表明模型具有一种 AI

排序理由 [lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

谁来评估痛苦?社区视角与大型语言模型在福祉帖子上的对齐研究

本文如何被排名

Signal score
7 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Andrew Aquilina, Xiang Lorraine Li, Yu-Ru Li ·

    谁来评估痛苦?社区视角与LLM在福祉帖子上的对齐

    arXiv:2608.29446v1 Announce Type: new Abstract: Judgments about psychological distress are socially situated: what counts as concerning hinges on community norms around emotional expression, vulnerability, and help-seeking. Yet large language models (LLMs) used for distress detec…