PulseAugur
实时 06:35:23
English(EN) SDARE-Bench: Evaluating Large Language Models on Conversational Stigma Detection and Response in Dyadic and Group Dialogue

新基准揭示大型语言模型难以应对对话污名

一个名为 SDARE-Bench 的新基准已被开发出来,用于评估大型语言模型 (LLM) 在检测和响应对话污名方面的能力。该基准包含 1,138 个二元查询和 1,388 个群体对话场景。对八个 LLM 的初步测试显示,它们在识别污名方面存在显著弱点,尤其是在群体对话中,模型还表现出更高的污名表达率,并提供不太真实的建议。在模拟的群体压力场景中,LLM 在 97.5% 的响应中表达了污名,凸显了关键的安全漏洞。 AI

影响 凸显了 LLM 在复杂对话环境中存在的关键安全漏洞,可能影响其在敏感应用中的部署。

排序理由 该集群描述了一个用于评估 LLM 安全性的新学术基准,发布在 arXiv 上。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准揭示大型语言模型难以应对对话污名

本文如何被排名

Signal score
29 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一个用于评估 LLM 安全性的新学术基准,发布在 arXiv 上。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Stephanie Fong, Yiwen Jiang, Zimu Wang, Hongxi Yang, Yaling Shen, Hiu Weh Naomi Chow, Heung Ying Lai, Xiangyu Zhao, Qingyang Xu, Zhongxing Xu, Jiahe Liu, Guilherme C. Oliveira, Vincent Lee, Zongyuan Ge, Dominic Dwyer ·

    SDARE-Bench:在二元和群体对话中评估大型语言模型在对话污名检测和响应方面的能力

    arXiv:2609.01548v1 Announce Type: new Abstract: Large Language Models (LLMs) are increasingly used in advice seeking and decision making that may affect social judgements. Despite stigma's profound effects on people and communities, benchmarks remain scarce. Existing general-doma…