PulseAugur
实时 18:19:39
English(EN) Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations

研究:“仇恨言论”检测器可能错误分类党派内容

一项新的研究论文提出,当前用于检测社交媒体影响行动中“仇恨言论”的方法可能将党派或地缘政治内容错误地归类为仇恨言论。该研究分析了来自七个政府归因活动中的 2508 万条推文,开发了一个基于 LLM 的检测器和一个可审计的规则,以区分基于身份的攻击、党派攻击和地缘政治的谩骂。研究结果表明,将所有分裂性内容报告为仇恨言论可能是一种高估,只有一小部分符合更严格的仇恨言论定义。 AI

影响 这项研究可以改进用于内容审核的 AI 模型,从而更准确地识别仇恨言论与政治言论。

排序理由 该集群包含一篇发表在 arXiv 上的研究论文,详细介绍了新的发现和方法论。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

研究:“仇恨言论”检测器可能错误分类党派内容

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Emilio Ferrara ·

    制造分裂:解构七个社交媒体影响行动中的敌对内容

    arXiv:2607.14491v1 Announce Type: cross Abstract: State-backed influence operations are routinely measured as high-prevalence sources of ``hate'' and ``toxicity.'' We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definit…

  2. arXiv cs.CL TIER_1 English(EN) · Emilio Ferrara ·

    制造分裂:分解七个社交媒体影响行动的敌对内容

    State-backed influence operations are routinely measured as high-prevalence sources of ``hate'' and ``toxicity.'' We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definition inclusive of hostility or divisiveness aimed a…