PulseAugur
实时 14:19:23
English(EN) When Hate Meets Facts: LLMs-in-the-Loop for Check-worthiness Detection in Hate Speech

新的大语言模型框架应对隐性及基于事实的仇恨言论检测 · 跟踪 2 个来源

研究人员开发了新的框架,通过整合大语言模型 (LLM) 来更有效地检测仇恨言论。一种方法 WSF-ARG+ 引入了一个数据集和一个循环中的大语言模型系统,用于识别使用事实类(尽管不正确)信息的仇恨言论,提高了检测准确性并减少了人工标注工作。另一个框架 FAID 则通过将隐性仇恨言论分为浅层、目标型和依赖上下文三种形式,然后对每种形式应用自适应检测策略,以提高效率和准确性。 AI

影响 这些基于大语言模型的框架在识别细微的仇恨言论方面提供了更高的准确性和效率,可能有助于在线内容审核工作。

排序理由 该集群包含两篇在 arXiv 上发表的学术论文,详细介绍了使用大语言模型检测仇恨言论的新颖框架。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的大语言模型框架应对隐性及基于事实的仇恨言论检测 · 跟踪 2 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含两篇在 arXiv 上发表的学术论文,详细介绍了使用大语言模型检测仇恨言论的新颖框架。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
11 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [3]

  1. arXiv cs.CL TIER_1 English(EN) · Manuel Tonneau, Dylan Thurgood, Diyi Liu, Niyati Malhotra, Victor Orozco-Olvera, Ralph Schroeder, Scott A. Hale, Manoel Horta Ribeiro, Paul R\"ottger, Samuel P. Fraiberger ·

    仇恨言论审核的执行与可行性

    arXiv:2604.12289v2 Announce Type: replace-cross Abstract: Online hate speech is associated with harms ranging from deteriorating mental health to violence, yet how consistently platforms moderate hate, and whether enforcement is feasible at scale, remain poorly understood. We aud…

  2. arXiv cs.CL TIER_1 English(EN) · Nicol\'as Benjam\'in Ocampo, Tommaso Caselli, Davide Ceolin ·

    当仇恨遭遇事实:在仇恨言论中利用“循环中的大型语言模型”进行可核查性检测

    arXiv:2603.25269v2 Announce Type: replace Abstract: Hateful content online is often expressed using fact-like, not necessarily correct information, especially in coordinated online harassment campaigns and extremist propaganda. Failing to jointly address hate speech (HS) and misi…

  3. arXiv cs.CL TIER_1 English(EN) · Han Wang, Yuhu Cheng, Xuesong Wang, Yi Zhu ·

    铁锤还是手术刀?用于隐性仇恨言论的细粒度自适应框架

    arXiv:2608.27462v1 Announce Type: new Abstract: Unlike explicit attacks with obvious profanity, implicit hate speech hides malice within seemingly compliant expressions through metaphors and contextual hints, making its detection in online content review challenging. While existi…