PulseAugur
实时 07:16:38
English(EN) SWARM: A Multilingual Human-Annotated Dataset for Russian Propaganda Detection in Search Engine Results

新的SWARM数据集针对多语言搜索结果中的俄罗斯宣传

研究人员推出SWARM,一个新推出的多语言数据集,旨在检测搜索引擎结果中的俄罗斯宣传。该数据集包含九种语言的2,183个搜索引擎结果,每个条目都由训练有素的编码员标注,以识别对反复出现的俄罗斯宣传叙事的支持。使用基于来源的阻止列表、监督分类器和大型语言模型(LLM)进行的基准测试显示,阻止列表不足以应对宣传出现在主流网站上,而不仅仅是标记的出口。虽然内容级分析更有效,但LLM的F1分数(0.73)高于监督分类器(约0.5),并且较小的LLM有时会混淆主题相关性与认可。 AI

影响 该数据集可以提高AI模型识别和减轻国家支持的虚假信息在不同语言和在线平台上传播的能力。

排序理由 该集群包含一篇介绍新数据集和评估方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的SWARM数据集针对多语言搜索结果中的俄罗斯宣传

本文如何被排名

Signal score
23 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇介绍新数据集和评估方法的论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Manuel Tonneau, Abhinav Dubey, Farhan Shaikh, Ilaria Vitulano, Martha Stolze, Hale Dedeoglu, Clara Riechert, Ella Kuka, Maryna Sydorova, Mykola Makhortykh, Elizaveta Kuznetsova ·

    SWARM:搜索引擎结果中俄语宣传检测的多语言人工标注数据集

    arXiv:2609.12653v1 Announce Type: new Abstract: Russian state propaganda spreads across many languages and online spaces. Yet, most computational work examines only one such space, usually social media, in one or two languages, and analyses sources rather than content. We introdu…