PulseAugur
实时 09:04:34
English(EN) ParsHate: A Benchmark Dataset for Hate and Target Detection in Persian

新的ParsHate数据集旨在检测波斯语仇恨言论

研究人员推出了ParsHate,这是一个新基准数据集,旨在检测波斯语中的仇恨言论并识别其目标。该数据集包含2013年至2022年间10,000条手动标注的波斯语推文,是首个涵盖十年内容的此类数据集。它包含31%的仇恨内容,并提供关于显式和隐式仇恨、七个目标类别以及跨度级别理由的详细标注。使用最先进模型的初步评估显示,仇恨言论检测的性能适中,而目标识别的性能则显著较低,这凸显了该数据集的挑战性。 AI

影响 该数据集旨在改善波斯语中仇恨言论及其目标的检测,有可能为波斯语使用者带来更好的审核工具和更安全的在线环境。

排序理由 该项目描述了一个在arXiv上发布的新NLP研究基准数据集。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的ParsHate数据集旨在检测波斯语仇恨言论

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目描述了一个在arXiv上发布的新NLP研究基准数据集。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, safety
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Zahra Bokaei, Walid Magdy, Bonnie Webber ·

    ParsHate:波斯语仇恨与目标检测基准数据集

    arXiv:2609.16393v1 Announce Type: new Abstract: We introduce ParsHate, a manually annotated dataset of 10,000 Persian tweets spanning 2013-2022, representing the first decade-long benchmark for hate speech detection in Persian. The dataset contains 31% hateful content and support…