Researchers have introduced ParsHate, a new benchmark dataset designed for detecting hate speech and identifying its targets within the Persian language. This dataset, comprising 10,000 manually annotated Persian tweets from 2013 to 2022, is the first of its kind to cover a decade of content. It includes 31% hateful material and offers detailed annotations for explicit and implicit hate, seven target categories, and span-level rationales. Initial evaluations using state-of-the-art models show moderate performance for hate speech detection and significantly lower performance for target identification, highlighting the dataset's challenging nature. AI
IMPACT This dataset aims to improve the detection of hate speech and its targets in Persian, potentially leading to better moderation tools and safer online environments for Persian speakers.
RANK_REASON The item describes a new benchmark dataset for NLP research published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ParsHate
- Persian
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →