PulseAugur
EN
LIVE 17:15:34

Research: "Hate speech" detectors may misclassify partisan content

A new research paper proposes that current methods for detecting "hate speech" in social media influence operations may be misclassifying partisan or geopolitical content as hate. The study analyzed 25.08 million tweets from seven government-attributed campaigns, developing an LLM-based detector and an auditable rule to differentiate between identity-based attacks, partisan attacks, and geopolitical invective. The findings suggest that reporting all divisive content as hate could be an overestimation, with only a fraction meeting a stricter definition of hate speech. AI

IMPACT This research could refine AI models used for content moderation, leading to more accurate identification of hate speech versus political discourse.

RANK_REASON The cluster contains a research paper published on arXiv detailing new findings and methodologies.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Research: "Hate speech" detectors may misclassify partisan content

COVERAGE [2]

  1. arXiv cs.CL TIER_1 English(EN) · Emilio Ferrara ·

    Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations

    arXiv:2607.14491v1 Announce Type: cross Abstract: State-backed influence operations are routinely measured as high-prevalence sources of ``hate'' and ``toxicity.'' We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definit…

  2. arXiv cs.CL TIER_1 English(EN) · Emilio Ferrara ·

    Manufactured Divisiveness: Decomposing the Hostile Content of Seven Social Media Influence Operations

    State-backed influence operations are routinely measured as high-prevalence sources of ``hate'' and ``toxicity.'' We argue those rates rest on a measurement error: the detectors behind them are validated to catch a broader definition inclusive of hostility or divisiveness aimed a…