A new research paper proposes that current methods for detecting "hate speech" in social media influence operations may be misclassifying partisan or geopolitical content as hate. The study analyzed 25.08 million tweets from seven government-attributed campaigns, developing an LLM-based detector and an auditable rule to differentiate between identity-based attacks, partisan attacks, and geopolitical invective. The findings suggest that reporting all divisive content as hate could be an overestimation, with only a fraction meeting a stricter definition of hate speech. AI
IMPACT This research could refine AI models used for content moderation, leading to more accurate identification of hate speech versus political discourse.
RANK_REASON The cluster contains a research paper published on arXiv detailing new findings and methodologies.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →