Researchers are exploring advanced methods to improve AI's ability to detect hate speech, particularly in multilingual and multimodal contexts. One study focuses on training-time explainability to align AI reasoning with human rationales for better accuracy and interpretability in detecting anti-Muslim hate speech in English and Hinglish. Another paper qualitatively analyzes state-of-the-art vision-language models like LLaVA-7B, Qwen-VL, GPT-4o mini, and Claude 3 Haiku for their effectiveness in identifying hate speech within memes, going beyond simple accuracy to evaluate their justifications. A third study investigates cross-script safety inconsistencies in LLMs for Urdu hate speech detection, revealing significant label instability between original script and English translations, and highlighting a gap in current safety evaluations for this language. AI
IMPACT Advances in multilingual and multimodal hate speech detection could lead to more nuanced and culturally aware content moderation systems.
RANK_REASON The cluster consists of multiple academic papers published on arXiv detailing research into AI safety and model capabilities for hate speech detection.
- arXiv
- LoRA
- Roman Urdu
- Claude Sonnet 4.5
- Gemini 2.5 Flash
- GPT-4o
- Llama-3.1
- Qwen-2.5
- Urdu
- BullySent
- Claude 3 Haiku
- GPT-4o mini
- Grad X Input
- HateXplain
- Integrated Gradients
- LLaVA-7B
- Qwen VL
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →