HateXplain
PulseAugur coverage of HateXplain — every cluster mentioning HateXplain across labs, papers, and developer communities, ranked by signal.
-
AI models struggle with multilingual and meme-based hate speech detection
Researchers are exploring advanced methods to improve AI's ability to detect hate speech, particularly in multilingual and multimodal contexts. One study focuses on training-time explainability to align AI reasoning wit…
-
Span-guided vs. unguided AI detoxification: a nuanced comparison
A new research paper explores the effectiveness of span-guided detoxification in AI language models, comparing it to unguided methods. The study found that while span-guided rewriting is preferred when it preserves mean…
-
New methods probe generative models for bias and improve performance
Researchers have developed new methods, Attribution Graphs (AGs) and Causal Probing, to analyze the internal workings of generative models. These techniques aim to identify and correct issues like spurious correlations,…
-
Hate speech annotation pipeline flaw silences minority values
A new research paper highlights a critical flaw in how hate speech datasets are annotated, specifically concerning the boundary between offensive and hateful content. The study reveals that annotator disagreement is not…