HateXplain
PulseAugur coverage of HateXplain — every cluster mentioning HateXplain across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Span-guided vs. unguided AI detoxification: a nuanced comparison
A new research paper explores the effectiveness of span-guided detoxification in AI language models, comparing it to unguided methods. The study found that while span-guided rewriting is preferred when it preserves mean…
-
New methods probe generative models for bias and improve performance
Researchers have developed new methods, Attribution Graphs (AGs) and Causal Probing, to analyze the internal workings of generative models. These techniques aim to identify and correct issues like spurious correlations,…
-
Hate speech annotation pipeline flaw silences minority values
A new research paper highlights a critical flaw in how hate speech datasets are annotated, specifically concerning the boundary between offensive and hateful content. The study reveals that annotator disagreement is not…