PulseAugur
EN
LIVE 14:50:12

ToxGate improves multilingual toxicity detection by conditioning signals

Researchers have developed ToxGate, a novel trust-fusion head designed to improve the reliability of toxicity detection systems, particularly for multilingual and code-mixed text. Traditional moderation tools struggle with variations like transliteration, slang, and language mismatches. ToxGate addresses this by conditioning auxiliary toxicity signals on encoder representations before integrating them into the prediction state. Experiments across multiple datasets and encoders demonstrated that ToxGate significantly enhances performance in high-risk moderation scenarios, including explicit slurs and violent threats, offering a more nuanced approach to evidence-based moderation. AI

IMPACT Enhances the reliability of AI-driven content moderation systems for diverse linguistic contexts.

RANK_REASON Academic paper detailing a new method for toxicity detection. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

ToxGate improves multilingual toxicity detection by conditioning signals

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Indraveni Chebolu, Rohan Singh, Arnab Mallick, Harmesh Rana ·

    Conditional Reliability of Toxicity Signals for Multilingual and Code-Mixed Abuse Detection

    arXiv:2607.15861v1 Announce Type: cross Abstract: Moderation systems increasingly rely on external toxicity tools, but those tools are unreliable under code-mixing, transliteration, slang, and language mismatch. We study the \emph{conditional reliability} of toxicity priors in In…