Researchers have developed a new framework called SeTox to detect implicit toxicity in Chinese neologisms. This framework uses search-augmented large language models (LLMs) to incorporate real-time web context, enabling static LLMs to identify toxic neologisms that have evolved in public consensus. Experiments demonstrated that SeTox, even with smaller 3B-scale models, outperforms larger models in detecting this type of nuanced toxicity. AI
IMPACT This research could improve content moderation systems by enabling them to detect subtle forms of toxicity in evolving language.
RANK_REASON The cluster contains an academic paper detailing a new method for toxicity detection. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →