A new study, HATEDECIDE, evaluates six structured decision models for hate-speech moderation, comparing them against specialized, zero-shot, commercial, and supervised baselines. The research found that while commercial LLMs performed better on only one dataset, the structured decision models offered significantly lower inference costs. Supplying explicit definitions of hate speech or decomposing the criteria into multiple questions did not consistently improve classification accuracy. AI
IMPACT Identifies opportunities for cost-effective hate-speech moderation using structured decision models.
RANK_REASON The cluster contains an academic paper detailing a new evaluation of models for a specific task. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →