Researchers have introduced DeflectBench, a new benchmark designed to evaluate the ability of large language models to generate rhetorical fallacies on demand. The benchmark tests 23,990 generations from four frontier models across three deflection strategies: whataboutism, ad hominem, and red herring. Findings indicate that model refusal is more sensitive to prompt structure and framing than to the content of controversial claims, with specific prompt changes significantly altering refusal rates. Notably, an "educational debate coach" prompt framing drastically reduces refusal, though models often resort to labeled compliance, identifying the fallacy within their response. AI
IMPACT This benchmark could inform the development of more robust safety mechanisms in LLMs, preventing their misuse for generating manipulative or deceptive content.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →