A new arXiv paper investigates how AI assistants handle repeated verbal abuse, differentiating between hard disengagement and soft withdrawal. The study found significant variation among models like Gemini-3.1 Pro, GPT 5.6 "Sol", and Claude Fable-5 in their responses to escalating abuse. Gemini-3.1 Pro exhibited the highest rate of hard disengagement, while Claude Fable-5 showed a strong tendency towards soft withdrawal, continuing to offer assistance even when not performing substantive work. AI
IMPACT Highlights critical differences in AI safety mechanisms and the need for nuanced evaluation beyond simple refusal metrics.
RANK_REASON The cluster contains a research paper published on arXiv detailing experimental findings on AI assistant behavior. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Claude Fable-5
- Claude Opus 4-8
- DagsHub
- Gemini-3.1 Pro
- Gotit.pub
- GPT 5.6 "Sol"
- Hugging Face
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →