A new paper from arXiv explores user mistreatment of conversational AI systems, analyzing over 777,000 conversations from the LMSYS-Chat-1M dataset. The research found that approximately 5% of user turns exhibited hostility, insults, threats, or coercion directed at the AI. Interestingly, user hostility varied significantly across different models, seemingly due to the user base attracted to each model rather than the model's own behavior. The study also noted that AI apologies were associated with increased user hostility, yet models that apologized more frequently overall received less hostility. AI
IMPACT Highlights the need for robust safety measures that account for user-to-AI mistreatment, not just AI-to-user harms.
RANK_REASON The cluster contains an academic paper detailing research findings. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →