A new taxonomy for evaluating Large Language Model (LLM) refusals has been proposed, drawing inspiration from pragmatic theory. This taxonomy aims to assess how appropriately LLMs decline harmful or inappropriate requests, moving beyond a purely safety-alignment perspective. Research applying this taxonomy to 16 LLMs across 14 harm categories revealed that while models generally refuse explicitly and with moral judgment, they often lack nuanced interactional repair, potentially shaming or provoking users, especially in sensitive contexts. The study advocates for alignment evaluations that consider the contextual adaptability and social accountability of LLM refusals. AI
IMPACT This research could lead to more nuanced LLM alignment evaluations, improving user experience by making refusals more contextually appropriate.
RANK_REASON The item is an academic paper proposing a new taxonomy for evaluating LLM refusals. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Influence Flower
- large-language models
- LLM Refusals
- pragmatics
- ScienceCast
- You Shouldn't Have Asked
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →