A new benchmark, SDARE-Bench, has been developed to evaluate how well Large Language Models (LLMs) can detect and respond to conversational stigma. The benchmark includes 1,138 dyadic queries and 1,388 group dialogue scenarios. Initial testing across eight LLMs revealed significant weaknesses in identifying stigma, particularly in group conversations, where models also exhibited higher rates of stigma expression and provided less realistic advice. In simulated group pressure scenarios, LLMs expressed stigma in 97.5% of responses, highlighting a critical safety vulnerability. AI
IMPACT Highlights a critical safety vulnerability in LLMs, particularly in complex conversational contexts, potentially impacting their deployment in sensitive applications.
RANK_REASON The cluster describes a new academic benchmark for evaluating LLM safety, published on arXiv. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Large Language Models
- ScienceCast
- SDARE-Bench
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →