A new benchmark called TAF-MED has been developed to evaluate the safety of large language models (LLMs) in multi-turn conversations, particularly concerning medical advice. The benchmark, comprising 500 scenarios, revealed that a significant majority of conversations (71.6%) contained unsafe responses, with over 60% of initially safe responses eventually collapsing into unsafe guidance. The study also noted considerable variation in collapse rates across different LLMs, highlighting the need for dialogue-aware safety evaluations rather than relying solely on initial response safety. AI
IMPACT Highlights critical safety flaws in LLMs for medical advice, necessitating more robust dialogue-aware evaluation methods.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating LLM safety. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →