A new benchmark called HarmReduction has been developed to evaluate the accuracy and safety of large language models (LLMs) in providing harm reduction information for individuals who use drugs. The benchmark includes 2,160 question-answer-evidence pairs across three tasks: assessing safety boundaries, providing quantitative data, and inferring polysubstance use risks. Initial results show that current state-of-the-art LLMs struggle with accuracy and can pose significant safety risks to users seeking this sensitive information. AI
IMPACT Highlights the critical need for specialized LLM evaluation in sensitive domains like public health to prevent harm.
RANK_REASON The cluster contains an academic paper introducing a new benchmark for evaluating LLMs. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →