A new psychometric audit of the HELM Safety benchmark, specifically focusing on the HarmBench dataset, suggests that it does not effectively measure a singular attribute like harmful refusal. The analysis employed multidimensional item response theory and differential item functioning, revealing that HarmBench scores do not consistently isolate a single attribute and that models from different developers can score differently despite similar refusal abilities. The study argues that aggregating scores across datasets and items can obscure distinct harm behaviors, and that safety scores should be validated as measuring a single attribute before being used for model comparisons. AI
IMPACT Highlights potential flaws in AI safety evaluation methods, suggesting current benchmarks may not accurately reflect model behavior.
RANK_REASON Academic paper analyzing an AI safety benchmark. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →