Two new benchmarks have been released to evaluate the safety capabilities of language models. TSHA focuses on assessing visual language models' ability to identify safety hazards in real-world indoor environments, using over 66,000 question-answer pairs. CAREBench, on the other hand, targets language models specifically, evaluating their recognition of upstream child-safety risks beyond explicit abuse material, with 500 prompts across twelve categories. Both benchmarks highlight significant shortcomings in current frontier models' safety awareness. AI
IMPACT These benchmarks will drive improvements in AI safety evaluations, pushing models to better recognize and mitigate risks in real-world scenarios.
RANK_REASON Two academic papers released on arXiv introducing new benchmarks for evaluating AI safety.
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →