PulseAugur
EN
LIVE 12:27:23

New benchmarks TSHA and CAREBench reveal LLM safety gaps

Two new benchmarks have been released to evaluate the safety capabilities of language models. TSHA focuses on assessing visual language models' ability to identify safety hazards in real-world indoor environments, using over 66,000 question-answer pairs. CAREBench, on the other hand, targets language models specifically, evaluating their recognition of upstream child-safety risks beyond explicit abuse material, with 500 prompts across twelve categories. Both benchmarks highlight significant shortcomings in current frontier models' safety awareness. AI

IMPACT These benchmarks will drive improvements in AI safety evaluations, pushing models to better recognize and mitigate risks in real-world scenarios.

RANK_REASON Two academic papers released on arXiv introducing new benchmarks for evaluating AI safety.

Read on arXiv cs.LG →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New benchmarks TSHA and CAREBench reveal LLM safety gaps

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Qiucheng Yu, Ruijie Xu, Mingang Chen Jianfeng Dong, Xin Tan ·

    TSHA: A Benchmark for Visual Language Models in Trustworthy Safety Hazard Assessment Scenarios

    arXiv:2603.29759v2 Announce Type: replace-cross Abstract: Recent advances in vision-language models (VLMs) have accelerated their application to indoor safety hazards assessment. However, existing benchmarks suffer from three fundamental limitations: (1) heavy reliance on synthet…

  2. arXiv cs.LG TIER_1 English(EN) · Kaavya Krishna-Kumar, Elaine Lau, Vaughn Robinson, Jay Caldwell, Sheriff Issaka, Skyler Wang, Francisco Guzm\'an, Steven Kelling, Jonas Mueller ·

    CAREBench: A Child-Safety Risk Benchmark for Language Models

    arXiv:2606.29685v1 Announce Type: new Abstract: How can we evaluate whether frontier AI systems recognize child-safety risks before they escalate into explicit harm? Existing child safety evaluations focus on child sexual abuse material, yet many child-safety failures begin earli…