PulseAugur
EN
LIVE 09:48:29

New LPS-Bench benchmark reveals safety flaws in AI agents

Researchers have developed LPS-Bench, a new benchmark designed to evaluate the safety awareness of computer-use agents (CUAs) in long-horizon planning tasks. The benchmark addresses the challenge of creating executable environments for diverse scenarios by using a template-guided pipeline for generating instructions, toolkits, and safety criteria, which are then reviewed by humans. LPS-Bench includes 570 cases across 7 task domains and 9 planning-risk types, and an LLM-based evaluator assesses tool choices and responses throughout execution. Initial evaluations of 13 LLM agents using LPS-Bench revealed significant safety failures in both benign and adversarial conditions, with prompt-based interventions showing only model-dependent improvements. AI

IMPACT Highlights persistent safety failures in AI agents, indicating a need for improved safety protocols and evaluation methods in agent development.

RANK_REASON The cluster contains a research paper detailing a new benchmark for evaluating AI safety. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New LPS-Bench benchmark reveals safety flaws in AI agents

How we ranked this

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster contains a research paper detailing a new benchmark for evaluating AI safety. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
safety, paper, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Tianyu Chen, Chujia Hu, Dongrui Liu, Xia Hu, Wenjie Wang ·

    LPS-Bench: Benchmarking Safety Awareness of Computer-Use Agents in Long-Horizon Planning under Benign and Adversarial Scenarios

    arXiv:2602.03255v2 Announce Type: replace Abstract: Computer-use agents (CUAs) execute multi-stage tasks through tools, where an early unsafe decision can propagate to consequential actions. Evaluating only final outcomes can miss such decisions, while constructing executable env…