Researchers have introduced StepJack, a new benchmark designed to test the safety of computer-use agents (CUAs) against multi-step indirect prompt injection attacks. These attacks involve distributing adversarial instructions across a chain of web pages, making them appear innocuous individually. Evaluations on the StepJack benchmark, which includes 480 test examples, revealed that multi-step attacks significantly increased the success rate for several state-of-the-art CUAs, including GPT-5.4 Mini, by up to 31.2 percentage points. AI
IMPACT This research highlights a new vulnerability in AI agents, potentially requiring developers to implement more robust safety measures against sophisticated prompt injection techniques.
RANK_REASON The cluster describes a new academic paper introducing a novel benchmark and attack class for AI safety research. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Computer Use Agents
- EvoCUA-32B
- GPT 5.4 Mini
- indirect prompt injection
- multi-step indirect prompt injection
- StepJack
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →