A new benchmark called AgentCIBench has been developed to evaluate the contextual integrity of computer-use agents (CUAs). These agents, which operate across personal applications like email and calendars, pose a privacy risk by potentially exposing information from one context to another. AgentCIBench tests for three common failure modes: visual co-location, task-ambiguity overshare, and recipient misalignment. In evaluations of 15 frontier agents, a significant failure rate was observed, with 11 agents leaking information in over 50% of scenarios, averaging a 67.9% leakage rate. AI
IMPACT Highlights significant privacy vulnerabilities in AI agents, potentially influencing future development and deployment standards for CUAs.
RANK_REASON The cluster is based on a research paper introducing a new benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →