Researchers have developed CRATE, a new two-stage framework for evaluating language-guided mobile agents. This framework uses a Visual-Language Model (VLM) as a judge, processing trajectories at a step-level to reason about consequences and state changes, rather than processing the entire trajectory at once. CRATE focuses on both task completion and operational safety, with an extension called CRATE-S specifically for safety assessments. Experiments show CRATE achieves high accuracy on task completion benchmarks, outperforming existing methods, and CRATE-S demonstrates strong alignment with safety benchmarks. AI
IMPACT This framework could improve the development and safety testing of AI agents used in mobile applications.
RANK_REASON The cluster contains a research paper detailing a new framework for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →