A new benchmark called GMA has been developed to evaluate general mobile AI assistants in complex real-world scenarios. GMA includes seven applications and 300 tasks across four difficulty levels, designed to capture more diversity and complexity than existing benchmarks like AndroidWorld and MobileWorld. Evaluations of eight frontier models revealed significant performance drops as task complexity increased, indicating current agents are far from reliably meeting user needs. The research also highlighted that appropriate harness design, such as context retention and explicit state tracking, can notably improve agent performance on demanding workflows, though effectiveness varies by foundation model. AI
IMPACT Highlights the need for more robust AI agents capable of handling complex, real-world mobile tasks.
RANK_REASON The cluster contains a research paper introducing a new benchmark for evaluating AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →