A new benchmark called AndroidLife tested an AI agent's ability to perform real-world tasks on a smartphone, with the Qwen3.8-27B model achieving a 56.7% success rate. The test involved 60 sequential tasks on a OnePlus phone, resulting in a peak chip temperature of 98.2°C and a 69% battery drain. While Qwen3.8-27B set a new benchmark for text-based models, it struggled with multi-step app interactions and accurately retrieving specific user information. AI
IMPACT Highlights the current limitations of AI agents in performing complex, multi-step tasks on mobile devices.
RANK_REASON The item describes a new benchmark for evaluating AI agents on real-world tasks. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →