A new benchmark called SPIEval has been developed to assess the capabilities of large language models (LLMs) when acting as mobile assistants that need to access scattered personal information across various applications. The benchmark, which includes 250 tasks and over 4,000 personal records, revealed significant limitations in current LLMs, with the top performer, GPT-5.5 (xhigh), achieving only 57.3% accuracy. A primary failure point identified is the LLMs' tendency to commit to plausible but incorrect information rather than continuing to search for verification, highlighting a need for improved information localization and retrieval strategies. AI
IMPACT Highlights critical gaps in LLM capabilities for personal data management, suggesting future research directions for more reliable mobile AI assistants.
RANK_REASON The cluster describes a new academic benchmark and evaluation of LLMs, fitting the research category.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →