Researchers have introduced ElderBench, a new benchmark designed to evaluate autonomous mobile agents intended to assist older adults with smartphone usage. Existing benchmarks often fail to capture the nuanced and indirect language patterns common among older users, leading to performance degradation in current agents. ElderBench is built upon 249 naturally elicited smartphone tasks from older adults, allowing for a more authentic assessment of agent capabilities. The findings highlight significant challenges for mainstream agents and Vision-Language Models when handling elderly-specific instructions, offering insights for developing more adaptive and age-inclusive AI assistants. AI
IMPACT This benchmark could lead to more effective AI assistants tailored to the needs of older adults.
RANK_REASON The item is a research paper introducing a new benchmark for AI agents. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- ElderBench
- Gotit.pub
- Hugging Face
- Influence Flower
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →