A new benchmark called EgoMonth has been introduced to evaluate the long-term spatiotemporal memory of multimodal large language models (MLLMs). This benchmark consists of over 300 hours of first-person video recordings from 20 participants, spanning up to 120 days, and includes 1,443 human-crafted questions. Current state-of-the-art models, including Gemini 2.5 Pro, show significant performance gaps compared to human baselines, particularly in tasks requiring route reasoning and spatial judgment. The findings suggest that existing MLLMs act more as lossy summarizers than as faithful memorizers, indicating a need for architectures with improved long-term memory capabilities. AI
IMPACT Highlights limitations in current MLLMs for tasks requiring sustained memory, indicating a need for new architectures.
RANK_REASON The cluster describes a new academic benchmark and evaluation of existing models. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Cross-view Spatial Reasoning
- Direction Judgement
- EgoMonth
- Gemini 2.5 Pro
- Multimodal Large Language Models
- Route Reasoning
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →