PulseAugur
EN
LIVE 10:48:01

New EgoMonth benchmark reveals MLLMs struggle with long-term memory

A new benchmark called EgoMonth has been introduced to evaluate the long-term spatiotemporal memory of multimodal large language models (MLLMs). This benchmark consists of over 300 hours of first-person video recordings from 20 participants, spanning up to 120 days, and includes 1,443 human-crafted questions. Current state-of-the-art models, including Gemini 2.5 Pro, show significant performance gaps compared to human baselines, particularly in tasks requiring route reasoning and spatial judgment. The findings suggest that existing MLLMs act more as lossy summarizers than as faithful memorizers, indicating a need for architectures with improved long-term memory capabilities. AI

IMPACT Highlights limitations in current MLLMs for tasks requiring sustained memory, indicating a need for new architectures.

RANK_REASON The cluster describes a new academic benchmark and evaluation of existing models. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New EgoMonth benchmark reveals MLLMs struggle with long-term memory

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Weitao Chen, Hu Jiaxin, Xie Tianyidan, Yang Li, Yuyi Qian, Banghao Xu, Ziheng Tang, Shenyi Wang, Mingyue Yu, Duo Li, Jiacheng Shi, Gao Wang, Zhan Xu, Zhicheng Qiu, Xuanfu Li, Jian Yang, Lanjun Wang, Zili Yi ·

    EgoMonth: A Month-Level Egocentric Video Benchmark for Long-Term Spatiotemporal Memory

    arXiv:2608.13113v1 Announce Type: cross Abstract: Recent advances in Multimodal Large Language Models (MLLMs) have led to substantial progress in video understanding, accompanied by a growing number of long video benchmarks. However, existing benchmarks rely predominantly on web-…