English(EN)Incentivizing Temporal-Awareness in Egocentric Video Understanding Models
新的基准和模型推动了人工智能中以自我为中心的视频理解能力
作者PulseAugur 编辑部·[10 个来源]·
研究人员正在开发新的方法和基准来提高多模态大语言模型(MLLMs)的时间和空间推理能力,特别是针对以自我为中心的视频理解。论文介绍了时间全局策略优化(TGPO)等技术来增强时间感知能力,以及Whareformer等模型用于跟踪长时以自我为中心的视频中的对象。新的基准,如EgoPolice和EgoExoMem,正在被创建,以在具有挑战性的数据集上评估这些模型,包括警方佩戴的摄像机录像和同步的以自我为中心/以外的视频对,突显了即使是像Gemini 2.5 Pro这样的先进模型也存在局限性。
AI
Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning depends on the correct ordering and evolution of events. This deficiency stems in p…
The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become h…
arXiv:2607.08537v1 Announce Type: new Abstract: The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video ev…
arXiv:2607.08514v1 Announce Type: new Abstract: Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, obje…
The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become h…
Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, objects, or environmental context, rather than reaso…
arXiv:2605.18734v2 Announce Type: replace Abstract: Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and observer perspectives, we introduce EgoExoMem, the …
arXiv:2501.05711v4 Announce Type: replace Abstract: Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant training limits their ability to infer egocentric properties, such as human-obje…
arXiv cs.CV
TIER_1English(EN)·Max Gonzalez Saez-Diez, Jihoon Chung, Adam D. Wolsky, Gregory Lanzalotto, Dean Knox, Jonathan Mummolo, Brandon M. Stewart, Olga Russakovsky·
arXiv:2607.06468v1 Announce Type: new Abstract: We introduce EgoPolice, a carefully curated dataset of real, egocentric police-civilian interactions, sourced from publicly available body-worn camera videos. We select police-civilian action labels that are critical for police beha…
We introduce EgoPolice, a carefully curated dataset of real, egocentric police-civilian interactions, sourced from publicly available body-worn camera videos. We select police-civilian action labels that are critical for police behavioral research and annotate them at a second-by…