PulseAugur
实时 11:44:18
English(EN) Incentivizing Temporal-Awareness in Egocentric Video Understanding Models

新的基准和模型推动了人工智能中以自我为中心的视频理解能力

研究人员正在开发新的方法和基准来提高多模态大语言模型(MLLMs)的时间和空间推理能力,特别是针对以自我为中心的视频理解。论文介绍了时间全局策略优化(TGPO)等技术来增强时间感知能力,以及Whareformer等模型用于跟踪长时以自我为中心的视频中的对象。新的基准,如EgoPolice和EgoExoMem,正在被创建,以在具有挑战性的数据集上评估这些模型,包括警方佩戴的摄像机录像和同步的以自我为中心/以外的视频对,突显了即使是像Gemini 2.5 Pro这样的先进模型也存在局限性。 AI

影响 以自我为中心的视频分析的进步可以通过对动态环境进行更细致的理解来改进具身人工智能、机器人技术和监控技术。

排序理由 多篇研究论文介绍了用于以自我为中心的视频理解的新模型、算法和基准。

在 Apple Machine Learning Research 阅读 →

AI 生成摘要 · Google Gemini · 来自 10 个来源。 我们如何撰写摘要 →

新的基准和模型推动了人工智能中以自我为中心的视频理解能力

报道来源 [10]

  1. Apple Machine Learning Research TIER_1 English(EN) ·

    激励以自我为中心的视频理解模型中的时间感知能力

    Multimodal large language models (MLLMs) have recently shown strong performance in visual understanding, yet they often lack temporal awareness, particularly in egocentric settings where reasoning depends on the correct ordering and evolution of events. This deficiency stems in p…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    Whareformer:学习追踪长片段自我中心视频中的“何物”与“何处”

    The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become h…

  3. arXiv cs.CV TIER_1 English(EN) · Jacob Chalk, Saptarshi Sinha, Dima Damen, Yannis Kalantidis, Diane Larlus ·

    Whareformer:学习追踪长篇主观视频中的“何物”与“何处”

    arXiv:2607.08537v1 Announce Type: new Abstract: The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video ev…

  4. arXiv cs.CV TIER_1 English(EN) · Masatoshi Tateno, Alexandros Stergiou, Risa Shinoda, Yoichi Sato, Dima Damen ·

    以自我为中心的视频语言模型能同时捕捉手部和物体线索吗?

    arXiv:2607.08514v1 Announce Type: new Abstract: Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, obje…

  5. arXiv cs.CV TIER_1 English(EN) · Diane Larlus ·

    Whareformer:学习追踪长片段自我中心视频中的“何物”与“何处”

    The recently established 'Out of Sight, Not out of Mind' (OSNOM) task for egocentric videos focuses on tracking objects that are moved by the camera wearer, online, maintaining knowledge of instance locations throughout the video even when they leave the field of view or become h…

  6. arXiv cs.CV TIER_1 English(EN) · Dima Damen ·

    以自我为中心的视频语言模型能同时捕捉手部和物体线索吗?

    Hand-object interaction (HOI) recognition requires capturing both hand manipulations and object transformations. However, existing video-language models often fall into shortcuts by relying on spurious correlations among hands, objects, or environmental context, rather than reaso…

  7. arXiv cs.CV TIER_1 English(EN) · Ruiping Liu, Junwei Zheng, Yufan Chen, Di Wen, Shaofang Quan, Chengzhi Wu, Jiaming Zhang, Kailun Yang, Kunyu Peng, Rainer Stiefelhagen ·

    EgoExoMem:跨视图同步的自我中心和外中心视频的记忆推理

    arXiv:2605.18734v2 Announce Type: replace Abstract: Egocentric memory is widely used in embodied intelligence, but it may be insufficient for comprehensive spatial-temporal reasoning. Inspired by human recall from both field and observer perspectives, we introduce EgoExoMem, the …

  8. arXiv cs.CV TIER_1 English(EN) · Dominick Reilly, Manish Kumar Govind, Le Xue, Srijan Das ·

    从我的视角到你的视角:利用特权式自我中心监督从外中心视频中学习自我中心线索

    arXiv:2501.05711v4 Announce Type: replace Abstract: Vision Language Models (VLMs) have achieved strong performance across a wide range of video understanding tasks. However, their viewpoint-invariant training limits their ability to infer egocentric properties, such as human-obje…

  9. arXiv cs.CV TIER_1 English(EN) · Max Gonzalez Saez-Diez, Jihoon Chung, Adam D. Wolsky, Gregory Lanzalotto, Dean Knox, Jonathan Mummolo, Brandon M. Stewart, Olga Russakovsky ·

    EgoPolice:高风险警用执法记录仪视频中的自我中心视频理解基准

    arXiv:2607.06468v1 Announce Type: new Abstract: We introduce EgoPolice, a carefully curated dataset of real, egocentric police-civilian interactions, sourced from publicly available body-worn camera videos. We select police-civilian action labels that are critical for police beha…

  10. arXiv cs.CV TIER_1 English(EN) · Olga Russakovsky ·

    EgoPolice:高风险警用执法记录仪视频中的自我中心视频理解基准

    We introduce EgoPolice, a carefully curated dataset of real, egocentric police-civilian interactions, sourced from publicly available body-worn camera videos. We select police-civilian action labels that are critical for police behavioral research and annotate them at a second-by…