PulseAugur
中
实时 00:41:09
English(EN) ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

ObjectStream 框架使用潜在对象进行流视频理解 · arXiv

研究人员推出了一种新颖的框架 ObjectStream,旨在通过使用潜在对象作为记忆锚点来增强流视频理解。这种无需训练的方法直接从冻结的 Video-LLM 表示中提取空间连贯的潜在对象,并将它们跨帧链接起来,在有限的内存预算内维护持久的历史记录。ObjectStream 保留了对象的历史、变化和最近的视觉上下文,使 Video-LLM 能够在不改变基础模型的情况下推理对象的身份、交互和状态变化。实验表明,在性能和效率方面都有显著提高,ObjectStream 在 OVO-Bench 实时视觉感知基准测试中将 Qwen2.5-VL-7B 的性能提高了 10.0 个点,同时将内存使用量和首次字节时间减少了约 50%。 AI

影响 增强了 Video-LLM 在实时分析和长视频理解方面的能力,可能改进监控、内容审核和自动视频摘要等应用。

排序理由 这是一篇描述视频理解新框架的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

ObjectStream 框架使用潜在对象进行流视频理解 · arXiv

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是一篇描述视频理解新框架的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
69 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.CV TIER_1 English(EN) · Mingkang Dong, Muxin Pu, Jie Li, Bohan Guo, Songruo Chen, Bin Ren, Xu Zheng, Chen Zhao, Tianwen Qian, Mohamed Elhoseiny, Yuqian Fu ·

    ObjectStream:将潜在对象作为流视频理解的记忆锚点

    arXiv:2607.28312v1 Announce Type: new Abstract: Streaming video understanding requires models to continuously retain useful visual evidence before future questions are known. Existing approaches primarily manage the growing visual context according to token importance, temporal r…