PulseAugur
实时 07:10:52
English(EN) Locate Anything in Videos: Rethinking Efficient Generative Spatio-Temporal Video Grounding

新的并行管解码方法可大幅缩短视频定位延迟

研究人员开发了一种名为并行管解码(PTD)的新方法,以提高时空视频定位的效率和准确性。该技术消除了自回归依赖,通过允许同时进行空间和时间定位,显著降低了延迟。PTD将定位过程分解为时间块和时间条件空间块,这些块被并行解码。该方法还引入了解耦块注意力,以在消除跨框依赖的同时保持上下文。实验表明,PTD在延迟方面取得了显著降低,吞吐量有所提高,并且还展示了对相关视频理解任务的泛化能力。 AI

影响 该方法可以显著加快视频分析任务的速度,并提高需要理解和定位视频中对象或事件的AI系统的性能。

排序理由 该集群描述了一篇详细介绍视频定位新方法的最新研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新的并行管解码方法可大幅缩短视频定位延迟

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一篇详细介绍视频定位新方法的最新研究论文。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
4 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+1 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准

报道来源 [3]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    在视频中定位一切:重新思考高效生成式时空视频定位

    Parallel Tube Decoding enables simultaneous spatial and temporal video grounding by removing autoregressive dependencies, drastically cutting latency while improving accuracy.

  2. arXiv cs.CV TIER_1 English(EN) · Xingjian Wang, Shijian Wang, Yibo Wang, Zihao Yu, Runhao Fu, Xuelian Cheng, Zongyuan Ge ·

    利用合成课程学习组合时空视频基础

    arXiv:2608.30584v1 Announce Type: new Abstract: Despite the impressive progress of recent MLLMs on spatio-temporal video grounding (STVG), existing evaluations and training data focus primarily on simple queries. They largely overlook the compositional queries prevalent in real-w…

  3. arXiv cs.CV TIER_1 English(EN) · Hanoona Rasheed, Haania Siddiqui, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan ·

    在视频中定位一切:重新思考高效生成式时空视频精确定位

    arXiv:2608.28192v1 Announce Type: new Abstract: Spatio-temporal video grounding (STVG) requires models to identify when a referred event occurs and localize the target entity throughout that interval. Existing multimodal large language models typically serialize dense localizatio…