PulseAugur
实时 09:19:06
English(EN) Temporally Grounded Compositional Camera Motion Understanding via Geometric Knowledge Distillation

新基准和模型推动视频相机运动理解

研究人员推出了 CamChoreo,这是一个用于理解视频中复杂相机运动的新基准数据集。该数据集包含 4,229 个带有详细时间标注的真实世界剪辑,其中近一半的片段包含多个同时发生的相机运动。为了解决当前多模态大语言模型 (MLLMs) 在识别这些细粒度运动方面的局限性,研究团队开发了 CamDistill。该方法将几何知识蒸馏成轻量级 token,从而在推理时无需单独的 3D 基础模型即可实现准确的相机运动识别。 AI

影响 推动细粒度的时间和组合式相机运动识别,可能改进视频生成和空间智能应用。

排序理由 该集群包含一篇学术论文,详细介绍了用于视频感知的基准和模型。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准和模型推动视频相机运动理解

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Dazhao Du, Shiyan Du, Jian Liu, Yongjian Yu, Bohai Gu, Tao Han, Hualuo Liu, Eric Liu, Yujia Zhang, Xi Chen, Song Guo ·

    通过几何知识蒸馏实现时间约束的组合式相机运动理解

    arXiv:2608.10932v1 Announce Type: cross Abstract: Understanding camera motion is fundamental to video perception, with applications in spatial intelligence and controllable video generation. Multimodal large language models (MLLMs) provide a natural interface for this task, but e…