PulseAugur
实时 18:20:11
English(EN) VideoSEMA: a scalable and efficient Mamba-like attention for video understanding

新的类Mamba注意力模型VideoSEMA提升视频理解性能

研究人员推出VideoSEMA,这是一种新颖的类Mamba注意力模型,专为高效可扩展的视频理解而设计。该模型采用分时空注意力机制,在其空间分量中结合了局部窗口注意力和全局平均,并在其时间分量中使用了softmax时间注意力。与现有的视觉Transformer和Mamba模型相比,VideoSEMA在K400和SSv2等基准数据集上表现出优越的性能,尤其是在图像分辨率增加时准确率的平稳下降。该研究还强调了通过扩张或稀疏时间注意力将VideoSEMA扩展到处理更长视频的潜力。 AI

影响 这种新的类Mamba架构为视频理解任务提供了更高的效率和性能,可能影响未来视频AI的发展。

排序理由 该集群包含详细介绍视频理解新模型架构的研究论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新的类Mamba注意力模型VideoSEMA提升视频理解性能

报道来源 [5]

  1. arXiv cs.AI TIER_1 English(EN) · Nhat Thanh Tran, Fanghui Xue andShuai Zhang, Jiancheng Lyu, Yunling Zheng, Yingyong Qi, Jack Xin ·

    VideoSEMA:一种类似Mamba的可扩展高效注意力机制,用于视频理解

    arXiv:2607.14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each…

  2. arXiv cs.AI TIER_1 English(EN) · Jack Xin ·

    VideoSEMA:一种类似Mamba的可扩展高效注意力机制,用于视频理解

    We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window atte…

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    VideoChat3:全开源视频多模态大模型,实现高效通用视频理解

    Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse…

  4. arXiv cs.CV TIER_1 English(EN) · Xinhao Li, Yuhan Zhu, Xiangyu Zeng, Yuhao Dong, Haoning Wu, Zhiqiu Zhang, Yuandong Yang, Changlian Ma, Qingyu Zhang, Yansong Shi, Xinyu Chen, Haoran Chen, Zizheng Huang, Jun Zhang, Kun Ouyang, Lin Sui, Ziang Yan, Yicheng Xu, Chenting Wang, Yinan He, Hong… ·

    VideoChat3:完全开源的视频多模态大模型,实现高效和通才的视频理解

    arXiv:2607.14935v1 Announce Type: new Abstract: Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. …

  5. arXiv cs.CV TIER_1 English(EN) · Limin Wang ·

    VideoChat3:完全开放的视频多模态大模型,实现高效和通才的视频理解

    Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse…