English(EN)VideoSEMA: a scalable and efficient Mamba-like attention for video understanding
新的类Mamba注意力模型VideoSEMA提升视频理解性能
作者PulseAugur 编辑部·[5 个来源]·
研究人员推出VideoSEMA,这是一种新颖的类Mamba注意力模型,专为高效可扩展的视频理解而设计。该模型采用分时空注意力机制,在其空间分量中结合了局部窗口注意力和全局平均,并在其时间分量中使用了softmax时间注意力。与现有的视觉Transformer和Mamba模型相比,VideoSEMA在K400和SSv2等基准数据集上表现出优越的性能,尤其是在图像分辨率增加时准确率的平稳下降。该研究还强调了通过扩张或稀疏时间注意力将VideoSEMA扩展到处理更长视频的潜力。
AI
arXiv:2607.14711v1 Announce Type: cross Abstract: We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each…
We present for video understanding (classification) a split space-time attention model, VideoSEMA, consisting of a scalable and efficient Mamba-like attention (SEMA) block in space and a softmax temporal attention in time. In each frame, SEMA attention applies a local window atte…
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse…
arXiv:2607.14935v1 Announce Type: new Abstract: Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. …
Recent advances in video understanding have spanned motion, long video, and streaming interaction, driving this field toward real-world applications. Despite this progress, current open-source models remain limited in several ways. They often struggle to generalize across diverse…