English(EN)Accelerating Text-to-Video Generation with Calibrated Sparse Attention
新方法通过优化注意力机制加速文本到视频生成 · 跟踪4个来源
作者PulseAugur 编辑部·[7 个来源]·
研究人员开发了新的方法来加速文本到视频生成,目前该过程受到大型Transformer模型中注意力机制的计算需求的瓶颈限制。Apple的CalibAtt和来自arXiv的HeadCast框架提出了无需训练的方法,可以识别并跳过可忽略的token-to-token连接,从而实现显著的加速。FVAttn是另一个无需训练的系统,它通过动态迁移注意力头来解决多GPU设置中的工作负载不平衡问题,在保持视频质量的同时实现了显著的推理加速。
AI
Recent diffusion models enable high-quality video generation, but suffer from slow runtimes. The large transformer-based backbones used in these models are bottlenecked by spatiotemporal attention. In this paper, we identify that a significant fraction of token-to-token connectio…
arXiv cs.LG
TIER_1English(EN)·Jinliang Shen, Lianghao Su, Zheming Li, Kang He, ZiLiang Lai, Yanbing Jiang, Chengru Song·
arXiv:2607.20125v1 Announce Type: cross Abstract: Autoregressive (AR) video diffusion models have become a promising paradigm for long and streaming video synthesis, but the continuously growing Key-Value (KV) cache makes attention the dominant inference cost, especially at high …
We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softmax video DiTs in quality while retaining the favo…
Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-p routing creates uneven per-head workloads under multi-GPU seque…
arXiv:2607.20940v1 Announce Type: new Abstract: Streaming video diffusion models have made substantial progress toward interactive and dynamic world simulation, but the nested autoregressive and denoising loops of conventional next-frame generation hinder real-time deployment. Re…
arXiv:2607.21553v1 Announce Type: new Abstract: We introduce SANA-Video 2.0, a hybrid video diffusion transformer instantiated at 5B and 14B scales under a unified architecture. Designed to generate high-quality video up to 720p on a single GPU, SANA-Video 2.0 matches full-softma…
arXiv:2607.16190v1 Announce Type: new Abstract: Video Diffusion Transformers process long spatio-temporal sequences, making self-attention the main bottleneck in high-resolution video generation. Training-free sparse attention reduces this cost, but adaptive Top-$p$ routing creat…