Researchers have introduced Video DeltaNet (VDN-H3), a hybrid-attention model designed to accelerate video generation while maintaining high quality. The model utilizes a combination of efficient linear attention and a softmax branch to achieve fast inference speeds, generating a 14.4-second clip in just over 11 seconds on 8 B200 GPUs. VDN-H3 is designed to be plug-and-play, adding a separate linear attention branch and LoRA adapters that can be merged into existing backbones without altering their weights. The project is fully open-source, with weights, training code, and an optimized inference stack made publicly available. AI
IMPACT This model's hybrid attention architecture could lead to more efficient video generation tools, potentially lowering barriers for content creation.
RANK_REASON The cluster describes a new model release with technical details and open-source availability. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →