A new model named Prism has been released, designed for native joint video-audio generation at resolutions up to 2K. Developed by researchers from Fudan University and Tencent Hunyuan, Prism utilizes a dynamic sparse attention framework to efficiently process high-resolution data. This approach allows the model to adapt its attention structure based on local content, leading to a 2.5x training speedup compared to full attention while improving generation quality. AI
IMPACT This model's dynamic sparse attention could influence future approaches to efficient high-resolution generative AI training.
RANK_REASON The item describes the release of a new model and its associated technical report, detailing its architecture and performance. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
- Accelerate
- diffusers
- FrancisRing/Prism
- Fudan University
- Jiangfeng Xiong
- Jian Wei Zhang
- Kaihang Pan
- Prism
- PyTorch
- Qi Tian
- Shuyuan Tu
- Tencent Hunyuan
- transformers
- Weijie Kong
- Xintong Han
- Yinming Huang
- Yue Wu
- Yu-Gang Jiang
- Zuxuan Wu
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →