Researchers have introduced a new family of tokenizers called KVAE, designed for multimodal generative models. These tokenizers, including KVAE-Audio, KVAE-3D, and KVAE-2D, are specifically engineered for text-conditioned generation tasks across audio, video, and image data. The paper details the development, training, and ablation studies for these models, sharing code and training specifics with the community. Evaluations indicate that KVAE tokenizers meet or exceed the performance of existing open-source alternatives on various reconstruction and generation metrics. AI
IMPACT These new tokenizers could improve the efficiency and quality of multimodal generative models for audio, video, and image tasks.
RANK_REASON The cluster describes a new family of tokenizers presented in an academic paper, detailing their architecture and performance.
- FLUX.2
- HunyuanVideo 1.5
- Kirill Chernyshev
- KVAE-2D
- KVAE-3D
- KVAE-Audio
- MMAudio
- MovieGen
- StableAudio
- variational auto-encoder
- Wan-2.2
- kandinskylab/kvae
- kandinskylab/KVAE-Audio
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →