Researchers have introduced AuK, an open-source foundational model designed for both speech generation and editing. This model integrates natural language instructions and audio context, utilizing a multimodal large language model, a joint variational auto-encoder, and a hybrid rectified-flow Transformer. To enhance efficiency, AuK has been distilled into AuK-Flash, which achieves a significant speedup in inference time without compromising performance on various speech-related tasks. AI
IMPACT This open-source model could accelerate research and development in speech synthesis and manipulation, enabling new applications in content creation and accessibility.
RANK_REASON The item describes a technical report detailing an open-source foundational model for speech generation and editing. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
- AuK-Flash
- Diffusion Transformer
- Hugging Face
- MMDiT
- tencent/AuK
- Transformer++
- variational auto-encoder
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →