The Hugging Face Transformers library has released version 5.17.0, introducing several new models and frameworks. Notable additions include HYV4, a 780B-parameter mixture-of-experts language model with a 1M token context window, and VibeVoice, a diffusion-based framework for high-fidelity speech synthesis. The release also features NeoMME, a multimodal-native multilingual foundation encoder, and Fun-ASR-Nano, an efficient end-to-end speech recognition model supporting multiple languages and dialects. Additionally, Kimi Linear, a hybrid linear attention architecture, and Canary, a multilingual ASR and speech-to-text translation model, are now available. AI
IMPACT Expands the toolkit for AI developers with new models for language, speech synthesis, multimodal understanding, and speech recognition.
RANK_REASON This is a software library release with new model additions, not a frontier model release from a primary lab.
Read on Transformers — Releases →
- Alibaba DAMO Academy
- ArthurZucker
- Canary
- DeepSeek-V3
- Fun-ASR-Nano
- FunAudioLLM
- gpt-oss
- H Company
- Hy4-Preview
- HYV4
- Kimi Linear
- Moonshot AI
- NeoMME
- NeoMME-Retriever
- Parakeet
- pengzhiliang
- VibeVoice
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →