The Hugging Face Transformers library has released version 5.12.0, introducing new models like MiniMax-M3-VL, a vision-language model with a CLIP-style vision tower and a sparse Mixture-of-Experts decoder. This update also includes improvements to PP-OCRv6, an efficient OCR system, and Parakeet-RNNT, a fast conformer encoder with an RNN-T decoder. Additionally, version 5.11.0 added DiffusionGemma, an encoder-decoder model for faster text generation, and DeepSeek-V3.2-Exp, which features a novel sparse attention mechanism for long-context efficiency. AI
RANK_REASON The cluster contains release notes for a software library detailing new model additions and improvements, which falls under research and development.
Read on Transformers — Releases →
- DeepSeek Sparse Attention
- DeepSeek-V3.2
- DiffusionGemma
- FalconMamba
- Hugging Face
- NemotronH
- Qwen2.5-VL
- Qwen2-VL
- Qwen3-VL
- Transformers
- v5.11.0
- Zamba2
- DeepSeek-V3.1-Terminus
- DeepSeek-V3.2-Exp
- Hugging Face Transformers
- MiniMax-M3
- MiniMax-M3-VL
- Parakeet-RNNT
- PP-OCRv6
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →