The Hugging Face Transformers library has released version 5.18.0, introducing several new models and significant updates. Key additions include Nemotron 3 Diarization for speaker identification, NemotronH Omni for multimodal reasoning across text, image, video, and sound, and HyperCLOVAX Vision V2, a multimodal model from Naver. The release also incorporates GTE, a new text representation model, and includes breaking changes and bug fixes related to ROCm, vLLM, and various model integrations. AI
IMPACT Expands the range of readily available multimodal and specialized audio models for developers.
RANK_REASON This is a software library release with new model integrations, not a frontier model release from a core AI lab.
Read on Transformers — Releases →
- Alibaba Group
- GTE
- Hugging Face Transformers
- HyperCLOVAX Vision V2
- NAVER
- Nemotron 3 Diarization
- NemotronH Omni
- Nvidia
- ROCm
- Snowflake
- vLLM
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →