A new version of the GLM-5.2 large language model, now with vision capabilities, has been released on Hugging Face. This update integrates the vision encoder from the Kimi 2.6 model, addressing a previous limitation of GLM-5.2. The model was developed by baseten, an inference provider, and made publicly available. AI
IMPACT Enhances the multimodal capabilities of open-source LLMs, potentially improving performance on tasks requiring visual understanding.
RANK_REASON This is a model update from a third-party provider, not a frontier lab release.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →