Ollama has released version v0.32.9, adding support for NVIDIA's Nemotron 3.5 Lightning, a 30B mixture-of-experts model designed for agentic workloads. This integration allows users to run the specialized model locally on consumer GPUs. Additionally, the vLLM inference engine has released v0.27.0, featuring extensive improvements and full support for the Kimi K3 model. Ollama also recently added initial support for Meta's Muse Glimmer multimodal model in v0.32.7, optimized for Apple Silicon. AI
IMPACT Expands local inference capabilities for specialized agentic and multimodal models on consumer hardware.
RANK_REASON Updates to local inference runtimes and support for new models.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →