Ollama has released version v0.32.4, introducing significant enhancements for local AI inference, particularly for users with Apple Silicon hardware. This update brings support for the Laguna model family via Apple's MLX engine, enabling accelerated inference on integrated GPUs. Additionally, the release refines speculative decoding by quantizing draft-model output heads for improved efficiency and accuracy, and fixes decoding issues for Qwen3 MoE models. The update also includes preliminary vision support for the Minimax-M3 model in llama.cpp, expanding its multimodal capabilities. AI
IMPACT Enhances local AI inference capabilities for Apple Silicon users and expands multimodal support in llama.cpp.
RANK_REASON This is a software release for a tool that facilitates local AI model inference, not a frontier model release or core research.
Read on Mastodon — fosstodon.org →
- Anthropic
- Claude
- Gemini
- Ollama
- Apple Inc.
- Hermes
- Laguna
- Mlx
- Qwen3
- v0.32.3
- v0.32.4
- AMD ROCm
- Apple MLX
- Apple Silicon
- llama.cpp
- Minimax-M3
- NVIDIA
- Nvidia Rubin Gpu
- TensorRT-LLM
AI-generated summary · Google Gemini · from 6 sources. How we write summaries →