Ollama has released version 0.32.6, significantly improving the performance of the Qwen 3.5 model on Apple Silicon Macs through the MLX engine and speculative decoding. This update also enhances compatibility with OpenAI's API by aligning the streaming format for the /v1/chat/completions endpoint. Additionally, other related projects like KataGo and llama.cpp have received bug fixes, with KataGo addressing TensorRT issues and llama.cpp resolving Vulkan device lost errors. AI
IMPACT Enhances local AI inference performance and compatibility for users of Apple Silicon devices running Qwen 3.5.
RANK_REASON This is a software update for a local AI inference tool, not a frontier model release from a major lab.
Read on Mastodon — fosstodon.org →
- AMD
- Apple GPUs
- ComfyUI
- Hugging Face
- INT8 quantization
- MLX
- NVIDIA
- Ollama
- OpenAI
- Qwen 3.5
- Qwen3-VL-32B
- CDNA 5
- FFmpeg 9.0
- MLX engine
- Qwen3-VL-32B-Ultra-Heretic-H3
- Apple Silicon
- KataGo
- llama.cpp
- MTP Head
- PyTorch
- TensorRT
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →