PulseAugur
EN
LIVE 13:50:48

Ollama v0.32.6 boosts Qwen 3.5 speed on Apple Silicon, improves OpenAI compatibility · 4 sources tracked

Ollama has released version 0.32.6, significantly improving the performance of the Qwen 3.5 model on Apple Silicon Macs through the MLX engine and speculative decoding. This update also enhances compatibility with OpenAI's API by aligning the streaming format for the /v1/chat/completions endpoint. Additionally, other related projects like KataGo and llama.cpp have received bug fixes, with KataGo addressing TensorRT issues and llama.cpp resolving Vulkan device lost errors. AI

IMPACT Enhances local AI inference performance and compatibility for users of Apple Silicon devices running Qwen 3.5.

RANK_REASON This is a software update for a local AI inference tool, not a frontier model release from a major lab.

Read on Mastodon — fosstodon.org →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

Ollama v0.32.6 boosts Qwen 3.5 speed on Apple Silicon, improves OpenAI compatibility · 4 sources tracked

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
This is a software update for a local AI inference tool, not a frontier model release from a major lab.
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
52 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [4]

  1. dev.to — LLM tag TIER_1 English(EN) · soy ·

    Ollama v0.32.6 Speeds Qwen3.5 on Apple Silicon — Plus PyTorch, Mesa 26.2 & GPU Fixes

    <p>Today's digest brings significant updates for local AI, with Ollama v0.32.6 boosting Qwen3.5 performance on Apple Silicon and adding OpenAI streaming compatibility. Additionally, Mesa 26.2 adds NVK Mesh Shader support, llama.cpp fixes Vulkan errors, KataGo resolves critical Te…

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    ⚙️ New Ollama Release! ⚙️ Version: v0.32.6 Release Notes: ## What's Changed - Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for

    ⚙️ New Ollama Release! ⚙️ Version: v0.32.6 Release Notes: ## What's Changed - Qwen3.5 is faster on Apple GPUs: the MLX engine now uses the model's MTP head for speculative decoding automatically - `/v1/chat/completions` streaming now matches OpenAI's wire format: `role` only on t…

  3. dev.to — LLM tag TIER_1 English(EN) · LucioLiu ·

    Ollama's Latest RC Improves Qwen3.5 on Apple GPUs and Fixes Streaming Compatibility

    <p>Ollama v0.32.6-rc0 adds a useful path for people running Qwen3.5 on Apple GPUs: the MLX engine now uses the model's MTP head automatically for speculative decoding.</p> <p>The practical idea is straightforward. The MTP head predicts upcoming tokens, and the main model verifies…

  4. dev.to — LLM tag TIER_1 English(EN) · soy ·

    Ollama v0.32.6 Boosts Qwen3.5 on Apple GPUs — Plus AMD CDNA 5, FFmpeg 9.0 & More

    <p>Ollama v0.32.6 ships with significant Qwen3.5 performance boosts for Apple GPUs, while AMD unveils its new CDNA 5 hardware and Helios Rackscale. This digest also covers the release of FFmpeg 9.0, NVIDIA's Alpamayo 2 Super for commercial use, and a trending Qwen3-VL model on Hu…