The llama.cpp project has released version b11528, which includes a fix for handling host views of tensors. This change, assisted by Qwen3.8 Flash-Next, addresses an assertion error that occurred with KV cache views when using tensor splitting with partial offloading. The update also re-enables SM tensor support for K2 Horizon and adds a TODO reference. AI
IMPACT Improves performance and stability for users of the llama.cpp inference engine, particularly those utilizing partial offloading and KV caching.
RANK_REASON This is a software release for a specific tool, not a frontier model release or significant industry event.
Read on llama.cpp — Releases →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →