PulseAugur
EN
LIVE 10:13:25

llama.cpp adds Qwen3-Next support; KataGo, Kimi-K3 models see updates

The latest release of llama.cpp (b10238) now supports Qwen3-Next with MTP, enabling optimized local inference on consumer hardware. Additionally, the Go AI engine KataGo has introduced an experimental evaluation cache in version v1.16.4 to speed up game state analysis. The Kimi-K3 multimodal model is also now accessible in the GGUF format, optimized by Unsloth for efficient local use on platforms like Hugging Face. AI

IMPACT Enhances local AI inference capabilities for developers and enthusiasts with improved model compatibility and performance optimizations.

RANK_REASON Updates to open-source AI inference engines and models.

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

llama.cpp adds Qwen3-Next support; KataGo, Kimi-K3 models see updates

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 Bahasa(ID) · soy ·

    llama.cpp b10238 Ships Qwen3-Next MTP — Plus KataGo, Kimi-K3 GGUF, & GPU AI

    <p>Today's digest highlights llama.cpp b10238's release with Qwen3-Next MTP support, alongside KataGo v1.16.4's experimental evaluation cache and Kimi-K3's GGUF format availability. We also delve into new AI infrastructure guidance from NVIDIA and AMD's AI Workbench, including Ne…