The llama.cpp project has released version b11477, which introduces optimizations for handling tensor flags across different models. This update centralizes the detection logic for trunk-only and MTP-only models into a helper function within llama_model_base. Additionally, several models including DeepSeek4, nemotron-h, Qwen35, qwen35moe, qwen3next, and qwen4exp now support trunk-only files, enhancing compatibility and efficiency. AI
IMPACT Improves efficiency and compatibility for running various LLMs locally.
RANK_REASON Software release for an open-source project focused on infrastructure for running LLMs.
Read on llama.cpp — Releases →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →