The latest release of llama.cpp, version 0.6.0, introduces a new extended batch API called llama_batch_ext. This update allows for mixed token and embedding batches, along with per-token state embeddings for Multi Token Prediction (MTP) and deepstack models. Additionally, the release adds support for the GLM-5.3-Flash (GLM5-Next) 320B text and vision hybrid model, as well as the Clef decision model which handles both text and vision inputs. AI
IMPACT Enhances inference capabilities for large language models with improved batch processing and support for new hybrid and decision models.
RANK_REASON This is a software release for an open-source project focused on AI model inference, including new features and model support. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →