The llama.cpp project has released an update (b11513) that significantly optimizes the CUDA implementation of the top-k algorithm. This update replaces the per-row DeviceTopKKernel with a more efficient grid-over-rows radix select for large row counts, drastically reducing processing time. For instance, on the qwen4exp model with 34,816 tokens, the top-k operation speed improved from over 5.7 seconds to approximately 941 milliseconds. The release also refines the selection of the optimal top-k implementation based on shape, utilizing bitonic, radix select, or DeviceTopK/CUB argsort depending on row length and configuration. AI
IMPACT Improves inference performance for models running on CUDA-enabled hardware via llama.cpp.
RANK_REASON This is a software update for an open-source project that improves performance of a specific algorithm on certain hardware, not a frontier release or significant industry event.
Read on llama.cpp — Releases →
- ARGSORT
- CUDA
- DeviceTopK
- DeviceTopKKernel
- GGML_CUDA_TOPK_RADIX_MIN_ROWS
- llama.cpp
- praneshgo
- Pranesh Gonegandla
- qwen4exp
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →