The landscape of serving large language models on Kubernetes is rapidly evolving, with many once-open challenges now addressed by well-funded projects and standards. Areas like KV-cache-aware routing, prefill/decode disaggregation, fractional GPU sharing, and basic GPU scheduling primitives are becoming crowded. However, significant gaps remain, particularly in treating model-weight distribution as a first-class Kubernetes primitive, which currently leads to long cold-start times. AI
IMPACT Highlights critical infrastructure gaps for efficient LLM deployment, particularly in reducing cold-start times for scaled applications.
RANK_REASON Article analyzes existing and emerging solutions for LLM serving on Kubernetes, referencing academic papers and projects. [lever_c_demoted from research: ic=1 ai=0.7]
- arXiv 2607.16596
- arXiv 2609.20874
- Gateway API Inference Extension
- GKE Inference Gateway
- Grove
- KAI Scheduler
- Kubernetes
- Kueue
- llm-d
- NVIDIA Dynamo
- NVIDIA GPU Operator
- Run:ai
- vCluster
- Volcano
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →