Llm D
PulseAugur coverage of Llm D — every cluster mentioning Llm D across labs, papers, and developer communities, ranked by signal.
- 2026-08-01 product_launch The open-source project llm-d has released version v0.7. source
1 day(s) with sentiment data
-
llm-d enhances Kubernetes for LLM inference
A new open-source project called llm-d aims to bridge the gap between Kubernetes' orchestration capabilities and the specific needs of large language model (LLM) inference. llm-d acts as a layer on top of existing tools…
-
Kubernetes-Native CXL Memory Enhances LLM Serving Efficiency
Researchers have developed a Kubernetes Dynamic Resource Allocation (DRA) driver that enables composable CXL memory to function as a schedulable cluster resource for LLM serving. This system allows for cross-node KV-cac…
-
Red Hat AI's vllm and llm-d teams lauded for inference expertise
SemiAnalysis highlighted the exceptional work of the vllm project and llm-d maintainers at Red Hat AI, recognizing them as leading experts in inference. The post specifically praised the helpfulness and kindness of indi…
-
Kimi-VL model scaled using heterogeneous E/PD on llm-d and SGLang
This item details the technical advancements in scaling the Kimi-VL model, which is designed for vision-heavy tasks. The scaling was achieved through heterogeneous E/PD (Execution/Processing Distribution) methods implem…
-
GLM-5.2 model released for agentic workloads on llm-d
A new model, GLM-5.2, is now available for agentic workloads on the llm-d platform. This release focuses on enhancing capabilities for AI agents.
-
llm-d v0.7 released with focus on production hardening
The open-source project llm-d has released version v0.7, focusing on the transition from initial feature implementation to robust production hardening. This update aims to improve the stability and reliability of the sy…
-
llm-d routing layer boosts Qwen 7B inference speed by 2.3x on AWS EKS
A new routing layer called llm-d has demonstrated a significant speedup for LLM inference, specifically with the Qwen2.5-7B-Instruct model on AWS EKS. By intelligently routing requests to vLLM replicas that are likely t…
-
Google enables OSS production Kubernetes inferencing for LLMs
Google has enhanced its open-source production Kubernetes inferencing capabilities by adding nightly CI for llm-d. This development is seen as a significant step towards enabling broader adoption of large language model…