DCGM
PulseAugur coverage of DCGM — every cluster mentioning DCGM across labs, papers, and developer communities, ranked by signal.
-
Anyscale launches GPU Health Observability to diagnose hardware failures
Anyscale has launched a private preview of its GPU Health Observability tool, designed to bridge the gap between application-level failures and underlying hardware issues in GPU clusters. This new layer of observability…
-
Kubernetes GPU Node Setup Crucial for LLM Deployment
This article details the complex process of preparing GPU nodes for large language models (LLMs) within a Kubernetes environment. It emphasizes that simply adding GPUs to a node is insufficient, as Kubernetes needs spec…
-
LLM serving observability: A layered approach for vLLM and TGI
This article details how to achieve end-to-end observability for large language model inference servers like vLLM and TGI. It highlights that standard observability tools fall short due to unique LLM serving characteris…