NVML
PulseAugur coverage of NVML — every cluster mentioning NVML across labs, papers, and developer communities, ranked by signal.
-
Developer builds local LLM server with auto VRAM model selection
A developer has created a local LLM server using FastAPI and llama.cpp that automatically selects the appropriate GGUF model based on available GPU VRAM. This setup allows users to run various models, from 7B to 70B par…
-
Local LLM server uses VRAM routing for efficient model selection
A technical guide demonstrates how to set up a local large language model server using llama.cpp and FastAPI. The system features VRAM-aware routing, allowing it to automatically select the most suitable LLM based on av…
-
New method detects hidden ML training using GPU telemetry
Researchers have developed a method to detect hidden machine learning training using zero-overhead telemetry from graphics processing units (GPUs). This approach utilizes privacy-preserving NVML telemetry, which observe…
-
Kubernetes operators enable scale-to-zero for LLM serving
New Kubernetes operators are emerging to address the cost of running large language models, particularly the issue of idle GPUs burning money. Hearth, an alpha-stage operator, allows users to declaratively serve open-so…
-
NVIDIA Edge AI Hardware Lacks Energy Attribution Capabilities
A new paper highlights a significant energy observability gap in NVIDIA's flagship edge AI hardware, specifically the GB10 SoC found in ASUS Ascent GX10 systems. The research demonstrates that current hardware lacks the…