NVIDIA GPUDirect Storage
PulseAugur coverage of NVIDIA GPUDirect Storage — every cluster mentioning NVIDIA GPUDirect Storage across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
LLM storage latency tolerance shifts stepwise with GPU utilization
Storage latency tolerance for LLM inference does not decrease linearly with GPU utilization, but rather in a stepwise manner. At higher GPU utilization levels (around 90%), the compute queue saturates, making storage la…
-
Mingxin FX100 storage solution accelerates video inference, reducing latency
Mingxin's FX100 storage solution addresses latency bottlenecks in real-time video inference, which are often caused by storage and data path limitations rather than GPU compute. The system employs a tiered KV cache appr…
-
NVIDIA advances AI storage with Vera CPU and open cuFile APIs
NVIDIA is advancing AI storage infrastructure to meet the escalating demands of large datasets and context windows. The company is introducing new storage solutions, including the NVIDIA Vera CPU, which offers significa…
-
AWS cuts LLM load times with GPUDirect Storage and FSx
AWS has introduced a new method to significantly speed up the loading of large language models onto GPU instances. By leveraging NVIDIA GPUDirect Storage (GDS) with Amazon FSx for Lustre, model weights can be loaded dir…