NVM Express
PulseAugur coverage of NVM Express — every cluster mentioning NVM Express across labs, papers, and developer communities, ranked by signal.
11 day(s) with sentiment data
-
AirLLM enables 70B model inference on 4GB GPU by streaming layers from disk
AirLLM is a new project that enables running large language models, such as a 70B parameter model, on hardware with very limited VRAM, like a 4GB GPU. It achieves this by loading model layers sequentially from disk to t…
-
SSD speed impact on gaming performance analyzed across 11 titles
A recent analysis of 11 gaming titles reveals that while faster SSDs like PCIe 5.0 offer some improvements, the real-world impact on gaming performance is often marginal. The study compared SATA and NVMe SSDs, noting th…
-
KV Cache Prefetching Slashes LLM Inference Latency
A new prefetching strategy for KV Cache data has been developed, significantly reducing storage latency during large model inference. This method, tested on the Mingxin FX100 with a 480B model, improves inference throug…
-
NVMe adds virtualization for SSDs, boosting storage in virtualized environments
The NVMe specification has been updated to include virtualization capabilities for locally attached Solid State Drives. This enhancement aims to improve storage management and performance in virtualized environments. Th…
-
KV Cache tiering boosts LLM inference speed and cuts costs
A new approach to managing KV Cache in large language model inference suggests treating it as a high-frequency access subset within the warm storage tier, rather than in the traditional hot or cold tiers. This strategy,…
-
Amazon offers 60% discount on Samsung 2TB SATA SSD
Amazon is offering a significant discount on a Samsung 2TB SATA SSD, specifically the 870 Evo model. This deal provides over 60% off the retail price, making it an attractive option for consumers seeking storage solutio…
-
AI presents major opportunity and threat to data storage sector
The storage industry faces a dual challenge and opportunity with the rise of AI. While AI drives demand for faster access and more secure data recovery, it also introduces significant threats through potential mishaps a…
-
Repurpose old drives for external storage amid AI-driven price hikes
With external SSD prices soaring due to AI-driven demand, a new article from Tom's Hardware explores repurposing older internal drives as external storage. The test involves using SATA SSDs, NVMe drives, and even tradit…
-
Developer refines MoE model file layout with community benchmarks
A developer optimized MoE model files by reordering expert weights based on measured co-activation, resulting in a 2.23x reduction in disk reads. When sharing this work, two of the initial optimization pitches were refu…
-
WekaFS vs Ceph for AI GPU Starvation: A Storage Showdown
WekaFS and Ceph are compared for their effectiveness in AI training pipelines, specifically addressing GPU starvation. WekaFS offers extreme DPDK speed but comes with high software licensing fees, while Ceph provides pe…
-
Colibri project enables 744B parameter models on consumer hardware via disk streaming
The Colibri project has developed a novel disk-streaming technique to run massive language models, such as Z.ai's GLM-5.2 with 744 billion parameters, on consumer hardware with limited RAM. This method separates the den…
-
Denver bare metal servers offer balanced US latency and high throughput
Deploying bare metal servers in Denver offers significant advantages for nationwide US user bases by balancing latency and providing robust infrastructure. Denver's central location ensures sub-30ms latency to the West …
-
Gigantic AI models now runnable on consumer laptops and PCs
New developments are making it possible to run large AI models on consumer hardware, significantly lowering the barrier to entry for local AI development. Projects like AirLLM enable 70-billion-parameter models to run o…
-
New HGA Method Enables Long-Context LLM Fine-Tuning on Limited VRAM
Researchers have developed a new method called Hierarchical Global Attention (HGA) to enable efficient fine-tuning of large language models with limited VRAM. This technique combines segment-wise backpropagation with ti…
-
Asus ROG Strix Aiolos M.2 SSD enclosure drops to $59
Asus' ROG Strix Aiolos M.2 SSD enclosure, capable of transfer speeds up to 20 Gbps, is currently available on Amazon for $59, a 14% discount. This tool-less enclosure supports both NVMe and SATA M.2 SSDs and features RG…
-
Dragon Age Creator Believes Series is Dead; Amazon SageMaker Enhances Inference
The creator of the Dragon Age series, David Gaider, believes the franchise is likely finished due to the poor performance of Dragon Age: The Veilguard. Despite this, he expressed openness to returning to the series if t…
-
AWS SageMaker HyperPod boosts enterprise AI inference with new features
Amazon SageMaker HyperPod has introduced new features to enhance enterprise inference for generative AI workloads. These updates include improved data capture capabilities at various points in the inference pipeline, of…
-
Sugon's storage system leads IO500, marking a first for Chinese vendors
Chinese company Sugon has achieved a significant milestone in high-performance computing storage by securing the top positions in both the production full-node and 10-node categories of the IO500 list. This marks the fi…
-
Rosewill M.2 SSD Cloner and Eraser Hits Record Low Price of $47
Rosewill's M.2 SSD Cloner and Eraser is currently available at its lowest price of $47, offering a convenient solution for IT professionals and home users alike. This device supports cloning and erasing NVMe drives both…
-
Rust engine streams Mixtral 8x7B on cheap VMs
A new Rust-based inference engine called MER allows for efficient streaming of large language models like Mixtral 8x7B from NVMe storage onto less powerful and cheaper virtual machines. This approach bypasses the need f…