Nccl
PulseAugur coverage of Nccl — every cluster mentioning Nccl across labs, papers, and developer communities, ranked by signal.
6 day(s) with sentiment data
-
New T-CCL library boosts multi-GPU AI model performance with TMA
Researchers have developed T-CCL, a new collective communication library designed for efficient multi-GPU execution in large transformer models. T-CCL leverages the Tensor Memory Accelerator (TMA) to offload data moveme…
-
Nvidia GPU Terms Explained for AI Engineers
This article defines essential GPU terms for AI engineers, covering concepts from CUDA cores to NVLink. It explains how these NVIDIA GPU concepts influence model fitting, execution speed, and scalability. The piece aims…
-
Google enhances GPU cluster with NVIDIA auto-activation, but lags top providers
Google has significantly improved its GPU cluster experience, moving from a "Bronze" to "Gold" tier in under two years, according to SemiAnalysis. A key enhancement is the auto-activation of NVIDIA's ConnectX-7/8 NCCL p…
-
New ThunderEP design boosts MoE inference on consumer GPUs
Researchers have developed ThunderEP, a new communication design for efficiently running large Mixture-of-Experts (MoE) models on consumer GPUs connected via PCIe. This system addresses the communication bottlenecks inh…
-
Modal launches serverless GPU clusters with RDMA support
Modal has announced the general availability of its new product, Modal Clusters. This feature allows organizations to access large-scale compute resources, including RDMA support for high-speed communication between nod…
-
NVIDIA NVRx enhances distributed AI training fault tolerance on Amazon EKS
NVIDIA has introduced NVRx, a Python library designed to enhance fault tolerance for large-scale distributed AI training on Amazon EKS. NVRx integrates with PyTorch's Fully Sharded Data Parallel (FSDP) to enable asynchr…
-
llama.cpp adds Ubuntu-CUDA builds and GCC 14 support in latest release
The llama.cpp project has released version b10969, which includes new build jobs for Ubuntu-CUDA, supporting CUDA versions 12.8 and 13.3 on both x64 and arm64 architectures. This update also incorporates GCC 14 for CUDA…
-
New REACT system alleviates AI cluster congestion at application layer
Researchers have developed a system called REACT that addresses congestion issues in shared AI clusters during distributed training. REACT operates at the application layer, detecting network congestion in real-time usi…
-
NVIDIA and Google Automate NCCL Plugin Setup for GCP
NVIDIA and Google have collaborated to implement auto-activation for Google's NCCL Plugin, specifically for ConnectX-7 and ConnectX-8 NCCL. This enhancement, driven by feedback from SemiAnalysis, aims to significantly i…
-
Developer details Qwen3.8-27B setup on dual RTX 3090s
A developer details the extensive troubleshooting required to run the Qwen3.8-27B model on a dual RTX 3090 setup without NVLink. Initial attempts with vLLM and SGLang encountered significant issues, including compilatio…
-
NVIDIA B300 fine-tuning of Qwen3-32B detailed in new research
A new paper details the operational challenges and solutions encountered when fine-tuning the Qwen3-32B model on NVIDIA's B300 accelerators. The research focuses on practical aspects of multi-node training, offering ins…
-
New CTA-pipelining method slashes multi-GPU latency for LLMs
Researchers have introduced CTA-pipelining, a novel execution paradigm for multi-GPU systems that optimizes for latency in serving large language models. This method exploits dependencies at the Cooperative Thread Array…
-
New HSAP framework enhances LLM training efficiency for hybrid-context models
Researchers have introduced HSAP, a Hierarchical Sequence-aware Parallelism framework designed to improve the efficiency of training large language models. This new approach addresses challenges in handling hybrid-conte…
-
Together AI releases open-source Parallel Kernel Builder for LLM inference
Together AI has released Parallel Kernel Builder (PKB), an open-source tool designed to optimize inference performance for large language models. PKB can identify and generate novel kernels, such as those for NeMo vocab…
-
LLMs struggle to generate multi-GPU kernels, researchers find
Researchers at Together have found that while large language models can efficiently generate single-GPU kernels, they struggle significantly with multi-GPU kernel generation. These models perform poorly when asked to cr…
-
Frontier LLMs struggle with multi-GPU kernel generation, new benchmark reveals
A new benchmark called ParallelKernelBench (PKB) has been developed to evaluate the ability of frontier large language models to generate efficient multi-GPU kernels. Testing models like GPT-5.5, Gemini 3 Pro, and Opus …
-
User seeks advice on optimizing dual-GPU inference with llama.cpp
A user on the r/LocalLLaMA subreddit is seeking advice on optimizing performance with an asymmetric dual-GPU setup. They have a 3080 Ti with 12GB VRAM and a 3080 with 20GB VRAM, and are experiencing significant speed dr…
-
New OptCC algorithm minimizes AllReduce slowdown from network failures
Researchers have developed OptCC, a new algorithm designed to improve the efficiency of AllReduce operations in large-scale GPU clusters, particularly when network failures occur. This algorithm approaches theoretical l…
-
Developer details verl RL framework internals and NCCL bug
A developer detailed their experience working with ByteDance's verl framework for RL post-training, including its internal workings and the challenges of forking the project. The write-up covers the framework's orchestr…
-
New framework HetCCL boosts LLM training on mixed-hardware clusters
Researchers have developed HetCCL, a new framework designed to improve collective communication efficiency in heterogeneous computing clusters used for training large language models. This framework addresses the limita…