A100 GPUs
PulseAugur coverage of A100 GPUs — every cluster mentioning A100 GPUs across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
Astrolabe system optimizes LLM serving with randomized prediction-guided scheduling
Researchers have developed Astrolabe, a novel scheduling system designed to optimize the serving of large language models (LLMs). This system employs a randomized prediction-guided approach to balance load across multip…
-
New TemporalSinkhorn method accelerates optimal transport calculations
Researchers have developed TemporalSinkhorn, a novel parallel-in-time execution method for dynamic entropic optimal transport problems, particularly benefiting applications like Flow Matching for generative modeling. Th…
-
Meta AI models power Genesis Mission projects at Lawrence Berkeley Lab
Meta's AI models, including the Segment Anything Model and those built on PyTorch, are being utilized in the initial projects of the Genesis Mission at Lawrence Berkeley National Laboratory. These projects focus on area…
-
New framework enhances LLM training by reducing noise in weaker models
Researchers have developed a new framework called Contrastive Weak-to-Strong Generalization (ConG) to improve the training of large language models. ConG addresses limitations in existing weak-to-strong generalization m…
-
New research explores controllable generalization failures and efficient RL distillation for LLMs
Researchers are exploring new methods to improve language model generalization and reasoning capabilities. One paper proposes a technique to construct models that exhibit controllable generalization failures by training…
-
LINE MAN Wongnai cuts AI server costs 9x, boosts app speed 4x
LINE MAN Wongnai has significantly reduced its AI server costs by 9 times and improved application speed by 4 times. This was achieved through a strategic shift in their MLOps approach, optimizing resource utilization a…
-
Mixed-Precision CA-SGD Accelerates Training on GPUs
Researchers have developed a mixed-precision communication-avoiding SGD (CA-SGD) method for generalized linear models on GPUs. This approach aims to reduce communication bottlenecks in distributed training by amortizing…
-
FlashSinkhorn solver accelerates optimal transport on GPUs
Researchers have developed FlashSinkhorn, a new GPU-accelerated solver for entropic optimal transport (EOT) that significantly reduces memory input/output operations. By rewriting stabilized log-domain Sinkhorn updates …
-
Nvidia chips smuggled to China and Russia despite US export controls
U.S. authorities are investigating multiple cases of advanced Nvidia GPUs and other semiconductor technology being illegally smuggled to China and Russia, circumventing export controls. These efforts involve sophisticat…