H200
PulseAugur coverage of H200 — every cluster mentioning H200 across labs, papers, and developer communities, ranked by signal.
- used by ByteDance 90%
- instance of Blackwell 90%
- instance of graphics processing unit 90%
- instance of NVIDIA H100 70%
- used by NVIDIA H100 70%
- used by Blackwell 70%
- competes with MI300X 70%
- used by vLLM 70%
- used by Supermicro 70%
- competes with H.1000 Gnome 70%
- competes with Kimi k3 60%
- competes with RTX 4090 60%
- 2026-08-19 product_launch First shipments of Nvidia's H200 AI chips have reached China, with ByteDance and Tencent as initial recipients. source
- 2026-05-14 product_launch US government approves sale of NVIDIA H200 AI chips to ten Chinese companies.
8 day(s) with sentiment data
-
Nebius and Nvidia launch 2026 Physical AI Awards with $750K in compute credits
Nebius Group, in partnership with Nvidia, has launched the 2026 Nebius Physical AI Awards, offering five distinct categories for startups with demonstrable physical AI products. Each of the five winners will receive $15…
-
Report: Billions in Nvidia AI chips reach China via illicit channels · 1 source tracked
A report from the Center for Advanced Defense Studies (C4ADS) details how Chinese firms are circumventing U.S. export restrictions to acquire billions of dollars worth of advanced Nvidia AI chips. The report identifies …
-
Nunchux AI unveils VC-Attention to speed up video diffusion transformers
Nunchux AI has developed VC-Attention, a novel training-free low-bit attention kernel designed to accelerate video diffusion transformers. This innovation addresses two key bottlenecks: value quantization errors and the…
-
New tools and research tackle GPU optimization for AI workloads
Several research papers and a new open-source tool address challenges in optimizing AI workloads on GPUs. COMPASS-ABS aims to reduce fragmentation in shared GPU clusters for deep learning training, improving resource ut…
-
VC-Attention framework speeds up video generation by optimizing low-bit attention
Researchers have introduced VC-Attention, a novel framework designed to enhance the efficiency and accuracy of attention mechanisms in Diffusion Transformers, which are crucial for state-of-the-art video generation. Thi…
-
New metric quantifies LLM training power elasticity for grid-responsive AI infrastructure
A new research paper introduces the concept of "job power elasticity" to characterize how LLM training performance is affected by reduced GPU power. The study proposes a "Power Flexibility Index" (PFI) to quantify this …
-
NVIDIA vLLM supports DeepSeekv4.1 Flash on release; AMD vLLM lags
NVIDIA's vLLM software is functioning seamlessly with the new DeepSeekv4.1 Flash model across all six of its hardware SKUs, including H100, H200, B200, B300, GB200, and GB300. In contrast, AMD's vLLM implementation is e…
-
Qwen 3 4B Base model sees 31% boost on MATH-500 after puzzle fine-tuning
A fine-tuned version of the Qwen 3 4B Base model demonstrated a 31% improvement on the MATH-500 benchmark after being trained on 100 zebra puzzles. The process for reproducing this result, which took approximately 6.5 m…
-
Self-hosting LLMs: Hidden costs and utilization challenges
The decision between using closed frontier LLM APIs, hosted open-weight APIs, or self-hosting open-weight models is complex. While self-hosting might seem cost-effective due to lower per-token costs, the actual savings …
-
Crypto GPU rental economics: Hosting LLMs vs. mining
Renting out GPUs for cryptocurrency mining and hosting AI models presents a complex economic landscape. While renting personal GPUs can yield modest daily returns, factors like electricity costs, depreciation, and low u…
-
LLMs accelerate Persian chat anonymization with efficient NER training
Researchers have developed a method for efficient anonymization of Persian customer chats using LLM-labeled data. They compared three instruction-tuned LLMs—DeepSeek-V3-0324, GPT-OSS-120B, and Qwen3-235B-A22B-Instruct-2…
-
GLM-5.3 Cybersecurity variant released with reduced refusals
A new cybersecurity-focused variant of the GLM-5.3 model, named GLM-5.3-CYBERSECURITY-FP8, has been released by dealignai. This model is designed to reduce refusals for offensive security tasks, red-teaming, and exploit…
-
New codebook layout enables massive self-organizing maps on single GPU
Researchers have developed a novel feature-major codebook layout for sparse-binary self-organizing maps, significantly improving memory efficiency and training speed. This optimization allows for the creation of much la…
-
iFlytek CEO: Domestic compute limits long context, lags NVIDIA H200
iFlytek CEO Liu Qingfeng stated that domestic computing power limitations are hindering the development of extremely long context windows for AI models. He also noted that the company's training efficiency lags behind N…
-
Top GPU Neoclouds 2026: CoreWeave, Nebius, Lambda, Crusoe, Groq Ranked
A 2026 ranking of GPU neocloud providers highlights CoreWeave as the sole Platinum-rated option, commanding premium pricing for high-end NVIDIA GPUs. Nebius leads in publishing on-demand pricing for the latest B300 chip…
-
NVIDIA Vera Rubin NVL72 boosts AI agent efficiency by 30x, expands ecosystem with MediaTek deal · 10 sources tracked
NVIDIA has unveiled its Vera Rubin NVL72 system, which reportedly offers up to 30 times greater throughput per megawatt for AI agent workloads compared to previous NVIDIA GB300 NVL72 systems. This significant efficiency…
-
China allows NVIDIA H200 chip imports to boost domestic AI competition with US
China has permitted limited shipments of NVIDIA's H200 chips to enter the country. This move aims to support domestic AI companies in their efforts to compete with the United States in the rapidly advancing AI sector. T…
-
Nvidia H200 AI chips reach China, but Hong Kong deployment faces hurdles · 3 sources tracked
Nvidia has begun shipping its H200 AI accelerators to China, with major tech companies ByteDance and Tencent reportedly receiving initial batches of approximately 10,000 units each. Despite these shipments, Beijing is d…
-
China permits ByteDance, Tencent to import 10,000 NVIDIA H200 AI chips
China has reportedly permitted ByteDance and Tencent to import 10,000 NVIDIA H200 chips each, a move aimed at bolstering their AI model development capabilities. This decision appears to relax previous US-imposed restri…
-
Frontier AI agents demand terabytes of VRAM, making self-hosting infeasible
Self-hosting a frontier-class AI agent requires substantial VRAM, far exceeding typical consumer hardware capabilities. A 685-billion-parameter model, for instance, needs approximately 800 GB of VRAM just for its weight…