NVIDIA B300
PulseAugur coverage of NVIDIA B300 — every cluster mentioning NVIDIA B300 across labs, papers, and developer communities, ranked by signal.
-
Kimi K3 LLM self-hosting costs $89.52/hr, offers 1M context
Kimi K3, a large language model developed by Moonshot AI, has provided a performance report from its own operational environment. Running on eight NVIDIA B300 SXM6 GPUs with a total of 2.2 TB of VRAM, the model boasts a…
-
Together AI partners with IBM and NVIDIA for enterprise AI inference
Together AI has announced a multi-year partnership with IBM and NVIDIA to enhance AI inference capabilities on IBM Cloud. This collaboration will leverage a dedicated cluster of NVIDIA B300 GPUs and Spectrum-X networkin…
-
NVIDIA B300 fine-tuning of Qwen3-32B detailed in new research
A new paper details the operational challenges and solutions encountered when fine-tuning the Qwen3-32B model on NVIDIA's B300 accelerators. The research focuses on practical aspects of multi-node training, offering ins…
-
AMD MI355X GPUs offer better performance per dollar for Kimi K3 model
A blog post from Wafer.ai details how they achieved better performance per dollar by running the Kimi K3 model on AMD's MI355X GPUs. Despite Kimi K3's massive 2.8T parameter size requiring significant VRAM, the MI355X, …
-
Thinking Machines releases Inkling-Small, outperforming larger predecessor
Thinking Machines Lab has launched Inkling-Small, a new open-weights multimodal model that prioritizes efficiency over sheer size. Despite being significantly smaller than its predecessor, Inkling, Inkling-Small demonst…
-
Thinking Machines releases Inkling, an efficient multimodal MoE model
Thinking Machines has released Inkling, a new open-weight, multimodal Mixture-of-Experts model with 975 billion total parameters and 41 billion active parameters. The model supports a 1 million token context window and …
-
TurboServe system optimizes streaming video generation serving
Researchers have developed TurboServe, a novel serving system specifically designed for streaming video generation. This system tackles challenges like session state management and dynamic resource allocation by integra…
-
Together AI adds thousands of NVIDIA B200/B300 chips for inference
Together AI has significantly expanded its cloud computing resources, adding thousands of new chips including NVIDIA's B200 and B300 accelerators. This move is aimed at bolstering their dedicated model inference service…