Together has announced advancements in their inference and OSS capabilities, leveraging new features from Nvidia's Rubin GPU. The company has updated its b200 gemms to incorporate Rubin's wider MMA steps, increased tensor and shared memory, and B-side collectors. These optimizations have led to significant performance gains, with their 16k nvfp4 gemm reaching 22.4 pflops and ThunderKittens achieving competitive speeds on the new hardware. AI
IMPACT Optimizations for Nvidia's Rubin GPU are likely to accelerate AI inference performance and enable more efficient training of large models.
RANK_REASON The cluster details hardware optimizations and performance benchmarks for AI inference, involving a major hardware vendor (Nvidia) and a significant AI software/infrastructure company (Together).
Read on X — Together (inference / OSS) →
- b200 gemms
- Cublas
- CuTe-DSL
- nvfp4 gemm
- Nvidia
- Nvidia Rubin Gpu
- NVL72
- Tensor Cores
- ThunderKittens
- Together
AI-generated summary · Google Gemini · from 4 sources. How we write summaries →