Together AI has gained access to NVIDIA's Vera Rubin NVL72 platform, which is based on the Blackwell architecture. Their team has updated their ThunderKittens software to leverage new features of the Vera Rubin chip, specifically focusing on optimizing NVFP4 and FP8 GEMMs (General Matrix Multiply operations). Initial tests show that while the new platform doubles the K dimension capacity for tensor cores and increases tensor memory, existing Blackwell-optimized kernels were not feeding the cores fast enough, achieving only about 42-44% of theoretical performance. Together AI's optimizations aim to push performance closer to the theoretical ceiling, targeting over 22 PFLOPS and competitive speeds with existing libraries like cuBLAS. AI
IMPACT Optimizations for NVIDIA's Blackwell architecture could improve AI training and inference performance on compatible hardware.
RANK_REASON Software optimization for existing hardware architecture.
- Blackwell
- cuBLAS
- CuTE DSL
- FP8
- GEMM
- Hopper
- NVFP4
- Nvidia
- NVIDIA Vera Rubin NVL72
- ThunderKittens
- Together AI
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →