PulseAugur
EN
LIVE 19:02:32

Together optimizes inference on Nvidia Rubin GPU with new features · 4 sources tracked

Together has announced advancements in their inference and OSS capabilities, leveraging new features from Nvidia's Rubin GPU. The company has updated its b200 gemms to incorporate Rubin's wider MMA steps, increased tensor and shared memory, and B-side collectors. These optimizations have led to significant performance gains, with their 16k nvfp4 gemm reaching 22.4 pflops and ThunderKittens achieving competitive speeds on the new hardware. AI

IMPACT Optimizations for Nvidia's Rubin GPU are likely to accelerate AI inference performance and enable more efficient training of large models.

RANK_REASON The cluster details hardware optimizations and performance benchmarks for AI inference, involving a major hardware vendor (Nvidia) and a significant AI software/infrastructure company (Together).

Read on X — Together (inference / OSS) →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

Together optimizes inference on Nvidia Rubin GPU with new features · 4 sources tracked

How we ranked this

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
The cluster details hardware optimizations and performance benchmarks for AI inference, involving a major hardware vendor (Nvidia) and a significant AI software/infrastructure company (Together).
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [4]

  1. X — Together (inference / OSS) TIER_1 Français(FR) · togethercompute ·

    Check out our blog for more details: https://t.co/PQ8OTa779h

    check out our blog for more details: https://t.co/PQ8OTa779h

  2. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    in the blog, we start with our original b200 gemms and progressively add rubin’s new features — wider mma steps, more tensor + shared memory, b-side collectors,

    in the blog, we start with our original b200 gemms and progressively add rubin’s new features — wider mma steps, more tensor + shared memory, b-side collectors, early a release, and more these kernels are still early. lut gemms, hardware-native megakernels, and new engine https:…

  3. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    one of the biggest changes on rubin is that tensor cores can consume operands twice as fast — but simply widening the mma wasn’t enough

    one of the biggest changes on rubin is that tensor cores can consume operands twice as fast — but simply widening the mma wasn’t enough we had to increase data reuse, deepen the pipeline, and more to keep them fed, ultimately pushing our 16k nvfp4 gemm to 22.4 pflops https://t.c…

  4. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    thunderkittens is now running on nvidia vera rubin!

    thunderkittens is now running on nvidia vera rubin! our kernels team got early access to nvl72 and spent the past few days digging through the new isa and bringing nvfp4 + fp8 gemms to life after reworking the kernels for rubin, we pushed them past 22 and 12 pflops respectively…