PulseAugur
实时 19:02:32
English(EN) thunderkittens is now running on nvidia vera rubin!

Together 利用 Nvidia Rubin GPU 的新功能优化推理 · 跟踪 4 个来源

Together 宣布了其推理和 OSS 能力的进步,利用了 Nvidia Rubin GPU 的新功能。该公司已更新其 b200 gemms 以整合 Rubin 更宽的 MMA 步长、增加的张量和共享内存以及 B 侧收集器。这些优化带来了显著的性能提升,其 16k nvfp4 gemm 达到了 22.4 pflops,ThunderKittens 在新硬件上实现了具有竞争力的速度。 AI

影响Nvidia Rubin GPU 的优化可能会加速 AI 推理性能并实现更高效的大型模型训练。

排序理由 该集群详细介绍了 AI 推理的硬件优化和性能基准测试,涉及主要的硬件供应商 (Nvidia) 和重要的 AI 软件/基础设施公司 (Together)。

在 X — Together (inference / OSS) 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

Together 利用 Nvidia Rubin GPU 的新功能优化推理 · 跟踪 4 个来源

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
该集群详细介绍了 AI 推理的硬件优化和性能基准测试,涉及主要的硬件供应商 (Nvidia) 和重要的 AI 软件/基础设施公司 (Together)。
Source corroboration
4 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [4]

  1. X — Together (inference / OSS) TIER_1 Français(FR) · togethercompute ·

    更多详情请查看我们的博客:https://t.co/PQ8OTa779h

    check out our blog for more details: https://t.co/PQ8OTa779h

  2. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    在博客中,我们从最初的 b200 gemms 开始,逐步添加 rubin 的新功能 — 更宽的 mma 步长、更多的 tensor + 共享内存、b 端收集器,

    in the blog, we start with our original b200 gemms and progressively add rubin’s new features — wider mma steps, more tensor + shared memory, b-side collectors, early a release, and more these kernels are still early. lut gemms, hardware-native megakernels, and new engine https:…

  3. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    rubin上最大的变化之一是张量核心可以两倍速消耗操作数——但仅仅扩大mma是不够的

    one of the biggest changes on rubin is that tensor cores can consume operands twice as fast — but simply widening the mma wasn’t enough we had to increase data reuse, deepen the pipeline, and more to keep them fed, ultimately pushing our 16k nvfp4 gemm to 22.4 pflops https://t.c…

  4. X — Together (inference / OSS) TIER_1 English(EN) · togethercompute ·

    thunderkittens 现已运行在 nvidia vera rubin 上!

    thunderkittens is now running on nvidia vera rubin! our kernels team got early access to nvl72 and spent the past few days digging through the new isa and bringing nvfp4 + fp8 gemms to life after reworking the kernels for rubin, we pushed them past 22 and 12 pflops respectively…