PulseAugur
实时 04:30:00
English(EN) To Infinity and Beyond: ThunderKittens Now on NVIDIA Vera Rubin NVL72!

Together AI 为 NVIDIA Vera Rubin Blackwell GPU 优化 ThunderKittens

Together AI 已获得基于 Blackwell 架构的 NVIDIA Vera Rubin NVL72 平台的访问权限。他们的团队已更新其 ThunderKittens 软件,以利用 Vera Rubin 芯片的新功能,特别专注于优化 NVFP4FP8 GEMM(通用矩阵乘法运算)。初步测试表明,虽然新平台将张量核心的 K 维度容量翻倍并增加了张量内存,但现有的 Blackwell 优化内核无法足够快地为核心提供数据,仅达到理论性能的 42-44% 左右。Together AI 的优化旨在将性能推向理论上限,目标是超过 22 PFLOPS,并达到与 cuBLAS 等现有库相媲美的速度。 AI

影响NVIDIA Blackwell 架构的优化可以提高兼容硬件上的 AI 训练和推理性能。

排序理由 针对现有硬件架构的软件优化。

在 Together AI blog 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Together AI 为 NVIDIA Vera Rubin Blackwell GPU 优化 ThunderKittens

本文如何被排名

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
针对现有硬件架构的软件优化。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
1 days old
Coverage has settled into its steady-state source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. Together AI blog TIER_1 English(EN) ·

    飞向无限远:ThunderKittens 现已登陆 NVIDIA Vera Rubin NVL72!

    We ported ThunderKittens to NVIDIA's Vera Rubin NVL72 and rebuilt our NVFP4 GEMM around the new hardware, taking it from 42% of roofline to over 22 PFLOPS — competitive with cuBLAS and CuTe DSL. Here is what changed in the ISA and how we used it.