ThunderKittens
PulseAugur coverage of ThunderKittens — every cluster mentioning ThunderKittens across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
Together optimizes inference on Nvidia Rubin GPU with new features · 4 sources tracked
Together has announced advancements in their inference and OSS capabilities, leveraging new features from Nvidia's Rubin GPU. The company has updated its b200 gemms to incorporate Rubin's wider MMA steps, increased tens…
-
Together AI optimizes ThunderKittens for NVIDIA Vera Rubin Blackwell GPUs
Together AI has gained access to NVIDIA's Vera Rubin NVL72 platform, which is based on the Blackwell architecture. Their team has updated their ThunderKittens software to leverage new features of the Vera Rubin chip, sp…
-
Stanford's ThunderKittens DSL optimizes AI kernel performance
A new article details ThunderKittens, a compact domain-specific language (DSL) developed at Stanford's Hazy Research Lab for creating high-performance AI kernels. The DSL aims to strike a balance between research produc…
-
Together AI kernels team optimizes GPUs with FlashAttention
The Together AI kernels team, including researchers Dan Fu and Tri Dao, developed FlashAttention, a software layer that significantly optimizes GPU performance for AI models. This breakthrough, achieved by applying data…