FlashAttention-3
PulseAugur coverage of FlashAttention-3 — every cluster mentioning FlashAttention-3 across labs, papers, and developer communities, ranked by signal.
4 day(s) with sentiment data
-
FlashAttention-2 & 3 Tackle GPU Memory Limits for LLMs
This article delves into the technical advancements of FlashAttention-2 and FlashAttention-3, explaining how they overcome the limitations of GPU memory bandwidth. It details the use of algorithmic tiling and asynchrono…
-
GRACE system accelerates real-time ad retrieval with generative recommenders
A new research paper introduces GRACE, a system designed to accelerate generative recommenders for real-time ad retrieval. GRACE addresses challenges in eligibility and compute by implementing Generative Target Matching…
-
New X-Stage pipeline optimization boosts DiT inference speed
Researchers have identified a new pipeline stage, termed X-Stage, that can optimize communication-computation overlap during the inference of Diffusion Transformers (DiTs). This stage focuses on the period after communi…
-
Sol-Attn speeds up video generation with efficient sparse attention
Researchers have developed Sol-Attn, a new training-free sparse attention method designed to accelerate inference for video generation models. Unlike previous methods that struggle with efficiency and accuracy due to ri…
-
New framework speeds up LLM inference on NVIDIA H20 GPUs
Researchers have developed FlashMLA-ETAP, a new framework designed to significantly speed up the inference of large language models on NVIDIA H20 GPUs. The framework introduces an Efficient Transpose Attention Pipeline …
-
ByteDance releases Bernini open-source video generation framework
ByteDance has released Bernini, an open-source framework for video generation and editing. The system combines a multimodal large language model for semantic planning with a DiT-based renderer. Bernini reportedly achiev…
-
Together AI releases FlashAttention-3 and -4 for faster LLM processing
Together AI has released FlashAttention-3 and FlashAttention-4, significant upgrades to their GPU-accelerated attention mechanism for large language models. FlashAttention-3, designed for Hopper GPUs, achieves up to 75%…
-
Together AI boosts AI training 90% with NVIDIA Blackwell
Together AI has launched new GPU clusters featuring NVIDIA's Blackwell platform, offering significant speedups for AI training and inference. These clusters, powered by the Together Kernel Collection, achieve up to 90% …