Researchers have developed a new method called Direct-P to optimize FlashAttention-4 for Blackwell's 4-bit floating-point (FP4) tensor cores, addressing performance bottlenecks caused by softmax conversion and on-chip dependencies. For noncausal inference, Direct-P maps scores directly to FP4 probabilities, achieving up to 2.13x the throughput of bfloat16 on NVIDIA GB200 hardware. A causal path for training reconstructs probabilities using FP8 gradient operands, accelerating an 8-billion-parameter update by up to 1.14x, though MXFP4 training trajectories were found to diverge. AI
IMPACT This research could lead to more efficient AI model training and inference by optimizing hardware utilization for lower-precision data types.
RANK_REASON The cluster describes a research paper detailing a new method for optimizing AI hardware performance.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →