AMD's MI355X graphics card has demonstrated superior performance over Nvidia's B200 in vLLM benchmarks for the Kimi K2.5 model, a significant achievement driven by community-developed kernels. This advancement stems from a $1.1 million kernel hackathon organized by AMD and GPU_MODE, which resulted in a more than 4x improvement in end-to-end MI355X performance through optimizations in MoE, Top-K, and tensor-parallel kernels. While AMD's vLLM performance on Kimi models still trails, these upstreamed kernel improvements to AMD's AITER library and the ATOM inference engine signal a positive trajectory towards parity with CUDA vLLM. AI
IMPACT Demonstrates the potential for community-driven optimization to close performance gaps in AI hardware, potentially influencing future hardware development and software integration.
RANK_REASON Community-driven kernel optimization leading to a benchmark improvement for AMD hardware against a competitor.
- AMD
- AnushElangovan
- ATOM
- CUDA
- GPUMODE Readonflow Team
- marksaroufim
- MI355X
- ROCm
- vLLM
- AITER kernel library
- ATOM inference engine
- Kimi K2.5
- Nvidia B200
- Readonflow Team
- ROCm stack
AI-generated summary · Google Gemini · from 7 sources. How we write summaries →