PulseAugur
EN
LIVE 16:54:24

AMD MI355X vLLM performance beats Nvidia B200 on Kimi K2.5 model

AMD's MI355X graphics card has demonstrated superior performance over Nvidia's B200 in vLLM benchmarks for the Kimi K2.5 model, a significant achievement driven by community-developed kernels. This advancement stems from a $1.1 million kernel hackathon organized by AMD and GPU_MODE, which resulted in a more than 4x improvement in end-to-end MI355X performance through optimizations in MoE, Top-K, and tensor-parallel kernels. While AMD's vLLM performance on Kimi models still trails, these upstreamed kernel improvements to AMD's AITER library and the ATOM inference engine signal a positive trajectory towards parity with CUDA vLLM. AI

IMPACT Demonstrates the potential for community-driven optimization to close performance gaps in AI hardware, potentially influencing future hardware development and software integration.

RANK_REASON Community-driven kernel optimization leading to a benchmark improvement for AMD hardware against a competitor.

Read on X — SemiAnalysis →

AI-generated summary · Google Gemini · from 7 sources. How we write summaries →

AMD MI355X vLLM performance beats Nvidia B200 on Kimi K2.5 model

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
Community-driven kernel optimization leading to a benchmark improvement for AMD hardware against a competitor.
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
58 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [7]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE For disaggregated upstream vLLM performance, AMD is still behind on Kimi models, but we are looking forward to improvements there as well, in addition

    @GPU_MODE For disaggregated upstream vLLM performance, AMD is still behind on Kimi models, but we are looking forward to improvements there as well, in addition to seeing AMD vLLM performance on Kimi K3. 5/6🧵

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE This work has been upstreamed to the main branch of AMD’s AITER kernel library and, excitingly, has now also been upstreamed to vLLM. 4/6🧵

    @GPU_MODE This work has been upstreamed to the main branch of AMD’s AITER kernel library and, excitingly, has now also been upstreamed to vLLM. 4/6🧵

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner. 3/6🧵 https://t.co

    @GPU_MODE They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner. 3/6🧵 https://t.co/fTcErEnzjF

  4. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE AMD launched a $1.1 mil kernel hackathon in collaboration with @GPU_MODE, and the Readonflow Team’s kernels improved end-to-end upstream MI355X perfor

    @GPU_MODE AMD launched a $1.1 mil kernel hackathon in collaboration with @GPU_MODE, and the Readonflow Team’s kernels improved end-to-end upstream MI355X performance by over 4x. 2/6🧵

  5. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE This work has been upstreamed to the main branch of @AIatAMD’s AITER kernel library and to the ATOM inference engine. We hope this work will also be u

    @GPU_MODE This work has been upstreamed to the main branch of @AIatAMD’s AITER kernel library and to the ATOM inference engine. We hope this work will also be upstreamed to vLLM so that AMD vLLM can reach performance parity with CUDA vLLM. 3/4🧵

  6. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner. 2/4🧵 https://t.co

    @GPU_MODE They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner. 2/4🧵 https://t.co/SWFbh16Op0

  7. X — SemiAnalysis TIER_1 Deutsch(DE) · SemiAnalysis_ ·

    GREAT WORK BY @GPU_MODE 🚨 FOR LAUNCHING THE $1.1mil AMD KERNEL HACKATHON.

    GREAT WORK BY @GPU_MODE 🚨 FOR LAUNCHING THE $1.1mil AMD KERNEL HACKATHON. The GPUMODE Readonflow Team’s kernels improved end-to-end MI355X performance by over 2x. We explain the optimizations below. 1/4🧵 https://t.co/2n1boLQi2j