PulseAugur
中
实时 02:26:55
English(EN) @GPU_MODE This work has been upstreamed to the main branch of AMD’s AITER kernel library and, excitingly, has now also been upstreamed to vLLM. 4/6🧵

AMD MI355X 在 Kimi K2.5 模型上的 vLLM 性能超越 Nvidia B200

AMD 的 MI355X 图形卡在 Kimi K2.5 模型的 vLLM 基准测试中表现优于 Nvidia 的 B200,这是社区开发的内核驱动的重大成就。这一进步源于 AMD 和 GPU_MODE 组织的一项价值 110 万美元的内核黑客马拉松,通过在 MoE、Top-K 和张量并行内核方面的优化,使 MI355X 的端到端性能提高了 4 倍以上。虽然 AMD 在 Kimi 模型上的 vLLM 性能仍有差距,但这些已合并到 AMD 的 AITER 库和 ATOM 推理引擎的内核改进预示着其朝着与 CUDA vLLM 相当的积极方向发展。 AI

影响 展示了社区驱动的优化在缩小人工智能硬件性能差距方面的潜力,可能影响未来的硬件开发和软件集成。

排序理由 社区驱动的内核优化导致 AMD 硬件在基准测试中相对于竞争对手有所提升。

在 X — SemiAnalysis 阅读 →

AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →

AMD MI355X 在 Kimi K2.5 模型上的 vLLM 性能超越 Nvidia B200

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
社区驱动的内核优化导致 AMD 硬件在基准测试中相对于竞争对手有所提升。
Source corroboration
7 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
63 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [7]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    对于分离式上游vLLM性能,@GPU_MODE AMD在Kimi模型上仍落后,但我们同样期待那里的改进,此外

    @GPU_MODE For disaggregated upstream vLLM performance, AMD is still behind on Kimi models, but we are looking forward to improvements there as well, in addition to seeing AMD vLLM performance on Kimi K3. 5/6🧵

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE 此项工作已合并到 AMD 的 AITER 内核库主分支,并且令人兴奋的是,现已合并到 vLLM。4/6🧵

    @GPU_MODE This work has been upstreamed to the main branch of AMD’s AITER kernel library and, excitingly, has now also been upstreamed to vLLM. 4/6🧵

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE 他们专注于优化 W4A4 MoE 内核、Top-K 内核、张量并行 all-reduce 内核以及 MLA 解码元数据规划器。3/6🧵 https://t.co

    @GPU_MODE They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner. 3/6🧵 https://t.co/fTcErEnzjF

  4. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    AMD 与 @GPU_MODE 合作启动了耗资 110 万美元的内核黑客松,Readonflow 团队的内核改进了端到端的上游 MI355X 性能

    @GPU_MODE AMD launched a $1.1 mil kernel hackathon in collaboration with @GPU_MODE, and the Readonflow Team’s kernels improved end-to-end upstream MI355X performance by over 4x. 2/6🧵

  5. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @GPU_MODE 此项工作已合并到@AIatAMD的AITER内核库主分支和ATOM推理引擎。我们希望此项工作也能被广泛应用

    @GPU_MODE This work has been upstreamed to the main branch of @AIatAMD’s AITER kernel library and to the ATOM inference engine. We hope this work will also be upstreamed to vLLM so that AMD vLLM can reach performance parity with CUDA vLLM. 3/4🧵

  6. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    GPU模式:他们专注于优化W4A4 MoE内核、Top-K内核、张量并行all-reduce内核以及MLA解码元数据规划器。2/4🧵 https://t.co

    @GPU_MODE They focused on optimizing W4A4 MoE kernels, Top-K kernels, tensor-parallel all-reduce kernels, and the MLA decode metadata planner. 2/4🧵 https://t.co/SWFbh16Op0

  7. X — SemiAnalysis TIER_1 Deutsch(DE) · SemiAnalysis_ ·

    GPU_MODE 发布的价值 110 万美元的 AMD 内核黑客马拉松,干得漂亮 🚨。

    GREAT WORK BY @GPU_MODE 🚨 FOR LAUNCHING THE $1.1mil AMD KERNEL HACKATHON. The GPUMODE Readonflow Team’s kernels improved end-to-end MI355X performance by over 2x. We explain the optimizations below. 1/4🧵 https://t.co/2n1boLQi2j