PulseAugur
EN
LIVE 02:13:28

AMD MI355X closes performance gap with GB300 via SGLang and UMBP integration · 3 sources tracked

SemiAnalysis reports that AMD's MI355X is rapidly improving in agentic inference performance and total cost of ownership, nearing parity with GB300. This advancement is attributed to AMD's SGLang team and their MoRI library, with significant contributions from Universitas Mandiri Bina Prestasi (UMBP). UMBP's integration into SGLang as a KV offloading backend, specifically through the KVCache Store Linker, optimizes the data path between High Bandwidth Memory (HBM) and DRAM, reducing data copies and improving efficiency. AI

IMPACT Accelerates agentic inference capabilities and potentially lowers costs for AI deployments by improving hardware efficiency.

RANK_REASON The cluster reports on a specific hardware product (AMD MI355X) achieving significant performance and TCO improvements, directly comparing it to a competitor (GB300) and detailing the underlying technological advancements.

Read on X — SemiAnalysis →

AI-generated summary · Google Gemini · from 3 sources. How we write summaries →

AMD MI355X closes performance gap with GB300 via SGLang and UMBP integration · 3 sources tracked

How we ranked this

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
The cluster reports on a specific hardware product (AMD MI355X) achieving significant performance and TCO improvements, directly comparing it to a competitor (GB300) and detailing the underlying te…
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [3]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    HiCache sits between HBM and the external DRAM KV store as a per-rank host-side L2 buffer, so every load and offload takes an extra copy, the L3-to-L2 fetch is

    HiCache sits between HBM and the external DRAM KV store as a per-rank host-side L2 buffer, so every load and offload takes an extra copy, the L3-to-L2 fetch is not pipelined with compute, and under MLA with TP each rank replicates the same KV across PCIe and host memory while

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    These rapid improvement can be attributed to both AMD's amazing SGLang team as well as their incredibly based, first-principles MoRI (Modular RDMA Interface) li

    These rapid improvement can be attributed to both AMD's amazing SGLang team as well as their incredibly based, first-principles MoRI (Modular RDMA Interface) library. These particular improvements also reflect UMBPs integration in SGLang as a KV offloading backend. (2/3) https://…

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    BREAKING: AMD MI355X is quickly closing the perf/TCO gap in agentic inference, when compared to GB300 apples-to-apples. (1/3)🧵 https://t.co/xmmgaeifDv

    BREAKING: AMD MI355X is quickly closing the perf/TCO gap in agentic inference, when compared to GB300 apples-to-apples. (1/3)🧵 https://t.co/xmmgaeifDv