PulseAugur
实时 02:13:43
English(EN) HiCache sits between HBM and the external DRAM KV store as a per-rank host-side L2 buffer, so every load and offload takes an extra copy, the L3-to-L2 fetch is

AMD MI355X 通过 SGLang 和 UMBP 集成缩小了与 GB300 的性能差距 · 跟踪 3 个来源

SemiAnalysis 报道称,AMD 的 MI355X 在代理推理性能和总拥有成本方面正在迅速提高,接近 GB300 的水平。这一进步归功于 AMD 的 SGLang 团队及其 MoRI 库,Universitas Mandiri Bina Prestasi (UMBP) 也做出了重要贡献。UMBP 作为 KV 卸载后端集成到 SGLang 中,特别是通过 KVCache Store Linker,优化了 High Bandwidth Memory (HBM) 和 DRAM 之间的数据路径,减少了数据复制并提高了效率。 AI

影响 通过提高硬件效率,加速代理推理能力,并可能降低 AI 部署的成本。

排序理由 该集群报道了一款特定的硬件产品 (AMD MI355X) 在性能和 TCO 方面取得了显著的改进,并将其与竞争对手 (GB300) 进行了直接比较,详细介绍了其技术进步。

在 X — SemiAnalysis 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

AMD MI355X 通过 SGLang 和 UMBP 集成缩小了与 GB300 的性能差距 · 跟踪 3 个来源

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
该集群报道了一款特定的硬件产品 (AMD MI355X) 在性能和 TCO 方面取得了显著的改进,并将其与竞争对手 (GB300) 进行了直接比较,详细介绍了其技术进步。
Source corroboration
3 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [3]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    HiCache 作为每路器端的 L2 缓存,位于 HBM 和外部 DRAM KV 存储之间,因此每次加载和卸载都需要额外复制一份,L3 到 L2 的获取是

    HiCache sits between HBM and the external DRAM KV store as a per-rank host-side L2 buffer, so every load and offload takes an extra copy, the L3-to-L2 fetch is not pipelined with compute, and under MLA with TP each rank replicates the same KV across PCIe and host memory while

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    这些快速的改进可归因于AMD出色的SGLang团队以及他们极其基于第一性原理的MoRI(模块化RDMA接口)li

    These rapid improvement can be attributed to both AMD's amazing SGLang team as well as their incredibly based, first-principles MoRI (Modular RDMA Interface) library. These particular improvements also reflect UMBPs integration in SGLang as a KV offloading backend. (2/3) https://…

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    突发:AMD MI355X 在代理推理方面正迅速缩小与 GB300 的性能/TCO 差距(一对一比较)。(1/3)🧵 https://t.co/xmmgaeifDv

    BREAKING: AMD MI355X is quickly closing the perf/TCO gap in agentic inference, when compared to GB300 apples-to-apples. (1/3)🧵 https://t.co/xmmgaeifDv