PulseAugur
实时 07:30:25
English(EN) @sgl_project The first AMD DI CI PR, #29084, was implemented three months after SemiAnalysis’s initial request and multiple meetings with @AnushElangovan and Va

SemiAnalysis 批评 AMD 领导层阻碍 vLLM 进展

SemiAnalysis 批评 AMD 领导层将计算资源从其内部 vLLM 团队转移,导致 vLLM 自动化测试的开发出现倒退。这种集群资源的转移正在阻碍 AMD 在实现与 CUDA 的 vLLM 测试能力相匹配方面取得进展。尽管面临这些挑战,AMD 的 DI CI 系统在最初的请求和会谈后得以实施,已成功识别出 DeepSeek V4 和 Kimi 等模型中的重大错误,提高了整体代码质量。 AI

影响 AMD 的内部资源分配决策可能会减缓其在开发具有竞争力的 AI 模型和测试基础设施方面的进展。

排序理由 该集群包含 SemiAnalysis 关于 AMD 内部资源分配及其对 AI 发展影响的关键推文,而不是 AMD 的官方公告或发布。

在 X — SemiAnalysis 阅读 →

AI 生成摘要 · Google Gemini · 来自 7 个来源。 我们如何撰写摘要 →

SemiAnalysis 批评 AMD 领导层阻碍 vLLM 进展

报道来源 [7]

  1. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @sgl_project @AnushElangovan @LisaSu 我们希望AMD领导层能够优先为内部vLLM团队提供稳定的集群,以便AMD的硬核实习生

    @sgl_project @AnushElangovan @LisaSu We hope AMD leadership can reprioritize providing its internal vLLM team with stable clusters so that AMD’s hardcore internal vLLM engineers can do the work required to reach 90%+ gating parity with CUDA vLLM. 7/7🧵

  2. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    门禁/阻塞测试意味着测试质量最高,因为除非测试通过,否则无法合并PR。虽然A

    @sgl_project @AnushElangovan @LisaSu Gating/blocking tests mean that the tests are of the highest quality, since PRs cannot merge unless the tests pass. While AMD leadership may distract non-technical folks by showing its non-gating pass rate, gating parity and the gating pass ra…

  3. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    vLLM 方面,由于 AMD 集群基础设施不稳定,自动 vLLM 门控测试的进展已大幅回退

    @sgl_project @AnushElangovan @LisaSu On the vLLM side, progress on automated vLLM gating tests has massively regressed due to AMD cluster infrastructure stability issues. AMD’s hardcore engineers had been making good progress on vLLM gating over the past couple of weeks until AMD…

  4. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    为使 AMD DI CI 达到与 CUDA SGLang 的同等水平,SGLang 项目、Anush Elangovan 和 Lisa Su 的团队还需要为 WideEP 解码优化实施夜间测试

    @sgl_project @AnushElangovan @LisaSu In order for AMD DI CI to reach parity with CUDA SGLang, it also needs to implement nightly tests for the WideEP decode optimization, on which it is making good progress in 31500. 4/7🧵

  5. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    @sgl_project @AnushElangovan @LisaSu 它捕获的第一个错误是 DeepSeek V4 上 #30336 中的 AMD MoRI 缓冲区 MR 错误。第二个是由 b 引起的 Kimi K2.6 错误

    @sgl_project @AnushElangovan @LisaSu The first bug it caught was an AMD MoRI buffer MR error in #30336 on DeepSeek V4. The second was a Kimi K2.6 error caused by an AMD AITER kernel regression in #30433 that made Kimi’s math accuracy massively lower. Through AMD’s nightly CI, the…

  6. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    AMD DI CI PR #29084 在 SemiAnalysis 最初请求并与 @AnushElangovan 和 Va 多次会面三个月后实现

    @sgl_project The first AMD DI CI PR, #29084, was implemented three months after SemiAnalysis’s initial request and multiple meetings with @AnushElangovan and Vamsi (Head of AI), as well as a meeting with @LisaSu early in the year regarding improvements to code quality. We will wa…

  7. X — SemiAnalysis TIER_1 English(EN) · SemiAnalysis_ ·

    AMD @sgl_project 团队在实现夜间分离服务 CI 以提高代码质量方面做得非常出色!它已经捕获并阻止了 2 个重大错误

    Great work by the AMD @sgl_project team on enabling nightly disaggregated serving CI to improve code quality! It has already caught and prevented 2 massive bugs from reaching customers, as we explained before 👇️ 1/7🧵 https://t.co/cUCr6niKEt