PulseAugur
实时 11:11:29
中文(ZH) 英伟达Vera Rubin的Agent吞吐暴涨最高提升30倍,却不只靠GPU | Hot Chips 2026

Nvidia Vera Rubin 通过异构系统架构将 Agent 通量提升 30 倍 · 跟踪到 1 个来源

Nvidia 正在为其 Vera Rubin NVL72 系统优化 Agent 工作负载,与前几代相比,每兆瓦的吞吐量提高了高达 30 倍。这一性能提升并非完全归功于更快的 GPU,也源于一种异构计算方法,该方法将 Agent 任务分配给专用硬件。该系统现在集成了 Vera Rubin GPU 用于大型模型计算,Groq 3 LPX 用于低延迟 token 生成,Vera CPU 用于工具编排和数据处理,以及 Spectrum-X 网络来连接这些组件。 AI

影响 这种用于 Agent 工作负载的异构架构有望为 AI 推理设定新标准,优化跨不同任务的延迟和吞吐量。

排序理由 文章详细介绍了 AI 推理系统的一次重大架构转变,从单一 GPU 优化转向涉及专用处理器和网络的异构方法,以应对 Agent 工作负载。[lever_c_demoted from significant: ic=1 ai=1.0]

在 雷峰网 (Leiphone) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Nvidia Vera Rubin 通过异构系统架构将 Agent 通量提升 30 倍 · 跟踪到 1 个来源

本文如何被排名

Signal score
18 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
文章详细介绍了 AI 推理系统的一次重大架构转变,从单一 GPU 优化转向涉及专用处理器和网络的异构方法,以应对 Agent 工作负载。[lever_c_demoted from significant: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Nvidia Vera Rubin的Agent吞吐量飙升高达30倍,但并非仅依赖GPU | Hot Chips 2026

    <p><span style="font-family: Arial; text-align: left; font-size: 16px;">在Agent推理上,英伟达又把性能往前推了一大步。</span></p><p><span style="font-size: 16px;"><span style="font-family: Arial; font-size: 16px;">Hot Chips 2026期间,英伟达披露了Vera Rubin NVL72最新Agent工作负载测试结果:在DeepSeek V4-Pro、单用户每秒160 Token的…