PulseAugur
实时 01:57:52
English(EN) Serving Gemma 4 E2B on a TPU v6e-1: what Trillium buys, and what it doesn't

Google 的 TPU v6e-1 提供内存升级但成本更高

一项技术分析显示,Google 新的 Cloud TPU v6e-1 (Trillium) 在性能上优于 v5e-1,但其更高的成本使其在某些工作负载下性价比不高。v6e-1 提供双倍内存和 3.6 倍的 KV 缓存容量,但价格是 v5e-1 的 2.25 倍。对于未能充分利用增加内存的工作负载,v6e-1 的每个输出 token 的成本仅略微降低,甚至更高。分析还强调了新硬件配置方面的问题,包括跨区域的可用性不一致以及 spot 和 flex-start 选项之间出乎意料的定价差异。 AI

影响 提供了关于专用 AI 硬件成本效益权衡的见解,影响了 LLM 部署的基础设施决策。

排序理由 硬件性能和成本效益的技术分析。[lever_c_demoted from research: ic=1 ai=0.7]

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Google 的 TPU v6e-1 提供内存升级但成本更高

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · xbill ·

    在 TPU v6e-1 上部署 Gemma 4 E2B:Trillium 的所得与所失

    <h1> Serving Gemma 4 E2B on a TPU v6e-1 </h1> <p>A Cloud TPU v6e-1 (Trillium) costs <strong>2.25×</strong> a v5e-1 and returns <strong>1.62–1.68×</strong> the throughput on workloads that fit in a v5e, and <strong>2.32–2.77×</strong> on workloads that do not. Per output token tha…