PulseAugur
中
实时 11:31:21
中文(ZH) “接力跑”盘活全国算力,PD分离终于破局:延迟砍半、成本直降近40%!

Infinigence 发布 PDD 架构,降低 LLM 推理时延与成本

Infinigence 发布了一种名为 PDD 的新型跨集群异构推理架构,旨在提高大型模型推理的效率并降低成本。该架构将传统的 Prefill-Decode (PD) 分离分解为三阶段的 Prefill-RelayDecode-MainDecode (P-RLD-MD) 系统。通过利用本地的 RelayDecode 实例,PDD 有效地掩盖了通过广域网传输 KV Cache 所带来的时延,据称首次 Token 时延降低了 51.5%,每个 Token 的成本降低了 37.5%。 AI

影响 该架构通过优化推理效率,有望显著降低大型语言模型的运营成本并改善用户体验。

排序理由 该集群描述了一种新的 LLM 推理技术架构,包括其设计原则和性能声明,并在会议上发布。[lever_c_demoted from research: ic=1 ai=1.0]

在 量子位 (QbitAI) 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Infinigence 发布 PDD 架构,降低 LLM 推理时延与成本

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一种新的 LLM 推理技术架构,包括其设计原则和性能声明,并在会议上发布。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
62 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 思邈 ·

    接力赛激活国家算力,PD分离终获突破:时延减半,成本降低近四成!

    最新完整技术报告出炉