PulseAugur
EN
LIVE 23:04:21
中文(ZH) “接力跑”盘活全国算力,PD分离终于破局:延迟砍半、成本直降近40%!

Infinigence unveils PDD architecture to cut LLM inference latency and cost

Infinigence has unveiled a new cross-cluster heterogeneous inference architecture called PDD, designed to improve the efficiency and reduce the cost of large model inference. This architecture breaks down the traditional Prefill-Decode (PD) separation into a three-stage Prefill-RelayDecode-MainDecode (P-RLD-MD) system. By utilizing local RelayDecode instances, PDD effectively masks the latency associated with transmitting KV Cache over wide-area networks, leading to a reported 51.5% reduction in first-token latency and a 37.5% decrease in per-token cost. AI

IMPACT This architecture could significantly reduce the operational costs and improve the user experience of large language models by optimizing inference efficiency.

RANK_REASON The cluster describes a new technical architecture for LLM inference, including its design principles and performance claims, presented at a conference. [lever_c_demoted from research: ic=1 ai=1.0]

Read on 量子位 (QbitAI) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Infinigence unveils PDD architecture to cut LLM inference latency and cost

How we ranked this

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The cluster describes a new technical architecture for LLM inference, including its design principles and performance claims, presented at a conference. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
58 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

Full methodology in our editorial standards.

COVERAGE [1]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 思邈 ·

    Relay race revitalizes national computing power, PD separation finally breaks through: latency halved, costs reduced by nearly 40%!

    最新完整技术报告出炉