PulseAugur
EN
LIVE 07:40:33
中文(ZH) “接力跑”盘活全国算力,PD分离终于破局:延迟砍半、成本直降近40%!

Infinigence unveils PDD architecture to cut LLM inference latency and cost

Infinigence has unveiled a new cross-cluster heterogeneous inference architecture called PDD, designed to improve the efficiency and reduce the cost of large model inference. This architecture breaks down the traditional Prefill-Decode (PD) separation into a three-stage Prefill-RelayDecode-MainDecode (P-RLD-MD) system. By utilizing local RelayDecode instances, PDD effectively masks the latency associated with transmitting KV Cache over wide-area networks, leading to a reported 51.5% reduction in first-token latency and a 37.5% decrease in per-token cost. AI

IMPACT This architecture could significantly reduce the operational costs and improve the user experience of large language models by optimizing inference efficiency.

RANK_REASON The cluster describes a new technical architecture for LLM inference, including its design principles and performance claims, presented at a conference. [lever_c_demoted from research: ic=1 ai=1.0]

Read on 量子位 (QbitAI) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Infinigence unveils PDD architecture to cut LLM inference latency and cost

COVERAGE [1]

  1. 量子位 (QbitAI) TIER_1 中文(ZH) · 思邈 ·

    Relay race revitalizes national computing power, PD separation finally breaks through: latency halved, costs reduced by nearly 40%!

    最新完整技术报告出炉