Infinigence has unveiled a new cross-cluster heterogeneous inference architecture called PDD, designed to improve the efficiency and reduce the cost of large model inference. This architecture breaks down the traditional Prefill-Decode (PD) separation into a three-stage Prefill-RelayDecode-MainDecode (P-RLD-MD) system. By utilizing local RelayDecode instances, PDD effectively masks the latency associated with transmitting KV Cache over wide-area networks, leading to a reported 51.5% reduction in first-token latency and a 37.5% decrease in per-token cost. AI
IMPACT This architecture could significantly reduce the operational costs and improve the user experience of large language models by optimizing inference efficiency.
RANK_REASON The cluster describes a new technical architecture for LLM inference, including its design principles and performance claims, presented at a conference. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →