Sarathi-Serve
PulseAugur coverage of Sarathi-Serve — every cluster mentioning Sarathi-Serve across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
LLM serving splits into two phases to double GPU throughput
Modern LLM serving architectures are evolving to handle requests more efficiently by splitting the process into two distinct phases: prefill and decode. The prefill phase, which processes the entire prompt, is compute-b…
-
LLM serving scheduler extended to support multi-tier SLAs
Researchers have extended the Llumnix LLM serving scheduler to support more than two priority tiers, enabling it to better manage heterogeneous service-level objectives (SLOs). The enhanced scheduler was evaluated using…
-
LLM inference and reasoning techniques advance with new research and hardware
Researchers are exploring novel methods to enhance the efficiency and reasoning capabilities of large language models (LLMs). Google Research is developing techniques to train LLMs to reason in a Bayesian manner, improv…