PulseAugur
中
实时 06:25:45
English(EN) Astrolabe: Balancing Load in LLM Serving with Randomized Prediction-Guided Scheduling

Astrolabe系统通过随机预测引导调度优化大语言模型服务

研究人员开发了Astrolabe,一个旨在优化大语言模型(LLM)服务的新型调度系统。该系统采用随机预测引导的方法,在无需昂贵的基于迁移的重新平衡的情况下,平衡多个LLM实例之间的负载。Astrolabe结合了响应长度估计、延迟预测和power-of-two-choices调度策略,以增强负载平衡并降低延迟。实验表明,Astrolabe可以匹配或超越现有基线的性能,显著降低平均和P99的首个token时间和端到端延迟,同时还减少了预测器的CPU使用率。 AI

影响 优化大语言模型服务基础设施,可能降低成本并改善AI应用的响应时间。

排序理由 该集群包含一篇详细介绍LLM服务新调度系统的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Astrolabe系统通过随机预测引导调度优化大语言模型服务

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇详细介绍LLM服务新调度系统的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
infra, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
57 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Wei Da, Evangelia Kalyvianaki ·

    Astrolabe:通过随机预测指导调度来平衡 LLM 服务中的负载

    arXiv:2508.03611v3 Announce Type: replace-cross Abstract: This paper presents Astrolabe, a randomized prediction-guided scheduler for one-shot request dispatch in multi-instance large language model (LLM) serving. Astrolabe improves load balancing without relying on migration-bas…