PulseAugur
实时 10:14:38
English(EN) SAGE: SLO-Aware Adaptive Retrieval for Production RAG Systems

SAGE系统优化RAG检索以降低延迟和成本 · 跟踪2个来源

研究人员开发了SAGE,一种面向生产检索增强生成(RAG)系统的新型自适应检索策略。SAGE根据估计的查询难度动态调整每个查询检索到的文档数量,旨在满足延迟和成本方面的严格服务水平目标(SLO)。该系统使用初始检索的轻量级特征,并进行离线训练,在推理时增加的开销极小。实验表明,SAGE在各种数据集和LLM家族中,在保持答案质量的同时,显著提高了SLO合规性,并降低了延迟和成本,与静态基线相比。 AI

影响 通过提高延迟和成本效率,优化了生产环境中的RAG系统,可能促进更广泛的应用。

排序理由 该集群包含一篇详细介绍检索增强生成系统新方法的论文。

在 arXiv cs.IR (Information Retrieval) 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

SAGE系统优化RAG检索以降低延迟和成本 · 跟踪2个来源

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Muhammad Faizan Raza (Luna), Shuo (Luna), Yang, Satish Mahadevan Srinivasan ·

    SAGE: 面向生产RAG系统的SLO感知自适应检索

    arXiv:2608.08237v1 Announce Type: cross Abstract: Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cost. However, standard retrieval pipelines rely on fixed retrieval budgets that ig…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Satish Mahadevan Srinivasan ·

    SAGE:面向生产RAG系统的SLO感知自适应检索

    Retrieval-Augmented Generation (RAG) systems in production operate under strict service level objectives (SLOs) on tail latency and infrastructure cost. However, standard retrieval pipelines rely on fixed retrieval budgets that ignore query difficulty, over-retrieving for easy qu…