Researchers have developed SAGE, a new adaptive retrieval policy for production Retrieval-Augmented Generation (RAG) systems. SAGE dynamically adjusts the number of passages retrieved per query based on estimated query difficulty, aiming to meet strict service level objectives (SLOs) for latency and cost. The system uses lightweight features from initial retrieval and is trained offline, adding minimal overhead at inference. Experiments show SAGE significantly improves SLO compliance and reduces latency and cost compared to static baselines, while maintaining answer quality across various datasets and LLM families. AI
IMPACT Optimizes RAG systems for production environments by improving latency and cost efficiency, potentially enabling wider adoption.
RANK_REASON The cluster contains a research paper detailing a new method for retrieval-augmented generation systems.
Read on arXiv cs.IR (Information Retrieval) →
- arXiv
- Gemma
- HotpotQA
- llama
- Mistral AI
- Muhammad Faizan Raza
- Natural Questions
- Qwen
- retrieval-augmented generation
- Sage
- UnSeenTimeQA
- SLO-Aware Adaptive Retrieval
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →