A new paper proposes that optimizing the order of retrieved text pieces in retrieval-augmented generation (RAG) systems can significantly improve Large Language Model (LLM) serving efficiency. The research demonstrates that current fixed ordering conventions are suboptimal, especially with multiple retrieved pieces, and introduces a structure theorem that equates optimal ordering to selecting a hierarchy over requests. An exact algorithm and a 1/2-approximation using agglomerative clustering are presented, showing substantial prefill reduction on benchmark datasets. AI
IMPACT Optimizing prompt ordering in RAG systems could lead to more efficient LLM serving and reduced computational costs.
RANK_REASON Academic paper detailing a novel theoretical approach to optimizing LLM serving. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- Beir
- BM25
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- retrieval-augmented generation
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →