A new research paper proposes a method for training small language models (SLMs) to act as multi-agent routers, improving the quality of search results. The approach uses supervised fine-tuning followed by reinforcement learning, incorporating retrieval relevance and query-agent topic alignment into a hierarchical reward function. This allows the SLM to learn when to select specific agents and when to redirect queries away from agents that produce low-relevance results, even if they appear topically aligned. The trained model achieved a significantly higher NDCG@10 score compared to LLM baselines that rely solely on intent-based routing, while also reducing selection latency. AI
IMPACT This research could lead to more efficient and accurate information retrieval systems by enabling smaller models to effectively manage complex multi-agent interactions.
RANK_REASON Research paper detailing a novel approach to training SLMs for multi-agent routing. [lever_c_demoted from research: ic=1 ai=1.0]
- Amazon Nova Lite
- arXiv
- Claude Haiku 4.5
- DagsHub
- Hugging Face
- Small Language Models
- Venkatashesha Gayathri Kondapalli
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →