Researchers have explored on-device tool routing for AI assistants, distinguishing between tool selection and abstention (when no tool applies). Traditional methods using a single language model are expensive in terms of latency and memory on local devices. An alternative approach replaces the language model with a retriever, which ranks local actions but cannot signal when no action is suitable. The study found that while BM25 can effectively select tools for lexically matched requests, a neural component is crucial for accurate abstention. Using a frozen encoder like multilingual-e5-base for abstention alone kept many requests local while correctly identifying those needing delegation, though a neural ranker improved overall quality at the cost of increased latency and memory. AI
IMPACT Highlights the trade-offs between neural models and retrievers for on-device AI, suggesting neural components are key for abstention.
RANK_REASON Academic paper detailing a novel approach to AI assistant tool routing. [lever_c_demoted from research: ic=1 ai=1.0]
- AI Assistant
- alphaXiv
- arXiv
- BM25
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- language model
- multilingual-e5-base
- Retriever
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →