New research explores advanced LLM routing techniques for efficiency and performance · 8 sources tracked
ByPulseAugur Editorial·[14 sources]·
Multiple research papers explore advanced routing techniques for large language models (LLMs) to improve efficiency and performance. VANE focuses on scoring the update rather than the token for better expert selection in Mixture-of-LoRA-experts. Semantic Routing Calibration (SRC) addresses LLM over-refusal by dynamically suppressing hypersensitive safety heads. FlexRouter optimizes for answer coverage by modeling model complementarity using Determinantal Point Processes, while Just Initialize provides a training-free component for large-scale routing optimization. RouteFM aims to be a foundation model for LLM routing, enabling reuse across different environments, and SaveRouter focuses on reducing supervision costs for economical LLM routing. Finally, market-aware routing is proposed for open-weight LLM inference, considering provider choice alongside model selection.
AI
IMPACT
These advancements in LLM routing could lead to more efficient and cost-effective AI inference, improving response quality and reducing computational overhead.
RANK_REASON
Multiple research papers published on arXiv detailing new methods for LLM routing.
arXiv:2610.00493v1 Announce Type: new Abstract: Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them rou…
arXiv cs.AI
TIER_1English(EN)·Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo·
arXiv:2609.25049v2 Announce Type: replace-cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely…
arXiv cs.AI
TIER_1English(EN)·Wang Wei, Harry Yang, Tiankai Yang, Samyadeep Basu, Hongjie Chen, Andy Zhao, Franck Dernoncourt, Ryan A. Rossi, Hoda Eldardiry·
arXiv:2609.38585v1 Announce Type: cross Abstract: Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundan…
arXiv:2609.35443v2 Announce Type: replace Abstract: Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational co…
arXiv cs.LG
TIER_1English(EN)·Yuchen Li, Mingyu Du, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong·
arXiv:2609.36724v1 Announce Type: new Abstract: Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose l…
arXiv:2609.37362v1 Announce Type: new Abstract: Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned throug…
arXiv cs.AI
TIER_1English(EN)·Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye·
arXiv:2609.37402v1 Announce Type: new Abstract: Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on histor…
arXiv:2609.37902v1 Announce Type: new Abstract: Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider s…
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping…
Existing Large Language Model (LLM) routing methods score LLMs independently to select top-k models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the over…
Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a par…
Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality fee…
dev.to — LLM tag
TIER_1English(EN)·Abdulsalam Abdulsalam·
<p>Every LLM tutorial ends the same way. You get back a blob of text, and now you have to parse it. You beg the model for JSON. It hands you JSON wrapped in an apology. You write a regex. The regex breaks on the next prompt. You add a retry. The retry costs you another second and…
dev.to — LLM tag
TIER_1English(EN)·Sanskriti Harmukh·
<p><a href="https://portkey.ai/" rel="noopener noreferrer">Portkey</a> is an open-source AI gateway that unifies access to 250+ Large Language Models (LLMs) across dozens of providers behind a single, OpenAI-compatible API. Instead of integrating with each provider's SDK, applica…