English(EN)FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing
新研究探索先进的LLM路由技术以提高效率和性能 · 已追踪8个来源
作者PulseAugur 编辑部·[14 个来源]·
多篇研究论文探讨了用于大型语言模型(LLM)的先进路由技术,以提高效率和性能。VANE专注于对更新而非token进行评分,以更好地选择Mixture-of-LoRA-experts中的专家。语义路由校准(SRC)通过动态抑制超敏感的安全头来解决LLM过度拒绝的问题。FlexRouter通过使用行列式点过程(Determinantal Point Processes)对模型互补性进行建模来优化答案覆盖率,而Just Initialize为大规模路由优化提供了一个无需训练的组件。RouteFM旨在成为LLM路由的基础模型,实现跨不同环境的重用,而SaveRouter则专注于降低监督成本以实现经济高效的LLM路由。最后,提出了面向市场的路由,用于开放权重LLM推理,在模型选择的同时考虑提供商选择。
AI
arXiv:2610.00493v1 Announce Type: new Abstract: Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them rou…
arXiv cs.AI
TIER_1English(EN)·Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo·
arXiv:2609.25049v2 Announce Type: replace-cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely…
arXiv cs.AI
TIER_1English(EN)·Wang Wei, Harry Yang, Tiankai Yang, Samyadeep Basu, Hongjie Chen, Andy Zhao, Franck Dernoncourt, Ryan A. Rossi, Hoda Eldardiry·
arXiv:2609.38585v1 Announce Type: cross Abstract: Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundan…
arXiv:2609.35443v2 Announce Type: replace Abstract: Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational co…
arXiv cs.LG
TIER_1English(EN)·Yuchen Li, Mingyu Du, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong·
arXiv:2609.36724v1 Announce Type: new Abstract: Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose l…
arXiv:2609.37362v1 Announce Type: new Abstract: Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned throug…
arXiv cs.AI
TIER_1English(EN)·Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye·
arXiv:2609.37402v1 Announce Type: new Abstract: Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on histor…
arXiv:2609.37902v1 Announce Type: new Abstract: Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider s…
Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping…
Existing Large Language Model (LLM) routing methods score LLMs independently to select top-k models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the over…
Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a par…
Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality fee…
dev.to — LLM tag
TIER_1English(EN)·Abdulsalam Abdulsalam·
<p>Every LLM tutorial ends the same way. You get back a blob of text, and now you have to parse it. You beg the model for JSON. It hands you JSON wrapped in an apology. You write a regex. The regex breaks on the next prompt. You add a retry. The retry costs you another second and…
dev.to — LLM tag
TIER_1English(EN)·Sanskriti Harmukh·
<p><a href="https://portkey.ai/" rel="noopener noreferrer">Portkey</a> is an open-source AI gateway that unifies access to 250+ Large Language Models (LLMs) across dozens of providers behind a single, OpenAI-compatible API. Instead of integrating with each provider's SDK, applica…