PulseAugur
中
实时 03:45:06
English(EN) FlexRouter: Learning Complementary Model Sets for Flexible LLM Routing

新研究探索先进的LLM路由技术以提高效率和性能 · 已追踪8个来源

多篇研究论文探讨了用于大型语言模型(LLM)的先进路由技术,以提高效率和性能。VANE专注于对更新而非token进行评分,以更好地选择Mixture-of-LoRA-experts中的专家。语义路由校准(SRC)通过动态抑制超敏感的安全头来解决LLM过度拒绝的问题。FlexRouter通过使用行列式点过程(Determinantal Point Processes)对模型互补性进行建模来优化答案覆盖率,而Just Initialize为大规模路由优化提供了一个无需训练的组件。RouteFM旨在成为LLM路由的基础模型,实现跨不同环境的重用,而SaveRouter则专注于降低监督成本以实现经济高效的LLM路由。最后,提出了面向市场的路由,用于开放权重LLM推理,在模型选择的同时考虑提供商选择。 AI

影响 LLM路由的这些进步可能带来更高效、更具成本效益的AI推理,提高响应质量并降低计算开销。

排序理由 多篇在arXiv上发表的研究论文,详细介绍了LLM路由的新方法。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 14 个来源。 我们如何撰写摘要 →

新研究探索先进的LLM路由技术以提高效率和性能 · 已追踪8个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇在arXiv上发表的研究论文,详细介绍了LLM路由的新方法。
Source corroboration
14 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
7 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [14]

  1. arXiv cs.LG TIER_1 English(EN) · Priya Nair, Lukas Brenner, Maya Lindqvist, Daniel Whitmore, Wen-Hsuan Liu, Tom Saliencro, Amara Okonkwo, Rohan Desai ·

    评分更新,而非评分Token:组合式LoRA专家模型的下降对齐路由

    arXiv:2610.00493v1 Announce Type: new Abstract: Mixture-of-LoRA-experts methods raise the capacity of low-rank adaptation by routing each token to a few low-rank experts. Nearly all of them tie one input-side factor to one output-side factor per expert, and nearly all of them rou…

  2. arXiv cs.AI TIER_1 English(EN) · Zixuan Wang, Bingjie Zhang, He Zhao, Dandan Guo ·

    通过动态语义路由校准缓解 LLM 过度拒绝

    arXiv:2609.25049v2 Announce Type: replace-cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely…

  3. arXiv cs.AI TIER_1 English(EN) · Wang Wei, Harry Yang, Tiankai Yang, Samyadeep Basu, Hongjie Chen, Andy Zhao, Franck Dernoncourt, Ryan A. Rossi, Hoda Eldardiry ·

    FlexRouter:为灵活的LLM路由学习互补模型集

    arXiv:2609.38585v1 Announce Type: cross Abstract: Existing Large Language Model (LLM) routing methods score LLMs independently to select top-$k$ models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundan…

  4. arXiv cs.AI TIER_1 English(EN) · Jiale Zhao, Sirui Mao, Zimu Chen, Wentao Yang, Zihan Wang, Xuefeng Huang, Junji Cheng, Liyuanjun Lai ·

    Just Initialize:一种用于大规模路由优化的无训练初始化组件

    arXiv:2609.35443v2 Announce Type: replace Abstract: Large-scale routing problems are difficult to solve efficiently as their search spaces grow rapidly with problem size. Existing approaches primarily improve the optimization procedure itself, often at increasing computational co…

  5. arXiv cs.LG TIER_1 English(EN) · Yuchen Li, Mingyu Du, Zongqi Fan, Nguyen H. Tran, Ken-Tye Yong ·

    梯度空间中的路由:均衡使用并非专家专业化

    arXiv:2609.36724v1 Announce Type: new Abstract: Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose l…

  6. arXiv cs.AI TIER_1 English(EN) · Guannan Lai, Han-Jia Ye ·

    一次预训练,处处路由:迈向 LLM 路由的基础模型

    arXiv:2609.37362v1 Announce Type: new Abstract: Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned throug…

  7. arXiv cs.AI TIER_1 English(EN) · Guannan Lai, Gelin Bian, Hao-Xuan Ma, Jun-Peng Jiang, Long Chen, Jian-Dong Liu, Zhi-Hao Tan, Han-Jia Ye ·

    路由应自负盈亏:稀疏监督助力经济型LLM路由

    arXiv:2609.37402v1 Announce Type: new Abstract: Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on histor…

  8. arXiv cs.AI TIER_1 English(EN) · Liang He, Jingbo Wen, Yixiong Chen, Yue Yang, Qizhen Lan, Kangning Cui, Xilu Wang ·

    无法从价格列表中选择提供商:面向开放权重 LLM 推理的市场感知路由

    arXiv:2609.37902v1 Announce Type: new Abstract: Existing LLM routers choose among models using static per-model costs. We show that open-weight inference markets introduce a second, largely ignored decision axis: after choosing a model, a client must still choose which provider s…

  9. Hugging Face Daily Papers TIER_1 English(EN) ·

    梯度空间中的路由:均衡使用并非专家专业化

    Sparse expert models can distribute traffic evenly while still grouping incompatible training signals within the same experts. We study routing as a gradient-partitioning problem and introduce gradient-aligned routing (GAR), whose load-normalized router objective rewards grouping…

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    FlexRouter:为灵活的LLM路由学习互补模型集

    Existing Large Language Model (LLM) routing methods score LLMs independently to select top-k models. However, this ignores model correlations and enforces a rigid computational budget. Consequently, routers often select redundant models that share failure modes, limiting the over…

  11. Hugging Face Daily Papers TIER_1 English(EN) ·

    一次预训练,处处路由:迈向 LLM 路由的基础模型

    Large language model (LLM) routing aims to assign each query to the most suitable model from a heterogeneous candidate pool, improving the quality--efficiency trade-off of LLM inference. Existing routers are typically learned through local fitting: a router is optimized for a par…

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    路由应自负盈亏:稀疏监督助力经济型LLM路由

    Large language model (LLM) routing reduces serving cost by assigning each query to an appropriate model while preserving response quality. Learning such a router, however, often requires executing multiple candidate models on historical queries to collect query--model quality fee…

  13. dev.to — LLM tag TIER_1 English(EN) · Abdulsalam Abdulsalam ·

    停止解析LLM输出:我构建了一个基于仅决策模型的路由管道

    <p>Every LLM tutorial ends the same way. You get back a blob of text, and now you have to parse it. You beg the model for JSON. It hands you JSON wrapped in an apology. You write a regex. The regex breaks on the next prompt. You add a retry. The retry costs you another second and…

  14. dev.to — LLM tag TIER_1 English(EN) · Sanskriti Harmukh ·

    部署 Portkey - 用于 LLM 路由的开源 AI 网关

    <p><a href="https://portkey.ai/" rel="noopener noreferrer">Portkey</a> is an open-source AI gateway that unifies access to 250+ Large Language Models (LLMs) across dozens of providers behind a single, OpenAI-compatible API. Instead of integrating with each provider's SDK, applica…