PulseAugur
实时 14:34:20
English(EN) I put my cost router on a neutral benchmark. It ranked near the bottom, and that's the interesting part

OmnisRouter 因成本高昂而在基准测试中排名靠后,而非路由能力

OmnisRouter 的开发者分享了基准测试结果,该工具旨在将 LLM 请求路由到成本最有效的模型,但其排名接近底部。尽管测试设置存在初始错误,但修正后的 OmnisRouter 在 RouterArena 上达到了 72.7% 的准确率,每千次查询成本为 3.71 美元。与使用更便宜的开源模型的其他路由器相比,其高昂的成本使其排名第 16 位(共 18 个路由器)。 AI

影响 此基准测试突显了在路由到高级 LLM 时成本与性能之间的权衡,表明当前的基准测试可能无法反映真实的代理使用情况。

排序理由 该条目详细介绍了特定 LLM 路由工具在独立基准测试中的表现。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

OmnisRouter 因成本高昂而在基准测试中排名靠后,而非路由能力

本文如何被排名

Signal score
50 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该条目详细介绍了特定 LLM 路由工具在独立基准测试中的表现。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Developer at Fortitude Omnis Group ·

    我将我的成本路由器放在一个中性基准上进行测试。它的排名接近底部,而这正是有趣之处

    <p>A few weeks back I repriced three months of my own Claude Code usage. Real traffic, not a demo: 39.5 billion tokens across 139,835 requests. At API rates that's about $30,000, and 91% of it went to Opus because that's what the default reaches for. Route the fraction a smaller …