PulseAugur
EN
LIVE 14:46:03

OmnisRouter benchmarked near bottom due to high cost, not routing skill

The developer of OmnisRouter, a tool designed to route LLM requests to the most cost-effective model, has shared benchmark results that place their router near the bottom. Despite initial bugs in the testing setup, the corrected OmnisRouter achieved 72.7% accuracy at a cost of $3.71 per thousand queries on RouterArena. This performance ranks it 16th out of 18 routers, primarily due to its high cost compared to other routers that utilize cheaper, open-source models. AI

IMPACT This benchmark highlights the trade-offs between cost and performance when routing to premium LLMs, suggesting current benchmarks may not reflect real-world agent usage.

RANK_REASON The item details the performance of a specific LLM routing tool on an independent benchmark.

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

OmnisRouter benchmarked near bottom due to high cost, not routing skill

How we ranked this

Signal score
45 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The item details the performance of a specific LLM routing tool on an independent benchmark.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, infra
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Developer at Fortitude Omnis Group ·

    I put my cost router on a neutral benchmark. It ranked near the bottom, and that's the interesting part

    <p>A few weeks back I repriced three months of my own Claude Code usage. Real traffic, not a demo: 39.5 billion tokens across 139,835 requests. At API rates that's about $30,000, and 91% of it went to Opus because that's what the default reaches for. Route the fraction a smaller …