Saltar al contenido
ES EN

Why you should stop using one model for everything: an AI model router

How to cut Claude Code costs with an AI model router, which model to use for coding each task, and what our own benchmark really measures.

If you code with Claude Code, you usually pick a model when you start and let it do everything: plan a migration, write a test, summarise a file or touch up a README. It is convenient, but each of those requests is billed at the price of the model you picked, and on simple tasks that price rarely buys a better result.

Why you should stop using one model for everything: an AI model router
Image: Unsplash — Tabs labeled "vibe coding" with code on bottom

To measure it we built Rebel Router, an open-source model router that sits between your agent and the providers, and put it through a benchmark of our own. These are the results, with their limits.

The problem: paying for the most expensive model every time

In our benchmark, the same 12 coding tasks cost 0.96 dollars with Claude Opus 5.5 and 0.37 with Claude Sonnet 5.5. GPT-6 Luna solved them for 0.011 dollars and passed the same 11 of 12 as Opus: roughly 80 times less money for the same result on this set. That does not mean Luna replaces Opus for any job; it means many everyday tasks do not need the most expensive model.

How Rebel Router decides

Rebel Router is a local proxy. Claude Code sends it every request through ANTHROPIC_BASE_URL, and OpenCode sees it as one more provider. For each request it does four things:

  1. Classifies the task into six classes (orchestration, complex code, bulk code, unit tests, QA or documentation) with a complexity score. It can use JEV, Venice's decision model (in beta), if you give it a key; without one, or if it takes longer than 800 ms, it uses local rules.
  2. Picks with the Broker: it drops models below a quality floor for that class and ranks the rest on quality, price and speed. It keeps three: the chosen one and two fallbacks.
  3. Controls spend with FinOps: daily and per-project budgets. From 80% it prefers cheaper models; at 100% it downgrades or blocks.
  4. Forwards and records: it translates between the Anthropic and OpenAI formats when needed and logs the model, tokens, cost and the reason for each pick to SQLite.

One important detail: in Claude Code no hook can change the session model. The proxy does the switching, so the interface keeps showing the model you asked for and Rebel Router's status line shows the one that answered.

What we measured

The benchmark has 12 real tasks, two per class, each scored deterministically: hidden tests, mutation testing, deliberately buggy servers and rubrics. No model grades another. We ran it on 11 October 2026 with 10 models, one attempt per task, for 1.78 dollars in total. The raw results are public.

  • Letting the router choose (auto, with the rule classifier) cost 0.17 dollars against 0.96 with Opus 5.5 for everything: 82% less.
  • In exchange it passed fewer: 9 of 12 tasks (75%) against 11 of 12 (92%). The three misses were two unit-test tasks and one QA task the router sent to Claude Haiku 5.5, whose answers ran past the 8,000-token limit.
  • Mean latency dropped from 31.7 to 19.5 seconds per task.
  • The rule classifier, without JEV, got the class right on 9 of the 12 tasks (75%).
by larebelion

Cost and pass rate of each model

Our benchmark: the same set of real tasks run with every model through Rebel Router. Updated daily.

Cost per task against pass rate

Loading data…

Pass rate by task class (mean across models)

Loading data…

Savings from automatic routing

Loading data…

Intelligence index: top 12

Data: Artificial Analysis

Loading data…

Key points On this data, routing saves a lot on documentation and bulk code and is risky on unit tests and QA. It is a small measurement: a first data point, not a leaderboard.

Which model to use for coding: the recommender

The same data feeds the model recommender. Pick the kind of task, the budget, the speed, the minimum context and whether you only want open or local models, and it returns the three that fit best, with their estimated cost and why. It also has its own page.

by larebelion

Model Recommender

Pick the kind of task and your limits: we return the three models that fit best, with their estimated cost per task and why.

Quality: pass rate in our benchmark when the model was measured, otherwise the Rebel Router catalogue score. Cost: measured, or estimated from official prices for a typical task.

Limits

  • Small sample: 12 tasks and one attempt per model. Per-class pass rates are indicative.
  • The savings have a price: the router lost two tasks Opus solved. You can keep on Opus whatever you ask Opus for (aliases.opus: passthrough) or pin a model per class.
  • Not covered: Grok, Gemini Pro, Ollama and GPT models other than Luna. DeepSeek V4 Pro, GLM 5.3 and Qwen 3.8 Max ran only some of the tasks. JEV was not measured: the benchmark used the rules.
  • Anthropic does not support routing Claude Code to non-Claude models: it works, but Claude-only features are dropped in translation.
  • Cost is list price: the tokens each provider reports times the official price in the catalogue.

Try it

The step-by-step lab shows how to install it from GitLab, bring your keys (or a single OpenCode Zen key), connect it to Claude Code and OpenCode, set budgets and check which model answers and what it costs. The code is on GitLab under the MIT licence.

Original source: Rebel Router (GitLab)

Produced with AI support and reviewed by the newsroom

Kernel

· Head of technology · Spain

“Saving 82% is worth little if the tests you lose are the ones that warn you.”

Comentarios

Publicar un comentario