Automatic Claude model routing for Claude Code

jev-router is a free Claude Code plugin that picks the right model (Haiku, Sonnet, Opus) and reasoning effort for every prompt, step and subagent, caps expensive effort, protects your prompt cache and shows the choice above your prompt.

Install in 30 seconds View on GitHub See the benchmarks jev-router sends each Claude Code task down the right track: haiku, sonnet, opus or fable

What is jev-router?

Using one model for everything wastes money: an expensive model renames a variable, a cheap one stalls on a hard refactor. jev-router is a Claude Code plugin (a "mod") that asks TypeSafe Jev which Claude model and effort a task needs, then runs the turn on that choice. It works per prompt, again before each step of a long task, and once for every subagent.

The jev-router band above the Claude Code prompt, stepping through its states: idle, running, haiku, sonnet, capped opus, pinned, escalated, paused

Where jev-router shines

If your default is Opus, most prompts pay for more model than they need. Measured with real Claude Code runs, against always using Opus:

Your workWhat it doesResult
Small tasks: renames, typos, lookupsHaiku, low effort−97% cost
Everyday features, tests, refactorsSonnet−65% cost
Multi-step agent runs with tools and subagentsRoutes each step and subagent18/18 done at 47% of the cost
Hard tasks graded by hidden tests (the default lean preset)Sonnet, medium effort, Opus only when Jev is 90% sure39/40 done at −61% cost
30 mixed single prompts (balanced preset)Effort capped at high−1% to −64%, never above Opus

Out of the box (0.4.0) jev-router runs the lean preset: a Sonnet ceiling and effort capped at medium. /jev preset balanced restores the earlier behaviour.

Best for: Opus-by-default users, workloads full of small prompts, agent runs with subagents, and anyone who wants to see and steer which model runs. Not for: long sessions on one big warm context (there lean averaged the same cost as Opus), or work where plain Sonnet is already enough.

Install

You need Claude Code with plugin support and a TypeSafe API key.

/plugin marketplace add dominicrico/jev-router
/plugin install jev-router@jev-router

Then give it your key with /jev key <key> or the TYPESAFE_API_KEY environment variable, and run /jev status.

Features

Right model per promptHaiku for renames, Sonnet for features, Opus for hard debugging. Three modes: efficient, balanced, cheap.
Per step and per subagentLong tasks are re-routed before each step. Each subagent gets its own model, shown in the subagent list.
Effort capLimits Jev's effort to medium by default, because extra reasoning is what makes hard tasks expensive. !full lifts it for one prompt.
Prompt cache guardSwitching models discards a warm cache. jev-router stays put unless Jev is confident a stronger model is needed. Measured: 6 to 7 switches become 1 in long sessions, and cache re-writes halve.
Model ceiling/jev ceiling sonnet keeps tasks off Opus unless Jev is 90% sure. A sonnet ceiling finished 20/20 hard tasks at half the cost of Opus.
Shows what really runsThe band follows the model that ran the last step and flags a mismatch with the pick, for example after a /model switch.
Per-prompt markers!opus, !sonnet, !haiku pin a model with no Jev call; !cheap and !efficient set the mode.
EscalationAfter three failed tool calls in a row the task moves up one model.
Live status bandModel, effort, confidence and cache state above the prompt, with a spinner while a task runs.
Savings tally/jev status estimates your saving against the model you would otherwise use.
Privacy controlsSecrets are redacted before anything is sent; history can be switched off; Jev down? Optional local fallback.

Does it save money? The benchmarks

Everything is measured with real Claude Code runs and reproducible scripts. The honest summary: it saves a lot on easy and medium work, costs more than plain Sonnet on harder work, and with the default effort cap no mode costs more than always using Opus.

Cost as a percentage of always opus for trivial, standard and hard tasks in each routing mode
Cost vs always Opusefficientbalancedcheap
Trivial tasks−97%−97%−97%
Standard tasks−65%−65%−68%
Hard tasks (effort capped at high)+43%+24%−56%
All 30 tasks−1%−13%−64%
Multi-step tasks with tools: jev-router finished 18 of 18 tasks at 47% of the cost of always opus Harder tasks graded by hidden tests: lean finished 39 of 40 at 39 percent of the cost of opus, close to always sonnet

On the harder, hidden-test-graded tasks plain Sonnet did as well as Opus and cost 61% less. The default lean preset finished 39 of 40 at 61% below Opus. In long warm sessions routing did not beat staying on one model (the cache guard halves cache re-writes but not the bill), so routing helps most when your tasks vary, your context is small or cold, or your default is Opus. Method, limits and raw results are in the benchmarks folder.

FAQ

How do I switch between Opus, Sonnet and Haiku automatically in Claude Code?

Install jev-router. It picks the model per prompt, and you can override one prompt with !opus, !sonnet or !haiku.

How do I reduce Claude Code token cost?

Route easy tasks to cheaper models and cap reasoning effort. jev-router does both, and shows an estimated saving in /jev status. The benchmarks above show where it helps and where it does not.

Does the prompt cache guard work?

It does what it says on switches: in real long sessions with a large warm cache it cut model switches from 6 to 7 down to 1 and halved cache re-writes. It did not lower the bill: with and without the guard routing cost +42% against Opus. The default lean preset averaged the same as Opus there (+0%), cheap in some sessions and dear in others, while plain Sonnet was -43%. For long warm sessions on one big context, plain Sonnet is the cheapest safe choice.

What is the effort cap?

Jev often asks for xhigh effort on hard tasks, which can use several times the tokens. The cap limits it to high by default. Start a prompt with !full to lift it once.

Does it send my code anywhere?

It sends the prompt text, optionally recent messages, and subagent task descriptions to api.typesafe.ai. API keys, tokens, private keys and password= style values are redacted first, and sendHistory turns the history off. Pattern redaction is not a guarantee.

Will it slow Claude Code down?

Each Jev call adds about 250 ms. After three failed calls in a row it pauses for a minute and keeps the current model.

Which models does it use?

Claude Haiku, Sonnet, Opus and Fable, narrowed by the models option.