Repository navigation
Releases: dominicrico/jev-router
Releases · dominicrico/jev-router
Release list
v0.4.1
[0.4.1] - 2026-10-10
Fixed (documentation, corrected measurements)
- Re-measured the long-session benchmark with 5 sessions per strategy and the new
leandefault. The guard still cuts model switches (6 to 7 down to 1) and halves cache re-writes, but it does not lower cost: guard on and off both cost +42% against Opus. The earlier "+24% vs +38%" (3 sessions) was noise and is retracted.leanaveraged the same cost as Opus in long warm sessions (range $0.26 to $0.63 per session); plain Sonnet was 43% cheaper. README, benchmarks, site and charts updated.
v0.4.0
[0.4.0] - 2026-10-10
Changed (behaviour change)
- The default is now the
leanpreset: a Sonnet ceiling and effort capped atmedium. Chosen by a pre-registered rule on 100 more graded runs: 39/40 hard tasks done against 38/40 for always-Opus and always-Sonnet, at $0.064 per run against $0.163 (Opus) and $0.062 (Sonnet); 18/18 on the multi-step set at $0.061 per run. Opus is still used when Jev is at least 90% sure (ceilingBreak). To keep the old behaviour setpreset: balanced(or/jev preset balanced).
Added
/jev preset lean|balanced|maxand apresetoption;/jev statusshows the active preset.- Benchmark harness:
ONLY=strategy filter, a lean strategy, and an opt-in Wilson interval for pass rates (CI=1). CI checks that the shipped defaults match the lean preset.
v0.3.0
[0.3.0] - 2026-10-09
Added
- The band shows the model that really ran the last step. A mismatch reads
sonnet (picked opus); after a Jev failure it shows the session model instead of "waiting for the first task"; it flips from warm to cold on its own. - Model ceiling and floor (
ceiling,floor,ceilingBreak,/jev ceiling,/jev floor), applied before the cache guard. Off by default. A sonnet ceiling finished 20/20 hard tasks at half the cost of opus. - A real long-session benchmark (
benchmarks/session.mts, one process per session via stream-json) and a debug line with the cache state per Jev call.
Changed
- README and charts now use the measured long-session numbers. The earlier simulation's "3x more" is replaced by +38% (guard off) and +24% (guard on) against opus, and the README says plainly that in a long warm session routing cost more than staying on opus.
v0.2.0
[0.2.0] - 2026-10-09
Added
- Per-step routing: Jev is asked again before every step of a task, not only at the prompt (
routeSteps). - Per-subagent routing: each subagent is routed at spawn from its full task prompt, sets the subagent's model and tags its description in the subagent list (
routeSubagents). Verified live. - Effort cap,
highby default, so thexhigh/maxthat Jev suggests on hard tasks no longer drives cost above always-opus. Lift it for one prompt with a!fullprefix or/jev full; change it with/jev cap <effort|none>or theeffortCapoption. The band shows the cap (⤓xhigh) and the unlock (🔓). - Circuit breaker: three failed Jev calls in a row pause routing, then it recovers; strict stickiness with a warm cache makes no step calls.
- Per-prompt markers:
!opus,!sonnet,!haiku,!fablepin a model and skip Jev;!cheap,!efficientset the mode; they combine with!full. - Escalation: after
escalateAfterfailed tool calls in a row the rest of the task moves up one model (verified live). - Estimated savings tally in
/jev statusagainstbaselineModel, kept across sessions;/jev stats clearresets it. - Privacy:
redact(on by default),sendHistory,maxTaskChars, and an opt-in localfallback: heuristicwhen Jev is down. - A debug log line per Jev call (
jev prompt|step N|agent), used to verify the above live. - CI workflow, band render test and an animated band GIF.
- Benchmark on harder tasks graded by hidden tests (100 real runs). On those tasks plain sonnet did as well as opus and cost 61% less; jev-router sits between them.
- Benchmarks for the effort cap and for multi-step and subagent routing on tasks with tools, with updated README images.
Changed
- Ponytail cleanup pass over the hook module: one subagent router, one usage tally (steps per model, in
/jev status), shared benchmark helpers, and the band GIF now renders the realsegments()so it cannot drift.
Fixed
- Secret redaction now catches lowercase names (
password=,client_secret:,api_key=), JSON keys ("apiKey": "...") anduser:password@in URLs, and no longer mangles type words liketoken: string. - A pinned model (
!opus) now works without a Jev key, and the marker no longer sticks to the next prompt. A prompt with no key no longer runs on the previous task's pick. - The hard-task benchmark's hidden tests and reference solutions are sealed in an archive and unpacked only into a random temp directory while grading, so an agent cannot read them from disk. (A first run was discarded because they were readable.)
- Background-task notifications no longer trigger a Jev call or overwrite the task text.
- Spinner leak: the band ticker now stops on
turn.completeand on/jevcommands, and cancels itself if a prompt never completes.
v0.1.0
First release: Jev model and effort routing per prompt, per step and per subagent, effort cap (default high) with !full unlock, cache stickiness, the band above the prompt, /jev commands and benchmarks.