Skip to content

Releases: dominicrico/jev-router

v0.4.1

Choose a tag to compare

@dominicrico dominicrico released this 10 Oct 20:35
8434b75

[0.4.1] - 2026-10-10

Fixed (documentation, corrected measurements)

  • Re-measured the long-session benchmark with 5 sessions per strategy and the new lean default. The guard still cuts model switches (6 to 7 down to 1) and halves cache re-writes, but it does not lower cost: guard on and off both cost +42% against Opus. The earlier "+24% vs +38%" (3 sessions) was noise and is retracted. lean averaged the same cost as Opus in long warm sessions (range $0.26 to $0.63 per session); plain Sonnet was 43% cheaper. README, benchmarks, site and charts updated.

v0.4.0

Choose a tag to compare

@dominicrico dominicrico released this 10 Oct 07:46
19cd162

[0.4.0] - 2026-10-10

Changed (behaviour change)

  • The default is now the lean preset: a Sonnet ceiling and effort capped at medium. Chosen by a pre-registered rule on 100 more graded runs: 39/40 hard tasks done against 38/40 for always-Opus and always-Sonnet, at $0.064 per run against $0.163 (Opus) and $0.062 (Sonnet); 18/18 on the multi-step set at $0.061 per run. Opus is still used when Jev is at least 90% sure (ceilingBreak). To keep the old behaviour set preset: balanced (or /jev preset balanced).

Added

  • /jev preset lean|balanced|max and a preset option; /jev status shows the active preset.
  • Benchmark harness: ONLY= strategy filter, a lean strategy, and an opt-in Wilson interval for pass rates (CI=1). CI checks that the shipped defaults match the lean preset.

v0.3.0

Choose a tag to compare

@dominicrico dominicrico released this 09 Oct 14:57
a38d8bf

[0.3.0] - 2026-10-09

Added

  • The band shows the model that really ran the last step. A mismatch reads sonnet (picked opus); after a Jev failure it shows the session model instead of "waiting for the first task"; it flips from warm to cold on its own.
  • Model ceiling and floor (ceiling, floor, ceilingBreak, /jev ceiling, /jev floor), applied before the cache guard. Off by default. A sonnet ceiling finished 20/20 hard tasks at half the cost of opus.
  • A real long-session benchmark (benchmarks/session.mts, one process per session via stream-json) and a debug line with the cache state per Jev call.

Changed

  • README and charts now use the measured long-session numbers. The earlier simulation's "3x more" is replaced by +38% (guard off) and +24% (guard on) against opus, and the README says plainly that in a long warm session routing cost more than staying on opus.

v0.2.0

Choose a tag to compare

@dominicrico dominicrico released this 09 Oct 13:20

[0.2.0] - 2026-10-09

Added

  • Per-step routing: Jev is asked again before every step of a task, not only at the prompt (routeSteps).
  • Per-subagent routing: each subagent is routed at spawn from its full task prompt, sets the subagent's model and tags its description in the subagent list (routeSubagents). Verified live.
  • Effort cap, high by default, so the xhigh/max that Jev suggests on hard tasks no longer drives cost above always-opus. Lift it for one prompt with a !full prefix or /jev full; change it with /jev cap <effort|none> or the effortCap option. The band shows the cap (⤓xhigh) and the unlock (🔓).
  • Circuit breaker: three failed Jev calls in a row pause routing, then it recovers; strict stickiness with a warm cache makes no step calls.
  • Per-prompt markers: !opus, !sonnet, !haiku, !fable pin a model and skip Jev; !cheap, !efficient set the mode; they combine with !full.
  • Escalation: after escalateAfter failed tool calls in a row the rest of the task moves up one model (verified live).
  • Estimated savings tally in /jev status against baselineModel, kept across sessions; /jev stats clear resets it.
  • Privacy: redact (on by default), sendHistory, maxTaskChars, and an opt-in local fallback: heuristic when Jev is down.
  • A debug log line per Jev call (jev prompt|step N|agent), used to verify the above live.
  • CI workflow, band render test and an animated band GIF.
  • Benchmark on harder tasks graded by hidden tests (100 real runs). On those tasks plain sonnet did as well as opus and cost 61% less; jev-router sits between them.
  • Benchmarks for the effort cap and for multi-step and subagent routing on tasks with tools, with updated README images.

Changed

  • Ponytail cleanup pass over the hook module: one subagent router, one usage tally (steps per model, in /jev status), shared benchmark helpers, and the band GIF now renders the real segments() so it cannot drift.

Fixed

  • Secret redaction now catches lowercase names (password=, client_secret:, api_key=), JSON keys ("apiKey": "...") and user:password@ in URLs, and no longer mangles type words like token: string.
  • A pinned model (!opus) now works without a Jev key, and the marker no longer sticks to the next prompt. A prompt with no key no longer runs on the previous task's pick.
  • The hard-task benchmark's hidden tests and reference solutions are sealed in an archive and unpacked only into a random temp directory while grading, so an agent cannot read them from disk. (A first run was discarded because they were readable.)
  • Background-task notifications no longer trigger a Jev call or overwrite the task text.
  • Spinner leak: the band ticker now stops on turn.complete and on /jev commands, and cancels itself if a prompt never completes.

v0.1.0

Choose a tag to compare

@dominicrico dominicrico released this 09 Oct 12:09

First release: Jev model and effort routing per prompt, per step and per subagent, effort cap (default high) with !full unlock, cache stickiness, the band above the prompt, /jev commands and benchmarks.