v3.1.0
The router gets graded, and the answer so far is that nothing can grade it yet. otito regret replays a repository's history and grades the tier each half of the router would have given against the same repaired outcome otito calibrate uses; --rescore grades a new arithmetic on frozen model answers. On three repositories no variant orders outcomes, and an offline audit of the join and of a same-session outcome says why, so the route stays advisory. The route-prompt hook now keeps every decision it makes, so the router can be graded once enough real requests exist. Attestation moves into the CLI as otito attest, with versioned records and a reusable workflow. No command, field or schema was removed.
Added
- The How It Works page is tied to every change in the CLI, the tool catalog and the release. Three new checks, each with a message that says where to make the fix. Every command
otito helplists must be a card on the page or a row inSUPPORTING_COMMANDS(scripts/how-it-works/content.js) with the reason it is not a stage of the loop, so a new command cannot ship without deciding where it belongs, and a command that leaves the CLI has to leave the table too. Each MCP tool now carries asummaryinsrc/lib/mcp.js, one or two plain sentences that the page andotito agent-tools --jsonprint; the MCP description is unchanged andsummaryis not sent overtools/list. The page stamps the package version (<meta name="otito:version">) and theversionlifecycle script re-renders it, so a release that skippednpm run docs:diagramfailsdocs:diagram:check. - "The gate never consults a model" is a test, not a sentence.
tests/gate-imports.test.jswalks the static import graph from every verdict-producing module (pass-local,pass-pr,pr-review,converge,review,impact) and fails if any path reachesjev.js,model-route.jsorcontext-read.js. The page's first guarantee is checked against the same test. otito attest/otito attest --verify. The post-merge attestation moves out ofaudit-pilot/attest.mjsand into the CLI.otito attest <repo> --verdict <file> --merge <sha> [--prev <sha>] [--pr <n>] [--author <name>] [--committed <iso>]appends a hash-chained record toaudit-pilot/ledger.jsonlunder the repository (or the file--ledgernames);--verifyrecomputes the chain and exits 1 if any record was altered. Both take--json.scripts/post-merge-attest.shandscripts/reconcile-attestations.shcall the command and honourOTITO_LEDGER.schemaVersionon the verdict JSON and on every ledger record.otito review --jsonnow reportsschemaVersion: 1besidereviewEngineVersion, and each attestation carriesschemaVersion: 1and theverdictSchemaVersionit was built from. Records written before these fields existed verify unchanged.- Jev calls identify Otito. Every call to TypeSafe goes out under
user-agent: otito/<version>, with the client name and version only. otito regret, the router graded against the repository's own history. For each non-fix commit it checks the first parent out into a temporary worktree, scores the commit subject as the request, and grades three tiers side by side: the deterministic half alone (AX, containment and bumps, every model term at zero), the shipped offline heuristic, and the Jev read whenTYPESAFE_API_KEYis set (--offlinekeeps a run keyless). Outcomes are the line-overlaprepairedjoinotito calibrateuses, now exported fromcalibrate.jsasjoinRepairsandreadHistoryso both grade the same thing. Per tier: n, repaired, rate, lift, with rates withheld below--min-sample; per variant: whether cheap < mid < premium orders outcomes, and the regret count, a commit routed cheap that was repaired within the window. Per question: mean, min, max and spread of the answers across the corpus, so a question that cannot separate its inputs shows as a flat line. Never a saving: otito does not know which model a host used. Offline the run is a pure function of repository state with a receipt; with a key the receipt says the model's answers are not replayable.--json,--out,--since,--max,--window,--quiet. Run on this repository (150 gradable commits, 47 younger ones censored, 30-day window,jev-1.13.0): no variant is shown to order outcomes, every interval overlaps every other, and no variant routes premium often enough to grade the top tier. The deterministic half sits at the 17% base rate in both tiers it uses; the keyless heuristic runs the wrong way within noise (cheap repaired 20.3% against 14.8% for mid); the Jev read runs the right way within noise (cheap 14.7%, mid 17.9%) and its answers now separate their inputs, but the arithmetic funnels three quarters of commits intomid. The table, intervals and receipts are in docs/18 under "Measured, otito, 2026-09-26".otito regret --rescore <run.json>, a change to the router's arithmetic graded on frozen model answers. Jev's answers are not replayable, so two arithmetics graded on two runs also differ by the model's drift between them. Each regret row now keeps the signalsscoreDecisionreads (containment,riskPathsbesideaxandcandidates) and the answers exactly as scored underinputs(Score expectation, confidence and level distribution; Noul probability), andrescoreRegret/--rescoreapplies the current arithmetic to a saved run with no checkout and no model call. The rescore carries its own receipt, names its source inmethod.rescoredFrom, and keeps the source'sreplayable: false. Runs saved by regret 0.2.0 are refused rather than rescored on rounded answers. docs/18 records the bar the next arithmetic has to clear, written before any candidate was scored, and the three frozen runs it is graded on: each repository's whole history,jev-1.13.0, 2,385 answered calls, $0.20, every saved tier reproduced by a rescore. Against that bar the shipped arithmetic fails in two places. On bashbop-api the keyless tier is inverted (cheap repaired 22.5% against 10.1% for mid, intervals apart) while the Jev tier is ordered (4.8%, 12.1%, 38.6%) and its escalations out of the cheap lane are right (13.5% repaired against 4.8% for the commits it left cheap). On bashbop-event-web the Jev read routes 8 of 1,016 commits cheap, too few to grade. The first candidate under the bar, which charges each Score term from a centre so a read at the easy end earns a share of AX back, passed the tuning runs at one centre only (0.2, with its neighbours failing inside the intervals) and failed the confirmation run on criterion 3; it does not ship, and docs/18 records the sweep, the confirmation, and that a cheap-ward push found almost nothing to lift in this backtest.regretEngineVersion0.3.0.otito regretgrades no release commits, and the bashbop-api findings above are withdrawn. A commit written by release tooling (chore(release): 2.26.5 [skip ci],chore: bump version to 1.4.0, a bare version, anything marked[skip ci]) is neither a request nor an outcome: no router saw it, and the join almost never reads one as repaired. They now leave the corpus the way fix commits do (RELEASE_SUBJECT), the corpus line prints how many, and--rescoreapplies the rule to runs saved before it, with a caveat naming how many rows left. On the frozen bashbop-api run they were 504 of 1,214 graded commits, 4 repaired, and 365 of the 474 in the deterministic cheap lane: that lane's 10.1% was the release commits, the keyless "inversion" was the heuristic moving them tomid, and the model's 186-commit cheap lane was 166 of them. Re-graded without them no variant orders outcomes on either bashbop repository, the human base rates are 41.7% (bashbop-api) and 50.5% (bashbop-event-web), and the centre candidate's tuning pass does not survive. docs/18 "Re-graded, 2026-09-26: without release commits" has the tables.regretEngineVersion0.4.0.- The
repairedjoin audited, and the fix rule widened. Before any further arithmetic, the question the re-grade left open was put to the frozen runs offline: 7 and 14 day windows, eight stricter joins (2 or 3 overlapping lines, fix commits capped by files or lines), the share of repairs owed to the ten largest fixes, the unlabelled fixes, and theDevelop (#N)and dependency-bump exclusions. The join reproduces every frozen row; on the bashbop repositories it is not a few sweeping fixes (204 and 284 distinct earliest fixes for 296 and 472 repairs); and no cut orders any graded variant. Commit history cannot grade the router on these repositories, and the next evidence has to come from live requests; docs/18 "Audited, 2026-09-26: the join" has the cuts and the record a live log should keep. One rule was wrong on its own terms:FIX_SUBJECTmissedhot-fix(...),hot-fit,bug(...),patch,fixes,fixedandfixing, so 79 bashbop-api and 40 bashbop-event-web fix commits were graded as requests and invisible as repairs (joined, they add 6 and 9 repairs and move no tier).calibrateandregretnow shareFIX_COMMIT_RULE, and--rescoredrops the rows the wider rule reads as fixes, counts them on the corpus line, and says their own repairs need a replay to join.regretEngineVersion0.4.1. - The route-prompt hook keeps its decision, and
route-outcomesgrades it against the session. A tier printed as context is gone when the turn ends, so every routed prompt now leaves one line in~/.otito/route-decisions.jsonl(OTITO_ROUTE_LOGmoves it,offdisables it): session id, timestamp, the prompt's hash and length (never the prompt), repository, branch and head at prompt time, the tier and route each half gave, the signals and the answers exactly as scored so--rescorecan re-tier a live corpus, and the subagent model.otito route --jsonreportsdeterministic: { tier, route }beside the scored tier;NEUTRAL_ANSWERSmoves tomodel-route.js.node scripts/hooks/route-outcomes.mjsjoins the log to the Claude Code transcripts by session and hash and grades every variant per tier on two same-session outcomes,corrected(the next prompt is an interruption or pushback) andreworked(the next turn edits the same file), with regret's minimum-sample rule and Wilson intervals. Both proxies were audited before they were written, on twenty days of one machine's transcripts: pushback (8% of follow-ups) is mostly the user's own typos, deploy failures and interruptions to add information; rework (17% of editing requests) is iteration; 84 tiered requests put every lane under the minimum sample, and at about seven editing requests a day the corpus grows no faster than commits do. docs/18 "Audited, 2026-09-26: a same-session outcome" has the numbers and the caveats the grader prints. Wire the hook once in~/.claude/settings.jsonrather than per repository. - A reusable post-merge attestation workflow.
.github/workflows/attest.ymltakesworkflow_callwithtarget_sha,ledger_path,ledger_branch,otito_refandreset_ledger, checks the tool out beside the repository under attestation, attests every missing first-parent commit and persists the ledger to the ledger branch. otito's own post-merge workflow now calls it with the merged commit as the tool ref.scripts/post-merge-attest.shandscripts/reconcile-attestations.shtakeOTITO_REPO,OTITO_BINandOTITO_LEDGERso they can attest a repository other than the one they ship in; the verdict is written beside the ledger.
Changed
- The How It Works page is redesigned around the loop, not the layers.
docs/assets/otito-how-it-works.htmlwas a dark SVG of circles in six bands, unreadable on a phone and silent about everything that happened after the gate:attest,calibrateandregretdid not appear, nor did the edit itself. It is now a responsive HTML timeline in the docs site's palette (light and dark) with six phases, Know the repository, Before the edit, The edit, Before the merge, After the merge and Over time, each naming who acts in it; otito is absent from exactly one. Every MCP tool is still a card generated fromgetAgentTools(), the CLI-only stages are cards the test checks againstotito help, and the page closes on the four guarantees the rest of the docs argue for: no model inside the gate, nothing leaves the machine by default, same inputs same output, evidence not approval.npm run docs:diagramregenerates it;--checkstill runs innpm run quality.
Fixed
- docs/18 printed a Jev cost that did not follow from its own token count, and a token count 2.6× lower than the call that ships. The "Cost and latency" section carried the pre-read measurement (728 to 845 input tokens) with a dollar range ($0.000019) that is not that count at $0.042/Mtok. Re-measured on 2026-09-26 against
jev-1.13.0with the request read folded in: 2,059 to 2,165 input tokens, $0.000086 to $0.000091, 291 to 323 ms. The same page still described the confidence floor as a routing fail-safe; it is a display threshold only since the double-count fix, and the paragraph now says so. docs/19 rounds the per-call cost to "about a hundredth of a cent" instead of "under", which was wrong for a 24-file context read.