Skip to content

v1.14.0

Choose a tag to compare

@github-actions github-actions released this 20 Sep 18:40
· 133 commits to main since this release
d9742ed

Added

  • scripts/hooks/route-prompt.mjs, a UserPromptSubmit hook that routes every request. A skill only routes when the model remembers to invoke it, so the request most worth routing — a quick one — is the one least likely to trigger it. The hook scores the request before any work starts and returns the tier as context, every time. It is explicit about the ceiling, because the ceiling is low: no hook output carries a model, PreModelSwitch may only block a switch and PostModelSwitch is read-only, and a session cannot re-price its own turns. So the hook advises the session and binds the subagent — the Task/Agent tool takes a model, so delegated work genuinely runs on the routed tier — and says as much in the context it injects, because claiming the host switched models when it only recommended a tier is an anti-pattern the skill names. It never eats a prompt: every failure path exits 0 and writes nothing, the routing call is killed at six seconds, turn-taking prompts and slash commands are skipped, and a subagent never routes so a routed agent cannot launch another. See model routing.

Fixed

  • An offline workspace search trusted every stored index without checking any of it. The capability signature above guards getCachedCodeMap, but otito search --offline never went through it: it read each catalogued repository's index file, took whatever map was inside, and searched it. So the fix landed for a single repository and not for the sweep across all of them — the surface where a stale index is hardest to notice, because one repository quietly contributing no Markdown records looks exactly like a repository that has none.

    • The acceptance rule is now one predicate shared by both readers, split along its cost. The repository fingerprint stays out of the offline path on purpose: it walks and stats every tracked file, and --offline is documented as using stored indexes without refreshing fingerprints, so re-fingerprinting every catalogued repository on every search is the regression this flag exists to avoid. The other three checks — cache version, capability signature, and the root the map names — are pure functions of an envelope already being parsed, they cost nothing, and they are the ones that catch an indexer mismatch. A repository whose index fails them is skipped with its reason recorded in the result's errors, naming the refresh, rather than being silently searched or unilaterally rebuilt.
  • A repository indexed before a capability change kept serving the old index. The Markdown fix above changed which files the indexer admits, but not cacheVersion, the hand-maintained counter that decides whether an on-disk index is still trustworthy. The repository fingerprint could not catch it either: it summarises the files on disk, and those had not changed. So an untouched repository indexed before that release went on being served a map with no Markdown records — route on "fix a typo in the README" found no candidates at all, the no evidence fail-safe fired on the empty set, and a one-line documentation edit was routed to the premium tier. The fail-safe was right; it was being handed a wrong answer.

    • The cache now also records a capability signature, derived by running the real eligibility and classification functions over a fixed corpus of representative paths and hashing the answers. Admitting a new extension, dropping one, or moving a path to a different kind each move the signature on their own, so an index written by an indexer that disagrees with today's is rebuilt without anyone remembering to bump anything. cacheVersion stays as the manual escape hatch for extraction changes the corpus cannot observe.
  • The code map never indexed Markdown, so no skill or docs page could be found. isSourceFilePath admitted a fixed list of code extensions plus a special case for CHANGELOG.md; every other .md file in the repository was invisible to impact analysis, route, and convergence. A request naming a skill therefore ranked whatever library files happened to share its vocabulary: "make the model-router skill layout-agnostic" returned integrations/herdr/runtime.mjs and src/lib/model-route.js at containment 18, and codex/skills/model-router/SKILL.md — the file the request names — appeared at no rank at all. The same blind spot hid every page under docs/. This is the failure class the empty-match fix addressed one release earlier: a confident answer resting on the wrong evidence.

    • Markdown is now indexed with two kinds. skill covers SKILL.md and its companion pages under a skills/ directory — these are instructions the repository ships and asks contributors to edit, so they are implementation, not commentary. doc covers the rest, and keeps the existing docs/ ranking demotion. Frontmatter keys and values, headings, and relative links become the file's symbols and imports; Markdown never reaches the TypeScript parser, and SQL quoted in a document is no longer mined as data access.
    • Both kinds own a change only when the request is about them. A skill can carry a coverage obligation for "change the model-router skill", but src/lib/model-route.js still owns "fix the model route scoring bug" — a skill whose path matches every term must not displace the code. The same request now ranks codex/skills/model-router/SKILL.md first at containment 90.
  • A document about a risky area escalated the tier as if it were that code. #175 made Markdown indexable, so prose became a routing candidate and, with it, risk evidence: docs/AUTH_TOKEN_VALIDATION.md classifies as auth/security on its "auth" and "token" path tokens, which fired the risk-path bump and forced premium for "fix a typo in the README". The bump itself is sound — a change touching auth code should escalate — but a document about token validation does not validate tokens, and matching it on a path substring prices a doc edit as an auth change. signalsFrom now skips prose candidates when collecting risk evidence, using the code map's kind where the ranking supplied one and falling back to isDocPath. Scoped to the risk signal: classifyPath is unchanged, because impact uses it for relevance, and a doc about auth should still rank for an auth request.

  • Post-merge attestation died on exit 128 and reported green when it had done nothing. The audit ledger chains against repoctx's merge commits, and otito's main is a fresh history, so none of the 97 durable records are reachable from it. reconcile-attestations.sh ran its coverage-gap check first, walking git rev-list FIRST..LAST before anything established those commits were present, so every run that resolved a merged commit failed with fatal: Invalid revision range — the script already held the right message for this, in the ancestry check below, but control never reached it. The runs that looked green were not passing: they resolved no merged commit, skipped every step and reported success, so the workflow alternated green and red while never once reconciling. Reachability is now established before any range is walked, and a ledger that fails it is reported as what it is with the recovery named. Recovery is opt-in via OTITO_ATTEST_RESET_LEDGER=1, never automatic, archives the superseded chain rather than deleting it, and starts the new chain at the tip instead of backfilling every ancestor, which would mint verdicts for changes the gate never ran on.

  • The route call priced output tokens at the input rate. JEV_PER_MTOK covers input tokens only, but askJev fell back to total_tokens — which also contains output — and generateRoute priced whatever landed there at the input rate, so a 900-total-token response and a 900-input-token response both printed $0.000038 and the inflated figure was indistinguishable from the correct one. The count and the billable quantity now stay apart: the response's token count is reported along with which quantity it is, only an input count is priced, and the terminal line says when a count cannot be priced rather than printing a number the rate does not support. priceRouteCall also separates a measured zero from an absent measurement, which the old truthiness check collapsed into null.

  • The router escalated two thirds of requests, sending one-file changes to the premium tier. confidence was Math.min of both Score answers, and a value under the floor bumped a tier. That discarded the confident answer: "rename the variable running to score in model-route.js" scored specificity 1.00 at confidence 1.00 — Jev put the entire distribution on "the request names the exact file, symbol, flag or user-visible string to change", which is exactly what that sentence does — while blast_radius put 0.80 on "contained to a single file" at confidence 0.54. Both answers say trivial. The minimum kept 0.54, missed the floor by one hundredth, and routed a one-file rename to the premium tier. It also double-counted, because score is the expectation over that question's own level distribution, so a spread answer already pays through its own term; bumping on the spread as well charged for it twice. Confidence is now reported for a reader and not read by the router. Measured on this repository over nine requests: escalation 67% → 0%, premium 6/9 → 1/9. The remaining bumps, no evidence and risk path, are otito's own deterministic repository signals, which is the half of this pairing entitled to overrule a model; a vendor's self-reported certainty is not. See model routing.