Repository navigation
Releases: giuliastro/token-harness
Release list
Token Harness 0.1.31 — Desktop application and guided benchmark recovery
Token Harness 0.1.31 — Desktop application and guided benchmark recovery
0.1.31 ships a native desktop application, retained configuration history and
reviewed benchmark controls. It also includes deterministic quality checks and
managed exploratory four-arm benchmarks added after 0.1.30.
Installation
Download the matching asset from the 0.1.31 release:
- Windows x64: install the setup EXE and open Token Harness.
- Linux x64: install the DEB or extract the TAR.GZ into a permanent directory and
runtoken-harness-desktop. - macOS Intel or Apple Silicon: open the matching DMG/ZIP and copy the application
into Applications.
Each desktop package includes its own Node runtime. Choose File → Open project
to select your coding project. Keep the installed runtime in a stable location while
owned hooks reference it. Windows packages are unsigned; macOS packages use ad-hoc
signing without notarization. WSL remains a separate CLI environment.
For the CLI:
npm install --global token-harness@0.1.31
token-harness --version
token-harnessThe version must be 0.1.31. Restart an already running application to load the
updated interface. Desktop updates use manual replacement of the published package.
What changed
- Native Windows x64, Linux x64 and macOS x64/arm64 applications share the guided UI
and transaction engine. Releases attach all native packages and SHA-256 checksums. - Results retains the latest 20 current-project operations across reopen. Reviewed
recovery targets the latest eligible configuration transaction; package operations
and arbitrary historical transactions remain outside this control. - Start or finish a comparison records manual baseline/optimized captures and
explicit quality outcomes, with retained progress. It executes no coding task or
check and does not infer quality from token counts. - Overview → Health and updates → Review benchmark recovery discovers the current
project's active prepared-arm lease, previews restoration and restores saved
configuration after a single-use approval. Drift, corrupt backups and invalid
checkpoints preserve recovery data. Recovery retains measurements without finishing
the capture or recording a quality result. - Advanced CLI benchmarks support approved direct executable checks, immutable quality
provenance, schema-1 reading, exploratory four-arm reports and reviewed temporary
configuration preparation/restoration. Each measurement class and quota window
remains separate; one four-arm set is not statistical evidence.
The Harness Remote test guide explains how
to use the same Machine/project, record a paired task and exercise prepared-arm
recovery. Factorial preparation, scheduling and handoff remain advanced workflows.
Verification scope
The feature PRs #378,
#380,
#381,
#382 and
#383 are merged. The latest
feature candidate passed native Windows, macOS and Linux CI and all four desktop
package/smoke targets. Local validation passed 2,525 tests with nine expected skips,
typecheck, source lint, formatting, build and bundle smoke.
Release-candidate CI,
merged-main CI and
exact-tag publication
passed. The tag workflow completed native builds, full tests, guided real-runtime
setup/update smoke, package installation checks, provenance/SBOM attestations,
npm Trusted Publishing and npm-latest verification.
The published npm tarball matches the GitHub asset byte-for-byte. Its isolated
installation reports 0.1.31, exposes the expected CLI and contains the same bundle
as the reviewed source build. The provenance signature, release workflow, tag and
source commit were verified. All seven desktop checksum entries match GitHub's
asset digests.
Tarball SHA-256: 4fd88149e58c50b6bc602186f3c31b5195b4ffd3524eaed002227ba11366dd80.
CLI bundle SHA-256: 658f80509bd88cce6826979eadadcdd1ebdf3f86837ac1eb61a1dbc8c5adb477.
Recovery verification is config-only. Unit fixtures, packaged smoke and desktop
builds do not prove native provider execution or savings. Windows published-artifact
verification (#255/#359), repeated isolated quality-gated measurements (#360), broader
desktop lifecycle/workflow coverage (#379), and broad promotion remain separate work.
Earlier runtime observations retain their original versions and scope; this release
makes no new token, subscription-quota or API savings claim.
Release source
Version preparation: PR #384. Immutable tag v0.1.31 points to merged main commit c9a90f0996a66addffa42f2ea4ab2aa5c3b62746.
Token Harness 0.1.30 — Attributed app evidence and paired comparisons
Token Harness 0.1.30 — Attributed app evidence and paired comparisons
0.1.30 makes Results show each coding app's attributed optimizer output, observed routing
callbacks and current allowance readings. It repairs paired comparisons so five-hour and weekly
quota evidence remain visible together with the actual quality gates.
Installation
npm install --global token-harness@0.1.30
token-harness --version
token-harnessThe version must be 0.1.30. These commands also work in PowerShell. Restart an already running
dashboard to load the updated Results interface.
What changed
- Coding-app rows use separately aggregated events with an explicit app identity. Shared and
unattributed history stays under its optimizer; configuration does not assign it to an app. - Runtime routing callbacks and current quota balances appear as separate evidence, with source,
observation and reset timestamps. Callback activity and remaining quota are not savings. - Paired benchmarks retain every comparable quota window, including simultaneous five-hour savings
and weekly growth. The existing primary comparison and verdict remain compatible. - Actual baseline and optimized quality gates remain visible even when efficiency usage is missing.
A positive allowance claim cannot borrow a different pair's passing quality gate; attributed
routing without comparable usage stays unmeasured for savings. - Results distinguish missing, unfinished and unreadable comparisons. Record a comparison opens
a read-only guide to the existing benchmark-start/finish collectors, without creating a baseline,
running a task or inventing a quality outcome. - Output history remains all-project and period-filtered. Paired quality/quota evidence remains
current-project; callbacks and current balances retain their own observation dates.
Verification scope
PR #371 and merged-main CI passed on Windows,
macOS and Linux. Local review validation passed 2,418 tests with nine expected skips, plus typecheck,
lint, formatting, build and bundle smoke. Regression coverage includes independent quota windows,
negative deltas, cached/reset-crossing rejection, quality-only and unknown-quality outcomes,
no-usage routing and shared/multi-app output attribution.
Release preparation also passed five version/packaging contract tests, formatting, lint, build,
bundle smoke, package staging, temporary tarball installation and exact version/tag validation.
Publication requires release-candidate cross-platform CI, exact-tag build, guided real-runtime smoke,
package installation checks, provenance/SBOM attestations, npm Trusted Publishing and npm-latest
verification.
Existing published 0.1.28 Linux callback/lifecycle observations
retain their exact build scope. Windows published-artifact checks, actual native child-model
selection and paired quality-gated savings measurements remain separate work. This release improves
evidence reporting and makes no new token, quota or API savings claim.
Full changelog: v0.1.29…v0.1.30
Publication completed through the exact-tag release workflow, including real-runtime smoke, package installation, provenance/SBOM attestations and npm-latest verification.
Token Harness 0.1.29 — Native subagent model ladders
Token Harness 0.1.29 — Native subagent model ladders
0.1.29 updates automatic prompt routing to select a cheaper native child by task difficulty and
the main agent's model tier. It fixes Codex fork guidance so explicit child-model overrides can apply
and asks Claude Code to set the child model explicitly.
Installation
npm install --global token-harness@0.1.29
token-harness --version
token-harnessThe version must be 0.1.29. These commands also work in PowerShell. Restart an already running
dashboard or agent session to load the updated interface and routing guidance.
What changed
- Codex reads the root model slug from the native hook payload. An Astra root may delegate bounded
implementation or tests togpt-6.1-sol, and read-only or mechanical work togpt-6-luna.
Sol/Terra roots request Luna; light-tier roots receive no routing context. - Codex guidance explicitly supplies the child model, reasoning effort and
fork_turns: "none".
Full-history forks inherit the root model and reject model overrides. - Claude guidance follows the Fable/Opus/Sonnet/Haiku ladder and sets the Agent tool's model
parameter explicitly. Model availability remains a native-harness requirement. - Eligible work includes read-only exploration, log/test triage, mechanical edits, spec-driven
tests or docs and a checked independent unit. Architecture, security, releases, migrations,
integration, unclear debugging and final review remain with the main agent. - One routed worker receives a self-contained brief. The main agent reviews its result and
finishes the work if the child fails its check. The portable skill, embedded copy, dashboard
explanation and RFC 0030 follow the same policy. - Hook processing remains local and deterministic, reads no prompt text and preserves the existing
callback receipts, trust requirements and transactional enable/disable lifecycle.
Verification scope
Local PR validation passed 2,403 tests with nine expected skips, plus typecheck, lint, build and
35 bundle smoke checks. Routing tests cover model tiers, malformed payloads, each ladder, the
light-tier no-op, root-only boundaries and the context character budget. Synthetic CLI probes cover
Sol/Astra output, empty Luna output, Claude's explicit model parameter, malformed input and
subagent-event output isolation.
Publication requires the release candidate's Windows, macOS and Linux CI, exact-tag build, guided
real-runtime smoke, package installation checks, provenance/SBOM attestations, npm Trusted Publishing
and npm-latest verification.
The native Linux routing audit records genuine
callbacks and lifecycle restoration with published 0.1.28. Those observations keep their exact
build scope and do not establish actual child-model selection under this revised policy. Windows
published-artifact checks, Claude Linux model authentication and paired quality-gated measurements
remain open. This release makes no new token, quota or API savings claim.
Publication evidence
Release-candidate CI passed on Windows, macOS and Linux. The exact-tag release workflow passed its build, tests, guided real-runtime smoke, packaging, provenance/SBOM attestations, npm Trusted Publishing and npm-latest checks.
The public npm and GitHub tarballs are byte-identical, with SHA-256 a27536e8e801e2bdc1492edf737d55cf74478e4044687c4b744947e78d4f37ee. npm integrity, both signed attestations against source commit 0985232ad20c081f874c78f66ade125e2a45550e, and an isolated installation of the public npm tarball (version and help) were independently verified. npm latest points to 0.1.29.
Token Harness 0.1.28 — Startup, Windows hooks and optimizer guidance
Token Harness 0.1.28 — Startup, Windows hooks and optimizer guidance
0.1.28 keeps automatic update checks from blocking dashboard actions, repairs reviewed Windows
hook paths, exposes read-only efficiency decisions and explains optimizer projects and benchmarks.
Installation
npm install --global token-harness@0.1.28
token-harness --version
token-harnessThe version must be 0.1.28. These commands also work in PowerShell. Restart an already running
dashboard to load the updated interface.
What changed
- Startup update checks are passive and coalesced. They preserve existing approval tickets and
discard results collected across a concurrent operation. Reviewing an update still creates an
exact-version approval before installation. - Native routing admits the reviewed Codex 0.146.0 hook format. Codex hook trust remains an explicit
native step; setup alone does not establish callback execution. - On Windows, setup can transactionally repair a missing or non-file standard HarnessTrim hook
executable. Custom commands stay under manual review. RTK's Codex launcher pins observed paths
and records to the agent's own database while preserving native command arguments. optimizeexposes advisory efficiency decisions with source-tagged evidence and exact-policy
capacity receipt references. Unavailable attempt and allowance budgets remain unknown.handoffaccepts repeatable--acceptanceand--factfields. Review caught and fixed a minimum
byte-budget regression: omission markers now compact before they can displace the mandatory
objective and next action. Per-section omission counts remain available in JSON.- Setup links each optimizer's source project and explains useful workloads, mechanisms and
limitations. Project-reported benchmark figures retain their scope and methodology and stay
separate from local measured results. The catalog remains available without a detected agent. - Routing explanations cover suitable delegation, the main agent's responsibility, native trust
and paired quality gates. Responsive disclosures support keyboard use and survive refresh. - RFC 0009 reconciles historical compatibility recordings with the shipped live-assignability
gates. No runtime guard or measurement gate is relaxed.
Verification scope
Local integration verification passed 2,395 tests with nine expected skips. Regression tests cover
mandatory checkpoint state and valid Unicode at the 256-byte minimum, update-check concurrency,
approval preservation and the reviewed hook repairs. The Windows launcher fixture exercises a PATH
containing Node alone and paths with shell-sensitive characters. UI interaction coverage checks
project links, claim isolation, the catalog without agents and disclosure persistence.
Publication requires the combined release candidate's Windows, macOS and Linux CI, exact-tag build,
guided real-runtime smoke, package installation checks, provenance/SBOM attestations, npm Trusted
Publishing and npm-latest verification.
The Windows audit and
six-pair pilot retain their pre-release scope, including
failed Codex workspace-write RTK commands, negative token/latency outcomes and excluded shared-account
quota. Package checks do not establish new runtime callbacks, marginal routing savings, successful
sandbox activation or broad promotion readiness. Issues #255, #359 and #360 retain their evidence gates.
v0.1.27
What's Changed
- docs: simplify onboarding and reconcile delivery roadmap by @gervaso-assistant in #363
- feat: simplify Overview and Results for 0.1.27 by @gervaso-assistant in #364
Full Changelog: v0.1.26...v0.1.27
v0.1.26
What's Changed
- Fix automatic routing and guided maintenance; release 0.1.26 by @gervaso-assistant in #358
Full Changelog: v0.1.25...v0.1.26
v0.1.25
What's Changed
- Fix Windows RTK and integration verification by @gervaso-assistant in #357
Full Changelog: v0.1.24...v0.1.25
v0.1.24
What's Changed
- release: Token Harness 0.1.24 — automatic native prompt routing by @gervaso-assistant in #356
Full Changelog: v0.1.23...v0.1.24
v0.1.23
What's Changed
- fix: verify Codex Desktop provider hooks by @gervaso-assistant in #355
Full Changelog: v0.1.22...v0.1.23
v0.1.22
What's Changed
- Remove obsolete CCR routing integration by @giuliastro in #353
- release: Token Harness 0.1.22 by @giuliastro in #354
Full Changelog: v0.1.21...v0.1.22