Releases: snapsynapse/harnessie
Release list
Harnessie 1.4.1: offline open-record AIDR export
Harnessie 1.4.1: offline open-record AIDR export
Published on GitHub on 2026-09-08 UTC at signed tag v1.4.1, pointing to 296deed2f91cd4c8eeecad82b83f137dd029ea26, with original-build assets attached. Both PyPI distributions match those assets, pass publisher-attestation verification and install from the public index in a fresh Python 3.13 environment. Package and assistant guide both use 1.4.1. Version 1.4.0 was intentionally skipped.
Harnessie 1.4.1 adds deterministic export of a supported open contested-phase record to an explicitly named AIDR file. It carries recorded dissent into a separate artifact while preserving the original decision and its run state.
Included
harnessie export-aidr RUN_ID PHASE --output decisions/AIDR-NNNN-short-slug.md --arbiter HUMAN_HANDLErequires an unused destination in the project's existingdecisions/directory and a declared human arbiter. It prints JSON and exits 0 for completed export or 2 for refusal.- The exporter accepts a strict subset of generated open records and their single emitted event-journal reference. Arbitration content, decided metadata, ambiguous Markdown/YAML, unresolved or broken evidence, unsafe paths, symlinks, and destination ID collisions refuse. Some records accepted by the runtime contain unsupported formatting and therefore cannot be exported; the exporter does not silently rewrite them.
- Repeated reviewer roles retain distinct participant-instance labels and original-role provenance. Recorded position prose and objections survive. Source and evidence hashes bind the consumed snapshot; validated output is created exclusively after the input bytes are rechecked.
- The executable example runs actual workflow code with scripted mock actors, invokes the installed export CLI, checks the result against the pinned AIDR 0.1.0 reference linter, and verifies source preservation and continued human-arbitration halt. It also demonstrates safe refusal of unsupported source formatting. Export itself has no Node dependency; reference-linter acceptance uses Node.
- Platforms without the required POSIX locking and file primitives return
unsupported_platformwith exit 2 before project reads or writes. Native Windows export is unsupported. - Engineering docs, user instructions, machine discovery, and the assistant guide describe the same export boundary and source limitations. Four deterministic exporter eval cases cover admissible open dissent, partial arbitration, namespace collision, and broken evidence.
Authority and evidence limits
Export is an operator-issued file write outside the runner, ownership ledger, tool registry, consent lock, and approval policy. It makes no model or provider calls, but it is not a read-only operation. An assistant invoking it needs authorization for the destination write. The command does not author or import arbitration, approve a decision, or resume the original run. The exported record remains open with empty Arbitration.
Model/provider metadata and arbiter declarations are reported identities, not authenticated authorship. Export-time hashes do not prove the source matches its original generation, that reviewers saw identical inputs, or that review was independent. Structural validation does not judge whether a question is decidable or an objection was adequately answered. Existing objection text is truncated upstream to 500 characters, event copies to 200; export cannot recover missing text or round attribution. Source and destination directories must remain operator-controlled. The export contract documents the complete supported subset and filesystem limits.
Verification status
The release execution audit records the local candidate gate, exact-commit CI, artifact validation, installed-consumer checks, and deployment evidence. CI, CodeQL, Scorecard, and Pages passed on release commit 296deed. The earlier completion and version preparation audits remain dated evidence for their own snapshots. All four GitHub assets passed provenance verification against the signed tag and release commit. Both PyPI distributions passed publisher verification; the fresh public-index consumer passed the installed export example. Verify Action 0.2.2 passed all seven exact-merge fixtures.
The final 8,036-byte guide earned hosted GuideCheck Level 4 under profile 2.0.0 with zero blocking findings and a qualifying DNS anchor. Its frozen SHA-256 is 7ab4c0a952109ea10257b1a9859533778261291605ce386708bccb305bbf09dc. The 1.3.1 receipt remains historical evidence for different guide bytes. Guide conformance does not establish software safety or accessibility conformance.
Deferred work and downstreams
Accessibility remains deferred for this session. The prior queue retains 35 table-contrast candidates, one video-caption applicability candidate, and manual checks; earlier evidence does not establish acceptance for newly changed routes.
Verify Action 0.2.2 and stable v0 pin core 1.4.1 at 9f18d70017f395ef4d15ac5746e30ac633f00f8e. Homebrew 1.4.1 passed strict audit, a real upgrade, formula and linkage tests, and the installed exporter demonstration; it is published in tap merge 953760f3968200f99658bfb068609532cb9fbf4d. Engine wrappers remain on their independent 0.1.0 release train.
Live review panels, arbitration import, broader source-format support, bidding, commentary, follow mode, automatic runner integration, model selection, and provider-policy changes remain deferred. No real human arbitration record was changed by the exporter delivery.
Original release assets
| Asset | SHA-256 |
|---|---|
harnessie-1.4.1-py3-none-any.whl |
2094c7793a0fde76226d5782b0785311d9dbc955998a0c7237580f43b1066dd5 |
harnessie-1.4.1.tar.gz |
4348ca226397a79d317dc619374b75019a43f8f660b30d3410d83d2cbf925d85 |
harnessie-1.4.1.SHA256SUMS |
11476d68a6e6f82f7e205457d11262e2eedb456cc90380ad25ec25796384b229 |
harnessie-1.4.1.cdx.json |
228a15e401cd337dc8acd7bc2264ea176162d02d2f13e0490e2a94cb9e1307c2 |
Final documentation verification
Closeout PR #21 merged at e2b7c80b7572945a7765bb8f7c1bf29e920b9133. Exact-commit CI, CodeQL, Scorecard and Pages all passed. All 23 checked live resources matched the merged bytes; the final ten-page production search check found zero defects and zero infrastructure failures. Scorecard's six previously recorded findings remain unchanged with the dispositions in the release audit. The frozen guide, immutable release tag and original package assets remain unchanged.
Accessibility remains deferred under the session instruction; this documentation verification is not an accessibility acceptance result.
Harnessie 1.3.1: offline observation and verified release integrity
Harnessie 1.3.1: offline observation and verified release integrity
Published on GitHub and PyPI on 2026-09-08 UTC, with verified original-build assets and PyPI attestations. Verify Action 0.2.1/stable v0 and Homebrew 1.3.1 are also published and tested.
Harnessie 1.3.1 adds offline run observation and strengthens release integrity while retaining the stable 1.x authoring and plugin contracts.
The initial 1.3.0 workflow stopped before asset upload and PyPI publication because of mutually exclusive provenance CLI flags. This corrective release fixes that invocation while preserving exact certificate/workflow/tag and commit verification. The original signed tag remains unchanged.
Included
harnessie observe RUN_IDverifies an existing local journal snapshot and writes cited JSON and Markdown without model calls, runner integration, source-journal changes or approval changes. Exit 0 means observation succeeded, not that the run passed.- Hash-verified runtime, development and release locks, including patched uv 0.11.15 tooling, with separate public-constraint consumer coverage.
- Original-build provenance enforcement for the wheel, source archive, CycloneDX SBOM and checksum record. Recovery refuses historical assets without qualifying original provenance.
- Bounded parser properties and refusal fixes for non-string claim statuses and NUL-containing evidence paths. Verdict parser identity advances to 3.
- Keyboard-focusable documentation scrolling regions, visible focus, homepage contrast improvements and preserved text alternatives.
- A current profile-2.0.0 assistant guide, synchronized trust artifacts and complete observer CLI documentation.
Verification
The full local locked gate passed: 572 tests, one environment-dependent skip, 28 expected failures for deliberately deferred features, 62/62 deterministic evals, both manifests, package inspection and a fresh-install observer smoke. Final source CI covers Linux sandboxed execution, fail-closed execution without a sandbox, macOS, public-constraint packaging and locked packaging.
The final 1.3.1 guide earned hosted GuideCheck Level 4 under profile 2.0.0 on 2026-09-08 UTC, with zero blocking findings. Served, sidecar, DNS and repository bytes agree on SHA-256 ef62ac6ecc50f1f277dda119299d1b145e0aae66a51bcf47ba233e23c072409b. This verifies guide provenance and structure, not software safety or runtime execution.
Live Siteline scored 97/100, grade A, on 2026-09-08 UTC. The production search contract passed across ten sitemap pages.
Scope and follow-up
AIDR-0009 continues to defer bidding, commentary, follow mode, automatic runner integration and model selection. Four confirmed serious accessibility violations and 22 review candidates were repaired. The maintainer explicitly deferred the remaining 35 table-contrast review candidates, one video-caption applicability candidate and manual accessibility acceptance as nonblocking follow-up. The automated audit remains inconclusive, with zero confirmed violations; this release does not claim complete accessibility conformance.
The release workflow builds and attests four final assets, verifies original build identity and publishes the same distributions through the protected PyPI Trusted Publisher. Verify Action and Homebrew propagate separately after core publication; engine wrappers retain their independent release train. Execution and delivery evidence records the exact outcomes.
Published assets
| Asset | SHA-256 |
|---|---|
| harnessie-1.3.1-py3-none-any.whl | 2de49487168fd58ac77c8a7aa44a024a0adb0239571fc2f0292966c9dd5d42cd |
| harnessie-1.3.1.tar.gz | fe250f5c86f993ec9af5723e85f42c57f0889c195ad2e57726d8782f827bb520 |
| harnessie-1.3.1.SHA256SUMS | 37b9497b392a61e2fb6c1df0e82303c97168b7adddf979766d3fa52ea5bd8fca |
| harnessie-1.3.1.cdx.json | 3b82707a4e654a969c969766eacb9e3a68134b5e1fa52c0e0dc75e609f278da6 |
Core tag v1.3.1 resolves to 6ddf84429ab8fbfa9ec96e59b283cf7cb341dd6f. Verify Action v0.2.1 and stable v0 resolve to 97ed2264818fc4c620c03de5cc9522965318079b; Homebrew formula commit is 630021462ef961669a0802d090c7c77c65190932. Engine wrappers remain 0.1.0 on their independent train.
Harnessie 1.3.0: offline observation and release integrity
Publication failed before asset upload or PyPI. The signed 1.3.0 tag is preserved. A provenance CLI argument conflict stopped the release workflow after building and attesting assets. Corrective package publication is proceeding as 1.3.1 with the same feature scope and retained provenance requirements.
Harnessie 1.3.0 adds offline run observation and strengthens release integrity while retaining the stable 1.x authoring and plugin contracts.
Included
harnessie observe RUN_IDverifies an existing local journal snapshot and writes cited JSON and Markdown without model calls, runner integration, source-journal changes or approval changes. Exit 0 means observation succeeded, not that the run passed.- Hash-verified runtime, development and release locks, including patched uv 0.11.15 tooling, with separate public-constraint consumer coverage.
- Original-build provenance enforcement for the wheel, source archive, CycloneDX SBOM and checksum record. Recovery refuses historical assets without qualifying original provenance.
- Bounded parser properties and refusal fixes for non-string claim statuses and NUL-containing evidence paths. Verdict parser identity advances to 3.
- Keyboard-focusable documentation scrolling regions, visible focus, homepage contrast improvements and preserved text alternatives.
- A current profile-2.0.0 assistant guide, synchronized trust artifacts and complete observer CLI documentation.
Verification
The full local locked gate passed: 572 tests, one environment-dependent skip, 28 expected failures for deliberately deferred features, 62/62 deterministic evals, both manifests, package inspection and a fresh-install observer smoke. Final source CI covers Linux sandboxed execution, fail-closed execution without a sandbox, macOS, public-constraint packaging and locked packaging.
Hosted GuideCheck verified the final 7,766-byte guide at Level 4 under profile 2.0.0 with zero blocking findings and an independent matching DNS anchor. Guide SHA-256: 5e00fa6903e0d42c49e9a25c08eec19116b04bf01ee0daf8c4d7d754b58b0793. The exact pre-publication receipt preserves warnings and limits; no Level 5 runtime claim is made.
Live Siteline scored 97/100, grade A, on 2026-09-08 UTC. The production search contract passed across ten sitemap pages.
Scope and follow-up
AIDR-0009 continues to defer bidding, commentary, follow mode, automatic runner integration and model selection. Four confirmed serious accessibility violations and 22 review candidates were repaired. The maintainer explicitly deferred the remaining 35 table-contrast review candidates, one video-caption applicability candidate and manual accessibility acceptance as nonblocking follow-up. The automated audit remains inconclusive, with zero confirmed violations; this release does not claim complete accessibility conformance.
The release workflow builds and attests four final assets, verifies original build identity and publishes the same distributions through the protected PyPI Trusted Publisher. Verify Action and Homebrew propagate separately after core publication; engine wrappers retain their independent release train. Execution and delivery evidence records the exact outcomes.
Harnessie 1.2.0: verify agent-produced changes
Harnessie 1.2.0: verify agent-produced changes
Harnessie 1.2.0 makes harnessie verify the smallest independently useful adoption surface. It binds claims to exact evidence, runs deterministic checks before model judgment, and derives a fail-closed exit from complete required-claim coverage. The full harness remains the growth path for consent, ownership lanes, containment, human arbitration, and tamper-evident audit.
Highlights
- A v1 evidence bundle binds stable claim IDs to an exact Git revision and dirty state, content-addressed diffs and proof files, and recorded deterministic checks. Unsafe paths, stale state, missing bindings, and hash drift refuse before model dispatch.
- Structured claim results classify every required claim as
reproduced,refuted, ornot_verifiable. Overall exit 0, 1, or 2 follows deterministically from complete claim coverage; legacy raw criteria remain compatible. - OpenAI Responses is now a first-class adapter for current reasoning models, stateless encrypted-reasoning replay, strict function tools, response validation, and usage accounting.
- Synthetic Ringer fixtures and trace metrics cover the adoption seam exposed by the first public Ringer cohort, including duplicate denials, repeated tool calls, work steps, token use, and claim coverage.
- Parallel copies of one denied tool call count as one failed turn, allowing a verifier to recover without weakening the tool allowlist.
Verification and supply chain
- The repository release gate composes the full test suite, 51 deterministic evals, authoring validation, inward and outward trust manifests, generated documentation, artifact inspection,
twine check, and fresh-install smoke. - The GitHub release workflow checks out the exact annotated tag, verifies tag/version/commit identity, runs the release gate under admitted Linux bubblewrap confinement, and builds the wheel and source distribution once.
- Those exact distributions are attached to this release with a reproducible CycloneDX 1.6 runtime SBOM and SHA-256 record, then sent unchanged to PyPI through Trusted Publishing and its protected
pypienvironment. PyPI attestations remain enabled. - The v1.2.0 tag is annotated but not locally GPG-signed. No project policy requires a maintainer-key tag signature, and adding a long-lived local key would create another identity and rotation surface. The compensating evidence is the exact protected tag commit, GitHub workflow identity, PyPI OIDC attestations, immutable release assets, SBOM, and recorded SHA-256 digests.
- The public OpenSSF Scorecard API returned no result for this repository on 2026-09-01. This is recorded as not yet measured, not as a zero or a pass. The absence is accepted for 1.2.0 because the release gate, full commit-pinned workflow, protected environment, OIDC publication, attestations, SBOM, and immutable digests are independently verified; a repository-owned Scorecard workflow remains a post-release measurement task.
- Exact counts, commit, workflow URLs, digests, PyPI attestation results, and downstream versions are verified during release closeout in
CHANGELOG.mdandNEXT.md; they are not invented in advance here.
Trust and downstream boundaries
- Manual keyboard-only, 200% zoom/reflow, and screen-reader testing was not completed before publication. On 2026-09-01 the release owner explicitly authorized a one-time 1.2.0 waiver after reviewing that gap. Automated Lighthouse 13.4.1 checks scored 1.00 with no failing accessibility audits on the homepage and all nine generated routes, but that evidence is not represented as a substitute for assistive-technology testing. The manual gate remains required for later releases.
- The 1.2.0 assistant guide, served copy, provenance sidecar, and repository trust pins move together. Its new DNS TXT anchor and hosted GuideCheck run are separate external gates. The 1.1.0 Level 4 receipt remains dated historical evidence and is not inherited by the new guide bytes.
- Harnessie Verify Action and Homebrew are separately released downstreams. Do not assume they expose the 1.2.0 evidence-bundle contract until their own pins and gates are verified.
- Engine wrappers retain their independent probe-gated release train because 1.2.0 consumes no new versioned wrapper seam.
The complete change record is in CHANGELOG.md.
Harnessie 1.1.0
Harnessie 1.1.0: the Golden Rule becomes inspectable
Harnessie 1.1.0 turns its ownership invariant into a memorable public contract and an executable proof:
Read together. Write only what you own.
Highlights
harnessie ownership PATH --agent AGENT [--json]explains the exact write decision used by the ownership ledger without claiming or changing the path. Human output names the governing lane, owner, pattern, reason, and remedy. JSON output uses schema version 1.- The zero-model, zero-network ownership-collision example performs a real overwrite attempt through the built-in
write_fileregistry and passes only when the second agent is denied, the original bytes survive, and the ledger still names the first writer. - The website, README, Guide,
llms.txt,agents.json, local CLI manifest, machine-readable changelog, assistant guide, roadmap, and project context now describe the same shipped 1.1.0 behavior. - Search discovery and generated-page contracts added since 1.0.0 remain part of the release gate.
Verification
- 433 tests passed with 1 environment-dependent skip.
- Deterministic eval scorecard: 50/50.
- Nine shipped authoring documents validated against schema v1.
- Outward trust manifest: 19 files. Inward manifest: 15 files.
- Nine generated documentation pages and ten-page search contract verified.
- Wheel and source distribution passed
twine check, structural inspection, private-surface scrubbing, and fresh-install smoke. - Ownership adversarial coverage includes workspace escapes, absolute paths, control characters, symlink resolution, invalid agents, declared-lane precedence, first-writer denial, allowed-decision false-positive checks, and proof that inspection does not mutate the ledger.
Trust boundaries and release residuals
harnessie ownershipis an explanation surface, not an authorization grant. Enforcement remains in the registry, runner preflight, ownership ledger, and OS sandbox.- Collaborative lanes deliberately allow co-editing, and operator-trusted in-process plugins remain outside child-process lane confinement.
- The 1.1.0 assistant guide and served copy are byte-identical and pinned by the provenance sidecar. After publication, the independently controlled DNS TXT and repository-file anchors matched the final hash and hosted GuideCheck re-earned Level 4 with zero blocking findings.
- Harnessie Verify v0.1.3, stable Action tag
v0, and the Homebrew formula carry core 1.1.0 after their separate gates passed. Engine wrappers remain independently released at v0.1.0 because 1.1.0 consumes no new wrapper seam.
The complete change record is in CHANGELOG.md.
1.0.0: extensibility earned
Harnessie 1.0.0: extensibility earned
Harnessie 1.0.0 freezes the authoring contract, closes the interpreter ownership gap, and admits installed tool extensions through one explicit trust boundary.
Highlights
- Six strict Draft 2020-12 authoring schemas cover models, cascade, boundary, approval policy, ownership, and workflows. Runtime startup and
harnessie validateuse the same fail-closed contract. - Worker shell calls, deterministic checks, and verifier commands receive agent-specific read-only ownership overlays. Unsupported nested profiles refuse instead of running unconfined.
- Installed tool plugins use only the
harnessie.tools.v1entry-point group and never auto-load. Explicit--plugin NAMEadmission validates and namespaces tools, applies registry policy, records loader-supplied provenance, and pins exact receipts across resume. - Anthropic and OpenAI-compatible adapters now normalize malformed provider responses into non-echoing error turns. Adversarial rebuttal agents receive complete peer positions without self-position leakage.
Verification
- 413 tests passed with 1 environment-dependent skip.
- Deterministic eval scorecard: 50/50.
- Nine shipped authoring documents validated against schema v1.
- Outward trust manifest: 19 files. Inward manifest: 15 files.
- Ecosystem manifest and 8 generated documentation pages verified.
- Wheel and source distribution passed
twine check, structural inspection, private-surface scrubbing, and fresh-install smoke. - Exact-commit GitHub Actions passed on macOS, Linux fail-closed, Linux bubblewrap, and package jobs.
- A clean install from the public PyPI index reported Harnessie 1.0.0.
- Adversarial plugin cases covered explicit-only admission, zero-selection non-discovery, malformed declarations, invalid parameter schemas, duplicate names, role denial, provenance attribution, and resume drift. No bypass or legitimate-corpus regression remained in the complete suite.
- Live provider scorecards were not run because the release changes harness mechanics and package contracts, not provider behavior. The deterministic mock-brain and real OS sandbox paths exercised the affected boundaries.
Release assets
harnessie-1.0.0-py3-none-any.whl:sha256:7a4bb95299eaae206010697ea2346b17024c6860480d4049b288f1ea75337076harnessie-1.0.0.tar.gz:sha256:40c2daa307d71a4687321205fac8fc1a24b6778c4412fb1f20cb2b20f89bd787
The same files and hashes are published on PyPI.
Trust boundaries and residuals
- Plugin implementations run in process as operator-trusted code. Registry mediation is not a sandbox and does not prove declared effects. Untrusted plugins remain unsupported pending a separately versioned out-of-process design.
- Docker remains admitted for the base workspace sandbox only. A nonempty lane profile requires bubblewrap, firejail, or Seatbelt until Docker has a truthful nested-mount admission probe.
harnessie-verify-actionv0.1.1 and the Homebrew formula remain on Harnessie 0.8.0 pending their separately authorized release trains.- The external
_assistant-guide.harnessie.comDNS TXT anchor must rotate to the 1.0.0 guide hash before hosted GuideCheck can re-earn its independently anchored level.
The complete change record is in CHANGELOG.md.
0.8.0: write-safety and self-integrity
0.8.0 (2026-08-04)
Theme: bound what a governed run may change and pin the harness inputs that define its behavior.
Highlights
- Blast-radius ceilings atomically roll back writes that exceed declared phase or workflow limits.
- Parallel groups may declare exact write paths; malformed, partial, aliased, or overlapping declarations refuse before dispatch.
- Maiden-voyage staging verifies new phase contracts without changing the target until explicit operator promotion.
- The inward manifest pins shipped prompts, configuration, and static ownership policy before model dispatch.
- The composed release gate now verifies source, generated docs, package metadata, archive safety, private-surface scrubbing, and a fresh install.
- Public machine handoffs truthfully expose local CLI capabilities, release history, support, and security contacts while explicitly denying a hosted API, hosted service, or MCP server.
The complete change record is in CHANGELOG.md.
Verification
- 352 tests passed with 1 environment-dependent skip.
- Deterministic eval scorecard: 47/47.
- Outward trust manifest: 13 files. Inward manifest: 9 files.
- Ecosystem manifest and 8 generated documentation pages verified.
- Exact release commit CI passed on Linux with bubblewrap, Linux without a sandbox backend, macOS, and the package job.
- Wheel and source distribution passed
twine check, structural inspection, private-surface scrubbing, and fresh-install smoke. - Fresh live Siteline rubric 2.3.0 scan: A, 97/100, Level 4 machine enablement at 16/18. The external response provenance and result-store residual are recorded in
audits/siteline-live-result-2026-08-05.json.
Artifact digests
harnessie-0.8.0-py3-none-any.whl:f51946a8f0fb33342627173f8e959b21ef54dd7b6a4b93b0f8e2cd7bd858602eharnessie-0.8.0.tar.gz:c9caffef61a8b1f9569cee36ede59f59c3dc8c66a47e500ecc090d445111f5e7
Ecosystem state and residuals
harnessie-verify-actionremains v0.1.0 and pins Harnessie 0.7.1 pending its separately authorized release train.- The Homebrew formula remains on Harnessie 0.7.1 pending its separately authorized release train.
harnessie-engine-wrappersremains independently released at v0.1.0; core 0.8.0 does not imply a wrapper bump.- The v0.8.0 assistant guide, sidecar, and repository-file anchor match SHA-256
19e8bae7ea3e66a69f484abf3c3bece469ce7e1645aa55b935f6fbdd1548a0a5. Hosted GuideCheck 0.7.1 currently achieves Level 3 with one blockinganchor.independent.mismatchbecause the external DNS TXT anchor still carries the older hashcf621fdb1adc036e6cbed7845d652fa1295338ebfdee4dc4053c0f32a924038c. Updating that record and re-running hosted verification remains an operator close-out; the two expected GitHub Pages header warnings are unchanged.
0.7.1: the verifier leaves the harness
0.7.1 (2026-07-09)
Theme: the verifier leaves the harness. One addition, adopted through the contested-decision process like everything before it.
Added
- Standalone verification surface
harnessie verify(harness/verify_standalone.py), adopted viadecisions/AIDR-0006(four-provider position sweep, human-arbitrated): point the VerificationGate's two layers at any workspace plus a claims file with no project scaffold, orchestrator, or run manifest. Deterministic checks run sandboxed and network-denied (opt-in--allow-networkfor artifacts whose own tests bind sockets; the verifier agent stays denied regardless), then a read-only fresh-context verifier tests the criteria claim by claim. Exit contract is scriptable and fail-closed: 0 verified, 1 failed, 2 cannot-verify (missing config, sandbox unavailable, provider error — neither pass nor fail was earned). Single pass by design: a foreign artifact's failure is the answer, not a prompt to reformulate and retry. The report carries the workspace git revision, criteria hash, verifier model, and network mode. First proving ground: independent verification of agent-produced pull requests. Proven bytests/test_verify_standalone.py.
v0.7.0: Sovereignty cascade routing and the containment boundary
Sovereignty cascade routing and the containment boundary. Route every task to the least-exposed environment that can complete it, and make containment a mechanical property of the run rather than an operator habit. The whole milestone was adopted through Harnessie's own contested-decision process: five arbitrated decision records (decisions/AIDR-0001 through AIDR-0005), three of them run on six- and three-model Ollama Cloud panels spanning eight providers.
A workflow that does not opt in to any of the new machinery routes byte-identically to 0.6.
Routing
- Cascade policies (
config/cascade.yaml): a phase opts in withcascade: <policy>instead of a fixed tier; the policy declares its tier ladder, which failure reasons climb it, a maximum climb, and a never-silent on-exhaust action. - Sideways provider fallback: refusals and availability failures move across providers at the same tier (a tier's
fallbacks:), never upward — up-tiering on a refusal is a containment leak. - Escalation headroom (AIDR-0005): a climb is refused before dispatch when the budget cannot cover the target tier's worst-case first turn; overspend is bounded to ceiling plus one in-flight turn, proven by a test.
- Sovereign tier: a fifth slot for any OpenAI-compatible controlled endpoint, deliberately off the default escalation walk — an exposure class, not a capability rung.
- Reserved pre-gate: work classes that never reach any model at any tier;
arbitrationis reserved by default. - routing_trace on every attempt (tier, effort, fallback, model, provider, outcome).
Containment boundary
harness/boundary.py, adapted with provenance from PAICE.work PBC production PII code released under Apache-2.0 (see NOTICE). Opt-in via config/boundary.yaml. Its guarantee is a per-data-class coverage table, not a blanket claim:
- Structured PII is stripped to stable placeholders before any content reaches a model, so no model call and no run artifact carries a raw value.
- Secrets are never mapped or rehydrated, and a secret in an egress payload halts the run (
secret_egress, kind labels only, no warn mode). - Unstructured free-text PII is explicitly not filter-caught — it is covered by contained routing, which keeps it on your local/sovereign tiers (never-egress, not imperfect filtering).
The strip map lives outside the run tree, reloads fail-closed on resume, and rehydration happens only at the operator boundary under deny-all per-tool grants.
Proof
Canary leak evals (zero seeded PII/secret bytes in any run artifact), gate-integrity canaries (phantom context, autonomy grab, counterfeit system frame), a per-brain placeholder-impact scorecard, and bundle identity (a scorecard pins model, provider, endpoint, prompt hash, parser version, sampling).
Verification: 255 passed / 1 skipped, 43/43 eval scenarios, trust manifest verified.
Install: pip install harnessie (or pipx install harnessie / uv tool install harnessie / brew install snapsynapse/tap/harnessie).
Full detail in CHANGELOG.md.
Harnessie 0.6.0 — first-harness readiness
First public release. Harnessie 0.6.0 is the "first-harness readiness" line: the release that had to make "the safest and easiest first AI harness for people" checkable before the repo, the site, and the package went public. It is on PyPI as of this release:
pip install harnessie # or: pipx install harnessie / uv tool install harnessie
harnessie init my-project # guided readiness check, ends on a green zero-dollar mock runSafest, as checkable claims
- Falsifiable threat model (docs/threat-model.md): an eleven-row table mapping Harnessie's structural properties against the failure modes of prevailing harness patterns (unsandboxed shell, prompt-level-only guardrails, self-verification, silent dissent-merging, and more). Every row cites the enforcing code and the test that proves it; all 25 cited test nodes pass. The honest residual is stated, not hidden.
- Default-deny posture audit: eleven assertions proving the shipped tool registry, ownership ledger, and CLI seams default closed — the orchestrator holds no side-effecting tool, the verifier never writes, network defaults off, and unknown-tool / wrong-role / unapproved / pre-consent dispatches all fail closed.
- Standing "break it" invitation: a vulnerability-disclosure path in SECURITY.md plus
evals/redteam.yamlpublished as red-team targets — three canary-exfiltration scenarios anyone can run against the exfil guard, the refusal grammar, and the shell allowlist. - Pre-run cost preview: every run states LIVE vs MOCK with ceilings and a worst-case dollar figure before any brain is built, and a live run with no spending ceiling refuses to start.
- Standards adopted, not just credited: Graceful Boundaries for the refusal surface (Level 1 grammar on every denial site, transport-adapted, test-proven) and GuideCheck for the assistant guide — served at /.well-known/assistant-guide.txt with a provenance sidecar and an independent DNS TXT anchor, verified end-to-end at GuideCheck Level 4 by the hosted verifier (zero blocking findings).
Easiest, for someone who has never identified as a developer
- Guided first run:
harnessie initchecks Python, detects the OS sandbox (framing a missing one as protection, not breakage), explains key handling (env vars, never files), and proves itself with a zero-dollar mock run before naming your next command. - Plain-language operator surface: runs end with a summary that leads with the outcome; every halt names one action (
harnessie resume ..., or the exact decision record to edit).harnessie reportreconstructs a readable summary even from a crashed run; the raw event dump moved behind--raw. - A quickstart that assumes nothing (harnessie.com/quickstart.html): no git or shell fluency required, a 19-term glossary in the order a newcomer meets each word, and an honest Windows page (bare Windows fails closed on shell steps; WSL2 is the supported path).
- Docs served on-site: every doc is a navigable HTML page on harnessie.com — no GitHub account needed to read the guide.
Also in this release
- Relicensed MIT → Apache-2.0 with NOTICE (trademark and PAICE-specification carveouts).
- Operability evals extended with risky/recovery coverage (invalid approval policies, parallel failure halting, audit-chain survival under concurrency); new stewardship evals for public-surface hygiene.
- The 0.7.0 milestone (sovereignty cascade routing and the containment boundary) is planned in the ROADMAP, gated on an adoption decision run through Harnessie's own contested-decision workflow.
Verify this release yourself
python3 -m pytest -q # 195 passed, 1 skipped
python3 -m harness.cli eval # 38/38 deterministic scenarios
python3 -m harness.cli verify-manifest # trust bundle, 9 filesFull detail: CHANGELOG.md. The safety claims are contestable in public — if you can break one, that is a bug report we want (SECURITY.md).