Repository navigation
Releases: rupeshpoojary9/OpenHarnX
Release list
OpenHarnX 0.1.1
OpenHarnX 0.1.1 is the first release on PyPI. The verifier is the same as 0.1.0; this release changes how you install it and how the project presents itself.
Commit: 2ea714c (signed tag v0.1.1), content verified READY under srt and by the gate in GitHub Actions (run 37576131380).
Install
uv tool install openharnx
npm install -g @anthropic-ai/sandbox-runtime@0.0.77 # provides srt, the sandbox
ohx doctorChanges
- On PyPI through trusted publishing: the upload comes from
.github/workflows/publish.ymlin this repository, approved by the maintainer, with no stored token; PyPI signs attestations of the files. - Package metadata: links to the repository, documentation, issues and releases; keywords and classifiers.
- README: opens with the problem, shows a 45-second scripted demo (no model ran) and a picture of a blocked report, states the evidence with what it missed and what it blocked, invites trial reports, and works as the PyPI page. The logo follows GitHub's own light or dark theme.
- Project files:
CONTRIBUTING.md,ROADMAP.md, a trial-report issue template.
Supported
Unchanged from 0.1.0: Python projects tested with pytest, on macOS (arm64) locally and Linux in CI, with srt as the sandbox. TypeScript, JavaScript and Go work but are experimental. Any coding agent can be checked with ohx verify or the CI gate; only Claude Code has a built-in integration and has been tested end to end.
What this release does not claim
It has no production users yet. It does not claim faster reviews or safe changes; a passing verdict is evidence about the checks that ran, not permission to merge. The evidence so far, with its limits, is in the README.
Security reports: through GitHub's private vulnerability reporting (see SECURITY.md).
OpenHarnX 0.1.0
OpenHarnX 0.1.0 is the first release: an open-source verifier for changes written by coding agents. Keep your agent. OpenHarnX locks the tests you agreed on, runs every check in a sandbox and gives a verdict with evidence, opened by a review brief: what was asked, what changed, what passed, what remains unverified and what needs your judgment. No model is called.
Commit: 0162666 (signed tag v0.1.0). The gate passed on this exact commit in GitHub Actions under the srt sandbox (run 37478954436), and its report carries a Sigstore attestation from that workflow: gh attestation verify report.json -R rupeshpoojary9/OpenHarnX.
Install
uv tool install git+https://github.com/rupeshpoojary9/OpenHarnX@01626665a341812ded02c0559f5428c2e71d5ef1 # pinned by commit
npm install -g @anthropic-ai/sandbox-runtime@0.0.77 # provides srt
ohx doctorThen follow the quick start in the README: the cheat demo, then ohx init --lock-tests on your own project.
Supported
Python projects tested with pytest, on macOS (arm64) locally and on Linux in CI, with srt as the sandbox. TypeScript, JavaScript and Go projects, the mutation check, ohx bug and ohx trace work but are experimental. Not supported: Linux outside CI, Windows. The full boundary, with the tests behind each criterion, is in docs/gate-release-criteria.md; the threat model, with what is prevented, detected and open, is in docs/threat-model.md.
What it does
- Locks the existing suite (
ohx init --lock-tests) or agreed acceptance tests (ohx contract new). Editing, skipping, deleting or weakening a locked test, loosening check configuration, or breaking a test that passed before is BLOCKED. - READY needs agreed acceptance tests; NO REGRESSIONS means nothing that passed before broke. A missing or crashed check is UNKNOWN, never a pass.
- Every report opens with a review brief, built from records, each verification claim linked to the check's output.
- A Claude Code Stop hook verifies when the agent says it is done and sends a BLOCKED agent back with what failed.
- CI:
ohx gatejudges a pull request against its base branch's tests and policy; a GitHub Action, GitLab and Jenkins examples. In GitHub Actions the report is signed by the pipeline (Sigstore). - Evidence in a hash-chained local store, signed with your SSH key.
What this release does not claim
It has been tested on its own development, on replays of public agent work and with simulated users, and the results are published with what it missed and what it blocked: on the impossible-tasks dataset, 80 of 102 runs replayed, 43 of 52 fake fixes caught with zero setup (all 26 that changed tests, 17 of 26 that changed only code; 5 of 11 on the held-out half), 8 legitimate test edits blocked until accepted; on 20 merged agent pull requests in github/spec-kit, 9 passed, 7 passed after approval of intended test changes, 4 failed in the replay's environment; on 31 agent commits in anthropics/claude-agent-sdk-python, 28 passed and 3 were intended test changes. Details: docs/tasks/T89.md, T90.md, T92.md. It has no production users yet. It does not claim faster reviews or safe changes; a passing verdict is evidence about the checks that ran, not permission to merge.
Security reports: through GitHub's private vulnerability reporting (see SECURITY.md).