Skip to content

v6.0.0 — SpecForge

Choose a tag to compare

@jefmonjor jefmonjor released this 05 Oct 15:49
· 4 commits to main since this release
4e70e7f

SpecForge 6 keeps the one rule, verify, don't trust, and makes the checking proportional: cheap for documentation, thorough for the code that handles money, secrets and permissions. Every new check is code, not a prompt.

Install

R=https://github.com/jefmonjor/specforge/releases/download/v6.0.0
F=specforge-darwin-arm64   # or darwin-amd64, linux-amd64, linux-arm64
curl -LO $R/$F && curl -LO $R/checksums.txt
sha256sum -c checksums.txt --ignore-missing   # macOS: shasum -a 256 -c …
chmod +x $F && sudo mv $F /usr/local/bin/specforge
specforge version
specforge doctor

On Windows, download specforge-windows-amd64.exe, rename it to specforge.exe and put it in a folder on your PATH. Then follow the quickstart and the user guide.

What's new in 6.0

⚖️ Proportional review

  • Known failures are evidence, not blockers. The loop runs the suite once before its first RED; tests that already fail are named from the runner's report, shown, kept out of REFACTOR's verdict and listed in the delivery.
  • Risk tiers from git. Passive, medium or high, from the paths and lines each scenario touched. The agent can raise the tier with a reason, never lower it. review: risk asks for your review only when it matters; mutation testing runs from medium up.
  • A model per phase with models.<phase>.

🔎 Review you can check

  • Review lenses (risk, reliability, readability, resilience): none for documentation, one for ordinary code, four for sensitive code. Each is read-only and schema-validated; a finding whose proof is not a changed line is discarded; inferential findings go to an independent refuter.
  • One correction, within a budget of half the change's lines, judged again by REFACTOR and validated on the corrected findings only. A regression goes to you, never to a second automatic correction.
  • Blind double review for high risk: what two independent passes prove on the same hunk skips the refuter.
  • specforge review runs the lenses over a whole branch.

🔍 An independent verifier

For high-risk scenarios, a second agent probes the specification in a disposable copy of your project, the only place an agent may run commands without asking. It must give a verdict for the scenario and the invariants it names, and SpecForge runs every failing command again before believing it. Your project must stay untouched. Proposed regression tests are offered, and added only if the suite still passes. specforge verify checks a whole specification.

🛡 Operational safety

  • A destructive-command guard installed by setup as Claude Code's or Gemini CLI's pre-tool hook: recursive deletes, git reset --hard, push --force, DROP TABLE, DELETE without WHERE, secrets… blocked with the reason the agent sees. Deleting /, ~ or the project is never allowed. It never fails open.
  • Edit surfaces from the plan: a file the agent changes outside the approved plan goes to you; a refused one must be put back.
  • specforge doctor: the agent, git, your runner and every gate's tool, with install hints.
  • Questions with options: in the loop, the plan and legacy turns, the agent must derive the candidate answers.

📦 Delivery and speed

  • Each scenario's size from its commit; above delivery.budget_lines, stacked slices with the command to branch each one (deliver --slices).
  • Scenarios side by side (loop.parallel): disjoint scenarios run in their own sandboxes, come back in order through a seam check, and get your review and their own commit in turn.

🧪 Proven with Claude Code

A real run on a Go payroll project with a broken legacy test: the baseline kept it apart, a medium and a high-risk scenario got one and four lenses, two scenarios ran side by side, and the guard stopped a real git reset --hard. What that run found (a verifier blaming a scenario for another's invariant, a "command" that was a code reading, probes waiting for an approval that never comes, a sandboxed scenario losing its gates) was fixed before release, and a quality review of the branch fixed three more problems in how parallel scenarios come back: see docs/DEMO.md. Coverage 78.9 %, 0 lint warnings, builds for Linux, macOS and Windows.

Full changelog: v5.0.0...v6.0.0