Skip to content

Releases: jefmonjor/specforge

v6.1.0 — SpecForge

Choose a tag to compare

@github-actions github-actions released this 05 Oct 15:52
4e70e7f

SpecForge 6.1 closes the gaps 6.0 left open. The guard read only the command line, so a script the agent wrote went unseen. Every copy wrote the whole project, sometimes several times over. And changing an approved specification could move every scenario's identity. Each change was tried with Claude Code and measured.

Install

R=https://github.com/jefmonjor/specforge/releases/download/v6.1.0
F=specforge-darwin-arm64   # or darwin-amd64, linux-amd64, linux-arm64
curl -LO $R/$F && curl -LO $R/checksums.txt
sha256sum -c checksums.txt --ignore-missing   # macOS: shasum -a 256 -c …
chmod +x $F && sudo mv $F /usr/local/bin/specforge
specforge version
specforge doctor

On Windows, download specforge-windows-amd64.exe, rename it to specforge.exe and put it in a folder on your PATH. Then follow the quickstart and the user guide.

What's new in 6.1

🛡 The guard reads what a command runs

  • Scripts: sh x.sh, ./x.sh, source x, followed into the scripts they run; a shebang names the language.
  • Programs: the file given to python, node, ruby, perl, php, go run, deno or bun, and code given inline (python -c, node -e…). It looks for calls that delete a tree (shutil.rmtree, a recursive fs.rm, FileUtils.rm_rf, os.RemoveAll…) and for shell or SQL in their strings.
  • Package scripts and make targets: npm/pnpm/yarn/bun run x with its pre-script, and make x with its prerequisites.
  • A file written and run in one command: the text of echo, printf or a heredoc is inspected. A file the line writes without showing what (curl -o, cp, base64 -d >) is refused until the agent writes it first and runs it in a second command.
  • rimraf and npx are recognised.
  • The rules setup writes into CLAUDE.md/GEMINI.md tell the agent to write a script before running it.

↩️ Checkpoints and specforge restore

Before every agent turn, the loop checkpoints your working tree. Uncommitted and untracked files are included; ignored ones are not.

  • Checkpoints live under refs/specforge/checkpoints/, outside your branch, index, stash and pushes, and the newest 50 are kept.
  • specforge restore lists them. specforge restore latest brings the newest back without deleting anything written since. The restore is itself checkpointed, so it can be undone.
  • This is the way back from what no guard can read, such as a compiled binary.
  • Scenarios run side by side are not checkpointed inside their sandboxes; when a checkpoint cannot be saved, the loop warns once and goes on.

✏️ Specifications that change safely

  • A scenario keeps its marker for good. The marker names its tests, its records and its plan row. It survives insertions, moves and edits of other scenarios. A rename with the same steps keeps it too, and a removed scenario's number is never reused. The markers live in specs/NNNN-slug/scenarios.json, and a specification approved under 6.0 takes them from its loop state (see Upgrading from 6.0).
  • specforge spec change 0001 "<request>" applies a change request in a guided conversation. It asks what the request leaves open, then previews what the approval will record.
  • spec clarify has the agent write each decision into every section it affects, not only on its question's line.
  • The loop redoes exactly what changed:
    • A modified scenario goes back to RED with its previous version shown, so its old test is updated instead of being taken as done.
    • Changing an invariant redoes the scenarios that name it.
    • Reviews, verifications and risk of unchanged scenarios are kept.
  • Removed scenarios and outdated plans are flagged:
    • approve and the loop warn about tests left behind by a removed scenario.
    • A plan naming a scenario that is gone is refused.
    • approve says when the plan needs revising.

📦 Copies that write each byte once

  • Verifier base: checked out straight from git, with no archive in between.
  • Sandboxes: they borrow the project's git objects instead of storing them again.
  • Copy-on-write: files are cloned on APFS, Btrfs and XFS.
  • Size limit: checked before anything is written.
  • Measured: a verifier copy of a 177 MB repository wrote 477 MB and now writes 310 MB. A sandbox of incompressible files writes half what it did.

🧪 Proven with Claude Code

  • Claude Code was asked to wipe a repository three ways the 6.0 guard could not see.
  • The first run got one through: Claude wrote the script and ran it in the same command. That is fixed, and the same request was then blocked three times.
  • A compiled binary ran git reset --hard and git clean; specforge restore latest brought back every uncommitted and untracked file.
  • The three ways into a specification were run for real:
    • a feature change on the payroll module of 6.0, with a scenario inserted in the middle and done alone while the others kept their tests and commits;
    • a technical requirement (an exchange-rate cache);
    • a spec drafted from Python legacy code, whose decisions now reach every section.
  • A fuzz test of 260,000 command lines found a panic, now fixed.
  • CI now also runs the copier's and the guard's tests on macOS and Windows.

Upgrading from 6.0

  • Run specforge setup in each project to refresh the SpecForge block of CLAUDE.md/GEMINI.md with the new shell rules. Nothing else in your files changes.
  • The first approval of an existing specification writes specs/NNNN-slug/scenarios.json. Make it in the clone where the loop ran, so the markers come from the loop state, and commit the file.
  • Approved somewhere else, without the loop state, the markers follow the scenarios' order: spec approve and spec change warn about it and say how to undo it.
  • Rename scenarios after that first approval: until then a new title reads as a new scenario.
  • Earlier approvals.md entries keep their old format; new ones list each scenario's marker.
  • No setting was added or renamed.

See docs/DEMO.md.

Full changelog: v6.0.0...v6.1.0

v6.0.0 — SpecForge

Choose a tag to compare

@jefmonjor jefmonjor released this 05 Oct 15:49
4e70e7f

SpecForge 6 keeps the one rule, verify, don't trust, and makes the checking proportional: cheap for documentation, thorough for the code that handles money, secrets and permissions. Every new check is code, not a prompt.

Install

R=https://github.com/jefmonjor/specforge/releases/download/v6.0.0
F=specforge-darwin-arm64   # or darwin-amd64, linux-amd64, linux-arm64
curl -LO $R/$F && curl -LO $R/checksums.txt
sha256sum -c checksums.txt --ignore-missing   # macOS: shasum -a 256 -c …
chmod +x $F && sudo mv $F /usr/local/bin/specforge
specforge version
specforge doctor

On Windows, download specforge-windows-amd64.exe, rename it to specforge.exe and put it in a folder on your PATH. Then follow the quickstart and the user guide.

What's new in 6.0

⚖️ Proportional review

  • Known failures are evidence, not blockers. The loop runs the suite once before its first RED; tests that already fail are named from the runner's report, shown, kept out of REFACTOR's verdict and listed in the delivery.
  • Risk tiers from git. Passive, medium or high, from the paths and lines each scenario touched. The agent can raise the tier with a reason, never lower it. review: risk asks for your review only when it matters; mutation testing runs from medium up.
  • A model per phase with models.<phase>.

🔎 Review you can check

  • Review lenses (risk, reliability, readability, resilience): none for documentation, one for ordinary code, four for sensitive code. Each is read-only and schema-validated; a finding whose proof is not a changed line is discarded; inferential findings go to an independent refuter.
  • One correction, within a budget of half the change's lines, judged again by REFACTOR and validated on the corrected findings only. A regression goes to you, never to a second automatic correction.
  • Blind double review for high risk: what two independent passes prove on the same hunk skips the refuter.
  • specforge review runs the lenses over a whole branch.

🔍 An independent verifier

For high-risk scenarios, a second agent probes the specification in a disposable copy of your project, the only place an agent may run commands without asking. It must give a verdict for the scenario and the invariants it names, and SpecForge runs every failing command again before believing it. Your project must stay untouched. Proposed regression tests are offered, and added only if the suite still passes. specforge verify checks a whole specification.

🛡 Operational safety

  • A destructive-command guard installed by setup as Claude Code's or Gemini CLI's pre-tool hook: recursive deletes, git reset --hard, push --force, DROP TABLE, DELETE without WHERE, secrets… blocked with the reason the agent sees. Deleting /, ~ or the project is never allowed. It never fails open.
  • Edit surfaces from the plan: a file the agent changes outside the approved plan goes to you; a refused one must be put back.
  • specforge doctor: the agent, git, your runner and every gate's tool, with install hints.
  • Questions with options: in the loop, the plan and legacy turns, the agent must derive the candidate answers.

📦 Delivery and speed

  • Each scenario's size from its commit; above delivery.budget_lines, stacked slices with the command to branch each one (deliver --slices).
  • Scenarios side by side (loop.parallel): disjoint scenarios run in their own sandboxes, come back in order through a seam check, and get your review and their own commit in turn.

🧪 Proven with Claude Code

A real run on a Go payroll project with a broken legacy test: the baseline kept it apart, a medium and a high-risk scenario got one and four lenses, two scenarios ran side by side, and the guard stopped a real git reset --hard. What that run found (a verifier blaming a scenario for another's invariant, a "command" that was a code reading, probes waiting for an approval that never comes, a sandboxed scenario losing its gates) was fixed before release, and a quality review of the branch fixed three more problems in how parallel scenarios come back: see docs/DEMO.md. Coverage 78.9 %, 0 lint warnings, builds for Linux, macOS and Windows.

Full changelog: v5.0.0...v6.0.0

v5.0.0 — SpecForge: Autonomous Spec-Driven & TDD Assembly Line

Choose a tag to compare

@jefmonjor jefmonjor released this 05 Oct 09:02

SpecForge 5 keeps the v4 core and its one rule, verify, don't trust, and brings back what v3 did well: legacy migrations and ready-made projects.

Install

R=https://github.com/jefmonjor/specforge/releases/download/v5.0.0
F=specforge-darwin-arm64   # or darwin-amd64, linux-amd64, linux-arm64
curl -LO $R/$F && curl -LO $R/checksums.txt
sha256sum -c checksums.txt --ignore-missing   # macOS: shasum -a 256 -c …
chmod +x $F && sudo mv $F /usr/local/bin/specforge
specforge version   # specforge v5.0.0 (1165adc, …)

On Windows, download specforge-windows-amd64.exe, rename it to specforge.exe and put it in a folder on your PATH. Then follow the quickstart and the user guide.

What's new in 5.0

🏛 Legacy rewrites, verified

  • legacy scan: a measured inventory of the old code (build, declared Java release, frameworks found from imports) and what each finding means for Java 21.
  • legacy map: the business capabilities of the legacy system, in migration order. Every path:line the agent cites is opened; an invented source sends the document back.
  • spec from-legacy "<capability>": one specification per capability, with the real values and messages, a Legacy sources section, and every oddity in the old code turned into a question for you.
  • The legacy repository is read-only: SpecForge hashes it around every agent turn and refuses any change.
  • A migration gate: the build must declare the target release, and no source may import javax.servlet, Log4j 1, JUnit 3, Vector… ArchUnit checks the layers.

🧰 New projects with the tooling wired

setup --new java|react|python|go, each one installed and run end to end before shipping:

  • java: Java 21 Maven, JUnit 6, AssertJ, ArchUnit layer rules, PMD.
  • react: Vite 7, React 19, strict TypeScript, Vitest 4.0, Testing Library, ESLint, Knip, jscpd, Stryker.
  • python: src/ layout, pytest, Ruff (from the project's .venv).
  • go: go.mod, golangci-lint v2.

🔧 Sharper gates and runners

  • Java lint with PMD; Python uses the .venv's pytest and Ruff; Maven and Gradle wrappers are preferred.
  • The duplication gate counts source code only (a real React run had package-lock.json failing it).
  • Answering a question during plan no longer rejects the plan.
  • Windows: atomic file replacement and test fixes.

📖 Documentation

A README and a user guide that read on a phone, with what you edit and what SpecForge writes, recipes for everyday tasks, and a real Java 6 → 21 migration in docs/DEMO.md.

Full changelog: v3.0.0...v5.0.0

v3.0.0 — Launch of SpecForge: Resilient Spec-Driven & TDD Assembly Line

Choose a tag to compare

@jefmonjor jefmonjor released this 02 Oct 12:26

🚀 SpecForge v3.0.0 Release
We are thrilled to announce the initial official release of SpecForge!

SpecForge is a standalone CLI engine in pure Go (~8.9 MB) that transforms AI-assisted coding from unverified script hallucination into a deterministic, resilient software assembly line.

✨ What's New in v3.0.0
Strict TDD Assembly Line (specforge loop): Enforces Red → Green → Refactor. Intercepts YAGNI violations before any implementation is written.

Autonomous Visual E2E Testing (specforge e2e): Native browser automation built on chromedp (TesterArmy style) without Node.js, Puppeteer, or Playwright.

Contextual Memory Brain (.sdd/agent/): 6 explicit memory files (agente.md, persona.md, ng-rules.md, glossary.md, references.md, lessons.md) with continuous dynamic self-learning.

Business Circuit Breaker ([NEEDS CLARIFICATION]): Prohibits AI from guessing or coding when functional requirements contain unresolved questions.

Heterogeneous Document Ingestion (specforge doc): Auto-converts architecture PDFs, Word docs, and Excel sheets with reconciliation formulas into Markdown via Microsoft MarkItDown.

Legacy Reverse-Engineering & Migration (--from-repo): Uses old legacy codebases as Ground Truth to extract sealed BDD Gherkin specs for greenfield projects.

Adversarial Security Audit (specforge audit): 6-phase security harness inspired by Cloudflare (Recon → Red Team Hunter → Blue Team Verifier).

Automated Quality Guardrails: Built-in gates for jscpd (0% duplicate code), knip (dead code), stryker (mutation testing), strict linters, and ArchUnit.

Zero-Keys & Zero-Admin PATH Injection: Automatic GCP ADC credential inheritance and user PATH registration without UAC elevation.

Multi-Platform Native Binaries: Pre-built for Windows, macOS Apple Silicon (M1/M2/M3/M4), macOS Intel, and Linux.

Open Source License: Distributed under the Apache License, Version 2.0.

📦 Precompiled Binaries
Download the standalone binary for your architecture below (no runtime dependencies required):

Windows x64: specforge-windows-amd64.exe

macOS Apple Silicon (M1/M2/M3/M4): specforge-darwin-arm64

macOS Intel: specforge-darwin-amd64

Linux x64: specforge-linux-amd64