Repository navigation
Releases: swingerman/engineer
Release list
v1.17.0 — engineer 0.33.0
engineer 0.33.0: role agents
Headline of this series: formal verification with TLA+ and Lean arrived in v1.16.0. TLA+ checks every interleaving, and Lean proves properties for every input. Each counterexample lands in your suite as a failing test.
Now you can:
- Fan out reviewers that can't touch your code. Six role agents ship with the plugin:
reviewer,panel-adviser,panel-advocate,gauntlet-critic,formal-verifierandverifier. Their tool limits and turn caps are enforced by the agent definition, not requested in a prompt. - Run parallel agents without flooding your context. Every subagent reply is capped at ~1,500 tokens. The detail goes to a file, and the reply returns its path.
- Override any role per project. Copy
agents/<name>.mdinto your.claude/agents/and edit it. Plugin agents have the lowest precedence, so your copy wins. - See what a fan-out costs before you agree to it. The parallelism offer now estimates agents × read set. Coding work that shares context stays sequential.
- Run
refinewith nothing extra installed. It dispatches threeengineer:revieweragents directly, with no external skill needed.
Earlier in the series: v1.16.1 (safety fixes).
/plugin marketplace add swingerman/engineer
v1.16.1 — engineer 0.32.1
engineer 0.32.1: safer hardening
Now you can:
- Leave
verify: autoon without it merging code that was never hardened. The merge-readiness bar now covers every checkpoint through CP8. - Let autonomy decide who picks the hardening. At high autonomy the refinement-advisor decides alone and tells you why. Below that, it asks.
- Make the advisor's picks mandatory with one line. Set
harden.required: truein the manifest. It replacesmutation.default_per_feature, which now triggers a validation warning. - Trust the mutation gate's numbers. Scores and thresholds are both percentages (0–100), and the CP7 handoff carries crap-analyzer's results for harden to read.
- See fixes harden on the board too. The fix lifecycle gains a
hardenedstage. - Run
leanandtlapluswithout surprise PRs. They ask before opening one. The TLA+ jar is pinned to v1.7.4.
v1.16.0 — engineer 0.32.0
engineer 0.32.0: formal verification with TLA+ and Lean
Now you can:
- Model-check your code with TLA+.
/engineer.tlaplusmodels retries, locks, async flows and state machines as written, and TLC checks every interleaving, not just the ones your tests happen to hit. - Prove all-inputs invariants with Lean.
/engineer.leanproves that a property holds for every input: a masker never leaks a password, a rounding rule never loses a cent.
How it keeps paying off in your project:
- Counterexamples become failing tests in your suite. Any violation is reproduced against your real code and pinned with a test that fails before the fix and passes after. The model is scaffolding; the test stays.
- It only runs where your code warrants it.
/engineer.refinement-advisorreads each diff and recommends TLA+, Lean, mutation testing, an introversion scan, or nothing, each with a reason. For formal checks it drafts the invariant. You pick with one click. - It's part of the pipeline, not a side quest.
/engineer.hardengives CP8 a skill of its own for features and fixes alike. It runs the selected checks and records a reason for anything skipped.
v1.15.1 — boards see the express lane
Marketplace v1.15.1 · engineer v0.31.1.
Tools that read DAE's lifecycles (Decarchy) now match how 0.31.0 actually runs.
You can now:
- Track express work on the board. A new
expresslifecycle: create the feature, then one/engineer.expresspass, stopping only for you to review the PR. - Stop being asked in the middle. Feature and prototype lifecycles now stop only where the gate profile does: plan approval after CP4 and review at CP7. A bundled feature no longer shows "waiting on you" at CP2 while the skills carry on.
autonomy_level: lowstill stops at every checkpoint, as before.
Install: /plugin marketplace add swingerman/engineer.
v1.15.0 — the express lane
Marketplace v1.15.0 · engineer v0.31.0.
The process now scales to the size of the change. Small work stops paying the big-feature tax.
You can now:
- Ship a small feature in one pass. A new XS express lane runs intake, build and verify in a single context, no checkpoint relay, code first. Still tracked, still gated by tests and the deterministic checks.
- Let the workflow size itself. It projects size and complexity from your idea and recommends the path (prototype / express / full pipeline) as a recommend-and-redirect choice point. You approve or bump one notch, no cold interview.
- Decide once, at the front. Size and path are a single choice up front, not a question in the middle of the build.
- Spend less on small work without dropping the discipline that keeps large features safe.
Prototype-first and full DAE are now two weights on one dial: a converged prototype lands as express or full. Install: /plugin marketplace add swingerman/engineer.
v1.14.1 — Collision-free feature numbers
engineer 0.30.0: feature-init allocates the feature number (NNN) across every git worktree and branch via the new dae_feature_number.py, so parallel unmerged features no longer collide.
v1.15.0-rc.1 — plugins publish their own lifecycles
Pre-release. Cut in preparation for Decarchy going public; the surfaces it replaces come out here rather than shipping twice.
DAE publishes its own pipeline
scripts/dae_lifecycle.py prints DAE's checkpoints as machine-readable lifecycle definitions on stdout, and plugin.json declares where to find it. A board, a dashboard or another agent can now ask what the pipeline is instead of keeping a hand-written copy that goes stale the day a checkpoint is renamed.
Stage order comes from dae_progress.CHECKPOINTS — already this repo's source of truth — so the published answer cannot drift from the one the skills use.
Three are emitted:
feature— spec-first: decide the criteria up front, then build to themprototype— prototype-first: build first, then derive the criteria from what convergedfix— reproduce, fix, and close the gap that let it through
Per references/two-paths.md, prototype-first is not a second process — it is the same checkpoints entered as derive-from-artifact, with iteration 0 in front and CP5 reading prototype_disposition.
Exit criteria are deliberately absent from the definitions: DAE's real criteria live in each project's charter and arrive on the handover at runtime.
Every agentic stage's brief now opens by setting a goal with /goal and an explicit definition of done — a stage whose agent has not said what done looks like has no way to know when to stop.
The dashboard and the control surface are gone
Both were local views of a feature-pipeline board plus an architecture map. Decarchy is that board now and /archify draws the map, so the plugin no longer carries its own.
Removed: dae_dashboard.py, dae_control.py, the control skill, and their tests. If you used /engineer.control, that is the change to know about — the control surface shipped on master after v1.14.0 and is withdrawn before it ever reached a release.
Kept deliberately: dae_arch.py and arch-check (a fitness gate, not a viewer — /archify draws diagrams, it does not gate a merge), dae_metrics.py (DORA), and dae_progress.py.
This also settles a drift that had already happened: dae_dashboard.CP_STAGES and dae_progress.CHECKPOINTS were two tables for one pipeline and disagreed on ids and labels — 1 Discuss vs 0 Onboard, Build vs Implement, Ship vs Harden. The divergent copy lived in the viewer and went with it.
Also in this release
- AI-native-SDLC adoptions: intent, the maintenance loop, and DORA/governance metrics (
f661a4f)
Verification
python3 -m unittest discover -p 'test_dae_*.py' in engineer/scripts — 500 tests, OK.
Plugin version: 0.28.0. The plugin's version line and the repo's release tags have always been independent; this does not change that.
v1.14.0: Less micromanaging your agents
The point of this release: less micromanaging your agents. Decide what to build and whether it's right, then let them do the rest without pulling you into the middle.
Repo renamed
disciplined-agentic-engineeringtoengineer. Old URLs redirect; the marketplace name and@disciplined-agentic-engineeringinstall refs are unchanged.
What you get
-
Stop deciding in the middle. Approve the spec and architecture once up front, verify once at the end. The build runs autonomously between, held to deterministic gates instead of your attention. The amount of ceremony scales to feature size on its own: small changes stay light, large ones get the upfront rigor. Risky or charter-capped work still gets full review, because safety always wins.
-
Ship without babysitting every PR. Opt a feature into auto-merge and it merges itself the moment an objective bar goes green (acceptance, coverage, mutation, architecture, CI), and never a moment before. Fail-closed, with hard stops for anything a machine shouldn't sign off: unproven visual/UX changes, payments and auth, canary rollouts.
-
Idea to working prototype in one step. Rough something out fast with zero ceremony. If it's worth keeping, that prototype becomes the reference your real build is graded against. The prototype is the spec.
Built on Uncle Bob's agent-gauntlet approach: put the human where judgment is required, and give the machine everything else.
v1.13.0 — Gauntlet loop at CP5
engineer 0.22.0 · atdd 0.8.4
DAE's objective graders only cover behavior — acceptance tests answer does it work, CRAP and mutation score answer is it tested. Anything they can't assert (visual fidelity to a ready design, interaction feel, output quality) fell to the human, which is why implementation got babysat even when the specs and designs were ready first.
The gauntlet loop replaces the human grader with a critic agent and a concrete bar: builder builds → a fresh critic A/Bs the output against the reference and names the single largest gap → builder closes it → repeat until ties-or-wins. Rounds run without a pause in between; the human sees them afterwards.
Where it fires
| bar | |
|---|---|
| CP5 exit, post-green | the declared reference (design export, screenshot, reference impl, golden output) |
| CP5 green loop | the failing acceptance tests — the no-human-between-rounds contract |
Not at CP2/CP4 (nothing to A/B; the review panel already covers judgment there), not at CP7/CP8 (already numeric bars with loops around them).
Opt-in by construction
The bar is declared at CP4 in plan.md's Test strategy:
gauntlet:
bar: [design/checkout-desktop.png]
capture: npm run screenshot -- --route /checkout --viewport 1440x900
max_rounds: 5No gauntlet: block → no loop, silently. Not a flag, not a default. A bar must be inspectable — prose and the ACs themselves are explicitly not bars.
Bounded four ways
clear · max_rounds · two identical gaps (no-progress) · a green stream regressing. Non-clear stops hand off with human_action_needed and the open gap named. Critics are plain subagents, never forks.
Changes
- new
engineer/references/gauntlet.md— the full contract planstep 5 — declare the bar when a real reference existsatdd-teamPhase 4 — run rounds without asking; post-green gauntlet; gateparallelism.md— CP5-exit row (loop-until, tier 3)handoff-summary.md—gauntlet_rounds[]schema + required-when ruledae_release.py— could not findatddat all (marketplace source./); now falls back to the marketplace root
490 tests pass. Credit: the Gauntlet Loop, somethingbig.ai.
v1.12.0 — review panel, ontology gate, model classes, host independence
Marketplace: 1.9.0 → 1.12.0 · engineer: 0.16.0 → 0.21.0 · atdd: 0.8.2 → 0.8.3
Covers everything since v1.9.0 — three engineer releases that shipped to master without a tag, plus this one.
Review panel — standing adversarial review at the artifact gates
The adviser + devil's-advocate pattern had been hand-rebuilt four times across four projects, each a fresh ~800-word prompt rediscovering the same lessons. It is now a standing pair, wired into discover-acs (CP2 exit) and plan (CP4 exit), autonomy-keyed.
- adviser (
frontierclass) — constructive: what's missing, underspecified, would bite later. - advocate (
inherit) — adversarial: assume a confident claim is false, find it.
Findings land as panel_findings[] in the handoff — accepted or rejected, both recorded, so the next agent doesn't re-litigate. An unaddressed error-severity finding blocks the checkpoint at autonomy medium/high.
Those manual runs had found an AC set fencing the wrong perimeter, a plan whose central factual claim was false, and an unfalsifiable AC. See engineer/references/review-panel.md.
Ontology gate — constraints at the ledger, not in prose
consistency-check's rule table was English evaluated by an LLM, and ran once in 121 skill invocations. A constraint that only holds when someone remembers to ask isn't a constraint.
The mechanical rows are now scripts/dae_ontology.py, run at every checkpoint exit. Constraint vocabulary borrowed from OWL, because naming a check precisely makes the missing ones obvious:
| Kind | Catches |
|---|---|
enumeration |
status: probably-shipped — invented values |
functional |
two features claiming one branch; ac_count ≠ actual |
inverse |
parent_feature ↔ child_features disagreement |
transitive |
parent chains that cycle |
disjoint |
Principle 7 — the verifier is the implementer |
closure |
an AC with no covering @AC-N scenario |
The judgment rows stay with the skill. No RDF, no triple store, no SPARQL — a set difference suffices at ten entity types.
First run found real defects: a CP8 handoff verified by the same agent that implemented CP5 (a shipped Principle 7 violation), uncovered ACs in two repos, six child_features naming features that don't exist, and a renamed feature with a stale back-reference.
Model classes — no product names to deprecate
Skills now say economy / inherit / frontier, never sonnet/opus/fable. Resolution rule: read the host's live model enum at dispatch time — the only model list that cannot go stale. inherit means omit the field, never name the parent's model.
Measured driver: 98% of input tokens are cache reads, so the lever is which class carries the long-context loop, not prompt length. The top class took 23% of spend for 4% of output — right for a bounded one-shot adviser, wrong for anything that loops.
Host independence
DAE is a methodology, not host features wearing a methodology's name. engineer/references/host-capabilities.md holds the seam: 3 required capabilities (dispatch, filesystem, script execution), 8 optional ones each with a specified degradation. Host tool names in skill prose are bindings, not the contract; porting is editing that table's binding column.
Adds an optional render capability — rendered boards for the roadmap view, next survey, panel findings, and ontology report, always alongside and never instead of the terminal text.
"Where are we" — roadmap breadcrumb
The breadcrumb answered which checkpoint but never is this done, what's next, where does this sit on the roadmap — and only rendered on entry, never at the boundary where that question actually gets asked. Now carries a ROADMAP line from feature.md frontmatter, and renders on checkpoint exit too.
Also in this release
session-summaryauto-invokes frompost-merge, plus an optionalPreCompacthook example. Deliberately not per-checkpoint — that would spamsession-log.md.- Release tooling fix:
dae_release.pynever syncedmarketplace.json, soengineeradvertised 0.18.0 while 0.21.0 was live. Now synced and verified. - Earlier, untagged: roadmap support (0.19.0), the parallelism reflex + session-mining fixes (0.20.0), test-introversion scanning (0.18.0).
- README reworked — host-neutral, 4548 → 3299 words, table of contents, stale facts corrected.
489 tests pass.