Releases: iammurtaza53/dev-agent-autopilot
Release list
v0.4.2: optional HostLatch trust-handoff gate
An optional trust gate for what agents leave behind.
An agent's sandbox ends when its work lands in your repository. The files it changed there can still run later with your authority: when you install dependencies, open the folder in your IDE, let CI run, or start the next agent session. v0.4.2 can run HostLatch on every task branch to catch those changes before you merge.
What's new
-
trustGate(off by default). With"trustGate": { "enabled": true }, everydev-autopilot checkends with a HostLatch scan of the task branch against the base branch. The scan flags:- package lifecycle scripts;
- IDE tasks;
- agent hooks and settings;
- MCP commands;
- CI workflows;
- Git attributes and dev containers.
-
One line when the scan is clean. Otherwise the output shows the decision, the risk score, the findings (rule, path, title) and the manifest path. Evidence stays in the stored manifest and isn't printed.
-
failOncontrols what fails the check:block(the default) orreview.dev-autopilot check --only trustruns only the scan. -
Fails closed. If the gate is on and HostLatch can't run, the check fails instead of passing silently, and
doctorreports it. -
Full review budget for flagged changes. Anything HostLatch flags is treated as high-risk, so the Codex review gets
reviewer.maxRounds. -
The Claude rule says what to do with a finding:
- remove a change the task doesn't need;
- list a needed one under "Trust handoff (HostLatch)" in the PR;
- treat a
blockas a human gate; - never rewrite or hide a change to pass the scan.
-
Install without cloning. Nothing is published to the npm registry. Install from this release instead:
npx --yes github:iammurtaza53/dev-agent-autopilot#v0.4.2 --help npm install -g github:iammurtaza53/dev-agent-autopilot#v0.4.2
Upgrading
cd dev-agent-autopilot && git pull && npm install # or: npm install -g github:iammurtaza53/dev-agent-autopilot#v0.4.2
cd your-project
dev-autopilot upgrade # refreshes the rule; your config is left exactly as it was
dev-autopilot doctor
git add .claude/rules/dev-autopilot.md && git commit -m "chore: upgrade Dev Agent Autopilot to v0.4.2"No config change is needed; the gate stays off until you switch it on. To use HostLatch without installing it:
"trustGate": { "enabled": true, "command": "npx --yes github:iammurtaza53/hostlatch#v0.2.0", "failOn": "block" }Tests: 215 pass on Linux, Windows and macOS with Node 22 and 24. The gate was also checked by hand against the real HostLatch 0.2.0 CLI. Full notes are in the CHANGELOG.
v0.4.1: LeanLoop, token-efficient orchestration
LeanLoop: send evidence, not history.
v0.4.1 cuts the context and output Autopilot puts in front of Claude Code and Codex, without weakening checks, reviews or safety rules. Nothing in LeanLoop calls a model: it works with hashes, Markdown sections, git diffs and exit codes, and your repository files stay the source of truth. No API keys, no model IDs.
On the benchmark, agent-facing text drops from 281 KB to 64 KB (77% less) against v0.4.0's exact prompt and rule, and Codex calls go from 8 to 6. All 21 required-information checks pass. These are orchestration-layer bytes, not provider-billed tokens.
What's new
- Context Capsule. One file per task instead of "read all the context files". It has the task verbatim, your instruction files and safety sections in full, and the sections the task points to, each as exact text with its path, lines and sha256. Everything else is indexed by line range.
- Delta Resume. A stopped session gets a short "nothing changed" message, or only the sections that changed. It never gets the whole context again and never keeps stale context.
- Quiet Checks.
dev-autopilot checkkeeps full logs on disk and shows PASS lines, or for a failure the exit code, the log path and a bounded excerpt. - Adaptive Codex review. The git diff sets the budget: docs-only changes skip review, small changes get 1 round, others 2. Security, payments, migrations, dependencies, CI, build and agent-instruction changes always get the full
reviewer.maxRounds. - No duplicate planning. Codex never plans when
planner.enabledis off, and never plans the same task twice. - Quota resume tickets (opt-in). When Claude or Codex states when an exhausted usage limit resets, Autopilot can resume the task 2 minutes after that time, after checking that it is still the same task. It never guesses reset times or buys credits.
dev-autopilot efficiencyreports the bytes kept out of agent context, on your machine only.- A tighter Claude rule, with subagent and chat discipline.
Upgrading
cd dev-agent-autopilot && git pull && npm install
cd your-project
dev-autopilot upgrade # refreshes the rule; your config is left exactly as it was
dev-autopilot doctor
git add .claude/rules/dev-autopilot.md && git commit -m "chore: upgrade Dev Agent Autopilot to v0.4.1"- No config changes are needed. LeanLoop is on with its defaults, quota auto-resume stays off, and your
reviewer.maxRoundsstays the cap. New projects default to 2 rounds. - Keep
dev-autopiloton PATH (npm link), because sessions call its LeanLoop helpers. - To keep v0.4.0 behaviour exactly, set
"leanloop": { "enabled": false }.
Full design, settings and limitations: docs/leanloop.md. All changes: CHANGELOG.
v0.4.0: official Codex plugin support and workflow hardening
Codex plans. Claude Code builds. Codex reviews. You merge.
OpenAI now publishes an official Codex plugin for Claude Code. v0.4 supports it as a companion for reviews you start yourself, keeps the native Codex CLI as Autopilot's unattended planner and reviewer, and hardens the workflow.
Official Codex plugin
- It's switched off inside Autopilot's background sessions. This is done through session-only
--settings: your own Claude Code sessions and settings files are untouched. The plugin's optional Stop-time review gate therefore can't add a second, unbounded review loop next toreviewer.maxRounds. Its write-capablecodex:codex-rescuesubagent,/codex:rescueand/codex:setupare always denied there. To load it anyway, set"codexPlugin": { "loadInAutopilotSessions": true }. dev-autopilot doctorreports whether the plugin is installed and enabled.dev-autopilot install-reviewer --install-plugininstalls the plugin only in an interactive terminal, after you confirm.- Why it isn't the automated reviewer: its review commands are user-invoked only (
disable-model-invocation). Use/codex:adversarial-review --base mainyourself for a second opinion before you merge. The full reasoning is in docs/codex-plugin.md.
Workflow hardening
- Codex failures stop the run. A failed Codex plan or review (login, usage or credits, network) makes the session stop and report a blocker. Claude never substitutes its own review.
- Any identifier works.
attach,logs,stopandresumeaccept the shortid, the fullsessionIdor the session name. - Session status and cleanup.
statuslabels sessionsactive,currentorstale. The newcleanupcommand removes finished Autopilot sessions from earlier tasks withclaude rm. It is a dry run by default, keeps transcripts, and never removes worktrees that hold uncommitted changes or unpushed commits. - Doctor and validation.
doctorchecks your Codex login.runanddoctorvalidate the reviewer settings. - Bug fixes: session names with long prefixes; case-sensitive project paths on Linux; unknown flags; duplicated command output.
Upgrading
Upgrading needs no config changes. upgrade leaves v0.3.x configs byte-for-byte unchanged, and the plugin is optional.
cd dev-agent-autopilot && git pull && npm install
cd your-project
dev-autopilot upgrade
dev-autopilot doctorCommit the refreshed .claude/rules/dev-autopilot.md.
Validation
- Tests: 93/93 Dev Agent Autopilot tests and 5/5 demo Todo tests pass.
- CI: passes on Windows, macOS and Linux with Node 22 and 24.
- Independent Codex review: 3 rounds, 7 findings, all fixed with regression tests.
- Real Claude Code background-session validation on Windows. Details are in #2.
See CHANGELOG.md for details.
v0.3.1: Claude worktree handling polish
Codex plans. Claude Code builds. Codex reviews. You merge.
A maintenance release from the first real end-to-end v0.3 demo run.
Fixed
init,upgradeandmigrate-v1now ignore.claude/worktrees/, so Claude background-session worktrees no longer make projects look dirty.- Existing
.gitignorecontent is preserved, and equivalent entries are not duplicated. .claude/rules/dev-autopilot.mdstays tracked. Autopilot does not ignore all of.claude/.- The human merge docs no longer assume a PR branch can be deleted while a Claude worktree still has it checked out.
Existing projects: run dev-autopilot upgrade to add the .gitignore entry.
Validation
- 20/20 Dev Agent Autopilot tests pass.
- 5/5 demo Todo tests pass.
- CI passes on Windows, macOS and Linux with Node 22 and 24.
- Independent Codex review found no blocking issues.
See CHANGELOG.md for details.
v0.3.0: first public release
Codex plans. Claude Code builds. Codex reviews. You merge.
The first public release of Dev Agent Autopilot: a project-agnostic development orchestrator for Claude Code + Codex + GitHub CI, designed to remove the manual copy/paste loop between implementation, review, testing and pull requests.
Write the task in NEXT_TASK.md, run dev-autopilot run, and it will:
- have Codex plan the architecture (read-only);
- have Claude Code implement it and run your real test/lint/build checks;
- have Codex review the branch, with Claude fixing findings;
- open the pull request and fix CI failures;
- stop for you to review and merge. It never auto-merges.
Highlights
- Works with any repository and stack
- Uses your existing
claude,codexandghlogins; no model IDs are hard-coded - Resumes background sessions after you close the terminal or reboot
- Destructive and production actions stay behind human gates
- Runnable demo project in
examples/demo-todo-app
Get started
git clone https://github.com/iammurtaza53/dev-agent-autopilot.git
cd dev-agent-autopilot && npm install && npm link
dev-autopilot --helpSee the README for the five-minute quick start, and CHANGELOG.md for the full list of changes.