Repository navigation
Releases: raditia/CraftKit
Release list
craftkit v1.58.0
Director mode reaps finished and stalled subagents
A finished background agent could sit idle instead of exiting, and an idle agent reads as
work in progress, so the director waited on agents that had nothing left to do.
- Rule 12a gains a Lifecycle step: at every turn boundary (each user prompt and each completion
notification) the director runsListAgentsand acts on every agent it spawned. Result
received:TaskStopit. Past its estimate with no notification:SendMessagefor status or
the final report, thenTaskStoponce it lands or when it is still silent at the next
boundary, reporting its work "not verified". Still progressing: left alone. The end-of-turn
status names every agent still running. - The delegation contract now ends with "your last message is the final report; stop after
sending it and wait for no reply", so agents stop waiting for a reply after they finish. - The routing hook's director line carries the short form, and
check.shcheck 23d fails if
the Lifecycle step is deleted. - The repo-local
CLAUDE.mdRelease section now says which number to bump: left for a major
update or revamp, middle for a minor update, right for a bug fix, with the numbers to the
right reset to 0 and the largest bump winning in a mixed release.
craftkit v1.57.0
Director mode: the main session directs, background agents build
A main session that edits ten files itself burns its context on mechanics and leaves nobody
to check the result, so larger work now goes to background agents with a verify contract.
- Codex users: the PreToolUse hook's matcher changed, so Codex treats it as untrusted until you re-approve it once in
/hooks; until then the Codex verify and delegate gates are skipped silently. Verified live on codex-cli 0.160.0: withCRAFTKIT_DELEGATE=onthe 3rd file was handed to a subagent. - Rule 12a sets a budget: one review pass and one gate run per change, with measurement runs and extra reviews only on request.
using-agent-skillsgains rule 12a (Claude Code): the main session assesses each prompt
first. Answers, lookups, edits touching up to two files and anything done in about 2 minutes
stay direct; long-running work, 3+ files, parallel pieces and orchestrator commands go to
background agents, and the call is stated in one line. The user's "do it here" or "in
background" overrides. The agent runs the verify command in the foreground. Refinements go through SendMessage, pivots through TaskStop and a respawn,
and integration requires the agent's verify result, else it is reported "not verified". It
lives in aCRAFTKIT-DIRECTORblock that Gemini and Cursor strip. Codex gets a delegation
paragraph in itsCRAFTKIT-CODEXblock: quick, coupled work of up to two files stays direct,
steering uses the available v1 or v2 agent tools, and every delegated result is collected
before replying. Passing verification is reused when the final tree is unchanged.- Codex verification runs through a PreToolUse command rewrite that observes the actual
process exit and matching start/finish snapshots. Codex 0.160 sends raw stdout to
PostToolUse without exit metadata, so yielded checks previously triggered repeated
verification requests even after success; stdout claims cannot grant verification credit. - New
gate-delegate.js(PreToolUse, Bash included): asks on the 3rd distinct source file a
main turn edits, and on each further one until the turn hands the work off (a background
spawn, or a profile that can edit). Unattendedclaude -psessions skip it, since an ask
there is auto-denied;CRAFTKIT_DELEGATE=offturns it off for any other automation. - Codex gets the same gate in
craftkit-codex.js, opt-in withCRAFTKIT_DELEGATE=on(PreToolUsenow onBash|apply_patch|spawn_agent): Codex rejectsaskand cannot tellcodex execfrom an interactive run, so it can only deny; it passes once the turn spawns an agent that can edit. gate-verify-on-stop.js: a notification turn measures dirty files from the finished agent's
spawn time, so a background agent's edits are verified; agent worktree dirs are ignored. It
passes when the agent's own transcript shows the verify command succeeding after its last
write, so the main session does not re-run it, and ignores edits outside the repo root.- Cursor now strips
CRAFTKIT-CODEX, which leaked into its rules before. - Sync migrates an installed hook off a matcher an earlier release registered; before, a
changed matcher never reached an existing install. A user-set matcher is left alone. - Measured: rule text alone delegated 0/5 multi-file tasks; with the gate, 2-3/5, every one
after a gate ask. Long-running single-file work and orchestrator commands are not gated and
stay rule-only guidance, a known gap.
craftkit v1.56.0
Execution units: pick the smallest one that fits
Agents had rules for when to spawn (parallel orchestrators, fusion panel) but none for which
unit to use, so coupled work got split across subagents and parallel edits could share a checkout.
using-agent-skillsgains core behavior #12: main task, subagent, custom agent, worktree, or
fork, each with its use. Delegation follows one flow: workstream contract, permission boundary,
read-only subagent or isolated worktree, main-task review, human decision.- The Codex runtime block carries a four-line version, since Codex loads only that section.
craftkit v1.55.0
Context sources: planning checks asks against connected team docs repos
Teams keep PRDs, specs, SQL, meeting notes and decisions in a GitHub repo, but planning never
looked at it, so a spec could quietly contradict a decision the team had already recorded.
- New
/context-sourceskill connects, replaces, disconnects and lists any number of docs repos
per project. Connections live in~/.craftkit/context-sources.json, keyed by project root;
each repo is a craftkit-managed shallow clone under~/.craftkit/context-cache/, shared across
projects and pruned when none uses it (ADR-0003). Nothing is written into the project except
citations in its planning files. /interview,/spec,/planand/defineinjectpartials/context-source.md: refresh once,
search the requirement's terms under one 8-file budget across all sources, cite
<source>@<sha>:<path>:<line>, and pause on a conflict with a fixed A/B/C prompt, including a
new cross-source case when two repos disagree. A helper that cannot run is reported as
context source not consulted, never as silence; an unconnected project is unchanged.scripts/context-source.{sh,js}install to~/.craftkit/bin. URLs carrying a credential are
refused; symlinks and filenames carrying control, line-separator or bidi characters are never
read, and every path is printed as one JSON string, so a filename cannot forge output fields
or lines; a failed listing is reported as an error, never as "nothing found"; a corrupt store is refused rather than read as empty; a read-only sandbox
(Codex default) falls back to the cached SHA ascannot-verify.partials/external-sources.mdscopes its "never write fetched content" rule to Figma and Lark
and learns thekind: gitsource row;partials/planning-resolve.mddocuments its shape.- Verified headless on Claude Code and Codex against fixture repos, with runs and pass bars in
docs/research/context-source-hosts.md: prompt-vs-doc and cross-source conflicts each cited
correctly 5/5 per host, unconnected baseline unchanged, and a real github.com repo without
credentials refused in about a second with no prompt. Cursor and Gemini are unverified.
craftkit v1.54.0
Codex-native agents and working-tree change scope
Codex ran CraftKit's agents by shelling out to codex exec, and every review and context
step compared only committed branch work, so unstaged and untracked edits went unreviewed.
- Named agents now install for Codex as TOML profiles under
${CODEX_HOME:-~/.codex}/agents/:
read-only sandbox, medium reasoning effort, the configured model. Profiles CraftKit does
not own are left alone. hooks/craftkit-codex.jsloads a short Codex runtime guide, taken from aCRAFTKIT-CODEX
block inusing-agent-skills, in place of the Claude routing text.- New
partials/change-scope.mdis injected into every review, ship, build, fix, context and
eval workflow. Change scope now covers staged, unstaged and untracked files as well as the
committed diff, and reportscannot-verifywhen no base resolves. scripts/test-codex.pycovers the Codex hook and adapter.
Claude-side regressions found in review and fixed
- Platform detection treated any
package.jsonas RN/web. Codex narrowed it to roots with a
React dependency, which routed RN monorepos (React only in a workspace package) as plain
Node and droppedfe-rules. Detection now also reads workspace packages from
workspaces,lerna.jsonorpnpm-workspace.yaml. - The Claude and Gemini adapters strip the
CRAFTKIT-CODEXblock, so Codex-only guidance
no longer loads in every session there. - Shared commands no longer hardcode this repo's
bash check.shas the verification command. - The parallel classifier again tells Claude to launch every agent and the background test
run in one message. check.sh23b gains a monorepo fixture, a Node-without-React fixture, and an assertion that
the Claude block carries no Codex section.
craftkit v1.53.0
Update notice on session start
npm never tells an installed package's users that a new version exists, and craftkit has
no CLI they run, so a release reached only the people who went looking for it.
hooks/craftkit-update-check.jsruns on Claude CodeSessionStart. It reads the
installed version from~/.craftkit-state/version(written by every sync), asks the npm
registry for the latest at most once a day, and shows a one-line notice with the update
command when the registry is newer. The answer is cached in
~/.craftkit-state/update-check; a failed fetch is cached too, so an offline machine
pays the 1.5s timeout once a day rather than every session.- Any failure stays silent.
CRAFTKIT_UPDATE_CHECK=offturns it off. check.shcheck 24a runs the hook against a seeded cache: a notice for a newer
version, none for an equal or older one, none when switched off.- Only reaches installs from this version on: anyone on v1.52.0 or older needs one manual
update before they see notices.
craftkit v1.52.0
Opt-in agent dashboard for Claude Code and Codex
A live terminal view of running agents: the main session, a box per running subagent that
appears when it starts and disappears when it finishes (elapsed time, tool calls, current
action), and a session log. Nothing showed which
subagents were running or what each was doing without attaching to them one by one.
CRAFTKIT_DASHBOARD=1 bash sync.shturns it on; the choice persists in
~/.craftkit-state/dashboard, because the post-merge hook syncs without the caller's
environment.CRAFTKIT_DASHBOARD=0turns it off, and that sync removes every piece.
Off by default: the logger writes each tool call's file path or command to disk, so the
log is0600in a0700directory, common credential shapes are masked before writing,
files older than 7 days are pruned by the logger itself, and off deletes the directory.
Values other than 1/0 and on/off words warn and leave the setting alone.hooks/craftkit-agent-log.jslogsSubagentStart,SubagentStop,PostToolUseandSessionEndon
Claude (through_CRAFTKIT_DASHBOARD_HOOKS, so the existing prune pass removes it when
off) and on Codex (its own registration, leaving other hooks alone).hooks/craftkit-statusline.jsbecomes the ClaudestatusLine. An existing one is saved to
~/.craftkit-state/statusline.jsonand wrapped: it runs first on the same stdin (2s cap,
failures ignored) and the dashboard fields are appended, so the numbers reach the dashboard
for the many users who already have a status line. Off restores it exactly, other keys
(padding, refreshInterval) included.scripts/dashboard.pyandscripts/ccdashinstall to~/.craftkit/bin(linked into
~/.local/bin). The dashboard keeps one state per session and reads only appended bytes,
redraws in place without wrapping, closes on Esc, restores the terminal on SIGTERM or
SIGHUP, and prints one frame when it has no terminal, which is what stops! ccdashfrom
stacking a new frame every second, so no instance can run unseen andccdashneeds no
stop command. It opens iTerm when that is the terminal, and a tmux side pane before 3.2.- Running subagents are tracked as one file each under
<session>.agents/: created only by
SubagentStart(Claude Code's internal helpers fire a bareSubagentStop, and a tool hook
can finish after the stop hook), deleted on stop, cleared onSessionEnd. A box idle for
30 minutes is labelled quiet rather than dropped, since one long tool call looks the same;
one whose stop and session end both never fired is hidden after a day. The status line
counts those files instead of re-reading the whole log. - Each subagent box shows its model and tokens (input, cache and output, summed per reply,
with output also on its own), and the tree line totals every subagent in the session.
Both come from Claude Code's subagent transcript, found by session and agent id and read
incrementally; a reply streamed over several lines counts once. Codex shows only the
model its events carry. - Several sessions at once: a numbered strip lists every Claude and Codex session active in
the last 30 minutes and not ended, in start order, with a short session id. Arrows or 1-9
pin one by session id,afollows the newest (holding the current one until it has been
quiet 5s, so two busy sessions do not flip the view), andccdash <n>resolves n to a
session id before opening its window, so a renumbered strip cannot retarget it.SessionEnd
leaves an.endedmarker the strip reads without opening the log; any later event clears
it. The strip takes at most a quarter of the screen. The logger records the project folder
name only, never the path. Sessions not viewed cost a first-line read and a listing. - The orphan staging-dir prune in
sync.shnow skips~/.craftkit/agent-tree; it had been
deleting the dashboard's logs on every sync while the dashboard was on. - Frames are fitted to the window height, dropping the oldest log lines first, because a
frame taller than the window scrolled the header off the top on every redraw. - Config writes go through symlinks and keep the file's mode, so a dotfiles link or a 0600
settings.jsonsurvives; Codexhooks.jsonshapes the sync does not recognise are left
untouched. - A malformed
settings.jsonorhooks.jsonis left alone with a warning instead of
aborting the sync, both are written atomically, and Codex removal filters inside a hook
group so a user's hook sharing it survives. check.shcheck 40 runs the real adapter functions in a throwaway HOME through off, on,
on again (no writes, no+/-lines) and off (no trace, logs included), plus a user's
ownstatusLine, a malformedhooks.json, credential masking, file modes, malformed log
lines, and everyCRAFTKIT_DASHBOARDvalue. Check 24 now reads the dashboard hook table.
craftkit v1.51.0
Codex loads CraftKit rules and checks verification at Stop
Codex previously installed full rules as files but loaded only a short guide into the
session. Skill and workflow routing depended on the agent choosing to read and follow them.
adapters/codex.shregisters a native Codex gateway in~/.codex/hooks.jsonwhile
preserving other hooks. Codex requires a/hooksreview and trust before it runs.SessionStartloads the full applicable rule bodies, including platform-scoped rules only
where their platform matches.UserPromptSubmitsupplies routing guidance and the
installed body of an explicitly requested$skillor leading/command.Stopcompares the working tree with its turn-start snapshot and continues a turn that
edited files but skipped the project's verification command. It preserves the baseline
across its automatic continuation and limits blocks to two. Pre-existing dirty files do
not trigger the gate.check.shexercises rule scoping, explicit command loading, pre-existing edits,
verification continuations, and idempotent hook registration. A live Codex smoke test
confirmed that skipping verification after an edit triggers the Stop continuation.
The runtime does not expose native skill activation as a stable hook event, so following a
skill's instructions remains an agent responsibility; the verification command is the
mechanically enforced part.
craftkit v1.50.0
Cold agents skip CLAUDE.md
Every agent spawn loaded every CLAUDE.md layer: the ~45 KB CRAFTKIT managed block, the
project CLAUDE.md and its imports. None of it was needed, since the rules an agent works
by already arrive through craftkitInject. A bulk-read probe that made no tool calls cost
32.6k tokens, and a third of that was the global block alone.
- All 16
agents/*.mdsetomitClaudeMd: true(Claude Code v2.1.271+). The same probe
now costs 6.6k, andfe-review9.4k. Org managed policy still reaches both agents, which
the go/no-go probe checked alongside the absence of three phrases found only in the managed block. - The field drops the project
CLAUDE.mdtoo, which a second probe confirmed. The
CONTEXT:payload inparallel-review,parallel-shipandparallel-buildnow carries a
PROJECT CONVENTIONS:entry: the rootCLAUDE.md, its@-imports, and the
.claude/rules/*.mdfiles withoutpaths:, ornot presentwhen there are none. Agents
spawned outside those three (plan-roaster,eval-judge,bulk-read) judge plans, scores or
one named file, so they go without it.CLAUDE.local.md, nestedCLAUDE.mdfiles and
path-scoped rules are not carried either, which is a known gap. check.shcheck 39 fails an agent without the field and a template without the entry.- Authoring rule #4 in the repo
CLAUDE.mdand the README agents notes say so. - To roll back, delete the field. Nothing on disk changes shape.
Stop gate stops blaming a turn for files dirty before it
When a turn wrote through the shell or spawned an agent, the verification gate added every
dirty file in the working tree to that turn's edits. So one untracked planning doc, last
touched the day before, blocked every turn that delegated work.
gitDirtycounts only files modified since the prompt that opened the turn (startedAt,
now read from the transcript bycraftkit-transcript.js). It comparesmax(mtime, ctime),
because a chmod moves only ctime. A deleted file, or a turn with no start time, still counts,
so the gate errs on the side of firing.git status --porcelain -zreplaces the line-based parse. A rename now resolves to its new
name instead ofold -> new, and a name with a space is no longer quoted.check.shcheck 23 gains fixtures for a stale delegating turn, a fresh one, agit mv,
a quoted name and a chmod-only change. The last three failed before the fix.
craftkit v1.49.0
/cross-review: Claude and Codex review the same diff, then check each other
craftkit's instructions already ran in every tool, but collaboration did not: /parallel-* and
/team-build spawn through Claude-only runtimes, so a second opinion was always the same
model twice. /cross-review puts two providers on one diff.
scripts/cross-review.sh(installed to~/.craftkit/bin/by a new[bin]sync step) runs
claude -p --restricted(Read/Grep/Glob only) andcodex exec -s read-only --ignore-user-config
in parallel on one verbatim prompt, then one critique round where each marks every one of the
other's findings AGREE, DISPUTE or CANNOT-VERIFY. One round only, because further rounds drift
toward agreement, not evidence.- Fail closed before anything is sent: both CLIs present, both login methods and the project
listed in~/.craftkit/cross-review-allowed-auth, tracked diff under 256 KiB. API-key and
endpoint override variables are cleared for the auth checks and the panelists. A miss stops
the run withcross-review could not run: <reason>; it never falls back to a same-model review. - Untracked files are never sent, only counted, so an unignored
.envcannot leave the machine. - Replies are validated layout-tolerant and content-strict: fences, preambles, wrapped lines and
lowercase severities pass; an unparseable line, or a critique that skips a peer finding, stops
the run with the raw replies kept. The run also stops if the tree changed during the review. - Prompts mark the diff and the peer's findings as untrusted data, keep panelists to files inside
the repository, and forbid invoking skills or agents. Panelists run withCRAFTKIT_PANELIST=1
(the script refuses to start under it, the routing hook stays silent) andCRAFTKIT_GATE=off. - Each run keeps the diff, exact prompts, commit, CLI versions and auth methods in a unique
owner-only directory under~/.craftkit-state/cross-review/, so a disagreement can be reproduced. commands/cross-review.mdis the host's adjudication table: consensus kept, disputes settled
by reading the cited lines, unverified claims capped at[WARNING]. README gains a flow diagram.check.shcheck 38 drives the script against stub CLIs: healthy runs, every fail-closed path,
untracked exclusion, tolerated layouts, and a critique that skips a finding.prune_orphan_stagingskips~/.craftkit/bin, which is not an adapter staging dir.