Skip to content

Releases: raditia/CraftKit

craftkit v1.58.0

Choose a tag to compare

@github-actions github-actions released this 07 Oct 07:05
66c128a

Director mode reaps finished and stalled subagents

A finished background agent could sit idle instead of exiting, and an idle agent reads as
work in progress, so the director waited on agents that had nothing left to do.

  • Rule 12a gains a Lifecycle step: at every turn boundary (each user prompt and each completion
    notification) the director runs ListAgents and acts on every agent it spawned. Result
    received: TaskStop it. Past its estimate with no notification: SendMessage for status or
    the final report, then TaskStop once it lands or when it is still silent at the next
    boundary, reporting its work "not verified". Still progressing: left alone. The end-of-turn
    status names every agent still running.
  • The delegation contract now ends with "your last message is the final report; stop after
    sending it and wait for no reply", so agents stop waiting for a reply after they finish.
  • The routing hook's director line carries the short form, and check.sh check 23d fails if
    the Lifecycle step is deleted.
  • The repo-local CLAUDE.md Release section now says which number to bump: left for a major
    update or revamp, middle for a minor update, right for a bug fix, with the numbers to the
    right reset to 0 and the largest bump winning in a mixed release.

craftkit v1.57.0

Choose a tag to compare

@github-actions github-actions released this 07 Oct 00:16
def8635

Director mode: the main session directs, background agents build

A main session that edits ten files itself burns its context on mechanics and leaves nobody
to check the result, so larger work now goes to background agents with a verify contract.

  • Codex users: the PreToolUse hook's matcher changed, so Codex treats it as untrusted until you re-approve it once in /hooks; until then the Codex verify and delegate gates are skipped silently. Verified live on codex-cli 0.160.0: with CRAFTKIT_DELEGATE=on the 3rd file was handed to a subagent.
  • Rule 12a sets a budget: one review pass and one gate run per change, with measurement runs and extra reviews only on request.
  • using-agent-skills gains rule 12a (Claude Code): the main session assesses each prompt
    first. Answers, lookups, edits touching up to two files and anything done in about 2 minutes
    stay direct; long-running work, 3+ files, parallel pieces and orchestrator commands go to
    background agents, and the call is stated in one line. The user's "do it here" or "in
    background" overrides. The agent runs the verify command in the foreground. Refinements go through SendMessage, pivots through TaskStop and a respawn,
    and integration requires the agent's verify result, else it is reported "not verified". It
    lives in a CRAFTKIT-DIRECTOR block that Gemini and Cursor strip. Codex gets a delegation
    paragraph in its CRAFTKIT-CODEX block: quick, coupled work of up to two files stays direct,
    steering uses the available v1 or v2 agent tools, and every delegated result is collected
    before replying. Passing verification is reused when the final tree is unchanged.
  • Codex verification runs through a PreToolUse command rewrite that observes the actual
    process exit and matching start/finish snapshots. Codex 0.160 sends raw stdout to
    PostToolUse without exit metadata, so yielded checks previously triggered repeated
    verification requests even after success; stdout claims cannot grant verification credit.
  • New gate-delegate.js (PreToolUse, Bash included): asks on the 3rd distinct source file a
    main turn edits, and on each further one until the turn hands the work off (a background
    spawn, or a profile that can edit). Unattended claude -p sessions skip it, since an ask
    there is auto-denied; CRAFTKIT_DELEGATE=off turns it off for any other automation.
  • Codex gets the same gate in craftkit-codex.js, opt-in with CRAFTKIT_DELEGATE=on (PreToolUse now on Bash|apply_patch|spawn_agent): Codex rejects ask and cannot tell codex exec from an interactive run, so it can only deny; it passes once the turn spawns an agent that can edit.
  • gate-verify-on-stop.js: a notification turn measures dirty files from the finished agent's
    spawn time, so a background agent's edits are verified; agent worktree dirs are ignored. It
    passes when the agent's own transcript shows the verify command succeeding after its last
    write, so the main session does not re-run it, and ignores edits outside the repo root.
  • Cursor now strips CRAFTKIT-CODEX, which leaked into its rules before.
  • Sync migrates an installed hook off a matcher an earlier release registered; before, a
    changed matcher never reached an existing install. A user-set matcher is left alone.
  • Measured: rule text alone delegated 0/5 multi-file tasks; with the gate, 2-3/5, every one
    after a gate ask. Long-running single-file work and orchestrator commands are not gated and
    stay rule-only guidance, a known gap.

craftkit v1.56.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 15:11
a3f4a67

Execution units: pick the smallest one that fits

Agents had rules for when to spawn (parallel orchestrators, fusion panel) but none for which
unit to use, so coupled work got split across subagents and parallel edits could share a checkout.

  • using-agent-skills gains core behavior #12: main task, subagent, custom agent, worktree, or
    fork, each with its use. Delegation follows one flow: workstream contract, permission boundary,
    read-only subagent or isolated worktree, main-task review, human decision.
  • The Codex runtime block carries a four-line version, since Codex loads only that section.

craftkit v1.55.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 12:30
3de090e

Context sources: planning checks asks against connected team docs repos

Teams keep PRDs, specs, SQL, meeting notes and decisions in a GitHub repo, but planning never
looked at it, so a spec could quietly contradict a decision the team had already recorded.

  • New /context-source skill connects, replaces, disconnects and lists any number of docs repos
    per project. Connections live in ~/.craftkit/context-sources.json, keyed by project root;
    each repo is a craftkit-managed shallow clone under ~/.craftkit/context-cache/, shared across
    projects and pruned when none uses it (ADR-0003). Nothing is written into the project except
    citations in its planning files.
  • /interview, /spec, /plan and /define inject partials/context-source.md: refresh once,
    search the requirement's terms under one 8-file budget across all sources, cite
    <source>@<sha>:<path>:<line>, and pause on a conflict with a fixed A/B/C prompt, including a
    new cross-source case when two repos disagree. A helper that cannot run is reported as
    context source not consulted, never as silence; an unconnected project is unchanged.
  • scripts/context-source.{sh,js} install to ~/.craftkit/bin. URLs carrying a credential are
    refused; symlinks and filenames carrying control, line-separator or bidi characters are never
    read, and every path is printed as one JSON string, so a filename cannot forge output fields
    or lines; a failed listing is reported as an error, never as "nothing found"; a corrupt store is refused rather than read as empty; a read-only sandbox
    (Codex default) falls back to the cached SHA as cannot-verify.
  • partials/external-sources.md scopes its "never write fetched content" rule to Figma and Lark
    and learns the kind: git source row; partials/planning-resolve.md documents its shape.
  • Verified headless on Claude Code and Codex against fixture repos, with runs and pass bars in
    docs/research/context-source-hosts.md: prompt-vs-doc and cross-source conflicts each cited
    correctly 5/5 per host, unconnected baseline unchanged, and a real github.com repo without
    credentials refused in about a second with no prompt. Cursor and Gemini are unverified.

craftkit v1.54.0

Choose a tag to compare

@github-actions github-actions released this 05 Oct 04:31
7137ba7

Codex-native agents and working-tree change scope

Codex ran CraftKit's agents by shelling out to codex exec, and every review and context
step compared only committed branch work, so unstaged and untracked edits went unreviewed.

  • Named agents now install for Codex as TOML profiles under ${CODEX_HOME:-~/.codex}/agents/:
    read-only sandbox, medium reasoning effort, the configured model. Profiles CraftKit does
    not own are left alone.
  • hooks/craftkit-codex.js loads a short Codex runtime guide, taken from a CRAFTKIT-CODEX
    block in using-agent-skills, in place of the Claude routing text.
  • New partials/change-scope.md is injected into every review, ship, build, fix, context and
    eval workflow. Change scope now covers staged, unstaged and untracked files as well as the
    committed diff, and reports cannot-verify when no base resolves.
  • scripts/test-codex.py covers the Codex hook and adapter.

Claude-side regressions found in review and fixed

  • Platform detection treated any package.json as RN/web. Codex narrowed it to roots with a
    React dependency, which routed RN monorepos (React only in a workspace package) as plain
    Node and dropped fe-rules. Detection now also reads workspace packages from
    workspaces, lerna.json or pnpm-workspace.yaml.
  • The Claude and Gemini adapters strip the CRAFTKIT-CODEX block, so Codex-only guidance
    no longer loads in every session there.
  • Shared commands no longer hardcode this repo's bash check.sh as the verification command.
  • The parallel classifier again tells Claude to launch every agent and the background test
    run in one message.
  • check.sh 23b gains a monorepo fixture, a Node-without-React fixture, and an assertion that
    the Claude block carries no Codex section.

craftkit v1.53.0

Choose a tag to compare

@github-actions github-actions released this 02 Oct 07:04
1e38553

Update notice on session start

npm never tells an installed package's users that a new version exists, and craftkit has
no CLI they run, so a release reached only the people who went looking for it.

  • hooks/craftkit-update-check.js runs on Claude Code SessionStart. It reads the
    installed version from ~/.craftkit-state/version (written by every sync), asks the npm
    registry for the latest at most once a day, and shows a one-line notice with the update
    command when the registry is newer. The answer is cached in
    ~/.craftkit-state/update-check; a failed fetch is cached too, so an offline machine
    pays the 1.5s timeout once a day rather than every session.
  • Any failure stays silent. CRAFTKIT_UPDATE_CHECK=off turns it off.
  • check.sh check 24a runs the hook against a seeded cache: a notice for a newer
    version, none for an equal or older one, none when switched off.
  • Only reaches installs from this version on: anyone on v1.52.0 or older needs one manual
    update before they see notices.

craftkit v1.52.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 15:19
8adee7e

Opt-in agent dashboard for Claude Code and Codex

A live terminal view of running agents: the main session, a box per running subagent that
appears when it starts and disappears when it finishes (elapsed time, tool calls, current
action), and a session log. Nothing showed which
subagents were running or what each was doing without attaching to them one by one.

  • CRAFTKIT_DASHBOARD=1 bash sync.sh turns it on; the choice persists in
    ~/.craftkit-state/dashboard, because the post-merge hook syncs without the caller's
    environment. CRAFTKIT_DASHBOARD=0 turns it off, and that sync removes every piece.
    Off by default: the logger writes each tool call's file path or command to disk, so the
    log is 0600 in a 0700 directory, common credential shapes are masked before writing,
    files older than 7 days are pruned by the logger itself, and off deletes the directory.
    Values other than 1/0 and on/off words warn and leave the setting alone.
  • hooks/craftkit-agent-log.js logs SubagentStart, SubagentStop, PostToolUse and SessionEnd on
    Claude (through _CRAFTKIT_DASHBOARD_HOOKS, so the existing prune pass removes it when
    off) and on Codex (its own registration, leaving other hooks alone).
  • hooks/craftkit-statusline.js becomes the Claude statusLine. An existing one is saved to
    ~/.craftkit-state/statusline.json and wrapped: it runs first on the same stdin (2s cap,
    failures ignored) and the dashboard fields are appended, so the numbers reach the dashboard
    for the many users who already have a status line. Off restores it exactly, other keys
    (padding, refreshInterval) included.
  • scripts/dashboard.py and scripts/ccdash install to ~/.craftkit/bin (linked into
    ~/.local/bin). The dashboard keeps one state per session and reads only appended bytes,
    redraws in place without wrapping, closes on Esc, restores the terminal on SIGTERM or
    SIGHUP, and prints one frame when it has no terminal, which is what stops ! ccdash from
    stacking a new frame every second, so no instance can run unseen and ccdash needs no
    stop command. It opens iTerm when that is the terminal, and a tmux side pane before 3.2.
  • Running subagents are tracked as one file each under <session>.agents/: created only by
    SubagentStart (Claude Code's internal helpers fire a bare SubagentStop, and a tool hook
    can finish after the stop hook), deleted on stop, cleared on SessionEnd. A box idle for
    30 minutes is labelled quiet rather than dropped, since one long tool call looks the same;
    one whose stop and session end both never fired is hidden after a day. The status line
    counts those files instead of re-reading the whole log.
  • Each subagent box shows its model and tokens (input, cache and output, summed per reply,
    with output also on its own), and the tree line totals every subagent in the session.
    Both come from Claude Code's subagent transcript, found by session and agent id and read
    incrementally; a reply streamed over several lines counts once. Codex shows only the
    model its events carry.
  • Several sessions at once: a numbered strip lists every Claude and Codex session active in
    the last 30 minutes and not ended, in start order, with a short session id. Arrows or 1-9
    pin one by session id, a follows the newest (holding the current one until it has been
    quiet 5s, so two busy sessions do not flip the view), and ccdash <n> resolves n to a
    session id before opening its window, so a renumbered strip cannot retarget it. SessionEnd
    leaves an .ended marker the strip reads without opening the log; any later event clears
    it. The strip takes at most a quarter of the screen. The logger records the project folder
    name only, never the path. Sessions not viewed cost a first-line read and a listing.
  • The orphan staging-dir prune in sync.sh now skips ~/.craftkit/agent-tree; it had been
    deleting the dashboard's logs on every sync while the dashboard was on.
  • Frames are fitted to the window height, dropping the oldest log lines first, because a
    frame taller than the window scrolled the header off the top on every redraw.
  • Config writes go through symlinks and keep the file's mode, so a dotfiles link or a 0600
    settings.json survives; Codex hooks.json shapes the sync does not recognise are left
    untouched.
  • A malformed settings.json or hooks.json is left alone with a warning instead of
    aborting the sync, both are written atomically, and Codex removal filters inside a hook
    group so a user's hook sharing it survives.
  • check.sh check 40 runs the real adapter functions in a throwaway HOME through off, on,
    on again (no writes, no +/- lines) and off (no trace, logs included), plus a user's
    own statusLine, a malformed hooks.json, credential masking, file modes, malformed log
    lines, and every CRAFTKIT_DASHBOARD value. Check 24 now reads the dashboard hook table.

craftkit v1.51.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 10:00
8f275a0

Codex loads CraftKit rules and checks verification at Stop

Codex previously installed full rules as files but loaded only a short guide into the
session. Skill and workflow routing depended on the agent choosing to read and follow them.

  • adapters/codex.sh registers a native Codex gateway in ~/.codex/hooks.json while
    preserving other hooks. Codex requires a /hooks review and trust before it runs.
  • SessionStart loads the full applicable rule bodies, including platform-scoped rules only
    where their platform matches. UserPromptSubmit supplies routing guidance and the
    installed body of an explicitly requested $skill or leading /command.
  • Stop compares the working tree with its turn-start snapshot and continues a turn that
    edited files but skipped the project's verification command. It preserves the baseline
    across its automatic continuation and limits blocks to two. Pre-existing dirty files do
    not trigger the gate.
  • check.sh exercises rule scoping, explicit command loading, pre-existing edits,
    verification continuations, and idempotent hook registration. A live Codex smoke test
    confirmed that skipping verification after an edit triggers the Stop continuation.

The runtime does not expose native skill activation as a stable hook event, so following a
skill's instructions remains an agent responsibility; the verification command is the
mechanically enforced part.

craftkit v1.50.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 08:55
3795d26

Cold agents skip CLAUDE.md

Every agent spawn loaded every CLAUDE.md layer: the ~45 KB CRAFTKIT managed block, the
project CLAUDE.md and its imports. None of it was needed, since the rules an agent works
by already arrive through craftkitInject. A bulk-read probe that made no tool calls cost
32.6k tokens, and a third of that was the global block alone.

  • All 16 agents/*.md set omitClaudeMd: true (Claude Code v2.1.271+). The same probe
    now costs 6.6k, and fe-review 9.4k. Org managed policy still reaches both agents, which
    the go/no-go probe checked alongside the absence of three phrases found only in the managed block.
  • The field drops the project CLAUDE.md too, which a second probe confirmed. The
    CONTEXT: payload in parallel-review, parallel-ship and parallel-build now carries a
    PROJECT CONVENTIONS: entry: the root CLAUDE.md, its @-imports, and the
    .claude/rules/*.md files without paths:, or not present when there are none. Agents
    spawned outside those three (plan-roaster, eval-judge, bulk-read) judge plans, scores or
    one named file, so they go without it. CLAUDE.local.md, nested CLAUDE.md files and
    path-scoped rules are not carried either, which is a known gap.
  • check.sh check 39 fails an agent without the field and a template without the entry.
  • Authoring rule #4 in the repo CLAUDE.md and the README agents notes say so.
  • To roll back, delete the field. Nothing on disk changes shape.

Stop gate stops blaming a turn for files dirty before it

When a turn wrote through the shell or spawned an agent, the verification gate added every
dirty file in the working tree to that turn's edits. So one untracked planning doc, last
touched the day before, blocked every turn that delegated work.

  • gitDirty counts only files modified since the prompt that opened the turn (startedAt,
    now read from the transcript by craftkit-transcript.js). It compares max(mtime, ctime),
    because a chmod moves only ctime. A deleted file, or a turn with no start time, still counts,
    so the gate errs on the side of firing.
  • git status --porcelain -z replaces the line-based parse. A rename now resolves to its new
    name instead of old -> new, and a name with a space is no longer quoted.
  • check.sh check 23 gains fixtures for a stale delegating turn, a fresh one, a git mv,
    a quoted name and a chmod-only change. The last three failed before the fix.

craftkit v1.49.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 10:11
fe08a5a

/cross-review: Claude and Codex review the same diff, then check each other

craftkit's instructions already ran in every tool, but collaboration did not: /parallel-* and
/team-build spawn through Claude-only runtimes, so a second opinion was always the same
model twice. /cross-review puts two providers on one diff.

  • scripts/cross-review.sh (installed to ~/.craftkit/bin/ by a new [bin] sync step) runs
    claude -p --restricted (Read/Grep/Glob only) and codex exec -s read-only --ignore-user-config
    in parallel on one verbatim prompt, then one critique round where each marks every one of the
    other's findings AGREE, DISPUTE or CANNOT-VERIFY. One round only, because further rounds drift
    toward agreement, not evidence.
  • Fail closed before anything is sent: both CLIs present, both login methods and the project
    listed in ~/.craftkit/cross-review-allowed-auth, tracked diff under 256 KiB. API-key and
    endpoint override variables are cleared for the auth checks and the panelists. A miss stops
    the run with cross-review could not run: <reason>; it never falls back to a same-model review.
  • Untracked files are never sent, only counted, so an unignored .env cannot leave the machine.
  • Replies are validated layout-tolerant and content-strict: fences, preambles, wrapped lines and
    lowercase severities pass; an unparseable line, or a critique that skips a peer finding, stops
    the run with the raw replies kept. The run also stops if the tree changed during the review.
  • Prompts mark the diff and the peer's findings as untrusted data, keep panelists to files inside
    the repository, and forbid invoking skills or agents. Panelists run with CRAFTKIT_PANELIST=1
    (the script refuses to start under it, the routing hook stays silent) and CRAFTKIT_GATE=off.
  • Each run keeps the diff, exact prompts, commit, CLI versions and auth methods in a unique
    owner-only directory under ~/.craftkit-state/cross-review/, so a disagreement can be reproduced.
  • commands/cross-review.md is the host's adjudication table: consensus kept, disputes settled
    by reading the cited lines, unverified claims capped at [WARNING]. README gains a flow diagram.
  • check.sh check 38 drives the script against stub CLIs: healthy runs, every fail-closed path,
    untracked exclusion, tolerated layouts, and a critique that skips a finding.
  • prune_orphan_staging skips ~/.craftkit/bin, which is not an adapter staging dir.