Skip to content

v0.1.2

Choose a tag to compare

@aviggiano aviggiano released this 30 Sep 07:00
6ac8c2b

This is a stability release: it fixes the orchestration and bookkeeping failures that stopped campaigns short of their report, makes recovery commands finish what they start, and removes a large amount of dead code, with stricter CI gates to keep it out. On the release commit, a fresh smoke campaign launched in about four minutes and finished succeeded with a verified, complete report in about 11 minutes.

  • Campaigns run to completion. The default, low-cost, exhaustive and invariant-only profiles launch again, the supervisor relaunches an engine that dies mid-run, and run synchronization, attempt and usage ledgers, and event journals no longer stop or strand a run over bookkeeping problems.
  • Recovery that works. resume --retry-failed reopens skipped work and judges a rerun on its own output, and a report whose agent reworded a finding verifies with a warning instead of being discarded. A tiny-vault smoke campaign that had failed twice with Report: unavailable recovered with resume --retry-failed to a verified, complete report.
  • Less code, stricter CI. The codebase ends about 14,000 lines smaller, mostly from deleting dead gates, pipelines and recovery code. Every package is now validated on pull requests, a complexity ceiling and an unused-export check block regressions, and a model-free end-to-end campaign kills and resumes a real engine on every pull request.

Breaking changes

  • [runtime] [artifacts] Finish or cancel every run launched on v0.1.1 before upgrading. Before each task, a v0.1.1 run's generated workflow checks that the validator build matches the one the run was planned with, and this release changes that build. A plain ultrafuzz resume of such a run therefore fails each remaining task with artifact-contract failure: planned schema binding changed; pause and cancel still work. Runs launched on this release validate each artifact against the schema it was planned with and no longer compare the validator build, so a later rebuild that changes only that build, such as an ajv bump, no longer strands them. report --require-verified now reports an invalid artifact as JSON_SCHEMA_VIOLATION instead of ARTIFACT_SCHEMA_INVALID. Thanks @aviggiano! (#1188)
  • [runtime] After upgrading, run ultrafuzz init in each existing project (no --force needed): it now replaces any stock .smithers/agents adapter that differs from its packaged copy, the only one planning accepts. Provider route IDs change for routes derived from a Codex config.toml, Claude settings.json or Kimi config.toml, for IDs that included a proxy or an unused Claude cloud variable, and for a Codex config that sets only a non-empty openai_base_url, which cloud planning now rejects. Update ULTRAFUZZ_DATA_GOVERNANCE_POLICY, re-acknowledge, and re-plan any run whose acknowledged ID changed. In return, a CLI rewriting its own config, a proxy change or an edit to an unused openai_base_url no longer fails every later task, DeepSeek tasks no longer hang, and Kimi and Pi no longer crash on unexpected output or usage data. Thanks @aviggiano! (#1173, #1210)
  • [config] [prompts] Review nodes in the default, exhaustive and invariant-only topologies and in new scaffolds now get 7,200 seconds, even where run.default_timeout_seconds or a model profile sets more. To keep a larger budget, ultrafuzz topology copy the topology, raise its review group's timeout_seconds, and point topology_path at the copy. A cloud config whose execution.resources.timeout_seconds is below 7,200 needs per-node overrides for the review nodes too. An existing .ultrafuzz/topology.yml keeps its old review budget, one hour by default, until you add defaults: { timeout_seconds: 7200 } to its review group. The stateful-invariant and Vyper setup prompts stop inviting production-source edits the workspace handoff rejects. Existing projects keep old copies of these prompts and of review/final-report.md, whose ## Audit context instruction fails report verification if followed: ultrafuzz validate now flags each as PROMPT_DIFFERS_FROM_BUILT_IN; delete the ones you have not customized and rerun ultrafuzz init. Thanks @aviggiano! (#1164)
  • [runtime] [prompts] Runs launched on v0.1.1 whose dynamic groups have fanned out, such as runs past goal-plan, also cannot be synchronized after upgrading. A prompt that waits on a dynamic group is re-rendered and compared byte for byte with its published copy, and this release changes the coverage section that the final-report prompt embeds. status and inspect stop synchronizing such a run, and resume, replay, fork, why and stats fail, most with runtime rendered prompt changed. The same change stops host-side checks rejecting output that a node's own verification accepted: a failed optional test-generation strategy, dynamic goal nodes, the triage, severity, test-aggregation and final-report nodes of any run whose goal plan named a goal, false-positive findings without severity fields, and ordinary coverage sentences in reports, which are now warnings. Thanks @aviggiano! (#1176)
  • [runtime] Final reports published before this release no longer verify: readers re-render them, and the restated run summary no longer matches. For such a run, ultrafuzz report and the dashboard show a complete run as PARTIAL with verification not-checked, report bundle falls back to a report-only archive, and --require-verified reads and the Modal public worker fail. Build any verified bundles you need before upgrading. In return, reports give elapsed time, models, tokens and spend for the whole run instead of as of the report agent's start. The renderer also no longer throws on link syntax, non-ASCII issue titles or label-like prose, which had left such runs with no report. Thanks @aviggiano! (#1167)
  • [config] [cli] validate and run reject [models] default = "<id>" naming any profile other than default (CONFIG_MODEL_DEFAULT_UNSUPPORTED), which produced an unparseable config.resolved.toml. Select the primary profile with [retry] agents = ["<id>"] instead. Commands on existing runs keep working. Prompt-authority selectors now sort the same way on every host instead of by host locale. That fixes task-preparation failures when producer attempt IDs share a prefix (node-1/node-10), and gives a custom multi-path selector that the old order sorted differently a new ID. It also accepts a symlinked --project, artifact path segments starting with .., a workflow_deadline_seconds of up to seven days, and clean, materialize and dashboard audit journals across a backward clock step. Thanks @aviggiano! (#1195)
  • [runtime] [cli] stats JSON gains a canceled node status and a canceled key in totals.status_counts without a schema version change, so consumers that validate stats against the previous closed schema must accept them. A node whose failed tasks were all cancelled, for example through ultrafuzz cancel, now reports canceled instead of failed; a node with any other failed task stays failed. An agent attempt interrupted by a controller crash is now recorded in attempts.jsonl as canceled once the resumed run restarts its task, so stats and status count the same attempts. On the first sync after upgrading, a finished run with such an attempt raises that node's retry_count and republishes its report once, still verified. Thanks @aviggiano! (#1200, #1213)
  • [evals] [cli] Removes ultrafuzz eval history --max-age-days, the six per-metric history SVGs the README never embedded, and json validate's support for the telemetry-cursor, publication-state and automatic history-publication schemas. Remove --max-age-days from any script that passes it; the CLI now rejects the flag. eval run stops running its unused telemetry forwarding, which aborted the whole suite before run-summary.json was written when a run's journal held a record it did not expect, and no longer writes telemetry/. Thanks @aviggiano! (#1174)
  • [runtime] Deletes about 2,350 lines of controller-refresh and recovery code that has had no caller since v0.0.23; resume --refresh-controller is unchanged. status, pause, cancel, replay and fork no longer replay the whole event journal, and run status now follows the workflow engine's result. A run that a pre-v0.0.23 build refreshed now fails those commands, and a run that such a build recovered can no longer re-read its final report. Both have been unsupported since v0.0.25, so start a new run. Thanks @aviggiano! (#1193)

Improvements

  • [runtime] [cli] Halves the per-file disk flushes a launch spends copying the run's execution snapshot, and a very large run can now read back the launch record it wrote. ultrafuzz doctor requires only the CLIs of agents the selected topology or retry chain can use, and warns when the temporary directory is a tmpfs or has under 2 GiB free, sizing its leftover controller directories for at most about a second. resume --refresh-controller no longer leaves untracked .smithers/continuations files that made the next private or cloud launch reject the target as dirty. Thanks @aviggiano! (#1179, #1208)
  • [runtime] Names the runner's run-level error, such as WORKFLOW_RENDER_FAILED: <what the workflow threw>, in the WORKFLOW_TERMINAL_WITHOUT_FAILED_NODE diagnostic of a run that ends failed with no failed node, redacted and capped at 1,000 characters, and adds a pinned-runner test that recovers such a run through a same-ID resume once its cause is removed. Thanks @aviggiano! (#1172)
  • [prompts] [config] Deletes code no production path calls: the prompt rename feature, the always-empty config prompt-metadata layer, redaction-restore helpers that docs/config.md described but nothing ran, unused Modal exports and two orphaned CI scripts. Output-contract templates now resolve packaged-first like the other prompt assets, and the dashboard's /api/flow drops seven always-false capabilities commands. Thanks @aviggiano! (#1175)
  • [cli] Records each file ultrafuzz report bundle had to skip in bundle-manifest.json as omitted_files, with its path, reason and size when known, so a recipient holding only the ZIP can tell which evidence is missing and why. Thanks @mrthankyou! (#1161)
  • [artifacts] Names the path, the offending segment and the reason, such as a segment over the 128-character limit or a disallowed character, when an unsafe workspace patch path is rejected, instead of a bare unsafe path segment message. Thanks @mrthankyou! (#1159)
  • [workflows] [config] Runs every release-validation lane on pull requests without waiting for the build and static-analysis jobs, and stops pushes to main cancelling each other. pnpm -w lint enforces a global complexity ceiling with no suppressions, and Vitest packages default to a 30 s test timeout. Config loading accepts the null-prototype tables smol-toml 1.9 returns, which had broken ultrafuzz init in packed installs. Thanks @aviggiano! (#1184)
  • [workflows] Makes pnpm -w knip fail on unused exports, unused exported types and duplicate exports, after the remaining module-private symbols are un-exported and duplicate aliases deleted. A contributor change that orphans an export must now delete or un-export it. Thanks @aviggiano! (#1219)
  • [tests] [workflows] Adds a required pull-request lane, cli-e2e, that runs a whole campaign without a model: separate ultrafuzz processes drive the pinned Smithers engine, and a stub codex writes every declared artifact. The test SIGKILLs the engine and supervisor mid-node, resumes, and requires a verified report with no finished task started again. Thanks @aviggiano! (#1187)
  • [tests] Deletes tests that could not catch a regression while keeping behavioural coverage. Prompt tests require the template variables that gates and the runtime depend on, now including the goal hunter's {{item.node_id}}, instead of pinning wording; source-text tests are gone; and lifecycle-inspection shares one launched run across its read-only tests (38 launches down to 5). Thanks @aviggiano! (#1196, #1203, #1212)
  • [tests] Gives dynamic-lifecycle test fixtures a one-hour controller lease so slow CI setup can no longer expire it, fixing an intermittent controller-loss failure in the runtime-supporting lane. Thanks @aviggiano! (#1214)
  • [docs] Documents that [run].workflow_deadline_seconds is checked only when a command synchronizes the run, so an unattended run keeps executing and spending past its deadline, and how to bound one with a periodic ultrafuzz status or ultrafuzz cancel. Thanks @mrthankyou! (#1157)
  • [docs] Explains the HTTP 403 Request blocked: prompt injection patterns detected that OpenRouter-backed agents get when a guardrail covering their key blocks prompt injection: set detection to Flag, or turn it off, on every guardrail that covers the key, and do not use Redact. Thanks @aviggiano! (#1182)
  • [runtime] Corrects the Claude adapter's subscription-auth comment: configDir is always forwarded, from the configured home, else $CLAUDE_CONFIG_DIR, else ~/.claude. No behavior change. Thanks @mrthankyou! (#1158)
  • [docs] [tests] Records this release's changes and upgrade steps in CHANGELOG.md, and reconciles the tests and lint that the back-to-back stability merges had invalidated for one another. Thanks @aviggiano! (#1191, #1222)

Bug fixes

  • [runtime] [artifacts] Fixes launches of the default, low-cost, exhaustive and invariant-only profiles, which in v0.1.1 failed at submission with optional dependency artifact directories do not match before any model ran. A launch that fails after creating its run directory but before sealing its controls is now recorded as failed and reported as RUN_LAUNCH_FAILED by status, resume and the other lifecycle commands. Thanks @mrthankyou and @aviggiano! (#1160, #1185)
  • [runtime] Stops a pnpm install anywhere on the host from failing a concurrent launch with workflow execution file … changed while reading: launch-time and run-document reads no longer treat a changed link count or ctime, which pnpm's hard-linked store produces, as a changed file. Thanks @aviggiano! (#1206)
  • [runtime] Stops a literal {{word}} in a task or operator prompt from failing every workflow render and every resume; goal-plan's own verification now rejects {{ in replacement values instead. The engine runs the validator CLI self-test only until it first succeeds in each process, and the ultrafuzz CLI, which every agent validator call starts, no longer loads the TypeScript compiler. Thanks @aviggiano! (#1168)
  • [runtime] Keeps a dynamic group's node list fixed once it has fanned out, such as the goal nodes that goal-plan plans, so a planning node that runs again after a reset keeps the original fan-out instead of failing every render, cancel, pause, why, fork and replay; resume --retry-failed still re-plans when the planning node's verifier failed. A lock left by a killed process no longer blocks every later render. Thanks @aviggiano! (#1163)
  • [runtime] Lets the Smithers supervisor relaunch a workflow engine that dies mid-run, so the run continues; before, a fresh run stayed orphaned until an operator resumed it, and a resumed run was marked failed after three relaunch attempts. Controller Bun processes no longer load a target repository's bunfig.toml, its preload scripts or .env. Thanks @aviggiano! (#1199)
  • [runtime] Gives agent retries a real wait, 60 s doubling to 300 s instead of 1 s and 2 s, and identical failures no longer end the planned chain before its later attempts and [retry].agents fallbacks. A task whose upstream outputs fail its input checks is no longer retried, since a retry reads the same files. Only typed deadline codes and heartbeat timeouts now mark a node timed out, so a failure whose text merely mentions a timeout is no longer mislabelled, and a Modal cloud-node deadline reports failed. Thanks @aviggiano! (#1171)
  • [runtime] Stops the workflow engine appending a TaskHeartbeat event after every heartbeat write, which flooded the event log and events --type node, and running an untimed git fetch origin and a failing git rebase when it creates or re-enters task worktrees. Thanks @aviggiano! (#1166)
  • [runtime] [artifacts] Stops the host re-checking invariant-campaign output with stricter copies of checks the node's own verification already ran, so Recon commands with cd/tee/echo framing, campaign results without the optional sequence_length, and nested campaign result paths no longer fail after the full fuzzing budget has been spent. A sequence length other than 100 now fails inside the attempt, where the agent can retry. Thanks @aviggiano! (#1178)
  • [security] Stops the publication secret gate rejecting Anvil's public test test … junk mnemonic, the 10 keys anvil prints, and long dotted identifiers mistaken for JWTs; its JWT rule now requires an eyJ header. A cloud run started before upgrading whose allowlisted variable holds that mnemonic or such a dotted value must drop it from ULTRAFUZZ_AGENT_ENV_ALLOWLIST to be replayed or forked. Thanks @aviggiano! (#1165)
  • [runtime] Lets pause and cancel stop a run whose sealed launch records no longer verify, for example after a rebuild or a hand-edited workflow, instead of refusing while the run kept executing. cancel still records canceled only when the runner confirms it, and a launch still publishing its execution snapshot is reported as incomplete instead of as a run to abandon. Thanks @aviggiano! (#1170)
  • [runtime] [cli] Makes run synchronization run one pass at a time: a pass that finds another running skips with WORKFLOW_SYNC_IN_PROGRESS, and status no longer rewrites state.json on every poll. Read-only runner queries time out after 120 s (ULTRAFUZZ_RUNNER_QUERY_TIMEOUT_MS, up to 600 s), and status reports runner, lock and bookkeeping failures as warnings instead of failing. A paused run is no longer cancelled at its deadline, a failed deadline cancel is retried, and an orphaned run's status suggests ultrafuzz resume <run-id>. Thanks @aviggiano! (#1180)
  • [runtime] Stops attempt-ledger and usage-accounting problems from blocking run synchronization: a reset that reuses an attempt number, a recorded attempt whose details later drift, a failed runner inspection, or a sync stopped between the usage.jsonl and run.json writes no longer stops the run's state from updating, and cancelled attempts are now recorded as canceled. Thanks @aviggiano! (#1186)
  • [runtime] [modal] Fixes model fan-out runs that stopped synchronizing after their first pass, and fan-out nodes that stayed skipped after their attempts recovered under resume --retry-failed (both need a custom topology that fans a node out over several model profiles). A private eval on Modal no longer fails once its run's dynamic goal lanes expand, and it still treats a model node with no run-state record as possible model work. Thanks @aviggiano! (#1192, #1202)
  • [artifacts] Fixes run event journals that stopped accepting events at 100,000 records, which then failed every sync, cancel and other lifecycle command; the 64 MiB size limit remains. Appends no longer re-parse the whole journal, so their cost grows far more slowly with its size, and new runs have no events.index/. The findings and property validators no longer throw internal schema parity invariant violated, and 48 artifact checks that never ran are deleted. Thanks @aviggiano! (#1181)
  • [artifacts] Keeps run journals accepting appends. After the host clock steps back behind a run's last event, a new event is stamped one millisecond after it instead of failing, so materialize --confirm, sync, resume and cancel keep working. Concurrent ultrafuzz commands can no longer tear or lose journal records: each append to a run or audit journal holds <journal>.lock. A waiter gives up after 30 s with a journal-locked error naming the holder, and the next append takes over a lock left by a killed process on the same host. Thanks @aviggiano! (#1211, #1221)
  • [runtime] Makes resume --retry-failed finish the recoveries it starts. It reopens the verifier and every node skipped behind a producer it resets, and a failed agent's descendants are skipped instead of each failing in preparation. A node whose verifier rejected its output is judged on the rerun's output, so the node and the run can end succeeded and publish their report; --reset-node recovers the same way. Retrying a dynamic source also re-renders the later prompts that wait on its group, which had stranded custom topologies whose join names the group's children. Thanks @aviggiano! (#1169, #1205, #1220)
  • [runtime] [cli] Makes a resumed run keep each agent's auth, api_key_env and config_dir; resumed Codex tasks had silently fallen back from API-key to subscription auth and now need the configured key. A resume that only attaches to an active run no longer extends the workflow deadline. A failed stale-worktree cleanup or trusted-CLI check is now a warning, and the run's own launcher stays first on PATH. Thanks @aviggiano! (#1177)
  • [runtime] [artifacts] Keeps the final report when its agent rewords or omits a finding's explanatory text, such as its proof of concept: the report verifies with a warning, while changed identity, location, evidence or classification still fail. A schema-valid report that verification still rejects is published as an unchecked PARTIAL report instead of ending with Report: unavailable, unless it fails the secret gate. Thanks @aviggiano! (#1204)
  • [runtime] [evals] Makes observers and harnesses survive transient failures: the EVMBench adapter retries failed status calls and gives up after five in a row, ultrafuzz dashboard starts despite a run directory without run.json, and eval run writes run-summary.json when one row's evidence is unreadable and stops a watch after ten identical sync failures. A crash during submodule hydration no longer bricks the task worktree, and diagnostics keep real .smithers/… paths. Thanks @aviggiano! (#1194)
  • [runtime] Makes ultrafuzz why give each recovery step as a runnable ultrafuzz command for the run, such as ultrafuzz resume <run-id> --retry-failed, instead of a workflow runner invocation, and status on a paused run suggests ultrafuzz resume <run-id>. A failed runner query, such as ultrafuzz node with an unknown node ID, reports the runner's reason instead of the command line it ran. Thanks @aviggiano! (#1207)
  • [cli] Prints a successful result's warnings after the output of plain-text status (including each --watch poll), inspect, why, validate and most other commands; before, only --json showed them. Exit codes are unchanged, and the end-to-end campaign test now also checks that the supervisor relaunches a SIGKILLed engine of a fresh run. Thanks @aviggiano! (#1209)
  • [modal] Reports the real error instead of an intermittent EBADF when the Modal provider rejects a deterministic archive it has just written; the same bug could also close an unrelated open file. Thanks @aviggiano! (#1215)

Full changelog: v0.1.1...v0.1.2