Skip to content

v0.1.3

Latest

Choose a tag to compare

@aviggiano aviggiano released this 01 Oct 14:47
a375d3b

This release makes a running campaign easier to keep going. You can now edit prompts between resumes, an upgrade that changes a schema no longer fails every resumed task, and one failed property lens no longer ends the campaign. Lifecycle commands now run the workflow engine from Ultrafuzz's own pnpm install instead of installing a copy each time, so on the test host the first engine command of a resume took about a second instead of about 43. This release also removes per-node cloud execution, which never completed a node.

  • Edit prompts mid-run. A run's prompt files are now records, not seals. Every resume applies the project's current prompts to the tasks that have not finished, so a prompt fix or an upgrade's prompt change reaches a campaign in flight. To keep hand edits instead, set [run] refresh_prompts_on_resume = false.
  • Fewer campaign-ending failures. A failed property lens no longer skips the rest of the campaign, and a schema change in an upgrade no longer fails every resumed task. The Forge guard and the provider homes now also work under Ubuntu's default umask.
  • Clearer agent failures. A failed Claude Code attempt now records the cause Claude Code states, such as a contended OAuth refresh, beside the generic Claude run failed.

Breaking changes

  • [runtime] [cli] resume, the pre-resume inspect, ps and the other commands that installed their own copy of the workflow engine now run it from Ultrafuzz's pnpm install, where pnpm applies the Smithers compatibility patches. They make no registry install and leave no ultrafuzz-controller-* directory in TMPDIR, and on the test host the first engine command of a resume took about 1 s instead of about 43 s. A pnpm checkout with those patches is now the only supported install: launch and resume refuse a packed or plain npm install, and an engine inside the target project. They also need bun on PATH. Run long campaigns from a dedicated checkout and leave its install alone while they run. A pnpm install that changes the engine's resolved dependency tree moves the engine, and the run then continues only after an ultrafuzz resume from the new install. Thanks @aviggiano! (#1201)
  • [config] [modal] Removes per-node cloud execution ([execution] mode = "cloud"), which never completed a node in any release. validate, doctor, run and eval run reject it with CONFIG_EXECUTION_CLOUD_REMOVED, and resume refuses a run planned with it, so start a new run. To use Modal, run the whole campaign inside one sandbox, as the ultrafuzz-modal eval runner does. Thanks @aviggiano! (#1197)
  • [topology] [config] A failed property lens no longer stops the campaign. In the packaged default, exhaustive and invariant-only topologies, the properties group now continues on failure, and the fan-in moves to a new property-catalog group, which still halts. The other strategies, specialists and review run without the failed lens's properties, and the report is PARTIAL or unverified unless a later resume --retry-failed reruns the lens successfully. The exhaustive and invariant-only profiles use the new topology after the upgrade. An existing project on the default or low-cost profile keeps the old behaviour until it updates .ultrafuzz/topology.yml, either with the three edits in CHANGELOG.md or, if you have not customized it, with ultrafuzz topology copy default .ultrafuzz/topology.yml --force. Every project should also delete .ultrafuzz/prompts/properties/property-specification-fanin.md and rerun ultrafuzz init. A failure_policy: continue group's results are now optional to every node outside that group, not only to review. Thanks @aviggiano! (#1198)
  • [runtime] A run's prompt files are records, not seals. artifacts/<attempt>/prompt.rendered.md is the prompt that attempt receives on every verb, and nothing compares it with its launch bytes any more, so an edit no longer strands the run. A prompt that cannot be rendered, or a missing prompt file, now fails only its task, instead of every render. A run launched by an earlier release whose prompts this release renders differently can be synchronized, resumed and reported again. What is given up: a process running as the operator can change a later task's prompt undetected, and replay and fork use the current prompt files instead of the launch bytes. Thanks @aviggiano! (#1230)
  • [runtime] [config] resume now applies the project's current prompts to the tasks that have not finished, by default. It renders each prompt the way the run's launch did, records every file it rewrites in prompt-history/<time>-<id>/refresh.json, and reports PROMPTS_REFRESHED. It never fails the resume: a prompt that run would reject keeps its files and is reported as PROMPT_REFRESH_REJECTED, and an edited topology makes it skip with PROMPT_REFRESH_SKIPPED. A hand edit of an unfinished task's prompt.rendered.md is replaced by the next resume. To keep such edits, set [run] refresh_prompts_on_resume = false, which resume reads from the project's current ultrafuzz.toml, so it also applies to runs already in flight. Thanks @aviggiano! (#1237)
  • [runtime] [artifacts] A run keeps validating its artifacts against the schemas it was planned with, so an upgrade that changes a schema, even by one $comment, no longer fails every resumed task. What is given up: a run no longer adopts an upgrade's schema fixes. Pause or finish running campaigns before you upgrade. Then continue a run launched before this release with ultrafuzz resume <run-id> --refresh-controller, every time; a plain resume of such a run fails with WORKFLOW_CONTROLLER_REFRESH_REQUIRED before Smithers starts. Thanks @aviggiano! (#1238)
  • [runtime] Without ULTRAFUZZ_PROVIDER_HOME_ROOT, the provider-home root is now ~/.ultrafuzz-provider-homes, created with mode 0700, instead of $XDG_STATE_HOME/ultrafuzz/provider-homes. Ubuntu's default umask makes ~/.local group writable, and the provider-home check refused the old root below it, so OpenRouterAgent, DeepSeekAgent and agents with a config_dir could not start. Nothing is moved: a ClaudeAgent, CodexAgent or KimiAgent with a config_dir finds its home empty, so with auth = "subscription" it is not logged in. To keep those logins, move the old root before an agent first runs on this release: mv "${XDG_STATE_HOME:-$HOME/.local/state}/ultrafuzz/provider-homes" ~/.ultrafuzz-provider-homes. Then rerun ultrafuzz init in each project: ultrafuzz run refuses stock agent adapters from an earlier release (CONTROLLER_SOURCE_UNTRUSTED), and this release changes several of them. Thanks @aviggiano! (#1242)

New features

  • [runtime] [config] Adds an opt-in friction log, [run] friction_log_enabled (default false). When it is on, agents are told to record roadblocks they hit with Ultrafuzz, its tooling or their instructions through a run-local command. That command runs the pinned Frog CLI and accepts only log and list. Entries land in <run>/friction/.agents/friction-log/<id>/friction.md and stay there until you review them. They are free-form agent text that the artifact secret scan does not cover, so check them for credentials before you publish anything. The command guards against mistakes, not against agents, which run unsandboxed. Frog is not copied into runs, so a disabled log adds nothing to a run. Thanks @mrthankyou and @aviggiano! (#1162, #1245)

Improvements

  • [runtime] A failed ClaudeAgent or DeepSeekAgent attempt now records the failure Claude Code states, such as a contended OAuth refresh or a rejected API key, beside the generic error, for example Claude run failed See https://smithers.sh/reference/errors (agent stated: Failed to refresh OAuth token: …). It appears in the node's last_error (in ultrafuzz inspect <run-id> --json and the dashboard), in the attempt ledger's failure_message and in public eval diagnostics, with the same redaction and 1,000-byte cap as other failure text. status and why do not show it yet. Retry, quota and auth handling are unchanged. Thanks @mrthankyou and @aviggiano! (#1224, #1225, #1246)
  • [docs] Removes the Evals section and its charts from the README; the eval suite reference stays linked. Thanks @aviggiano! (#1228)

Bug fixes

  • [runtime] On a host whose umask leaves new directories group writable, such as Ubuntu's default 0002, the Forge guard now limits Forge. Tasks had run the real forge without forge_vmem_limit_kb and forge_rayon_threads while run.json recorded the guard as active. run.json now records whether the engine keeps the wrapper on PATH, and a FORGE_GUARD_INACTIVE warning says when it does not, such as with a custom run.output_dir. ultrafuzz clean of a run planned for per-node Modal execution now warns with CLEAN_CLOUD_STORAGE_RETAINED and names the Modal volume and sandboxes it no longer removes. Thanks @aviggiano! (#1235)
  • [runtime] A final-report producer retry or verifier in a restarted controller (resume, quota park, supervisor relaunch) no longer needs a smithers executable on the operator's PATH. Before, it failed every remaining attempt and the run had no report. The workflow now records each producer attempt's selections in the run directory. For a run launched by an earlier release whose verifier fails with report producer selection was never recorded, run ultrafuzz resume <run-id> --refresh-controller --reset-node <report node> (node:final-report in the packaged topologies). Thanks @aviggiano! (#1183)
  • [runtime] [modal] Moves production dependencies past new advisories: brace-expansion and undici past the High advisories published on 2026-09-29, including inside the patched operator npm, and @grpc/grpc-js to 1.14.5, past GHSA-m9gg-hp2v-232j. The dependency-advisory gate now merges the registry's per-range duplicates of one advisory. Thanks @aviggiano! (#1229, #1241)

Full changelog: v0.1.2...v0.1.3