Releases: monad-developers/ultrafuzz
Release list
v0.1.3
This release makes a running campaign easier to keep going. You can now edit prompts between resumes, an upgrade that changes a schema no longer fails every resumed task, and one failed property lens no longer ends the campaign. Lifecycle commands now run the workflow engine from Ultrafuzz's own pnpm install instead of installing a copy each time, so on the test host the first engine command of a resume took about a second instead of about 43. This release also removes per-node cloud execution, which never completed a node.
- Edit prompts mid-run. A run's prompt files are now records, not seals. Every
resumeapplies the project's current prompts to the tasks that have not finished, so a prompt fix or an upgrade's prompt change reaches a campaign in flight. To keep hand edits instead, set[run] refresh_prompts_on_resume = false. - Fewer campaign-ending failures. A failed property lens no longer skips the rest of the campaign, and a schema change in an upgrade no longer fails every resumed task. The Forge guard and the provider homes now also work under Ubuntu's default umask.
- Clearer agent failures. A failed Claude Code attempt now records the cause Claude Code states, such as a contended OAuth refresh, beside the generic
Claude run failed.
Breaking changes
- [runtime] [cli]
resume, the pre-resumeinspect,psand the other commands that installed their own copy of the workflow engine now run it from Ultrafuzz's pnpm install, where pnpm applies the Smithers compatibility patches. They make no registry install and leave noultrafuzz-controller-*directory in TMPDIR, and on the test host the first engine command of a resume took about 1 s instead of about 43 s. A pnpm checkout with those patches is now the only supported install: launch andresumerefuse a packed or plain npm install, and an engine inside the target project. They also needbunonPATH. Run long campaigns from a dedicated checkout and leave its install alone while they run. Apnpm installthat changes the engine's resolved dependency tree moves the engine, and the run then continues only after anultrafuzz resumefrom the new install. Thanks @aviggiano! (#1201) - [config] [modal] Removes per-node cloud execution (
[execution] mode = "cloud"), which never completed a node in any release.validate,doctor,runandeval runreject it withCONFIG_EXECUTION_CLOUD_REMOVED, andresumerefuses a run planned with it, so start a new run. To use Modal, run the whole campaign inside one sandbox, as theultrafuzz-modaleval runner does. Thanks @aviggiano! (#1197) - [topology] [config] A failed property lens no longer stops the campaign. In the packaged
default,exhaustiveandinvariant-onlytopologies, thepropertiesgroup now continues on failure, and the fan-in moves to a newproperty-cataloggroup, which still halts. The other strategies, specialists and review run without the failed lens's properties, and the report is PARTIAL or unverified unless a laterresume --retry-failedreruns the lens successfully. Theexhaustiveandinvariant-onlyprofiles use the new topology after the upgrade. An existing project on thedefaultorlow-costprofile keeps the old behaviour until it updates.ultrafuzz/topology.yml, either with the three edits inCHANGELOG.mdor, if you have not customized it, withultrafuzz topology copy default .ultrafuzz/topology.yml --force. Every project should also delete.ultrafuzz/prompts/properties/property-specification-fanin.mdand rerunultrafuzz init. Afailure_policy: continuegroup's results are now optional to every node outside that group, not only toreview. Thanks @aviggiano! (#1198) - [runtime] A run's prompt files are records, not seals.
artifacts/<attempt>/prompt.rendered.mdis the prompt that attempt receives on every verb, and nothing compares it with its launch bytes any more, so an edit no longer strands the run. A prompt that cannot be rendered, or a missing prompt file, now fails only its task, instead of every render. A run launched by an earlier release whose prompts this release renders differently can be synchronized, resumed and reported again. What is given up: a process running as the operator can change a later task's prompt undetected, andreplayandforkuse the current prompt files instead of the launch bytes. Thanks @aviggiano! (#1230) - [runtime] [config]
resumenow applies the project's current prompts to the tasks that have not finished, by default. It renders each prompt the way the run's launch did, records every file it rewrites inprompt-history/<time>-<id>/refresh.json, and reportsPROMPTS_REFRESHED. It never fails the resume: a prompt thatrunwould reject keeps its files and is reported asPROMPT_REFRESH_REJECTED, and an edited topology makes it skip withPROMPT_REFRESH_SKIPPED. A hand edit of an unfinished task'sprompt.rendered.mdis replaced by the nextresume. To keep such edits, set[run] refresh_prompts_on_resume = false, whichresumereads from the project's currentultrafuzz.toml, so it also applies to runs already in flight. Thanks @aviggiano! (#1237) - [runtime] [artifacts] A run keeps validating its artifacts against the schemas it was planned with, so an upgrade that changes a schema, even by one
$comment, no longer fails every resumed task. What is given up: a run no longer adopts an upgrade's schema fixes. Pause or finish running campaigns before you upgrade. Then continue a run launched before this release withultrafuzz resume <run-id> --refresh-controller, every time; a plainresumeof such a run fails withWORKFLOW_CONTROLLER_REFRESH_REQUIREDbefore Smithers starts. Thanks @aviggiano! (#1238) - [runtime] Without
ULTRAFUZZ_PROVIDER_HOME_ROOT, the provider-home root is now~/.ultrafuzz-provider-homes, created with mode0700, instead of$XDG_STATE_HOME/ultrafuzz/provider-homes. Ubuntu's default umask makes~/.localgroup writable, and the provider-home check refused the old root below it, soOpenRouterAgent,DeepSeekAgentand agents with aconfig_dircould not start. Nothing is moved: aClaudeAgent,CodexAgentorKimiAgentwith aconfig_dirfinds its home empty, so withauth = "subscription"it is not logged in. To keep those logins, move the old root before an agent first runs on this release:mv "${XDG_STATE_HOME:-$HOME/.local/state}/ultrafuzz/provider-homes" ~/.ultrafuzz-provider-homes. Then rerunultrafuzz initin each project:ultrafuzz runrefuses stock agent adapters from an earlier release (CONTROLLER_SOURCE_UNTRUSTED), and this release changes several of them. Thanks @aviggiano! (#1242)
New features
- [runtime] [config] Adds an opt-in friction log,
[run] friction_log_enabled(defaultfalse). When it is on, agents are told to record roadblocks they hit with Ultrafuzz, its tooling or their instructions through a run-local command. That command runs the pinned Frog CLI and accepts onlylogandlist. Entries land in<run>/friction/.agents/friction-log/<id>/friction.mdand stay there until you review them. They are free-form agent text that the artifact secret scan does not cover, so check them for credentials before you publish anything. The command guards against mistakes, not against agents, which run unsandboxed. Frog is not copied into runs, so a disabled log adds nothing to a run. Thanks @mrthankyou and @aviggiano! (#1162, #1245)
Improvements
- [runtime] A failed
ClaudeAgentorDeepSeekAgentattempt now records the failure Claude Code states, such as a contended OAuth refresh or a rejected API key, beside the generic error, for exampleClaude run failed See https://smithers.sh/reference/errors (agent stated: Failed to refresh OAuth token: …). It appears in the node'slast_error(inultrafuzz inspect <run-id> --jsonand the dashboard), in the attempt ledger'sfailure_messageand in public eval diagnostics, with the same redaction and 1,000-byte cap as other failure text.statusandwhydo not show it yet. Retry, quota and auth handling are unchanged. Thanks @mrthankyou and @aviggiano! (#1224, #1225, #1246) - [docs] Removes the Evals section and its charts from the README; the eval suite reference stays linked. Thanks @aviggiano! (#1228)
Bug fixes
- [runtime] On a host whose umask leaves new directories group writable, such as Ubuntu's default
0002, the Forge guard now limits Forge. Tasks had run the realforgewithoutforge_vmem_limit_kbandforge_rayon_threadswhilerun.jsonrecorded the guard as active.run.jsonnow records whether the engine keeps the wrapper onPATH, and aFORGE_GUARD_INACTIVEwarning says when it does not, such as with a customrun.output_dir.ultrafuzz cleanof a run planned for per-node Modal execution now warns withCLEAN_CLOUD_STORAGE_RETAINEDand names the Modal volume and sandboxes it no longer removes. Thanks @aviggiano! (#1235) - [runtime] A final-report producer retry or verifier in a restarted controller (resume, quota park, supervisor relaunch) no longer needs a
smithersexecutable on the operator'sPATH. Before, it failed every remaining attempt and the run had no report. The workflow now records each producer attempt's selections in the run directory. For a run launched by an earlier release whose verifier fails withreport producer selection was never recorded, runultrafuzz resume <run-id> --refresh-controller --reset-node <report node>(node:final-reportin the packaged topologies). Thanks @aviggiano! (#1183) - [runtime] [modal] Moves production dependencies past new advisories:
brace-expansionandundicipast the High advisories published on 2026-09-29, including inside the patched operator npm, and@grpc/grpc-jsto 1.14.5, past GHSA-m9gg-hp2v-232j. The dependency-advisory gate now merges the registry's per-range du...
v0.1.2
This is a stability release: it fixes the orchestration and bookkeeping failures that stopped campaigns short of their report, makes recovery commands finish what they start, and removes a large amount of dead code, with stricter CI gates to keep it out. On the release commit, a fresh smoke campaign launched in about four minutes and finished succeeded with a verified, complete report in about 11 minutes.
- Campaigns run to completion. The
default,low-cost,exhaustiveandinvariant-onlyprofiles launch again, the supervisor relaunches an engine that dies mid-run, and run synchronization, attempt and usage ledgers, and event journals no longer stop or strand a run over bookkeeping problems. - Recovery that works.
resume --retry-failedreopens skipped work and judges a rerun on its own output, and a report whose agent reworded a finding verifies with a warning instead of being discarded. A tiny-vault smoke campaign that had failed twice withReport: unavailablerecovered withresume --retry-failedto a verified, complete report. - Less code, stricter CI. The codebase ends about 14,000 lines smaller, mostly from deleting dead gates, pipelines and recovery code. Every package is now validated on pull requests, a complexity ceiling and an unused-export check block regressions, and a model-free end-to-end campaign kills and resumes a real engine on every pull request.
Breaking changes
- [runtime] [artifacts] Finish or cancel every run launched on v0.1.1 before upgrading. Before each task, a v0.1.1 run's generated workflow checks that the validator build matches the one the run was planned with, and this release changes that build. A plain
ultrafuzz resumeof such a run therefore fails each remaining task withartifact-contract failure: planned schema binding changed;pauseandcancelstill work. Runs launched on this release validate each artifact against the schema it was planned with and no longer compare the validator build, so a later rebuild that changes only that build, such as an ajv bump, no longer strands them.report --require-verifiednow reports an invalid artifact asJSON_SCHEMA_VIOLATIONinstead ofARTIFACT_SCHEMA_INVALID. Thanks @aviggiano! (#1188) - [runtime] After upgrading, run
ultrafuzz initin each existing project (no--forceneeded): it now replaces any stock.smithers/agentsadapter that differs from its packaged copy, the only one planning accepts. Provider route IDs change for routes derived from a Codexconfig.toml, Claudesettings.jsonor Kimiconfig.toml, for IDs that included a proxy or an unused Claude cloud variable, and for a Codex config that sets only a non-emptyopenai_base_url, which cloud planning now rejects. UpdateULTRAFUZZ_DATA_GOVERNANCE_POLICY, re-acknowledge, and re-plan any run whose acknowledged ID changed. In return, a CLI rewriting its own config, a proxy change or an edit to an unusedopenai_base_urlno longer fails every later task, DeepSeek tasks no longer hang, and Kimi and Pi no longer crash on unexpected output or usage data. Thanks @aviggiano! (#1173, #1210) - [config] [prompts] Review nodes in the default, exhaustive and invariant-only topologies and in new scaffolds now get 7,200 seconds, even where
run.default_timeout_secondsor a model profile sets more. To keep a larger budget,ultrafuzz topology copythe topology, raise itsreviewgroup'stimeout_seconds, and pointtopology_pathat the copy. A cloud config whoseexecution.resources.timeout_secondsis below 7,200 needs per-node overrides for the review nodes too. An existing.ultrafuzz/topology.ymlkeeps its old review budget, one hour by default, until you adddefaults: { timeout_seconds: 7200 }to itsreviewgroup. The stateful-invariant and Vyper setup prompts stop inviting production-source edits the workspace handoff rejects. Existing projects keep old copies of these prompts and ofreview/final-report.md, whose## Audit contextinstruction fails report verification if followed:ultrafuzz validatenow flags each asPROMPT_DIFFERS_FROM_BUILT_IN; delete the ones you have not customized and rerunultrafuzz init. Thanks @aviggiano! (#1164) - [runtime] [prompts] Runs launched on v0.1.1 whose dynamic groups have fanned out, such as runs past
goal-plan, also cannot be synchronized after upgrading. A prompt that waits on a dynamic group is re-rendered and compared byte for byte with its published copy, and this release changes the coverage section that the final-report prompt embeds.statusandinspectstop synchronizing such a run, andresume,replay,fork,whyandstatsfail, most withruntime rendered prompt changed. The same change stops host-side checks rejecting output that a node's own verification accepted: a failed optional test-generation strategy, dynamic goal nodes, the triage, severity, test-aggregation and final-report nodes of any run whose goal plan named a goal, false-positive findings without severity fields, and ordinary coverage sentences in reports, which are now warnings. Thanks @aviggiano! (#1176) - [runtime] Final reports published before this release no longer verify: readers re-render them, and the restated run summary no longer matches. For such a run,
ultrafuzz reportand the dashboard show a complete run asPARTIALwith verificationnot-checked,report bundlefalls back to a report-only archive, and--require-verifiedreads and the Modal public worker fail. Build any verified bundles you need before upgrading. In return, reports give elapsed time, models, tokens and spend for the whole run instead of as of the report agent's start. The renderer also no longer throws on link syntax, non-ASCII issue titles or label-like prose, which had left such runs with no report. Thanks @aviggiano! (#1167) - [config] [cli]
validateandrunreject[models] default = "<id>"naming any profile other thandefault(CONFIG_MODEL_DEFAULT_UNSUPPORTED), which produced an unparseableconfig.resolved.toml. Select the primary profile with[retry] agents = ["<id>"]instead. Commands on existing runs keep working. Prompt-authority selectors now sort the same way on every host instead of by host locale. That fixes task-preparation failures when producer attempt IDs share a prefix (node-1/node-10), and gives a custom multi-path selector that the old order sorted differently a new ID. It also accepts a symlinked--project, artifact path segments starting with.., aworkflow_deadline_secondsof up to seven days, and clean, materialize and dashboard audit journals across a backward clock step. Thanks @aviggiano! (#1195) - [runtime] [cli]
statsJSON gains acancelednode status and acanceledkey intotals.status_countswithout a schema version change, so consumers that validate stats against the previous closed schema must accept them. A node whose failed tasks were all cancelled, for example throughultrafuzz cancel, now reportscanceledinstead offailed; a node with any other failed task staysfailed. An agent attempt interrupted by a controller crash is now recorded inattempts.jsonlascanceledonce the resumed run restarts its task, sostatsandstatuscount the same attempts. On the first sync after upgrading, a finished run with such an attempt raises that node'sretry_countand republishes its report once, still verified. Thanks @aviggiano! (#1200, #1213) - [evals] [cli] Removes
ultrafuzz eval history --max-age-days, the six per-metric history SVGs the README never embedded, andjson validate's support for the telemetry-cursor, publication-state and automatic history-publication schemas. Remove--max-age-daysfrom any script that passes it; the CLI now rejects the flag.eval runstops running its unused telemetry forwarding, which aborted the whole suite beforerun-summary.jsonwas written when a run's journal held a record it did not expect, and no longer writestelemetry/. Thanks @aviggiano! (#1174) - [runtime] Deletes about 2,350 lines of controller-refresh and recovery code that has had no caller since v0.0.23;
resume --refresh-controlleris unchanged.status,pause,cancel,replayandforkno longer replay the whole event journal, and run status now follows the workflow engine's result. A run that a pre-v0.0.23 build refreshed now fails those commands, and a run that such a build recovered can no longer re-read its final report. Both have been unsupported since v0.0.25, so start a new run. Thanks @aviggiano! (#1193)
Improvements
- [runtime] [cli] Halves the per-file disk flushes a launch spends copying the run's execution snapshot, and a very large run can now read back the launch record it wrote.
ultrafuzz doctorrequires only the CLIs of agents the selected topology or retry chain can use, and warns when the temporary directory is a tmpfs or has under 2 GiB free, sizing its leftover controller directories for at most about a second.resume --refresh-controllerno longer leaves untracked.smithers/continuationsfiles that made the next private or cloud launch reject the target as dirty. Thanks @aviggiano! (#1179, #1208) - [runtime] Names the runner's run-level error, such as
WORKFLOW_RENDER_FAILED: <what the workflow threw>, in theWORKFLOW_TERMINAL_WITHOUT_FAILED_NODEdiagnostic of a run that endsfailedwith no failed node, redacted and capped at 1,000 characters, and adds a pinned-runner test that recovers such a run through a same-IDresumeonce its cause is removed. Thanks @aviggiano! (#1172) - [prompts] [config] Deletes code no production path calls: the prompt rename feature, the always-empty config prompt-metadata layer, redaction-restore helpers that
docs/config.mddescribed but nothing ran, unused Modal exports and two orphaned CI scripts. Output-contract templates now resolve packaged-first like the other prompt assets, and ...
v0.1.1
This release strengthens default audits with stateful invariant fuzzing, removes implicit third-party evaluation reporting, and updates setup guidance and CI reliability.
- Stateful audits by default. Fresh projects now run all five invariant stages, while exhaustive campaigns can spend four hours fuzzing selected properties.
- More control over evaluation data. Braintrust reporting and the implicit judge gateway are removed; optional judging now requires explicit endpoint and credential configuration.
- Clearer setup and safer maintenance. Documentation matches current configuration and agent behavior, and the ZIP reader update includes upstream archive handling fixes.
Breaking changes
- [evals] [modal] Removes Braintrust reporting and the
eval publishAPI. Optional LLM judging now requiresULTRAFUZZ_EVAL_JUDGE_URLand a dedicated judge key. Modal benchmark configuration moves toultrafuzz.modal.benchmark.v3; existing v2 configurations are rejected and must be regenerated. Thanks @aviggiano! (#1134)
New features
- [config] [runtime] Adds all five invariant stages to fresh default audits and extends exhaustive campaigns to four hours of fuzzing across selected properties, while preserving explicit project and runtime overrides. Thanks @aviggiano! (#1132)
Improvements
- [cli] Updates
adm-zipto 0.6.1, incorporating upstream archive parsing and extraction hardening. Thanks @dependabot! (#1127) - [workflows] [docs] Removes paid benchmark automation and its write-capable history publisher while retaining manual benchmark tools and credential-free CI. Thanks @aviggiano! (#1131)
- [docs] Aligns setup and configuration guidance with current schemas, agents, prerequisites, and Smithers migrations. Thanks @aviggiano! (#1123)
- [docs] [security] Adds private vulnerability reporting guidance and security ownership, and updates the license attribution to Monad Foundation. Thanks @aviggiano! (#1125)
- [evals] Updates the workspace to Vitest 4.1.11 and adapts package tests and fixtures for the new version. Thanks @aviggiano and @dependabot! (#1135, #1121)
Bug fixes
- [workflows] Gives runtime integration and CLI CI lanes up to 120 minutes to finish on slow runners, without changing their coverage or failure gates. Thanks @aviggiano! (#1137)
Full changelog: v0.1.0...v0.1.1
v0.1.0
Ultrafuzz v0.1.0 is the first public release of the AI-powered orchestrator for smart-contract fuzzing and threat hunting. It brings property specification, threat modeling, parallel investigation, stateful testing, and report generation into one configurable workflow for Solidity and Vyper projects.
- Investigate from multiple angles. Specialized agents examine a protocol through its actors, user flows, properties, and potential failure modes.
- Combine reasoning with execution. Property-guided investigation works alongside dedicated stateful fuzzing for behavior that depends on transaction sequences and changing state.
- Adapt the workflow to your project. Choose an audit profile and models, edit the prompts and workflow graph, and monitor the campaign from the CLI or dashboard.
Features
- [prompts] [topology] Property-guided threat hunting. Build a shared protocol context, derive concrete properties and invariants, and investigate them through parallel specialist strategies. Threat modeling and generated investigation goals draw on the OWASP Smart Contract Security reference material.
- [runtime] Stateful fuzzing. The
exhaustiveandinvariant-onlyprofiles include dedicated invariant campaigns. Generate and run stateful tests, retaining test artifacts, failure reproductions, and available coverage evidence for review. - [config] Ready-made audit profiles. Start with
smoke,low-cost,default,exhaustive, orinvariant-only. Adjust concurrency, strategy breadth, retry limits, and time budgets to fit the project. - [runtime] [modal] Model and execution choice. Use frontier or open-weight models through Claude, Codex, DeepSeek, Kimi, OpenCode, OpenRouter, and Pi adapters. Run locally or use Modal for supported cloud tasks.
- [cli] [dashboard] Campaign visibility and recovery. Follow the workflow graph, node progress, attempts, logs, token usage, and cost estimates. Inspect blockers and use resume, replay, or fork commands to continue or revisit work.
- [artifacts] Reports with explicit coverage. Deduplicate and review findings, then produce agent-written Markdown and JSON reports. Best-effort execution lets independent work continue after ordinary task failures. Successful reporting marks incomplete coverage as PARTIAL; if reporting fails or cannot start, saved results remain available and the report is marked unavailable. Strict completion and report verification are separate options.
- [prompts] [topology] An editable workflow. Prompts, workflow topology, model profiles, and references live in project files. Extend the specialist strategies, change their handoffs, or reuse the prompts in another orchestrator. Generated test files can be reviewed and explicitly copied into the target repository.
- [evals] [evmbench] Evaluation tooling. Use UltrafuzzBench, EVMBench integration, and the focused ScFuzzBench workflow to evaluate campaign variants. Track precision, recall, F1, time, token usage, and cost, with retained run evidence and published benchmark history.
Get started
Read the launch post, follow the first-campaign guide, or explore the prompt catalog.
Run campaigns on disposable, isolated virtual machines. Agents execute commands without permission prompts; review the security guidance before starting.
v0.0.25
v0.0.25 makes new runs use best-effort completion and gives operators a clear report outcome when tasks fail. Final report content remains agent-written.
- Keep useful work moving. Independent tasks continue after ordinary failures. Review uses successful, checked inputs, and incomplete coverage produces a PARTIAL report when the report agent succeeds.
- Know when reporting is unavailable.
status --watch --jsonexposes whether execution ended and the report's availability, completion, verification, and paths. Failed reporting keeps the saved results without creating a replacement report. - Choose the required checks.
run --require-completerequires complete execution coverage.report --require-verifiedseparately requires report verification. Neither option adds retries.
Breaking changes
- [runtime] [config] Adds default best-effort completion, agent-written PARTIAL reporting, and report availability in the existing status command. Local report verification is optional; strict bundles, benchmark scoring, and public publication retain their verification requirements. The sealed resolved-config contract advances to v4. This release applies to new runs; migration and resumption of older runs under the new contract are unsupported. Custom topologies retain their declared failure policies, and configured attempts and time limits still apply. Updates the TOML parser to 1.7.1 to address GHSA-7w5x-hrqm-74c2. Thanks @aviggiano! (#1120)
Improvements
- [benchmarks] Refreshes published benchmark history and charts through the automatic publisher. Direct commit: 46ae7dd2.
Bug fixes
- [evals] Fixes the required-backend preflight test so clean CI runners verify the intended rejection without depending on reference caches. Thanks @aviggiano! (#1122)
Published without waiting for the final PR CI rerun or CI on the release commit, as authorized by the maintainer. Local validation includes all 187 CLI tests and all 450 evaluation tests passing.
Full changelog: v0.0.24...v0.0.25
v0.0.24
v0.0.24 focuses on more reliable campaign continuation and report publication. It strengthens continuation, retry history, workspace preparation, and live diagnostics, improves macOS startup, and moves the shipped vulnerability reference to pinned OWASP Smart Contract Security data.
- More reliable continuation. Resumed and reset campaigns retain their original policy, task identities, failed-attempt history, and coherent workspace evidence.
- Dependable results and diagnostics. Reports survive controller restarts and public redaction, while status and worker failures preserve the evidence operators need.
- Pinned OWASP references. The new reference adapter validates source provenance and derived planning data; existing configurations must migrate from the former database format.
Breaking changes
- [references] [runtime] Replaces the former vulnerability database with commit-pinned OWASP Smart Contract Security records and a digest-verified planner catalog. Existing projects must review reference, topology, setup-prompt and schema changes, sync
vulnerability-database.owasp-scs, and start a fresh run; retain the previous CLI for historical database artifacts. The former database formats are no longer accepted. Optional SCSVS hints cannot exclude weaknesses, so this pin plans 156 class goals plus threat goals and roaming. See the migration guide. Thanks @aviggiano! (#1118)
New features
- [runtime] Enables campaign submission on macOS with inode-rooted snapshot paths that preserve file identity across renames and symlink replacement. Thanks @mrthankyou! (#1016)
Improvements
- [docs] Places disposable, isolated-host guidance and the unrestricted-agent warning before the first setup and campaign commands. Thanks @mrthankyou! (#1023)
- [runtime] [cli] Adds an opt-in synchronization deadline for status, inspection, stats, and diagnostics, with a stale-snapshot warning when the deadline expires. Observers continue to wait for complete synchronization by default. Thanks @aviggiano! (#1059, #1089)
- [workflows] [benchmarks] Removes the age-only benchmark-history monitor that opened outage issues when no benchmark-triggering changes had occurred. Thanks @aviggiano! (#1112)
- [benchmarks] Refreshes published benchmark history and its documentation charts through the automated publisher. These automated commits have no associated pull requests:
e876e555andc1229d04.
Bug fixes
- [runtime] [modal] Creates temporary roots beneath a canonical system directory, fixing macOS path validation in runtime, cloud, reference, and packaged-install operations, and cleans up read-only runtime test snapshots. Thanks @mrthankyou! (#1019, #1025)
- [runtime] Matches descriptor-backed macOS workflow files by device and inode so valid snapshot aliases retain their workflow identity. Thanks @aviggiano! (#1032)
- [workflows] [security] Marks known synthetic credential fixtures so secret-scanning checks stop rejecting their intentional test values while continuing to scan surrounding code. Thanks @mrthankyou! (#1021)
- [modal] [cli] Aligns report and publication fixtures with authenticated detection provenance, restoring the intended cloud and CLI validation checks. Thanks @aviggiano! (#1033)
- [runtime] [modal] Passes the engine-owned task runtime directly into cloud workflows so final reports can read durable workflow metrics across separately loaded package copies. Thanks @aviggiano! (#1034)
- [runtime] Preserves the admitted executable PATH and isolated Pi home through detached continuations so the generated adapter can find Pi and retain its approved environment. Thanks @aviggiano! (#1037)
- [cli] Projects only documented public metadata into inspect results, preventing persisted source provenance or private fields from invalidating the CLI response. Thanks @aviggiano! (#1040)
- [runtime] Keeps detached-resume startup helpers within their module scope and migrates the previous generated patch during controller refresh. Thanks @aviggiano! (#1042)
- [runtime] Normalizes Pi reasoning-token breakdowns without reducing the provider-reported output total. Thanks @aviggiano! (#1044)
- [runtime] Retains a process-owned copy of inherited execution-snapshot descriptors so later engine, supervisor, and resume handoffs survive closure or reuse of the startup descriptor. Thanks @aviggiano! (#1046)
- [prompts] Makes final-report producers preserve dedupe-owned descriptions and final lifecycle severity, preventing complete smoke runs from failing report publication on rewritten authority. Thanks @aviggiano! (#1047)
- [runtime] Restores missing static presentation prompts from verified immutable plan snapshots before resume or reset, preserving completed nodes. Thanks @aviggiano! (#1049)
- [cli] Marks stats results unsuccessful when synchronization reports an error while retaining the statistics snapshot and original diagnostics. Thanks @aviggiano! (#1051)
- [runtime] Retries recognized atomic-replacement races when observing live run documents, while continuing to reject stable integrity failures. Thanks @aviggiano! (#1055)
- [runtime] Retries partially written final event records with a bounded coherent-snapshot budget; stable malformed tails and malformed earlier records still fail. Thanks @aviggiano! (#1057)
- [runtime] Keeps lifecycle event queries and watchers readable during authenticated dynamic retries and exposes control divergence as an observation warning. Thanks @aviggiano! (#1062)
- [runtime] Archives complete dynamic-expansion generations before producer retry, retaining attempt artifacts and workspace evidence while safely releasing stale worktree registrations. Thanks @aviggiano! (#1064)
- [runtime] Records workflow failures before agent selection without inventing executed model attempts or immutable attempt provenance. Thanks @aviggiano! (#1066)
- [runtime] Repairs absent, prunable worktree registrations owned by the current run before ordinary resume, preserving live, locked, and unrelated registrations. Thanks @aviggiano! (#1070)
- [runtime] Recognizes a reset node’s retained current attempt even when its attempt number is lower than a previous failed attempt. Thanks @aviggiano! (#1074)
- [cli] Includes authenticated dynamically generated tasks in stats without admitting unrelated task additions or dropping the sealed baseline. Thanks @aviggiano! (#1076)
- [cli] [runtime] Keeps run listings available when one run directory cannot be read, showing the unreadable run and its diagnostic so operators can identify and clean it. Thanks @mrthankyou! (#1080)
- [runtime] Accepts blank lines inside invariant-ledger and canonical-property Markdown blocks while retaining exact field, ordering, and JSON parity checks. Thanks @mrthankyou! (#1083)
- [dependencies] Updates the transitive fast-uri resolution to 3.1.7, addressing four high-severity advisories without adding dependency overrides. Thanks @aviggiano! (#1085)
- [runtime] Corrects the exhausted-read-race regression to expect an unreadable run listing and warning, matching the combined live-read and listing behavior. Thanks @aviggiano! (#1088)
- [workflows] Retries transient registry failures in the dependency-advisory gate and runs deterministic checks first, so temporary registry outages do not hide formatting, lint, or build failures. Thanks @aviggiano! (#1090)
- [modal] Retains bounded, sanitized eval-score and eval-report failure envelopes in public worker diagnostics, including the underlying error code and message. Thanks @aviggiano! (#1092)
- [cli] [evals] Preserves eval error codes and bounded, redacted details in score and report failure messages without changing the CLI envelope contract. Thanks @aviggiano! (#1093)
- [evals] Uses an OpenAI-compatible strict judge response schema and bounded retries for transient or invalid judge responses, while retaining full local result validation. Thanks @aviggiano! (#1094)
- [artifacts] [runtime] Allows selected nonessential artifact metadata omissions with attributable warnings, preserves authenticated finding identity through review, and carries warnings into reports and public companions. Missing substantive evidence and conflicting identities remain fatal. Thanks @aviggiano! (#1095)
- [modal] Names and logs otherwise unhandled worker and public-bundle failures, preserving bounded, redacted causes instead of reporting unexplained sandbox exits. Thanks @aviggiano! (#1096)
- [runtime] [cli] Reports an incomplete launch without inventing workflow identity or launcher liveness, allowing read-only diagnosis while execution still requires valid control evidence. Thanks @mrthankyou! (#1097)
- [runtime] Makes public final-report redaction stable across repeated publication checks and preserves valid run identifiers without weakening positive credential detection. Thanks @aviggiano! (#1098)
- [runtime] Collapses only duplicate string values created by public-report redaction so valid reports continue to satisfy unique-array constraints. Thanks @aviggiano! (#1107)
- [artifacts] Uses the dedicated authenticated lifecycle ledger as the final report’s lifecycle authority, avoiding contradictory requirements from stale embedded copies. Thanks @aviggiano! (#1103)
- [runtime] Preserves run, logical-node, and Smithers-node identity in initial and reconstructed task specifications so fresh verifiers can reconcile recorded producer attempts. Thanks @aviggiano! (#1104, #1113)
- [runtime] Recovers final-report retry provenance from...
v0.0.23
v0.0.23 makes long-running audits easier to resume and trust. It moves ordinary continuation onto Smithers 0.35, repairs authenticated controller refresh across dynamic runs, and restores honest benchmark and report publication with tighter workflow and artifact validation.
- Native continuation. Stopped runs resume under their existing Smithers identity, with current engine support and safer ownership, session, and dependency handling.
- Authenticated recovery. Controller refresh, sealed snapshots, dynamic manifests, prompts, schemas, and workspace replay remain recoverable without relaxing integrity checks.
- Honest results. Benchmark freshness, smoke readiness, report metrics, coverage, and release validation now fail loudly and explain what blocked publication.
Breaking changes
- [runtime] [cli] Restores ordinary
resumeto native same-ID Smithers continuation and supersedes the short-lived controller re-finalization path, while keeping explicit reset, retry, refresh, ownership, and execution protections. Thanks @aviggiano! (#898, #924, #961) - [config] [topology] Removes the
fullaudit profile and folds its packaged specialist topology intoexhaustive; configurations that still selectfullnow fail with the standard unknown-profile diagnostic. Thanks @aviggiano! (#807) - [config] [runtime] Removes the unused
run.max_parallel_nodeskey andULTRAFUZZ_MAX_PARALLEL_NODESoverride;run.max_parallel_agentsremains the effective concurrency control. Thanks @aviggiano! (#936)
New features
- [runtime] [cli] Upgrades the pinned workflow engine to Smithers 0.35.0 and admits its new stalled states, lifecycle events, warnings, usage fields, migrations, and retry behavior across Ultrafuzz readers and adapters. Thanks @aviggiano! (#996)
- [config] [runtime] Adds Pi's
maxthinking level to model profiles, sealed configuration, and adapter execution while preserving terminal failure semantics. Thanks @aviggiano! (#992) - [benchmarks] [workflows] Adds pre-compute cohort reachability, scheduled benchmark-history freshness monitoring, explicit publication annotations, and opt-in
eval history --max-age-dayschecks so stale history cannot look healthy. Thanks @aviggiano! (#1013) - [prompts] [docs] Adds a generated 67-entry prompt catalog and freshness checks while polishing shipped setup, review, and property prompts. Thanks @aviggiano! (#1015)
- [security] [artifacts] Replaces hand-rolled vendor-secret matching and redaction with pinned Secretlint rules while preserving synchronous, bounded artifact gates. Thanks @aviggiano! (#865)
- [workflows] [tooling] Adds strict typed ESLint, Knip, Size Limit, and SHA-pinned Super-Linter gates and removes the unused code and dependencies they expose. Thanks @aviggiano! (#1005)
- [docs] Adds Contributor Covenant 2.1 to the repository's community documentation. Thanks @aviggiano! (#885)
- [runtime] [security] Seals the trusted CLI and its transitive package graph into a run-owned content-addressed closure, with confined module loading that works under both Node and Bun. Thanks @aviggiano! (#900, #938)
Improvements
- [artifacts] [runtime] Consolidates widely duplicated language helpers and standard-library utilities in a one-off simplification pass without changing validation or security boundaries. Thanks @aviggiano! (#812)
- [runtime] Reduces controller overhead by caching authenticated execution and workflow snapshots, indexing event probes, validating execution files linearly, bounding recovery memory, and preflighting each trusted CLI content address once. Thanks @aviggiano! (#905, #917, #920, #953, #959, #964)
- [runtime] [cli] Makes status and other read-only lifecycle inspection observe already-published evidence without taking the workflow control lock. Thanks @aviggiano! (#957)
- [runtime] [prompts] Replaces model-authored task summaries with a runtime-owned completion marker so artifact verification, not terminal prose or arbitrary JSON, remains the success contract. Thanks @aviggiano! (#947)
- [docs] Strengthens the VPS recommendation and unrestricted-agent warning so operators understand the host-level risk of model and prompt choices. Thanks @aviggiano! (#984)
- [prompts] [artifacts] Clarifies Foundry dependency materialization and removes contradictory goal-plan instructions about runtime-derived counts. Thanks @aviggiano! (#840, #862, #848)
- [workflows] [runtime] Strengthens release validation with cloud-worker and descriptor-confinement probes, CI-safe test budgets, the missing benchmark-history command, split release lanes, pull-request runtime coverage, and repaired post-merge fixtures. Thanks @aviggiano! (#836, #845, #851, #860, #864, #869, #954, #967, #994, #1001, #1029)
- [changelog] Removes stale entries for profiles that were added and removed before release, keeping the Unreleased narrative internally consistent. Thanks @aviggiano! (#884)
- [workspaces] Records the no-op macOS Smithers
jjpostinstall decision so clean pnpm installs no longer stop on anallowBuildsplaceholder. Thanks @mrthankyou! (#1009)
Bug fixes
- [runtime] [security] Redacts lifecycle command payloads while retaining bounded, sanitized process diagnostics and controller-recovery guidance. Thanks @aviggiano! (#800)
- [runtime] [modal] Resolves sealed snapshot entrypoints through descriptor aliases, retained retired paths, and native package resolution without admitting cwd decoys or outside-snapshot modules. Thanks @aviggiano! (#794, #832, #841)
- [runtime] Makes repeated controller refreshes migrate legacy Bun startup controls and predecessor runner patches, recover published preparations, and preserve fresh operator execution authority. Thanks @aviggiano! (#806, #809, #811, #843, #849, #853)
- [runtime] [artifacts] Sources current adapters and modules from the invoking package closure, retains authenticated prompts, refreshes schema bundles, exempts deferred prompts, compiles sealed pre-expansion tasks, and republishes sealed prompt paths for dynamic-run refreshes. Thanks @aviggiano! (#903, #913, #981, #983, #990, #991, #1004)
- [runtime] [workspaces] Bounds and serializes Git-backed workspace capture, excludes runtime, Foundry, dependency, and submodule roots, recovers stale locks, and restores reopened producer workspaces before replay. Thanks @aviggiano! (#825, #827, #829, #830, #834, #912, #950)
- [runtime] Repairs workflow launch and rerender admission by treating the engine database as controller-owned, importing the runtime budget helper, accepting SQL-null optional inputs, retaining source identity, and recovering dependency preparation. Thanks @aviggiano! (#798, #802, #874, #873, #877)
- [artifacts] [runtime] Aligns dynamic dependency, template, attempt, prerequisite, validator-command, and generated-test authority with the compiled and materialized graph while keeping partial or forged evidence fail-closed. Thanks @aviggiano! (#855, #857, #893, #907, #916, #946, #948)
- [artifacts] [reporting] Aligns goal-search coverage, recon-fuzzer paths, bounded-report enrichment, severity fields, and durable runtime accounting so canonical reports remain publishable and retain authoritative metrics. Thanks @aviggiano! (#826, #828, #919, #932, #1007)
- [modal] [benchmarks] Restores smoke and eval publication by using positive-only secret scans for fail-on-hit paths, preserving exact event and workflow inputs, digesting only route-affecting Codex config, and giving contended validators and final-report prompts viable semantics and budgets. Thanks @aviggiano! (#886, #892, #901, #909, #1027)
- [runtime] [cli] Classifies provider HTTP 402 responses as quota exhaustion, parks affected nodes without spending retry attempts, and shows actionable resume guidance. Thanks @aviggiano! (#824)
- [artifacts] Detects atomic schema-path replacement by link count instead of timestamp granularity, closing a race in snapshot reads. Thanks @aviggiano! (#887)
- [runtime] [pi] Preserves text-free Pi terminal answers, streams oversized prompts over stdin, and interprets structured event output without selecting unrelated tool payloads. Thanks @aviggiano! (#890, #896, #928)
- [runtime] Preserves wrapped continuation metadata, reconciles superseded attempt occurrences, passes sealed fork input to preflight, guards active resumes, resolves native dependencies, recovers exact OpenRouter sessions, admits native workflow identities, and scrubs stale local session pointers after restart. Thanks @aviggiano! (#930, #952, #966, #969, #974, #976, #979, #995)
- [runtime] [artifacts] Reopens the producer dependency closure when a zero-retry artifact verifier fails, allowing
resume --retry-failedto regenerate missing or malformed outputs. Thanks @aviggiano! (#926) - [runtime] [cli] Lets
ultrafuzz cleanremove read-only sealed execution snapshots while preserving source references on genuine cleanup failure. Thanks @mrthankyou! (#985) - [runtime] Lets read-only health and status survive expected manifest or checkout divergence and reports fresh divergences even when another control snapshot is reused. Thanks @aviggiano! (#871, #878, #989)
- [prompts] Builds packaged prompt assets into the directory the runtime resolver and package validator actually read. Thanks @aviggiano! (#796)
Full changelog: v0.0.22...v0.0.23
v0.0.22
v0.0.22 brings threat-aware, dynamically planned audits to Ultrafuzz and strengthens the authenticated runtime and cloud lifecycle that keeps long-running campaigns recoverable.
- Threat-driven audit planning. The default workflow can build a threat model, combine it with a pinned vulnerability database, and expand deterministic goal lanes with attributable findings.
- Authenticated controller recovery. Operators can refresh controller-only code for an existing compatible run while preserving sealed inputs, durable identities, and fail-closed recovery checks.
- More dependable cloud execution. Modal handoffs, restored workspaces, runtime-rendered prompts, dependency closures, schema bundles, and retry state remain bound to authenticated release evidence.
New features
- [runtime] [topology] Adds threat-model and goal-planning workflows, dynamic topology fanout, pinned vulnerability-database inputs, opt-in Pi and OpenCode adapters, expanded artifact provenance, a structural threat-model benchmark lane, and source-release package validation. Thanks @aviggiano! (#744)
- [runtime] [cli] Adds
resume --refresh-controllerfor compatible controller-only fixes, rebuilding stock controller files from installed packages while authenticating sealed run semantics and retaining the existing run and workflow identities. Thanks @aviggiano! (#726)
Improvements
- [workflows] Keeps the benchmark-history aggregation gate on
mainpushes and manual release validation while skipping it on pull requests where its prerequisites do not run. Thanks @aviggiano! (#773) - [docs] Refreshes the README's product description, threat-hunting guidance, provider-neutral introduction, VPS recommendation, and getting-started prompt. Thanks @aviggiano! (#790)
Bug fixes
- [runtime] [modal] Makes authenticated controller refresh durable across cloud continuations, interrupted snapshot publication, dynamic base-task reconstruction, sealed module and schema resolution, missing-run recovery, portable snapshot paths, retry recreation, and trusted CLI rotation. Thanks @aviggiano! (#757, #766, #777, #779, #781, #783, #785, #787, #789)
- [modal] [runtime] Preserves canonical cloud handoff and publication identity by deduplicating identical authenticated inputs, repairing controller-local Git metadata, limiting static reconstruction to direct dependencies, retaining immutable dependency evidence, and admitting only authenticated runtime-rendered prompts under the run root. Thanks @aviggiano! (#759, #771, #774, #775, #776)
Full changelog: v0.0.21...v0.0.22
v0.0.21
v0.0.21 strengthens Ultrafuzz's authenticated strategy and reporting chain while making cloud execution, recovery, and status reporting more resilient.
- Trustworthy strategy outputs. Strategy intake and report rows are now bound to sealed authorities, with clearer search roles and stricter lifecycle reconciliation.
- Safer cloud execution. Modal handoffs preserve committed inputs, lifecycle budgets leave room for setup and collection, and archive and download paths are bounded against stalls.
- More dependable recovery. Pending retries resume from terminal state, status remains useful around control drift, and Bun-backed adapter contracts honor conditional skips.
Improvements
- [runtime] [workflows] Adds a source-fingerprinted responsibility audit for shipped agent adapters, moves DeepSeek reasoning effort to the supported runtime option, and requires the boundary contracts in pull-request CI. Thanks @aviggiano! (#706)
Bug fixes
- [runtime] Lets read-only status reporting surface sealed control-file divergence as warnings while execution and lifecycle paths continue to fail closed. Thanks @aviggiano! (#681)
- [runtime] [cli] Recovers terminal runs with pending retry work, avoids duplicate active-workflow submissions, renews deadlines, and reports durable-state lifecycle divergence. Thanks @aviggiano! (#707)
- [config] [topology] Aligns audit-profile retry budgets across expanded agentic nodes, preserves explicit project and topology overrides, and removes the overlapping
thoroughprofile. Thanks @aviggiano! (#708) - [runtime] Makes conditional runtime tests use Bun-compatible skips and gives generated-adapter contracts stable selection and timeout behavior. Thanks @aviggiano! (#709)
- [modal] Preserves committed paths matched by
.gitignorein cloud handoff baselines while continuing to exclude untracked ignored files. Thanks @aviggiano! (#717) - [prompts] [artifacts] Binds strategy ancestor intake and published report rows to authenticated authorities, clarifies strategy search roles, and closes optional-prerequisite, recovery, archive, and report-lifecycle gaps. Thanks @aviggiano! (#621)
- [config] [modal] Preserves the configured inner agent timeout while reserving 30 minutes for the enclosing cloud lifecycle and rejects budgets beyond Modal's maximum. Thanks @aviggiano! (#723)
- [modal] [security] Bounds archive reads, writes, backpressure, and handle cleanup with a no-progress deadline so safe extraction cannot hang or finish partially. Thanks @aviggiano! (#729)
- [modal] Drains Modal result downloads concurrently with remote completion, preventing pipe backpressure deadlocks while preserving atomic publication and failure cleanup. Thanks @aviggiano! (#739)
Full changelog: v0.0.20...v0.0.21
v0.0.20
v0.0.20 hardens Ultrafuzz's execution and artifact boundaries, improves OpenRouter reliability, binds worktrees to immutable launch revisions, and formalizes MIT licensing across the workspace.
- Safer security boundaries. Canonical artifacts now fail closed on detected secrets, while cloud execution, provider environments, archive handling, and governance controls receive defense-in-depth hardening.
- More reliable runs. OpenRouter sessions recover from rate limits without duplicating work, and worktrees stay bound to the exact revision that launched a run.
- Clearer project defaults. The workspace now carries MIT license metadata, a one-hour default agent timeout, and aligned operational documentation.
New features
- [artifacts] [security] Adds a fail-closed publication gate for secret-bearing canonical artifacts, expands credential detection, and makes Kimi credential handling safer. Thanks @aviggiano! (#622)
- [docs] [security] Adds the MIT license, package license metadata across the workspace, licensing documentation, and CI policy checks for release readiness. Thanks @aviggiano! (#657)
Improvements
- [runtime] [security] Hardens governance, cloud execution, provider environment isolation, and runtime-integrity boundaries across high-priority AppSec controls. Thanks @aviggiano! (#635)
- [evals] [security] Bounds evaluation matrices and concurrency, secures remote pricing retrieval against private-address and redirect abuse, and aligns the documented ceilings with the enforced 32-target and 80-row limits. Thanks @aviggiano! (#640, #643)
- [runtime] [modal] Seals governance data used for cloud handoff, strengthens secret isolation, and makes security classification changes fail closed. Thanks @aviggiano! (#644)
- [config] [runtime] Raises the default local and cloud agent timeout from 30 minutes to one hour and documents configuration precedence. Thanks @aviggiano! (#646)
- [runtime] [artifacts] Captures the exact launch revision and binds local, retry, cloud, and workflow worktrees to that immutable source commit. Thanks @aviggiano! (#655)
- [docs] Simplifies the README landing page while retaining the dedicated licensing and security policy documentation. Thanks @aviggiano! (#666)
- [workflows] [runtime] Tiers CI by event so pull requests use a curated runtime smoke suite while pushes to
mainand manual runs retain the full release-validation matrix. Thanks @aviggiano! (#669)
Bug fixes
- [runtime] [openrouter] Recovers safely from initial and mid-run OpenRouter rate limits, resumes the exact session after substantive work, and prevents duplicate or conflicting replay output. Thanks @aviggiano! (#611, #626, #631)
- [cli] [security] Upgrades and hardens benchmark ZIP parsing with bounded archive handling and production dependency-advisory gates. Thanks @aviggiano! (#623)
- [runtime] Declares the generated plain-record validator required by invariant recovery paths and removes the test-only fallback that masked its absence. Thanks @aviggiano! (#668)
Full changelog: v0.0.19...v0.0.20