Skip to content

v0.0.24

Choose a tag to compare

@aviggiano aviggiano released this 08 Sep 22:21

v0.0.24 focuses on more reliable campaign continuation and report publication. It strengthens continuation, retry history, workspace preparation, and live diagnostics, improves macOS startup, and moves the shipped vulnerability reference to pinned OWASP Smart Contract Security data.

  • More reliable continuation. Resumed and reset campaigns retain their original policy, task identities, failed-attempt history, and coherent workspace evidence.
  • Dependable results and diagnostics. Reports survive controller restarts and public redaction, while status and worker failures preserve the evidence operators need.
  • Pinned OWASP references. The new reference adapter validates source provenance and derived planning data; existing configurations must migrate from the former database format.

Breaking changes

  • [references] [runtime] Replaces the former vulnerability database with commit-pinned OWASP Smart Contract Security records and a digest-verified planner catalog. Existing projects must review reference, topology, setup-prompt and schema changes, sync vulnerability-database.owasp-scs, and start a fresh run; retain the previous CLI for historical database artifacts. The former database formats are no longer accepted. Optional SCSVS hints cannot exclude weaknesses, so this pin plans 156 class goals plus threat goals and roaming. See the migration guide. Thanks @aviggiano! (#1118)

New features

  • [runtime] Enables campaign submission on macOS with inode-rooted snapshot paths that preserve file identity across renames and symlink replacement. Thanks @mrthankyou! (#1016)

Improvements

  • [docs] Places disposable, isolated-host guidance and the unrestricted-agent warning before the first setup and campaign commands. Thanks @mrthankyou! (#1023)
  • [runtime] [cli] Adds an opt-in synchronization deadline for status, inspection, stats, and diagnostics, with a stale-snapshot warning when the deadline expires. Observers continue to wait for complete synchronization by default. Thanks @aviggiano! (#1059, #1089)
  • [workflows] [benchmarks] Removes the age-only benchmark-history monitor that opened outage issues when no benchmark-triggering changes had occurred. Thanks @aviggiano! (#1112)
  • [benchmarks] Refreshes published benchmark history and its documentation charts through the automated publisher. These automated commits have no associated pull requests: e876e555 and c1229d04.

Bug fixes

  • [runtime] [modal] Creates temporary roots beneath a canonical system directory, fixing macOS path validation in runtime, cloud, reference, and packaged-install operations, and cleans up read-only runtime test snapshots. Thanks @mrthankyou! (#1019, #1025)
  • [runtime] Matches descriptor-backed macOS workflow files by device and inode so valid snapshot aliases retain their workflow identity. Thanks @aviggiano! (#1032)
  • [workflows] [security] Marks known synthetic credential fixtures so secret-scanning checks stop rejecting their intentional test values while continuing to scan surrounding code. Thanks @mrthankyou! (#1021)
  • [modal] [cli] Aligns report and publication fixtures with authenticated detection provenance, restoring the intended cloud and CLI validation checks. Thanks @aviggiano! (#1033)
  • [runtime] [modal] Passes the engine-owned task runtime directly into cloud workflows so final reports can read durable workflow metrics across separately loaded package copies. Thanks @aviggiano! (#1034)
  • [runtime] Preserves the admitted executable PATH and isolated Pi home through detached continuations so the generated adapter can find Pi and retain its approved environment. Thanks @aviggiano! (#1037)
  • [cli] Projects only documented public metadata into inspect results, preventing persisted source provenance or private fields from invalidating the CLI response. Thanks @aviggiano! (#1040)
  • [runtime] Keeps detached-resume startup helpers within their module scope and migrates the previous generated patch during controller refresh. Thanks @aviggiano! (#1042)
  • [runtime] Normalizes Pi reasoning-token breakdowns without reducing the provider-reported output total. Thanks @aviggiano! (#1044)
  • [runtime] Retains a process-owned copy of inherited execution-snapshot descriptors so later engine, supervisor, and resume handoffs survive closure or reuse of the startup descriptor. Thanks @aviggiano! (#1046)
  • [prompts] Makes final-report producers preserve dedupe-owned descriptions and final lifecycle severity, preventing complete smoke runs from failing report publication on rewritten authority. Thanks @aviggiano! (#1047)
  • [runtime] Restores missing static presentation prompts from verified immutable plan snapshots before resume or reset, preserving completed nodes. Thanks @aviggiano! (#1049)
  • [cli] Marks stats results unsuccessful when synchronization reports an error while retaining the statistics snapshot and original diagnostics. Thanks @aviggiano! (#1051)
  • [runtime] Retries recognized atomic-replacement races when observing live run documents, while continuing to reject stable integrity failures. Thanks @aviggiano! (#1055)
  • [runtime] Retries partially written final event records with a bounded coherent-snapshot budget; stable malformed tails and malformed earlier records still fail. Thanks @aviggiano! (#1057)
  • [runtime] Keeps lifecycle event queries and watchers readable during authenticated dynamic retries and exposes control divergence as an observation warning. Thanks @aviggiano! (#1062)
  • [runtime] Archives complete dynamic-expansion generations before producer retry, retaining attempt artifacts and workspace evidence while safely releasing stale worktree registrations. Thanks @aviggiano! (#1064)
  • [runtime] Records workflow failures before agent selection without inventing executed model attempts or immutable attempt provenance. Thanks @aviggiano! (#1066)
  • [runtime] Repairs absent, prunable worktree registrations owned by the current run before ordinary resume, preserving live, locked, and unrelated registrations. Thanks @aviggiano! (#1070)
  • [runtime] Recognizes a reset node’s retained current attempt even when its attempt number is lower than a previous failed attempt. Thanks @aviggiano! (#1074)
  • [cli] Includes authenticated dynamically generated tasks in stats without admitting unrelated task additions or dropping the sealed baseline. Thanks @aviggiano! (#1076)
  • [cli] [runtime] Keeps run listings available when one run directory cannot be read, showing the unreadable run and its diagnostic so operators can identify and clean it. Thanks @mrthankyou! (#1080)
  • [runtime] Accepts blank lines inside invariant-ledger and canonical-property Markdown blocks while retaining exact field, ordering, and JSON parity checks. Thanks @mrthankyou! (#1083)
  • [dependencies] Updates the transitive fast-uri resolution to 3.1.7, addressing four high-severity advisories without adding dependency overrides. Thanks @aviggiano! (#1085)
  • [runtime] Corrects the exhausted-read-race regression to expect an unreadable run listing and warning, matching the combined live-read and listing behavior. Thanks @aviggiano! (#1088)
  • [workflows] Retries transient registry failures in the dependency-advisory gate and runs deterministic checks first, so temporary registry outages do not hide formatting, lint, or build failures. Thanks @aviggiano! (#1090)
  • [modal] Retains bounded, sanitized eval-score and eval-report failure envelopes in public worker diagnostics, including the underlying error code and message. Thanks @aviggiano! (#1092)
  • [cli] [evals] Preserves eval error codes and bounded, redacted details in score and report failure messages without changing the CLI envelope contract. Thanks @aviggiano! (#1093)
  • [evals] Uses an OpenAI-compatible strict judge response schema and bounded retries for transient or invalid judge responses, while retaining full local result validation. Thanks @aviggiano! (#1094)
  • [artifacts] [runtime] Allows selected nonessential artifact metadata omissions with attributable warnings, preserves authenticated finding identity through review, and carries warnings into reports and public companions. Missing substantive evidence and conflicting identities remain fatal. Thanks @aviggiano! (#1095)
  • [modal] Names and logs otherwise unhandled worker and public-bundle failures, preserving bounded, redacted causes instead of reporting unexplained sandbox exits. Thanks @aviggiano! (#1096)
  • [runtime] [cli] Reports an incomplete launch without inventing workflow identity or launcher liveness, allowing read-only diagnosis while execution still requires valid control evidence. Thanks @mrthankyou! (#1097)
  • [runtime] Makes public final-report redaction stable across repeated publication checks and preserves valid run identifiers without weakening positive credential detection. Thanks @aviggiano! (#1098)
  • [runtime] Collapses only duplicate string values created by public-report redaction so valid reports continue to satisfy unique-array constraints. Thanks @aviggiano! (#1107)
  • [artifacts] Uses the dedicated authenticated lifecycle ledger as the final report’s lifecycle authority, avoiding contradictory requirements from stale embedded copies. Thanks @aviggiano! (#1103)
  • [runtime] Preserves run, logical-node, and Smithers-node identity in initial and reconstructed task specifications so fresh verifiers can reconcile recorded producer attempts. Thanks @aviggiano! (#1104, #1113)
  • [runtime] Recovers final-report retry provenance from durable runner history after controller restart and keeps physical attempt numbers distinct from retry-chain positions. Thanks @aviggiano! (#1114)
  • [runtime] Regenerates preparation, invariant, and workspace evidence coherently after authenticated linear dependency replacement, with interruption-safe recovery and archived original bytes. Changed legacy preparations without protected provenance still require a fresh run. Thanks @aviggiano! (#1115)
  • [runtime] Preserves OpenRouter stdout event order so terminal failures quarantine trailing output even within one chunk, while retaining valid earlier output and resuming the exact session. Thanks @aviggiano! (#1116)
  • [runtime] Restores the original authenticated launch policy for native continuation and preserves recordable failed executor attempts before supported stopped-run resets. Repeated status and stats retain failed and successful attempts even when reset reuses a physical attempt number; already lost historical authority is not reconstructed. Thanks @aviggiano! (#1117)

Full changelog: v0.0.23...v0.0.24