Skip to content

v5.15.0 — what a run loses when it stops and starts again

Latest

Choose a tag to compare

@Masterplanner25 Masterplanner25 released this 25 Sep 05:20
· 1 commit to main since this release
Immutable release. Only release title and notes can be modified.
d3e9c07

What a run loses when it stops and starts again.

A workflow that parks at workflow_wait and resumes in another process crosses a
boundary, and on the far side of it the runtime had been quietly dropping things.
Seven entries — five fixes and two changes — and every one was reported by a
project building on Nodus rather than found here.

#868 a derived VM inherits the host state it works for — twelve attributes across four derivation sites, not the two reported. Plus a module function's agent_call reaching the process-global registry, undoing per-tenant isolation at every module boundary
#869 a run parked longer than the store's 30-day scan bound stopped existing — invisible to nodus workflow runs, to the sweep that would have resumed it, to the migration that would have carried it, and to the warning whose job is to say these runs will be stranded
#870 reject-and-revise works again. #482 refused it on the stated grounds that the payload was "silently discarded"; measured on 4.0.8, it reached the replayed step
#871 rehydrated data comes back in the order it was built, so anything derived from a step result — a draft, a hash, an @exactly_once key — no longer differs after a restart
#873 a resume inherits the caller's bounds. Five were lost, so a guest escaped its instruction budget and its deadline by parking: measured, 8 resumes and ~800 ms inside a 300 ms budget, and ~270,000 instructions inside 5,000
#875 max_terminal_runs can see the records it exists to delete. A cap of 2 left 6 files, and the survivors were the oldest

Three things that stopped working

  • Two or more resume_workflow calls under the nodus run default now time
    out
    (#873). A resume costs ~99 ms against the 200 ms default, so one still
    fits and two do not. --time-limit N (seconds) is the fix, as it already is
    for anything else needing more than 200 ms.
  • max_terminal_runs deletes more than it used to (#875). Opt-in — the
    default is None, an unset cap still deletes nothing, and live runs are
    untouched at any age. Unset the cap to keep the old retention.
  • resume_workflow(id, "checkpoint", {payload}) replays instead of being
    refused
    (#870). A restoration, not a restriction; nothing can have depended
    on an error. resume_workflow(id, "checkpoint") with no payload is still
    refused, which is the part #482 earned.

New API

WorkflowStore.list_all_runs() — every record with no bound applied, which is
what a migration must enumerate. Concrete, not abstract, so no out-of-tree
store breaks at construction.

Validation

  • CI green on the release PR, both unittest discover and pytest
  • Gate 10a: all 8 dependent suites pass, 953 tests. It went red twice first —
    once on a checkout that had moved, once on a real version desync in nodus-mcp
    that had nothing to do with this release. Both fixed before publishing
  • Gate 10b: 136/136 probes against the wheel in a clean venv, resolved from
    site-packages under --require-installed
  • Stage 5: 8/8 through the published package

Full detail: CHANGELOG ·
eval record

pip install nodus-lang==5.15.0