Releases: ben-vargas/flowition
Release list
v0.7.1 — attempt-scoped Timeline, steered-claude teardown fix, textless-reasoning rendering
Combined release carrying 0.7.0 and 0.7.1 (0.7.0 was never tagged or published; its changelog section is preserved below as part of this release).
0.7.1 — Fixed
- Steered-claude teardown hang (#3). The claude CLI emits one
resultper turn — a mid-turn injected message coalesces into the running turn and never gets its own result — butAgentJobcounted one expected result per injection, sooutstandingstuck above zero, stdin never closed, and CLI (waiting for EOF) deadlocked against engine (waiting for exit) until the 30-minute stall SIGKILL discarded an already-valid result and expensively re-ran the agent. A parsed terminal result now settles the outstanding count and closes stdin immediately; a steer landing after the result queues for a--resumefollow-up turn. - A live-steer process exiting with outstanding messages but a valid result post-dating every injection now has that result accepted (unanswered messages requeued) instead of refused as
truncatedand re-run. True early death still refuses. - Empty-reasoning handling is deliberate and uniform across every stream parser. Claude Code ≥2.1 headless redacts thinking to empty text + signature; parsers now record
{kind:"reasoning", text:"", redacted:true}for final textless blocks and drop empty incremental deltas, keeping the honest record (and the Thinking… liveness signal) without inventing text. - The viewer renders textless reasoning as a compact non-expandable row ("text withheld by the CLI" / "no reasoning text recorded") instead of an expandable block that opens to nothing.
0.7.0 — Added
- Attempt-scoped Timeline in the viewer cockpit: the fold archives each agent's per-attempt view into the closing attempt scope on resume (
AttemptScope.agents, before the round-11 clock clear), so selecting an earlier attempt renders that attempt's real execution bars on its own window — no replay badges for agents that executed there, an explicit absence state where nothing was archived, and no now-line on a closed attempt. Server re-folds deterministically backfill archives for runs recorded before this release.
0.7.0 — Fixed
- The Timeline no longer ignores the lineage attempt selector (previously attempt 1 showed attempt 2's replay ticks).
- Archived attempts derive their capability verdict from their own opening-event engine (
AttemptScope.engine), not the run-level caps a resume overwrites. - Same-millisecond resume boundaries can no longer leak a previous attempt's lane: byte-order participation (
AgentView.inAttempt) is recorded and read ahead of timestamp inference.
Full details in CHANGELOG.md. CI: 8/8 green (root+viewer suites × Node 18.17.1/LTS).
v0.6.0 — grok adapter (Grok Build CLI)
New adapter: grok — orchestrate the Grok Build CLI
agent(prompt, { adapter: 'grok', model: '…', effort: 'xhigh' }) now runs the Grok Build CLI as an eighth real CLI agent.
- Invocation.
claude-streamprotocol. Prompt via a 0600--prompt-file(.md).FLOWITION_GROK_BINoverrides the executable. - Capabilities. Turn steering via
--resume(never create-only-s). Native--json-schema+ streaming-messages-json.--rulesforsystem(appends; do not use--system-prompt-override). - Yolo.
--always-approve --permission-mode bypassPermissions(do not also pass--yolo). - Effort. grok 1.0.3 accepts only
low|medium|high|xhigh. Omitted effort always passes--reasoning-effort high.none/minimal→low;low/medium/high/xhighpass through;max→xhigh. - cwd. Grok does not get
--cwd.AgentJobalreadyspawn()s withspec.cwd, and grok 1.0.3 resolves--cwdagainst process cwd, so a relativeagent({ cwd: 'packages/app' })would double-resolve. - Everywhere it should be.
flowition doctor,flowition guide, MCP enum,index.d.tsunions, viewer GK badge, and the README/ARCHITECTURE/skill docs.
Full changelog: v0.5.0...v0.6.0
v0.5.0 — cursor adapter (Cursor CLI / cursor-agent)
New adapter: cursor — orchestrate the Cursor CLI
agent(prompt, { adapter: 'cursor', model: 'gpt-5.6-sol-xhigh' }) now runs cursor-agent as a seventh real CLI agent, giving workflows access to Cursor's full model catalog (Opus/Sonnet/Fable, GPT-5.x families, Grok, Gemini, Composer, and more — cursor-agent --list-models).
- Invocation. Headless print mode with full permissions:
cursor-agent -p --output-format stream-json --force, prompt delivered on stdin (no argv limits).FLOWITION_CURSOR_BINoverrides the executable. - Capabilities. Turn steering and session resume via
cursor-agent -p --resume <session_id>; schema by prompt contract; provider session ids journaled for cross-run continuation. - New
cursor-jsonlparser. Normalizes Cursor's stream dialect: session capture from the init event, per-delta thinking → reasoning, tool calls paired by wirecall_id(ids can contain literal newlines), generic wrapper-key tool naming (shellToolCall→shell— no hard-coded tool list), camelCase usage tokens, and last-assistant-text result preference (the terminal event'sresultfield concatenates every message of the turn — critical for schema mode). Truncated streams are refused via the standardterminalRequiredgate. - The cursor quirk: effort lives in the model id. cursor has no effort flag — reasoning effort is encoded in the model id itself (
gpt-5.6-sol-xhigh,claude-fable-5-thinking-high, or bracket overrides likeclaude-opus-4-8[effort=high]), soefforton a cursor agent is rejected loudly atagent()time rather than silently dropped. - Everywhere it should be.
flowition doctor,flowition guide, MCP enum,index.d.tsunions, viewer tool-id registry, and the README/ARCHITECTURE/skill docs.
Full changelog: v0.4.0...v0.5.0
v0.4.0 — tailnet viewer access via Tailscale Serve
Viewer: tailnet access via Tailscale Serve — --tailscale-origin
The viewer can now be reached from other machines on your tailnet without ever leaving loopback. The server stays bound to 127.0.0.1; Tailscale Serve terminates tailnet TLS and proxies to the fixed local port:
flowition viewer --port 4646 --tailscale-origin https://machine.tailnet-name.ts.net
tailscale serve --bg 4646The flag teaches the existing request gates about exactly one extra path, and nothing else changes:
- One validated origin.
--tailscale-originrequires an explicit fixed--portand accepts exactly one canonical HTTPS*.ts.netauthority, which joins the closed Host→Origin allowlist. Nonstandard Serve ports work too (https://….ts.net:8443↔tailscale serve --bg --https=8443 4646). - Provenance, byte-exact. Requests for that authority must carry Serve's
X-Forwarded-Proto: httpsexactly; anything marked as public Funnel traffic (Tailscale-Funnel-Request) is refused before the Host gate — usetailscale serve, nevertailscale funnel. - No port walking when the flag is set; the port is part of the policy.
- Same auth everywhere. Tailnet requests need the same bearer token, mutations still need
--controlplus the control token, and all tokens travel in the URL fragment, which never crosses the network. - Rendezvous carries the policy. Reuse matches origin+port after the HMAC challenge; every mismatch — including corrupted records — refuses loudly and never echoes corrupt bytes. Every exit path prints a serve-teardown reminder.
- Without the flag, request handling is byte-for-byte unchanged.
The startup banner prints both the local and the tailnet URL; open the tailnet one from another machine on your tailnet.
Reviewed by Oracle (3 rounds, final ship) and an independent two-model panel (claude-fable-5 high + gpt-5.6-sol xhigh, both ship) via flowition runs flo_17686490 and flo_9f0b65dd. 512/512 tests, zero-deps gate green. New 565-line viewer-tailscale test suite pins the origin validation, provenance gate, Funnel refusal, and rendezvous policy.
Full changelog: v0.3.2...v0.4.0
v0.3.2 — quiescent cache TTL jitter
Viewer: quiescent cache TTL jitter — no more 30-second expiry herd
The run-list cache classifies quiescent runs (stale/unknown/corrupt-result) once and re-derives them on a ~30 s TTL. But a viewer opened on a home full of stale runs classifies all of them in the same cold request — and a single fixed TTL made the whole population expire together, dumping every re-derive on one later poll. At 5,000 runs with 500 stale, that one request spiked to 143.8 ms against the 120 ms P2 budget while its neighbours ran ~65–115 ms. Operators with many stale runs felt the same periodic stutter every 30 seconds; under concurrent load it surfaced as a flaky P2 perf gate.
Each run's quiescent TTL is now deterministically jittered ±25% around 30 s (FNV-1a over the run directory — not Math.random(), so a run keeps the same deadline across requests and restarts, and tests can compute it exactly):
- Any 5 s poll window now inherits roughly a third of the herd at most, instead of all of it.
- The population's amortized ~1/30 s probe rate — the documented contract — is unchanged.
- Resume signals (
.resumingmarker,run.lock) still re-derive immediately, TTL notwithstanding; nothing about resume detection is delayed. - Server-only change:
viewer/distis byte-identical.
Post-fix P2 under the full concurrent root suite: 67.6/87.1/67.4/79.5/91.6 ms — worst sample 91.6 ms, down from the 143.8 ms herd spike.
Pinned by new unit coverage: a cold-classified 500-run herd re-derives as a strict subset per poll and exactly once each; per-entry deadlines hold until their own expiry; jitter is bounded, deterministic, and leaves no dominant 5 s poll window. DESIGN.md §5.4.2/§10 updated from the fixed-30 s contract; MEASUREMENTS.md rebound with the new P2 row. Root suite 485 tests green.
v0.3.1 — cross-run result seeding
Cross-run result seeding — --seed-from <runId>
When you edit a workflow file, resume correctly refuses the changed hash — and until now, every completed agent in the old run had to be re-run from scratch. Their results were sitting in the old run's journal. Now:
flowition run edited.workflow.js --seed-from flo_a1b2c3d4- Loads the settled source run's journal read-only (never repaired) and reuses its final completed agent results as a candidate cache. Derived agent keys hash branch position + prompt + canonical keyed spec — no run id, seed, or file hash — so unchanged calls match across runs while edited calls miss and run fresh.
- Weaker than resume, on purpose: key equality identifies the same call shape, not the same world. It's operator-authorized cache reuse for research/pure-result agents, never a freshness guarantee.
- Never seeded:
step()results (a completed step proves a side effect in the old world), explicit-keyagents (an explicit key matches even a rewritten call),ask()answers, provider sessions, usage baselines, and any source key that ever accepted steering mail. Target-side steering around a seeded call drops with a warning, like same-run cache replay. - Durable: a hit is written into the target journal (with
seededprovenance: source run id, index, usage) before it's returned — the target resumes normally even after the source run is deleted. Chained seeding (source → target → target) names the immediate source but keeps surfacing the original provider cost. - Zero double-billing: seeded hits add nothing to the new run's provider spend or
--budget; imported source usage is exposed separately as provenance. - Refused loudly for missing/live/corrupt/self/key-version-mismatched sources, and cannot combine with
--resume(fresh runs only). Detached runs prevalidate the source before spawning. - Surfaces everywhere:
flowition run --seed-from, MCPflowition_runseedFrom,seeded from run …in status output, and aseededFromannotation in the viewer's shared fold.
Fit zoom fix — the timeline no longer scrolls at "Fit"
Fit zoom promised no horizontal scrollbar but grew one on any run with a real span: the trailing duration/replay tag deliberately hangs past the bar it describes, so a lane ending at ~100% of the track pushed its tag beyond the edge. The timeline plot now reserves a 130px trailing meta gutter — the bar and its tag always land inside the viewport at Fit; fixed zooms (200%, 400%, …) keep scrolling by design. Pinned by a Playwright regression that drives a real-span run in a real browser.
Both changes ship with full evidence: 11 acceptance tests for seeding, a browser regression for Fit proven to bite (fails with the gutter removed), regenerated §3.7 screenshot captures, and rebound perf fingerprints. Root suite 484 tests, viewer suite 1137, Playwright 21.
v0.3.0 — input contracts + durable steps
Workflow input contracts — meta.argsSchema
A workflow module can now declare a JSON Schema for its --args:
export const meta = {
argsSchema: {
type: 'object',
required: ['target'],
properties: { target: { type: 'string', minLength: 1 } },
additionalProperties: false,
},
}- Validates the effective args verbatim at admission time, on fresh runs and resumes alike — a violation terminates the run before any agent or step executes.
- Failures always produce full terminal artifacts (journal
endrecord,result.json, no agent events) — including malformed schemas, unsupported keywords, and validator throws. Nothing escapes as a bare crash. - The validator's supported subset (
type,required,properties,additionalProperties: false,items,enum,const, min/max bounds,anyOf) rejects anything it doesn't enforce loudly — no silently ignored constraints.
Durable side-effect nodes — step(name, args?, fn)
Journaled local work (git commands, file writes, API calls) with replay-on-resume:
const sha = await step('commit', { msg }, async () => {
await exec(`git commit -m "${msg}"`)
return exec('git rev-parse HEAD')
})- A completed callback's JSON result is journaled and replayed on resume without re-executing; failed or crash-window (start-only) attempts re-run.
- The contract is durable memoization, not exactly-once — callbacks should be idempotent or carry idempotency keys.
- Steps use an independent per-branch counter and a separate key domain, so adding or removing a
step()never shifts agent resume keys (existing journals stay resumable; key version unchanged). - The completed journal record seals the outcome: telemetry failures after completion can never cause a re-run.
- Strict JSON normalization for args/results:
undefined, functions,BigInt,NaN, cycles, and sparse arrays are rejected loudly;-0normalizes to0; a void callback resolves tonull.
Observability
Steps are first-class everywhere: CLI status/tail, the MCP surface, and the web viewer cockpit (step lifecycle, replay markers, failures, durations).
Docs
New examples/durable-steps.workflow.js, plus README, guide, architecture, skill, and TypeScript declaration updates for both features.
v0.2.0 — web observability + control cockpit
Web viewer: observability + control cockpit
flowition view now serves a browser cockpit for watching and steering runs:
- Live run cockpit — phases, agents, and logs folded from the event journal in real time, with per-agent output previews, token/cost usage, and stall indicators.
- Run list & detail views — every run under
~/.flowitionbrowsable with status, duration, and result; stale/dead runs detected and marked. - Control surface — answer
ask()questions and steer live agents (sendTo) from the browser, gated behind a control token (--control). - Zero new runtime dependencies — the SPA ships prebuilt in
viewer/dist; the server is plainnode:httpwith strict CSP, single-decode path validation, and 0600 token files.
Authoring surface
- TypeScript declarations (
index.d.ts) for the whole workflow authoring surface — toolkit, agent options, adapters, schema types. - flowition agent skill (
skills/flowition/SKILL.md) so coding agents can author and operate workflows correctly. - Clearer error when a workflow file fails to parse as ESM from a CommonJS scope.
Hardening
- Favicon CSP compliance and first-contact Linux portability fixes.
- CI walkthrough stabilized (steady-state windows, memory headroom, isolation); the live-session walkthrough is a local release gate.