Skip to content

Releases: ben-vargas/flowition

v0.7.1 — attempt-scoped Timeline, steered-claude teardown fix, textless-reasoning rendering

Choose a tag to compare

@ben-vargas ben-vargas released this 19 Aug 06:20
1ebdac2

Combined release carrying 0.7.0 and 0.7.1 (0.7.0 was never tagged or published; its changelog section is preserved below as part of this release).

0.7.1 — Fixed

  • Steered-claude teardown hang (#3). The claude CLI emits one result per turn — a mid-turn injected message coalesces into the running turn and never gets its own result — but AgentJob counted one expected result per injection, so outstanding stuck above zero, stdin never closed, and CLI (waiting for EOF) deadlocked against engine (waiting for exit) until the 30-minute stall SIGKILL discarded an already-valid result and expensively re-ran the agent. A parsed terminal result now settles the outstanding count and closes stdin immediately; a steer landing after the result queues for a --resume follow-up turn.
  • A live-steer process exiting with outstanding messages but a valid result post-dating every injection now has that result accepted (unanswered messages requeued) instead of refused as truncated and re-run. True early death still refuses.
  • Empty-reasoning handling is deliberate and uniform across every stream parser. Claude Code ≥2.1 headless redacts thinking to empty text + signature; parsers now record {kind:"reasoning", text:"", redacted:true} for final textless blocks and drop empty incremental deltas, keeping the honest record (and the Thinking… liveness signal) without inventing text.
  • The viewer renders textless reasoning as a compact non-expandable row ("text withheld by the CLI" / "no reasoning text recorded") instead of an expandable block that opens to nothing.

0.7.0 — Added

  • Attempt-scoped Timeline in the viewer cockpit: the fold archives each agent's per-attempt view into the closing attempt scope on resume (AttemptScope.agents, before the round-11 clock clear), so selecting an earlier attempt renders that attempt's real execution bars on its own window — no replay badges for agents that executed there, an explicit absence state where nothing was archived, and no now-line on a closed attempt. Server re-folds deterministically backfill archives for runs recorded before this release.

0.7.0 — Fixed

  • The Timeline no longer ignores the lineage attempt selector (previously attempt 1 showed attempt 2's replay ticks).
  • Archived attempts derive their capability verdict from their own opening-event engine (AttemptScope.engine), not the run-level caps a resume overwrites.
  • Same-millisecond resume boundaries can no longer leak a previous attempt's lane: byte-order participation (AgentView.inAttempt) is recorded and read ahead of timestamp inference.

Full details in CHANGELOG.md. CI: 8/8 green (root+viewer suites × Node 18.17.1/LTS).

v0.6.0 — grok adapter (Grok Build CLI)

Choose a tag to compare

@ben-vargas ben-vargas released this 14 Aug 00:57
18f4d5c

New adapter: grok — orchestrate the Grok Build CLI

agent(prompt, { adapter: 'grok', model: '…', effort: 'xhigh' }) now runs the Grok Build CLI as an eighth real CLI agent.

  • Invocation. claude-stream protocol. Prompt via a 0600 --prompt-file (.md). FLOWITION_GROK_BIN overrides the executable.
  • Capabilities. Turn steering via --resume (never create-only -s). Native --json-schema + streaming-messages-json. --rules for system (appends; do not use --system-prompt-override).
  • Yolo. --always-approve --permission-mode bypassPermissions (do not also pass --yolo).
  • Effort. grok 1.0.3 accepts only low|medium|high|xhigh. Omitted effort always passes --reasoning-effort high. none/minimal → low; low/medium/high/xhigh pass through; max → xhigh.
  • cwd. Grok does not get --cwd. AgentJob already spawn()s with spec.cwd, and grok 1.0.3 resolves --cwd against process cwd, so a relative agent({ cwd: 'packages/app' }) would double-resolve.
  • Everywhere it should be. flowition doctor, flowition guide, MCP enum, index.d.ts unions, viewer GK badge, and the README/ARCHITECTURE/skill docs.

Full changelog: v0.5.0...v0.6.0

v0.5.0 — cursor adapter (Cursor CLI / cursor-agent)

Choose a tag to compare

@ben-vargas ben-vargas released this 13 Aug 01:45

New adapter: cursor — orchestrate the Cursor CLI

agent(prompt, { adapter: 'cursor', model: 'gpt-5.6-sol-xhigh' }) now runs cursor-agent as a seventh real CLI agent, giving workflows access to Cursor's full model catalog (Opus/Sonnet/Fable, GPT-5.x families, Grok, Gemini, Composer, and more — cursor-agent --list-models).

  • Invocation. Headless print mode with full permissions: cursor-agent -p --output-format stream-json --force, prompt delivered on stdin (no argv limits). FLOWITION_CURSOR_BIN overrides the executable.
  • Capabilities. Turn steering and session resume via cursor-agent -p --resume <session_id>; schema by prompt contract; provider session ids journaled for cross-run continuation.
  • New cursor-jsonl parser. Normalizes Cursor's stream dialect: session capture from the init event, per-delta thinking → reasoning, tool calls paired by wire call_id (ids can contain literal newlines), generic wrapper-key tool naming (shellToolCall → shell — no hard-coded tool list), camelCase usage tokens, and last-assistant-text result preference (the terminal event's result field concatenates every message of the turn — critical for schema mode). Truncated streams are refused via the standard terminalRequired gate.
  • The cursor quirk: effort lives in the model id. cursor has no effort flag — reasoning effort is encoded in the model id itself (gpt-5.6-sol-xhigh, claude-fable-5-thinking-high, or bracket overrides like claude-opus-4-8[effort=high]), so effort on a cursor agent is rejected loudly at agent() time rather than silently dropped.
  • Everywhere it should be. flowition doctor, flowition guide, MCP enum, index.d.ts unions, viewer tool-id registry, and the README/ARCHITECTURE/skill docs.

Full changelog: v0.4.0...v0.5.0

v0.4.0 — tailnet viewer access via Tailscale Serve

Choose a tag to compare

@ben-vargas ben-vargas released this 08 Aug 08:36

Viewer: tailnet access via Tailscale Serve — --tailscale-origin

The viewer can now be reached from other machines on your tailnet without ever leaving loopback. The server stays bound to 127.0.0.1; Tailscale Serve terminates tailnet TLS and proxies to the fixed local port:

flowition viewer --port 4646 --tailscale-origin https://machine.tailnet-name.ts.net
tailscale serve --bg 4646

The flag teaches the existing request gates about exactly one extra path, and nothing else changes:

  • One validated origin. --tailscale-origin requires an explicit fixed --port and accepts exactly one canonical HTTPS *.ts.net authority, which joins the closed Host→Origin allowlist. Nonstandard Serve ports work too (https://….ts.net:8443 ↔ tailscale serve --bg --https=8443 4646).
  • Provenance, byte-exact. Requests for that authority must carry Serve's X-Forwarded-Proto: https exactly; anything marked as public Funnel traffic (Tailscale-Funnel-Request) is refused before the Host gate — use tailscale serve, never tailscale funnel.
  • No port walking when the flag is set; the port is part of the policy.
  • Same auth everywhere. Tailnet requests need the same bearer token, mutations still need --control plus the control token, and all tokens travel in the URL fragment, which never crosses the network.
  • Rendezvous carries the policy. Reuse matches origin+port after the HMAC challenge; every mismatch — including corrupted records — refuses loudly and never echoes corrupt bytes. Every exit path prints a serve-teardown reminder.
  • Without the flag, request handling is byte-for-byte unchanged.

The startup banner prints both the local and the tailnet URL; open the tailnet one from another machine on your tailnet.

Reviewed by Oracle (3 rounds, final ship) and an independent two-model panel (claude-fable-5 high + gpt-5.6-sol xhigh, both ship) via flowition runs flo_17686490 and flo_9f0b65dd. 512/512 tests, zero-deps gate green. New 565-line viewer-tailscale test suite pins the origin validation, provenance gate, Funnel refusal, and rendezvous policy.

Full changelog: v0.3.2...v0.4.0

v0.3.2 — quiescent cache TTL jitter

Choose a tag to compare

@ben-vargas ben-vargas released this 05 Aug 15:04

Viewer: quiescent cache TTL jitter — no more 30-second expiry herd

The run-list cache classifies quiescent runs (stale/unknown/corrupt-result) once and re-derives them on a ~30 s TTL. But a viewer opened on a home full of stale runs classifies all of them in the same cold request — and a single fixed TTL made the whole population expire together, dumping every re-derive on one later poll. At 5,000 runs with 500 stale, that one request spiked to 143.8 ms against the 120 ms P2 budget while its neighbours ran ~65–115 ms. Operators with many stale runs felt the same periodic stutter every 30 seconds; under concurrent load it surfaced as a flaky P2 perf gate.

Each run's quiescent TTL is now deterministically jittered ±25% around 30 s (FNV-1a over the run directory — not Math.random(), so a run keeps the same deadline across requests and restarts, and tests can compute it exactly):

  • Any 5 s poll window now inherits roughly a third of the herd at most, instead of all of it.
  • The population's amortized ~1/30 s probe rate — the documented contract — is unchanged.
  • Resume signals (.resuming marker, run.lock) still re-derive immediately, TTL notwithstanding; nothing about resume detection is delayed.
  • Server-only change: viewer/dist is byte-identical.

Post-fix P2 under the full concurrent root suite: 67.6/87.1/67.4/79.5/91.6 ms — worst sample 91.6 ms, down from the 143.8 ms herd spike.

Pinned by new unit coverage: a cold-classified 500-run herd re-derives as a strict subset per poll and exactly once each; per-entry deadlines hold until their own expiry; jitter is bounded, deterministic, and leaves no dominant 5 s poll window. DESIGN.md §5.4.2/§10 updated from the fixed-30 s contract; MEASUREMENTS.md rebound with the new P2 row. Root suite 485 tests green.

v0.3.1 — cross-run result seeding

Choose a tag to compare

@ben-vargas ben-vargas released this 05 Aug 04:31

Cross-run result seeding — --seed-from <runId>

When you edit a workflow file, resume correctly refuses the changed hash — and until now, every completed agent in the old run had to be re-run from scratch. Their results were sitting in the old run's journal. Now:

flowition run edited.workflow.js --seed-from flo_a1b2c3d4
  • Loads the settled source run's journal read-only (never repaired) and reuses its final completed agent results as a candidate cache. Derived agent keys hash branch position + prompt + canonical keyed spec — no run id, seed, or file hash — so unchanged calls match across runs while edited calls miss and run fresh.
  • Weaker than resume, on purpose: key equality identifies the same call shape, not the same world. It's operator-authorized cache reuse for research/pure-result agents, never a freshness guarantee.
  • Never seeded: step() results (a completed step proves a side effect in the old world), explicit-key agents (an explicit key matches even a rewritten call), ask() answers, provider sessions, usage baselines, and any source key that ever accepted steering mail. Target-side steering around a seeded call drops with a warning, like same-run cache replay.
  • Durable: a hit is written into the target journal (with seeded provenance: source run id, index, usage) before it's returned — the target resumes normally even after the source run is deleted. Chained seeding (source → target → target) names the immediate source but keeps surfacing the original provider cost.
  • Zero double-billing: seeded hits add nothing to the new run's provider spend or --budget; imported source usage is exposed separately as provenance.
  • Refused loudly for missing/live/corrupt/self/key-version-mismatched sources, and cannot combine with --resume (fresh runs only). Detached runs prevalidate the source before spawning.
  • Surfaces everywhere: flowition run --seed-from, MCP flowition_run seedFrom, seeded from run … in status output, and a seededFrom annotation in the viewer's shared fold.

Fit zoom fix — the timeline no longer scrolls at "Fit"

Fit zoom promised no horizontal scrollbar but grew one on any run with a real span: the trailing duration/replay tag deliberately hangs past the bar it describes, so a lane ending at ~100% of the track pushed its tag beyond the edge. The timeline plot now reserves a 130px trailing meta gutter — the bar and its tag always land inside the viewport at Fit; fixed zooms (200%, 400%, …) keep scrolling by design. Pinned by a Playwright regression that drives a real-span run in a real browser.


Both changes ship with full evidence: 11 acceptance tests for seeding, a browser regression for Fit proven to bite (fails with the gutter removed), regenerated §3.7 screenshot captures, and rebound perf fingerprints. Root suite 484 tests, viewer suite 1137, Playwright 21.

v0.3.0 — input contracts + durable steps

Choose a tag to compare

@ben-vargas ben-vargas released this 04 Aug 20:13

Workflow input contracts — meta.argsSchema

A workflow module can now declare a JSON Schema for its --args:

export const meta = {
  argsSchema: {
    type: 'object',
    required: ['target'],
    properties: { target: { type: 'string', minLength: 1 } },
    additionalProperties: false,
  },
}
  • Validates the effective args verbatim at admission time, on fresh runs and resumes alike — a violation terminates the run before any agent or step executes.
  • Failures always produce full terminal artifacts (journal end record, result.json, no agent events) — including malformed schemas, unsupported keywords, and validator throws. Nothing escapes as a bare crash.
  • The validator's supported subset (type, required, properties, additionalProperties: false, items, enum, const, min/max bounds, anyOf) rejects anything it doesn't enforce loudly — no silently ignored constraints.

Durable side-effect nodes — step(name, args?, fn)

Journaled local work (git commands, file writes, API calls) with replay-on-resume:

const sha = await step('commit', { msg }, async () => {
  await exec(`git commit -m "${msg}"`)
  return exec('git rev-parse HEAD')
})
  • A completed callback's JSON result is journaled and replayed on resume without re-executing; failed or crash-window (start-only) attempts re-run.
  • The contract is durable memoization, not exactly-once — callbacks should be idempotent or carry idempotency keys.
  • Steps use an independent per-branch counter and a separate key domain, so adding or removing a step() never shifts agent resume keys (existing journals stay resumable; key version unchanged).
  • The completed journal record seals the outcome: telemetry failures after completion can never cause a re-run.
  • Strict JSON normalization for args/results: undefined, functions, BigInt, NaN, cycles, and sparse arrays are rejected loudly; -0 normalizes to 0; a void callback resolves to null.

Observability

Steps are first-class everywhere: CLI status/tail, the MCP surface, and the web viewer cockpit (step lifecycle, replay markers, failures, durations).

Docs

New examples/durable-steps.workflow.js, plus README, guide, architecture, skill, and TypeScript declaration updates for both features.

v0.2.0 — web observability + control cockpit

Choose a tag to compare

@ben-vargas ben-vargas released this 04 Aug 20:13

Web viewer: observability + control cockpit

flowition view now serves a browser cockpit for watching and steering runs:

  • Live run cockpit — phases, agents, and logs folded from the event journal in real time, with per-agent output previews, token/cost usage, and stall indicators.
  • Run list & detail views — every run under ~/.flowition browsable with status, duration, and result; stale/dead runs detected and marked.
  • Control surface — answer ask() questions and steer live agents (sendTo) from the browser, gated behind a control token (--control).
  • Zero new runtime dependencies — the SPA ships prebuilt in viewer/dist; the server is plain node:http with strict CSP, single-decode path validation, and 0600 token files.

Authoring surface

  • TypeScript declarations (index.d.ts) for the whole workflow authoring surface — toolkit, agent options, adapters, schema types.
  • flowition agent skill (skills/flowition/SKILL.md) so coding agents can author and operate workflows correctly.
  • Clearer error when a workflow file fails to parse as ESM from a CommonJS scope.

Hardening

  • Favicon CSP compliance and first-contact Linux portability fixes.
  • CI walkthrough stabilized (steady-state windows, memory headroom, isolation); the live-session walkthrough is a local release gate.