Releases: beastoin/agent-flow
Release list
v0.5.11: fix stale HTML report on re-push
Bug fix: findRunFile returned the oldest timestamped report.html instead of the latest. When report was run twice (e.g., after adding backend logs), push uploaded the stale first HTML while run.json had the current data.
Root cause: readdirSync().find() returns the first match. Timestamped files sort chronologically, so the first match is the oldest.
Fix: Changed to filter().sort() and return the last (latest) match.
Test: 5 new tests for findRunFile (333 total).
Fixes the L3c3ZDTrZt report where HTML showed 651 entries (app only) but JSON had 1000 entries (app + backend + step).
v0.5.5: log timeline pipeline + timestamp-based file naming
What's new
Log timeline pipeline
- Log parser (
src/log-parser.ts): auto-extract timestamps from any common format (ISO 8601, Python comma-ms, logcat, space-separated), detect log levels - Timeline in HTML reports: auto-discover
.logfiles in run directory → parse → filter to run time window → correlate with steps → render chronological timeline with source badges and clickable citations to raw log lines - Raw log viewer: collapsible sections with line numbers and anchor IDs for citation linking
Timestamp-based file naming
- All run files now get a compact timestamp prefix:
20260330T100001200Z-events.jsonl,20260330T100005000Z-run.json, etc. run.meta.jsonis the only fixed-name file — serves as directory index storing all timestamped filenamesfindRunFile()resolver (src/run-files.ts): backwards-compatible lookup that checks fixed name first, then scans for timestamped variant
Stats
- 322 tests (18 new)
- 18 files changed, +725 / -37 lines
v0.5.2
Fix 15 design principle violations found by codex audit
Docs: update CLAUDE.md command list (8 not 6), fix README command table, event types, run.json example, architecture. Remove stale references to run/migrate commands and non-existent files.
Code: fix exit codes (no-subcommand=2, unverified=2), add input validation for all --run-dir inputs, remove raw ADB instructions from schema, fix error hint text.
v0.5.1
Bug fix: step outcome ignoring tier verification results
Step outcome was derived solely from step.end event outcome field (defaulting to fail when absent). Two-tier verification results were computed but never fed back into step outcome. Now: tier failures override step.end pass, and when step.end has no explicit outcome, tiers determine the result. 291 tests pass (3 new).
v0.5.0: Agent-Friendly Verification
What's new
Honest results — Reports no longer show misleading green PASS when nothing was verified. New unverified result when all automated checks have no evidence and all agent reviews are pending. Audit mode shows AUDIT badge instead of PASS.
Step recipes — record init --flow <yaml> now returns per-step event recipes telling agents exactly what to stream:
{"recipe": [{"id":"S1","events":["step.start","action","artifact (screenshot: step-S1.webp)","assert (text_visible: Counter: 2, milestone: counter)","agent-review (prompt_idx: 0, verdict: pass|fail)","step.end"]}]}agent-review events — New event type lets agents resolve tier 2 prompts:
{"type":"agent-review","step_id":"S1","prompt_idx":0,"verdict":"pass","reason":"Home tab bar visible"}Event gap warnings — record finish now warns when expected events are missing (e.g., step has expect but no assert was streamed).
Install / upgrade
npm install -g flow-walker-cli@0.5.0
288 tests, 18/18 eval gates.
v0.4.1
Bug fix
- Fix multi-step flows with judge blocks collapsing to 1 step — The flow-parser's
!inJudgeguard prevented subsequent step IDs from being recognized after ajudge:block. Uses indentation-aware detection now. - Filter empty prompt cards in reports
- Footer shows package version (0.4.1) not schema version (3.0.0)
--versionshows both:flow-walker 0.4.1 (schema 3.0.0)
Install / upgrade
npm install -g flow-walker-cli
v0.3.2 — v2-only, all v1 code removed
What's Changed
Breaking: v1 support completely removed. All agents must use v2 flows.
Removed
parseFlowV1()parser andLEGACY_STEP_KEYSmigratecommand (v1→v2 migration)- v1 types:
Flow,FlowStep - v1 reporter:
generateReport(),buildHtml() - v1 run-schema:
RunResult,StepResult,validateRunResult() - v1 yaml-writer:
generateFlows(),toYaml(),writeFlows()
Enforced
parseFlowFile()rejects non-v2 flows with clear errorreportcommand validates v2 run.json (requiresmode,steps[].outcome,steps[].do)- Agents must use the pipeline:
record init → stream → finish → verify → report → push
Stats
- 253 tests passing
- ~1,300 lines of v1 code removed
- Schema version: 2.1.0
Full Changelog: v0.3.1...v0.3.2
v0.3.1 — v2 enforcement + desktop video fix
What's new
v2 Schema Enforcement
reportcommand now rejects non-v2 run.json — requiresmode,steps[].outcome, andsteps[].do- Clear error message tells agents to run
flow-walker verifyfirst - No more silent "undefined" in reports from hand-crafted data
Desktop Video Recording Fix
- Replaced
screencapture -v(broken by TCC) with ffmpeg via Terminal.app - Terminal.app has Screen Recording permission — bypasses TCC entirely
- Tested on macOS Tahoe M4: 1920x1080, 10fps, h264, ~150KB for 20s
Pipeline Documentation
recordschema documents the full 6-step mandatory pipelineverifyschema says "REQUIRED after record finish"reportschema says "Requires verify — rejects v1/non-v2 data"
Stats
- 321 tests, typecheck clean
- Synced from beastoin/autoloop Phase 10+11
v0.3.0 — Desktop Support + Verify Fixes
What's New
Phase 10 — Desktop Support
- agent-swift bridge:
AgentType(flutter|swift),detectAgentType(), platform-specificexec/textPress/bringToForeground - CLI flags:
--agent flutter|swiftand--agent-pathfor desktop testing - Desktop video:
screencapture -vfor macOS recording - Schema v2.1.0: Updated command schema with new flags
Phase 11 — Verify Correctness + Polish
- Audit mode fix:
resultnow reflects actual step outcomes (was hardcoded to"pass") - Outcome normalization:
"skip"→"skipped","partial"→"fail", unknown →"fail" - Expectation checking:
text_visibleandinteractive_countnow check assert eventpassedfield - YAML multi-line scalars: folded (
>) and literal (|) support in flow parser - Event schema validation: warns on non-standard
step.endoutcome values
Stats
- 320 tests passing (37 new)
- 15 files changed, 627 insertions
Validated
- 6 desktop flows run by sora on macOS via agent-swift
- Navigation (6/6), Dashboard (3/6), Chat (5/5), Memories (5/6), Tasks (4/5), Settings (9/9)
- Remaining failures are agent-swift AX gaps, not flow-walker bugs
v0.2.1
What's new in v0.2.1
Agent-first v2 architecture (Phases 7-9)
- 3-phase recording:
record init→record stream→record finish - NDJSON event streaming: step.start, action, assert, artifact, step.end, run.start, run.end, note
- Verify command: strict, balanced, and audit modes with VerifyResult schema
- Snapshot & replay system: auto-save coordinates after successful runs, 1.85x speedup on replays
verify: truestep marking: agents control which steps always get verified
Report improvements
- Dark-themed v2 HTML reports with embedded base64 screenshots
- Duration display from real wall-clock timestamps
- Auto-detect screenshots by step ID
- Video embedding support
- Landing page v2 pass count fix
Other
- Walk command generates v2 flows
- Auto video recording in record init/finish (use
--no-videoto disable) - Comprehensive README rewrite covering full v2 pipeline
283 tests passing. Zero external dependencies.
Full changelog: v0.2.0...v0.2.1