Skip to content

plan(roadmap): roadmap-expansion — Fable 5 planning run (5 backlog features → Road-to-0.0.1-stable)#397

Merged
rickylabs merged 21 commits into
mainfrom
plan/roadmap-expansion
Jul 5, 2026
Merged

plan(roadmap): roadmap-expansion — Fable 5 planning run (5 backlog features → Road-to-0.0.1-stable)#397
rickylabs merged 21 commits into
mainfrom
plan/roadmap-expansion

Conversation

@rickylabs

@rickylabs rickylabs commented Jul 4, 2026

Copy link
Copy Markdown
Owner

Roadmap Expansion — Fable 5 planning run

Steering input + harness run for the Fable 5 roadmap supervisor to fold five backlog features
into Road to 0.0.1-stable (#301). This PR is the planning charter only — it opens no
epics/issues and touches no framework code until the owner ratifies.

The five topics

  • A — Dev Dashboard (killer feature; ships as a plugin at beta.6) — specs/topic-A-dashboard.md
  • B — Telemetry production-grade revamp (feeds A) — specs/topic-B-telemetry.md
  • C — Tutorial complete rewrites + minimal eis-chat tutorial (beta.7 docs cut) — specs/topic-C-tutorials.md
  • D — Per-feature storytelling / positioning docs (beta.7 docs cut) — specs/topic-D-positioning-docs.md
  • E — Deno-desktop + unified single-process deployment (beta.8/stable; extends #327) — specs/topic-E-desktop-deploy.md

Charter

  • specs/00-mission-and-flow.mdauthoritative: mission, the A→G delegation flow (Fable →
    Sonnet-5 deep-search workflow → Opus-4.8 deep-dive agents → WSL Codex adversarial → OpenHands
    PLAN-EVAL), the B1–B4 output contract, the overnight-system + memory directive, hard boundaries.
  • specs/01-ratified-decisions.md — milestone train, owner decisions D1–D4, R1/R4, positioning
    lock, delegated calls (D-NSONE, telemetry flow).
  • specs/02-eis-chat-reference.md — eis-chat as the working reference; per-topic reading map.
  • FABLE-PROMPT.md — the thin launch prompt + PR-discipline mandate.

Each topic-* spec opens with the owner's original bullets preserved verbatim, then the agreed
decisions, eis-chat pointers, delegated calls, dependencies, "what B researches", "what Fable
produces".

Status — ✅ Plan-Gate PASSED (OpenHands run-28716441078-1, minimax M3, 2026-07-04) · owner ratified · filing deferred to Phase 2

  • Harness run scaffolded (.llm/runs/plan-roadmap-expansion--seed/) + specs seeded
  • A — Supervisor online: charter + specs read; worklog + phase registry committed
  • B — Sonnet-5 deep-search corpus (5 concurrent agents; 75 files across 20 topic/folder cells; eis-chat reference staged)
  • C — Fable synthesis + both delegated decisions resolved with byte-evidence (analysis/FABLE-STAGE-C-SYNTHESIS.md) — D-NSONE = promote the missing fresh-ui L3 blocks/ layer; grouped-trace flow = two-tier (beta.6 Flow-B framework-native / stable Flow-A duckdb)
  • D — 4 Opus-4.8 deep-dive design proposals (design/{A-dashboard,B-telemetry,E-desktop,CD-docs}/): telemetry T1–T9, dashboard DDX-0…19, docs S0+C1–6+D1–9+V, desktop #E1–E8
  • D+ — Owner-expanded Topic-A source set folded (see note below)
  • E (locked design)research.md (Plan-Gate checklist + 14-row Findings + jsr-audit surface scan), plan.md (12 locked decisions, 13-fork open-decision sweep, cross-epic DAG, milestone train, risk register, per-epic gate matrix), worklog ## Design
  • F1/F2 — WSL Codex adversarial review (verdict FAIL_PLAN, 10 findings) + F2 fixes all 10 (6b12225d)
  • G — OpenHands PLAN-EVAL (minimax M3, separate session) — ✅ PASS; verdict-of-record plan-eval.md landed at d22df217 (comment)
  • Owner ratification — plan ratified 2026-07-04; forks locked: OF-5 opt-in OTel-SDK on fan-in · OF-10 per-capability sections · OF-11 contract in plugin-dashboard-core (not core axis)
  • Filing (Phase 2, owner-authorized) — milestones 0.0.1-beta.5/6/7/8 + epics/sub-issues filed with netscript-pr taxonomy; supersession map owner-approved before any close

Owner expanded the Topic-A source set (mid-run, 2026-07-04). At the owner's direction the
dashboard competitor corpus was extended beyond dev-consoles to the "manage framework features
through the UI" category: Appwrite Console (north-star), Directus (extensibility /
plugin-panel model), Strapi (codegen-from-UI mirroring the CLI + in-dashboard AI). Teardown at
research/A-dashboard/04-baas-admin-console-teardown.md + 17 new matrix rows. This sharpened the
Topic-A IA (flat "Plugin Control list" → cross-cutting panels + per-capability
create→configure→monitor sections), added DDX-17 (DashboardPanelContribution seam /
.withDashboardPanel), DDX-18a-d (per-capability sections), DDX-19 (codegen-from-UI), and
reframed D-NSONE via the Directus panel-contribution precedent. New owner forks OF-10…OF-13.

Draft until the roadmap is complete and PLAN-EVAL passes; the owner ratifies and cuts. Fable
pushes + updates this PR after every stage.

Refs #301

🤖 Generated with Claude Code

https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN

rickylabs and others added 2 commits July 4, 2026 19:10
… Fable delegation charter

Structured steering input for the Fable 5 roadmap supervisor to fold the 5
backlog features (dev-dashboard, telemetry revamp, tutorial rewrites,
positioning docs, deno-desktop+single-process deploy) into Road-to-0.0.1-stable.

- specs/00 authoritative: mission + A-to-G delegation flow (Fable -> Sonnet-5
  workflow -> Opus-4.8 deep-dives -> WSL Codex adversarial -> OpenHands PLAN-EVAL)
  + B1-B4 output contract + overnight/memory directive + hard boundaries
- specs/01 ratified decisions: milestone train, D1-D4, R1/R4, positioning lock,
  delegated calls (D-NSONE, telemetry flow)
- specs/02 eis-chat reference: per-topic reading map
- specs/topic-A..E: each preserves the owner's original bullets verbatim
- matrix/ analysis/ research/ context/: B-output folder contracts
- FABLE-PROMPT.md: thin launch prompt + PR discipline mandate
- harness artifact skeletons for Fable to fill

Planning only — no GitHub mutation, no framework code until the owner ratifies.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…h prompt

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
@rickylabs

Copy link
Copy Markdown
Owner Author

Supervisor online — run taken (stage A → B)

Fable 5 roadmap supervisor has taken this run (plan-roadmap-expansion--seed).

Done so far

  • Charter (FABLE-PROMPT.md) + all specs read in full: 00-mission-and-flow (authoritative), 01-ratified-decisions, 02-eis-chat-reference, and all five topic specs (A dashboard, B telemetry, C tutorials, D positioning, E desktop/single-process). Owner original bullets preserved verbatim in the specs — they remain the source of truth.
  • Harness skill activated (netscript-harness); run dir + B1–B4 folder scaffolding verified; worktree on plan/roadmap-expansion confirmed; WSL gh auth verified.

Entering stage B: launching the Sonnet 5 (high) deep-search workflow to fill matrix/ analysis/ research/ context/ with one sub-folder per topic (B1 resource matrix + index, B2 exhaustive codebase analysis, B3 external-source research distilled to .md, B4 other useful context), per the B output contract in specs/00.

Flow ahead: B corpus → Fable analysis → Opus 4.8 per-topic deep-dive design proposals → locked design + harness run docs → WSL Codex adversarial review → fixes → OpenHands PLAN-EVAL (separate session, hard stop).

Boundaries honored: planning only — no issue/label/milestone mutations, no framework code; PR stays draft; owner ratifies and cuts.

Delegated decisions I will resolve with rationale: D-NSONE (NS One promotion into @netscript/fresh-ui vs borrowed reference) and the telemetry grouped-trace showcase flow (derived from eis-chat's real pipeline, incl. cross-language-hop milestone choice).

rickylabs and others added 3 commits July 4, 2026 19:25
… B (Sonnet 5 deep-search x5) launched

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…ics (75 files, B1-B4)

Five concurrent Sonnet-5 agents (one per topic) filled the B1-B4 output contract:
matrix/ (resource matrices), analysis/ (exhaustive codebase surface), research/
(distilled external sources), context/ (diagrams, candidate flows, open questions).
Every topic sub-folder carries an INDEX.md. Facts-over-verdicts preserved on all
delegated decisions (D-NSONE, telemetry grouped-trace flow, E sequencing).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…ift candidates

Records the load-bearing divergences the B corpus surfaced (E1 172a-2 misattribution,
E2 #327 WATCH placement, A1 D-NSONE reframe, A2 no TS interaction-service, B1 thin
cross-language boundary, B2 telemetry-worse-than-leveling + triggers bug, CD1 #232
disjoint scope, C1 5-tracks + missing milestones, D1 two-IA blocker). Worklog +
phase-registry advanced to stage C.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
@rickylabs

Copy link
Copy Markdown
Owner Author

Stage B complete — deep-search corpus landed (entering C → D)

The Sonnet-5 deep-search stage is done: 5 concurrent agents (one per topic), 75 files across the B1–B4 contract (matrix/ analysis/ research/ context/, one sub-folder per topic, every cell + INDEX.md). Committed 3d70ff5a, pushed.

What the corpus changed vs the specs (9 drift candidates → drift.md)

Several are decision-shaping and I'll carry them into the locked design; the ones that touch scope or your delegated calls are flagged for you:

  • E1 (significant) — the topic-E "172a-2 service-base-seam" prerequisite is a misattribution. PR Re-architect plugin scaffold surface (#157): thin, typesafe, no plugin-source copy #172 is merged; its 172a-2 sub-slices are CLI plugin type-soundness, unrelated to single-process. The real precursor is much narrower: the server-side in-process .fetch() mount seam (ServiceApp in @netscript/service) already ships; only a client-side ClientLinkPort adapter is missing from @netscript/sdk. E is smaller than framed.
  • E2 (significant) — epic: NetScript enterprise deployment framework (cloud-agnostic + bare-metal, CLI + Aspire) #327 currently lists deno desktop as WATCH/reference-only, unscheduled, and feat(aspire): first-party deno desktop app support in the generator (native window as a dev-stack resource) #375 is p3/Backlog. The E deliverable (schedule it as the 4th tier + promote feat(aspire): first-party deno desktop app support in the generator (native window as a dev-stack resource) #375) is a real rescope of the issue's current state, not a note.
  • A1 (significant) — D-NSONE premise is weaker than stated. fresh-ui and NS One share a byte-identical L0–L2 layer (NS One's primitives are fresh-ui's own copy-source output). The only real gap is fresh-ui has no L3 "blocks" layer (eis-chat has 9). So the decision is really "promote the missing blocks layer into fresh-ui" vs "borrow per-dashboard" — your promotion lean stays valid, re-costed much cheaper.
  • A2 — Aspire pin verified 13.4.6 (clears the ≥9.4 WithCommand gate). But IInteractionService is not exposed in the TS AppHost SDK at any version — the dashboard's interactive prompts must route through command arguments, not interaction-service.
  • B1 (significant) — telemetry showcase reality: eis-chat's SigNoz "archaeology ⋈ live telemetry" join is aspirational (not provisioned); the only genuine non-Deno boundary today is a telemetry-dark duckdb.exe subprocess. NetScript already injects TRACEPARENT into subprocesses (7-runtime polyglot) but no example stitches a non-Deno child span back into the parent. So the hardest cross-language hop needs a built demonstration, not just instrumentation — this shapes my flow + milestone pick.
  • B2 (significant) — telemetry is worse than "level everyone up to workers": streams F (zero wiring), ai F (seam never invoked), triggers has a real correctness bug (inbound W3C context captured but never used to parent spans → severed trace), and the package itself has a tracked "Refactor" arch-debt mandating a ports/adapters restructure + OTEL-adapter split. Real span-links exist in exactly one place (database Prisma bridge). The revamp includes a restructure + a bugfix, not just parity leveling.
  • CD1 (significant) — epic: docs — march to 0.0.1-stable (coverage & accuracy) #232 is a 100% accuracy/coverage umbrella with zero overlap with tutorial-rewrite or positioning scope. Landing C+D "under epic: docs — march to 0.0.1-stable (coverage & accuracy) #232" needs an explicit rescope (or a new docs-cut child epic) — I'll draft both options for your call.
  • C1 — there are 5 live tutorial tracks, not 4 (chat landed separately and teaches against mid-flight @netscript/ai, publish:false today); tutorial chapter URLs are wired into 8 capability-hub nav sections. No beta.6/beta.7 GitHub milestones exist yet — issue-filing per AGENTS.md will need you to create them (no mutations this run).
  • D1 — two unreconciled docs IAs (capabilities/ ~15 pages vs the 9 pillar folders) block per-feature authoring; an IA-reconciliation slice must precede the D fan-out. Only 2 named-competitor mentions exist site-wide today.

Next

Stage C (my synthesis; verifying the decision-critical B2 files) → stage D (Opus 4.8 per-topic deep-dive design proposals). Still planning-only; PR stays draft; no issue/milestone mutations. I'll resolve D-NSONE and the telemetry grouped-trace flow in the locked design with rationale, and bring the scope forks above (CD1 rescope choice, missing milestones) back to you.

rickylabs and others added 7 commits July 4, 2026 19:53
…both delegated decisions

D-NSONE: promote the missing L3 blocks layer into fresh-ui (L0-L2 already byte-identical
copy-source); leave MCP-specific components out unless dashboard IA needs them. Telemetry
grouped-trace flow: two-tier — beta.6 flagship = Flow B framework-native multi-process
pipeline (workers->oRPC callback->streams fan-in with span-links); stable = Flow A
cross-language duckdb.exe subprocess hop (needs net-new span + env-carrier + language shim).
Records telemetry-revamp true scope (package restructure + triggers bugfix + streams/ai
from zero) and E scope correction (172a-2 misattribution; ClientLinkPort is the real precursor).
Owner-facing forks flagged (missing milestones, #232 rescope, #327 rescope).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
… stage C marked done

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…cision resolutions + jsr-audit surface scan

Written to the Plan-Gate checklist: re-baseline vs main@eeaff336, 14 verifiable load-bearing
findings, planned public-surface deltas (fresh-ui L3 blocks, telemetry OTEL-adapter subpath,
sdk ClientLinkPort, new plugin-dashboard-core) with slow-type risks, delegated decisions +
owner forks. Independent of the in-flight Opus deep-dives.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…16 files)

dev-dashboard (DDX-0..16): thin plugins/dashboard + packages/plugin-dashboard-core; Aspire
command-kind seam HARD beta.6; 7-panel IA; D-NSONE promote 7 L3 blocks (drop data-grid
collision, MCP out); DDX-8 flagship trace co-land gate on telemetry.
telemetry-revamp (T1..T9): TC-1..14 convention + ports/adapters restructure; thin-default
+ opt-in-SDK-on-fan-in; triggers W3C bugfix + oRPC span-creation + streams/ai from zero;
@netscript/telemetry/query surface; real Flow-B e2e; Flow-A duckdb + AI adapter at stable.
docs-cut (S0 + C1..C6 + D1..D9 + V): IA-reconciliation precursor; 5 tutorial rewrites +
minimal on-ramp; per-pillar positioning; #232 Option-2 (new child epic) recommended.
#327 rescope (E1..E8): ClientLinkPort in-process adapter precursor; tursodb single-writer
relocation; option b->c ladder; #375 folded; #349 kept WATCH-sibling (not merged).

Cross-topic edges reconciled (query contract, flagship co-land gate). Planning only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…BaaS admin consoles)

Owner steered the dashboard research to add the manage-through-UI category the first
dev-console teardown missed: Appwrite Console (north-star — per-capability create/configure/
monitor loop for every backend primitive), Directus (extension taxonomy: plugin-contributes-a-
panel SDK contract; schema-driven UI generation), Strapi (codegen-from-UI mirroring the CLI +
in-dashboard AI-on-codegen). New research/A-dashboard/04-baas-admin-console-teardown.md +
matrix/_draft-competitor-rows-baas.md (17 rows, kinds manage/extensibility/codegen-ui/ai-iterate).
Opus-A resumed to fold these into the panel IA, a .withDashboardPanel extensibility axis, and
the D-NSONE call. Planning only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
… (Stage E)

Integrated program plan folding owner topics A-E into Road-to-0.0.1-stable (#301):
12 locked decisions (Spine-1 co-land, D-NSONE=fresh-ui L3, two-tier trace flow,
telemetry restructure, dashboard thin-plugin+core, Aspire command-kind seam,
E ClientLinkPort precursor, docs-only wave). 9-fork open-decision sweep (only OF-5
must-resolve-now, resolved technically by LD-4). Cross-epic DAG + beta.5->stable
milestone train + risk register + per-epic gate matrix + deferred scope.
Commit-slice framing satisfies <30 at both run and per-epic levels.

Planning-only: no GitHub mutation, no framework code. PR #397 stays draft.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…oard design (Stage D+/E)

Opus-A integrated the manage-through-UI corpus (Appwrite/Directus/Strapi) into
design/A-dashboard/: IA shift from a flat "Plugin Control list" to cross-cutting
panels + per-capability create->configure->monitor sections; new DDX-17
DashboardPanelContribution seam (.withDashboardPanel, thinness-correct contract in
plugin-dashboard-core, Aspire-style discovery, @netscript/plugin stays
dashboard-agnostic), DDX-18a-d per-capability sections (workers/sagas/triggers/
streams, thin, beta.6), DDX-19 codegen-from-UI "Add resource" (stable, #157-safe,
#238 AI handshake). D-NSONE sharpened via the Directus panel-contribution precedent.

Reconciled into plan.md: dashboard slices DDX-0..19; DAG IA-shift note; new owner
forks OF-10..OF-13 (=OQ-11..14) added to the open-decision sweep (per-capability-vs-flat,
sugar-vs-core-axis, codegen milestone/#238, schema-driven db tab) — none FAIL_PLAN.
DDX-0..16 unchanged, so plan.md's DDX-8 flagship co-land references remain valid.
pr-body.md status refreshed to current reality (A-E done) + owner-expanded-source-set note.

Planning-only: no GitHub mutation, no framework code. PR #397 stays draft.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
@rickylabs

Copy link
Copy Markdown
Owner Author

Stage E locked — integrated roadmap ready for adversarial review (F1)

Design is locked. The five owner topics are folded into Road-to-0.0.1-stable (#301) as 2 NEW epics + 2 rescopes, fully decomposed with acceptance criteria, labels, milestones, a cross-epic DAG, a risk register, and per-slice agent briefs.

Both supervisor-delegated decisions resolved with byte-level evidence:

  • D-NSONE → promote the missing @netscript/fresh-ui L3 blocks/ layer (L0–L2 are byte-identical copy-source; the only real gap is fresh-ui has no L3 at all). data-grid NOT promoted (collides with the existing DataGrid<T>); MCP components out for beta.6. Reframed as a panel-contribution registry via the Directus precedent.
  • Telemetry grouped-trace flowtwo-tier: beta.6 flagship = Flow B (framework-native multi-process: eischat→workers-api→workers→oRPC callback→streams fan-in, real span-links); stable = Flow A cross-language duckdb.exe hop.

The four epics

Epic Milestones Slices Merge-gate
telemetry-revamp (NEW, Spine-1 enabler) beta.5 → stable T1–T9 T8 real Flow-B e2e
dev-dashboard (NEW, Spine-1 headline, ships as a plugin) beta.6 core → stable depth DDX-0…19 DDX-16 scaffold.runtime join
docs-cut (C+D) beta.7 S0 + C1–6 + D1–9 + V OpenHands per-domain
#327 rescope (E, desktop/single-process) beta.8 → stable #E1–#E8 #E7 deploy-e2e

Owner expanded the Topic-A source set mid-run (Appwrite/Directus/Strapi "manage-through-UI"): sharpened the dashboard IA (flat list → cross-cutting panels + per-capability create→configure→monitor sections) and added DDX-17 (DashboardPanelContribution / .withDashboardPanel), DDX-18a-d, DDX-19 (codegen-from-UI). Teardown + 17 matrix rows committed.

Consolidated decision register: 12 locked decisions (LD-1…12) + 13 owner forks (OF-1…OF-13). Only OF-5 (telemetry SDK-dep posture) is "must-resolve-now" in the rework sense, and it is technically resolved (LD-4) — it needs only the owner's dependency-posture ratification. No FAIL_PLAN open decision.

Next: F1 WSL Codex adversarial review of the locked design → F2 fixes → G OpenHands PLAN-EVAL (minimax M3, separate session). PR stays draft; the owner ratifies (creates milestones, picks forks) and cuts.

Artifacts: plan.md, research.md, design/{A-dashboard,B-telemetry,E-desktop,CD-docs}/, worklog.md § Design.

rickylabs and others added 2 commits July 4, 2026 20:42
Resolve all 10 WSL-Codex F1 findings in the roadmap-expansion planning
artifacts (planning-only; no framework code, no GitHub mutation):

- F1-01 BLOCKER: LD-2 + research#1 no longer overstate fresh-ui evidence
  (5/37 pairs sampled-identical; remaining 32 = DDX-0 full-tree-diff gate).
- F1-02 BLOCKER: OF-10 reclassified must-resolve-now (drafted issue graph
  already assumes per-capability); closing note corrected to OF-5 + OF-10.
- F1-03 BLOCKER: DDX-8 + DDX-16 hard-depend telemetry T4/T5/T6/T7 by ID;
  acceptance fails if the oRPC callback span is mock-only.
- F1-04 MAJOR: DDX-16 given an explicit full beta.6 dependency list.
- F1-05 MAJOR: wave:* label-sync folded into OF-1; DDX-19 wave:v2->wave:defer.
- F1-06 MAJOR: gate:jsr + publish acceptance on DDX-0/2/4/17 and #E1.
- F1-07 MAJOR: dashboard slice count corrected 20 -> 23 (< 30 holds).
- F1-08 MAJOR: DDX-2/proposal folder vocab fixed (telemetry/->middleware/;
  public/ reframed as harness-observed #305/#306 divergence).
- F1-09 MINOR: #E2 Closes #375 moved from issue body to resolving-PR body.
- F1-10 NIT: research#14 extended to beta.5/beta.8 milestones.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
@rickylabs

Copy link
Copy Markdown
Owner Author

Stage F2 — adversarial-review fixes applied (commit 6b12225d)

F1 verdict was FAIL_PLAN (WSL-Codex, gpt-5.5, thread 019f2e6c…): 3 BLOCKER, 5 MAJOR, 1 MINOR, 1 NIT. All 10 findings are now resolved in the planning artifacts — planning-only: no framework code, no epic/issue/label mutation, PR stays draft.

# Sev Finding F2 fix
F1-01 BLOCKER LD-2 claimed fresh-ui/eis-chat L0–L2 universally byte-identical, but only 5/37 pairs were sampled LD-2 + research #1 softened to "5/37 sampled-identical; remaining 32 = DDX-0 scripted full-tree-diff gate"; DDX-0 acceptance is the proving gate
F1-02 BLOCKER OF-10 marked safe-to-defer, but the drafted beta.6 issue graph already bakes in per-capability → deferring forces rework OF-10 reclassified must-resolve-now (with a documented flat-list fallback re-draft path); closing note corrected to "OF-5 and OF-10 must be ratified"
F1-03 BLOCKER DDX-8 child issue named only fan-in + triggers bugfix; omitted T6 (oRPC span-creation) + T7 (query surface) DDX-8 and DDX-16 now hard-depend telemetry T4/T5/T6/T7 by ID; acceptance fails if the oRPC callback span is mock-only / not queryable
F1-04 MAJOR DDX-16 dependency line used a range that omitted DDX-0/1/2/4/14/15/17/18a-d DDX-16 given an explicit full beta.6 dependency list + cross-epic T4–T7
F1-05 MAJOR Drafts use wave:* labels but .github/labels.yml has no wave: block; wave:v2 non-canonical Owner action to sync wave:v1/v1-min/defer folded into OF-1; DDX-19 wave:v2wave:defer
F1-06 MAJOR gate:jsr missing on public-surface slices Added gate:jsr + doc:lint/publish --dry-run acceptance to DDX-0, DDX-2, DDX-4, DDX-17, #E1
F1-07 MAJOR "20 slices" undercount (DDX-18a-d counted as 1) Corrected to 23 (17 core + DDX-17 + 4×DDX-18 + DDX-19); still < 30
F1-08 MAJOR DDX-2 called public/ + telemetry/ "doctrine layering" — not doctrine-05 roles telemetry/middleware/ (drift corrected); public/ reframed as harness-observed CLI-scaffolder divergence, tracked #305/#306 — dropped the "doctrine-clean" claim
F1-09 MINOR #E2 put Closes #375 in the issue body Moved to the resolving-PR body; issue draft = "Part of #327; folds #375"
F1-10 NIT research #14 covered only beta.6/beta.7 Extended to beta.5/beta.8; owner creates any missing milestones (OF-1)

Consolidated decision register after F2: LD-1…12 (unchanged) + OF-1…13. Two forks are ratify-now (not silently deferrable): OF-5 (telemetry SDK-adapter dependency posture) and OF-10 (per-capability IA shape) — both resolved technically with a locked recommendation and a documented override/fallback, so no FAIL_PLAN open decision remains.

Next: dispatching Stage G — OpenHands PLAN-EVAL (minimax M3, separate session, hard stop before implementation).

@rickylabs

Copy link
Copy Markdown
Owner Author

@openhands-agent model=openrouter/minimax/minimax-m3 provider=openrouter output=pr-comment iterations=800

use harness — run PLAN-EVAL (planning gate only) for the roadmap-expansion program plan. Hard stop before implementation: read and judge the planning artifacts only. Do NOT write framework code, do NOT scaffold, do NOT mutate any GitHub issue/label/milestone, do NOT run deno task/deno check/deno publish and do NOT touch deno.lock. You are the independent evaluator in a separate session; the plan was authored by a different (Fable) session — evaluate it adversarially, do not defend it.

SKILL

Activate these repo skills before judging (skill-first is mandatory; skipping is the #1 eval-fail):

  • netscript-harnessprimary. The Plan-Gate checklist, PLAN-EVAL protocol, verdict format (.llm/runs/<run-id>/plan-eval.md), and evaluator separation live here. Follow it exactly.
  • openhands-handoff — your own run/output contract: write the verdict to OPENHANDS_RUN_DIR and post one pr-comment; do not reuse legacy scratch; never mutate the lock.
  • netscript-doctrine — to check the DDX-2 package folder vocabulary against doctrine-05 role folders and the ARCHETYPE-2/5 archetype claims, plus the thin-plugin/fat-core direction.
  • netscript-pr — to verify the label taxonomy (namespaced colon labels, exactly one status:, wave: mapping), milestone assignment, and the closing-keyword-in-PR-body obligation used throughout the drafts.
  • jsr-audit — to sanity-check the gate:jsr surface-scan claims (public surface has doc:lint + publish --dry-run acceptance) on DDX-0/2/4/17 and #E1.
  • netscript-tools — for gate-evidence and lock-hygiene expectations you assert against (you must NOT run the gates yourself; judge that the plan names the right ones).

What to evaluate

Target: branch plan/roadmap-expansion (this PR, #397). Run artifacts under .llm/runs/plan-roadmap-expansion--seed/:

  • plan.md — the integrated program plan: locked decisions LD-1…12, open-decision sweep OF-1…13, cross-epic DAG, milestone train (beta.5→beta.6→beta.7→beta.8→stable), risk register, per-epic gate matrix, jsr-audit surface scan, commit-slice framing.
  • research.md — the re-baselined findings the plan depends on (esp. S0: initial public repo genesis #1 fresh-ui evidence, Wave 3 — @netscript/plugin (A4 plugin host) [umbrella · integration branch] #14 milestones).
  • design/A-dashboard/{proposal.md,epic-and-issues.md,open-questions.md} — Topic A (dev-dashboard, 23 slices DDX-0…19).
  • design/B-telemetry/, design/CD-docs/, design/E-desktop/epic-and-issues.md — the other epics.
  • worklog.md — the design section + progress log (includes the F1→F2 history).

This plan already went through an adversarial review (F1, WSL-Codex) that returned FAIL_PLAN with 3 BLOCKER + 5 MAJOR + 1 MINOR + 1 NIT; commit 6b12225d (F2) claims to fix all ten. Verify the F2 fixes actually hold — in particular:

  • (F1-01) LD-2 / research S0: initial public repo genesis #1 no longer overstate the fresh-ui byte-identical evidence and DDX-0 carries the full-tree-diff proving gate;
  • (F1-02) OF-10 is correctly classified so no rework-forcing decision is left silently deferred;
  • (F1-03) DDX-8 and DDX-16 hard-depend telemetry T4/T5/T6/T7 by ID with acceptance that fails on a mock-only oRPC span;
  • (F1-05/06/07/08) wave-label gap, gate:jsr coverage, the 23-slice count (<30), and the DDX-2 doctrine folder vocabulary.

Method

Run the harness Plan-Gate checklist and fill the plan-eval.md template:

| Plan-Gate item | PASS/FAIL | evidence |
| research present and current | | |
| decisions locked | | |
| open-decision sweep (any deferred decision that FORCES REWORK ⇒ FAIL_PLAN) | | |
| commit slices (<30, each names gate + files) | | |
| risk register | | |
| gate set selected | | |
| deferred scope explicit | | |
| jsr-audit surface scan (pkg/plugin) | | |

Then run your own open-decision sweep — independently look for any decision that, if deferred, forces rework (do not just trust the plan's own OF table). Give a final verdict: PASS or FAIL_PLAN. If FAIL_PLAN, enumerate the required fixes precisely (finding → fix), each tied to a file+location.

Output

  • Write the completed evaluator artifact to .llm/runs/plan-roadmap-expansion--seed/plan-eval.md (the verdict of record — OPENHANDS_RUN_DIR), committed on this PR branch. Do not commit anything else; do not touch deno.lock (if any tool dirties it, git checkout -- deno.lock before committing).
  • Also write OPENHANDS_SUMMARY_PATH and post a single pr-comment summarizing the verdict, the per-item checklist results, and (if FAIL_PLAN) the required fixes.
  • Keep PR plan(roadmap): roadmap-expansion — Fable 5 planning run (5 backlog features → Road-to-0.0.1-stable) #397 in draft. This is planning-gate evaluation only.

@rickylabs

rickylabs commented Jul 4, 2026

Copy link
Copy Markdown
Owner Author

OpenHands Agent — Completed

Model: openrouter/minimax/minimax-m3
Provider: OPENROUTER
Job status: failure
Trace: .llm/tmp/run/openhands/pr-397/run-28716441078-1 (committed to the PR branch)
Artifact: openhands-agent-28716441078-1

PLAN-EVAL — plan-roadmap-expansion--seed

Summary

Independent PLAN-EVAL pass on the plan/roadmap-expansion branch (PR #397, draft). Verified
all 8 Plan-Gate checklist boxes PASS and ran an independent open-decision sweep that found no
additional rework-forcing deferrals beyond what the plan already names. Verdict: PASS.

All ten F1 adversarial findings (3 BLOCKER + 5 MAJOR + 1 MINOR + 1 NIT) from WSL-Codex
F1-adversarial-review.md have been verified fixed in the F2 artifacts (commit 6b12225d),
without introducing new gaps. The two ratify-now forks in the OF table (OF-5 SDK-adapter
dep posture, OF-10 per-capability IA) are both technically pre-resolved in the drafted
artifacts with documented fallback re-draft paths — the plan does not ship an unresolved
decision that would force rework.

Changes

  • .llm/runs/plan-roadmap-expansion--seed/plan-eval.md — full PLAN-EVAL verdict
    artifact (overwrote the harness stub template with the completed 8-row checklist +
    independent evaluator-run open-decision sweep + PASS verdict + 4 non-blocking
    observations).

No other files modified. No framework code, no plugin/package surface, no deno task /
deno check / deno publish runs. deno.lock untouched. No GitHub issue / label /
milestone / PR-comment mutations (the PR-comment will be posted via the workflow owner per
the trigger contract; the runner did not post one itself).

Validation

This was a planning-only evaluator pass; no implementation validation is in scope. The
evaluation was driven by reading:

  • .llm/harness/gates/plan-gate.md (the checklist contract)
  • .llm/harness/evaluator/plan-protocol.md + verdict-definitions.md (evaluator discipline)
  • .llm/harness/gates/archetype-gate-matrix.md (gate family authority)
  • .llm/runs/plan-roadmap-expansion--seed/plan.md (LD-1…12 + OF-1…13 + DAG + milestones
    • risk register + per-epic gate matrix + deferred scope)
  • .llm/runs/plan-roadmap-expansion--seed/research.md (14-row Findings table, post-F2
    softened Finding S0: initial public repo genesis #1 and post-F2 extended Finding Wave 3 — @netscript/plugin (A4 plugin host) [umbrella · integration branch] #14)
  • .llm/runs/plan-roadmap-expansion--seed/worklog.md (Design section + F1→F2 history)
  • .llm/runs/plan-roadmap-expansion--seed/F1-adversarial-review.md (10 F1 findings
    enumerated; the basis for the F2 verification)
  • .llm/runs/plan-roadmap-expansion--seed/design/A-dashboard/{proposal,epic-and-issues, open-questions}.md (Topic-A — DDX-0…19 = 23 slices; DDX-8/16 hard-dep T4/T5/T6/T7 by
    ID; DDX-16 explicit full beta.6 dep list; gate:jsr on DDX-0/2/4/17; per-capability IA
    thesis; .withDashboardPanel seam; codegen-from-UI DDX-19 stable-tier)
  • .llm/runs/plan-roadmap-expansion--seed/design/B-telemetry/epic-and-issues.md (T1…T9;
    beta.6 crit path T1→T2→T3→T5→T8; T6 pulled onto flagship gate)
  • .llm/runs/plan-roadmap-expansion--seed/design/CD-docs/epic-and-issues.md (S0 + C1–6
  • .llm/runs/plan-roadmap-expansion--seed/design/E-desktop/epic-and-issues.md (#E1…E8;
    #E1 gate:jsr + #E2 Closes #375 in PR body not issue body; [Deploy-S13] WATCH: RFC-14 unified-mode + Nitro deno_server (tier-3 serverless) #349 WATCH)
  • .github/labels.yml (verified wave:* block is absent — OF-1 owner action; verified
    the four wave:* draft labels reduce to three canonical wave:v1 / wave:v1-min /
    wave:defer per netscript-pr skill)
  • .agents/skills/netscript-pr/SKILL.md (label taxonomy + closing-keyword obligation
    • wave→milestone map)

The runner used grep / cat -n / find only — no deno task / deno check /
deno publish / deno doc / deno fmt / git push / gh mutations. deno.lock is
unchanged (mtime + size identical to the rest of the unmodified tree).

Responses to review comments or issue comments when relevant

The trigger comment on PR #397 (@openhands-agent ... use harness — run PLAN-EVAL ...) is
the only request. The single PR-comment summary required by the trigger's
output_mode=pr-comment is being staged at OPENHANDS_COMMENT_PATH (the workflow owner
posts it; the runner did not post one). The verdict of record lives at
.llm/runs/plan-roadmap-expansion--seed/plan-eval.md and was not authored or defended by
this evaluator — it is an independent adversarial judgment.

Remaining risks

Non-blocking observations recorded in plan-eval.md § "Notes (non-blocking)" — none of
these affects the Plan-Gate verdict but a future F-cycle may want to:

  1. Drift Watch R2 only names a fallback if the Aspire app kind slips; it does not
    name a fallback if the command kind itself becomes unavailable. Current mitigation =
    Aspire pin 13.4.6 ceiling. Low-likelihood; only act if Aspire's roadmap signals
    command-kind removal.
  2. epic: NetScript AI Stack — first-class AI runtime, chat & plugin seams (anchor #219) #238 AI-on-codegen is a cross-epic edge to a flagship AI plugin. The plan correctly
    defers DDX-19 to stable; IMPL-EVAL should verify the epic: NetScript AI Stack — first-class AI runtime, chat & plugin seams (anchor #219) #238 owner is looped in before
    any beta.6 stretch attempt.
  3. .github/labels.yml wave:* block absence is correctly tracked as OF-1 owner
    action (post-F2 fix to F1-05). The plan is not in a position to mutate labels.yml
    from this evaluator session; the obligation is on the owner at ratification time.
  4. Per-slice agent-briefs.md are referenced in the per-epic drafts but live one
    level deeper (design/*/agent-briefs.md) than this evaluator's read. Lane discipline
    (WSL Codex / Opus workflows / OpenHands validation) is consistent across LD-11,
    worklog.md, and the DDX Lane columns; the per-slice handoffs are not re-examined
    here.

PR #397 stays draft per the trigger contract.

Run: https://github.com/rickylabs/netscript/actions/runs/28716441078

…ed forks

Restore the OpenHands run-28716441078-1 (minimax M3) filled plan-eval.md that
the CI commit-back steps failed to push (branch previously held the empty
template). Content recovered verbatim from the run log; provenance noted in-file.
Record owner ratification (2026-07-04): OF-5 opt-in OTel-SDK on fan-in; OF-10
per-capability sections; OF-11 DashboardPanelContribution contract owned by
plugin-dashboard-core (not a core axis).

Planning-only; PR #397 stays draft.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
@rickylabs

Copy link
Copy Markdown
Owner Author

Stage G — PLAN-EVAL verdict: ✅ PASS

Run: OpenHands run-28716441078-1 (minimax M3 via OpenRouter) — a separate session from the Fable author and the WSL-Codex F1 adversarial reviewer, per evaluator-separation discipline.

Verdict of record: .llm/runs/plan-roadmap-expansion--seed/plan-eval.md (landed at d22df217).

Result — all 8 Plan-Gate boxes PASS

Plan-Gate item Result
Research present and current ✅ PASS
Decisions locked (LD-1…LD-12) ✅ PASS
Open-decision sweep (OF-1…OF-13 + independent evaluator sweep) ✅ PASS
Commit slices < 30 (gate + files each) ✅ PASS
Risk register ✅ PASS
Gate set selected ✅ PASS
Deferred scope explicit ✅ PASS
jsr-audit surface scan (pkg/plugin) ✅ PASS
  • All 10 F1 findings (3 BLOCKER + 5 MAJOR + 1 MINOR + 1 NIT) verified fixed at F2 (6b12225d) with no new gaps.
  • The evaluator ran an independent open-decision sweep (8 candidate deferrals beyond the OF table) and found no additional rework-forcing decisions. No FAIL_PLAN open decision.
  • 4 non-blocking observations recorded (Aspire command-kind fallback, #238 AI-on-codegen cross-edge, .github/labels.yml missing wave:* block = OF-1, per-slice agent-briefs out of eval scope).

Job-status note (corrected)

The workflow reports "Job status: failure", but this is not an evaluation failure — the evaluator agent exited 0 and produced the PASS verdict above. The failing steps are "Commit changes back to PR branch" and "Commit run trace to PR branch" (a CI commit-back failure), which is why the branch previously carried the empty plan-eval.md harness template. I recovered the evaluator's filled artifact verbatim from the run log (provenance noted in-file) and landed it at d22df217, restoring the verdict of record the failed CI step should have pushed.

Owner ratification (2026-07-04)

Owner has ratified the plan ("absolute bar", authorized). The three locked forks are recorded in the register:

  • OF-5 → allow opt-in @opentelemetry/sdk-* on fan-in (recommendation adopted).
  • OF-10 → per-capability sections at beta.6 (recommendation adopted).
  • OF-11DashboardPanelContribution contract owned by plugin-dashboard-core, not a core definePlugin axis (overrides the earlier core-axis lean; now locked).

PR #397 stays draft. Issue-filing (milestones + taxonomy) is deferred to the owner-authorized filing phase.

rickylabs and others added 4 commits July 4, 2026 23:31
Topic F-ai (6th roadmap topic) folded in post-ratification. Stage B: 4 Sonnet-5
deep-search agents, 14 files across research/context/analysis/matrix/F-ai.
Verdict: F-ai is evaluate-and-harden, not rebuild — the five-home #238 split is
correctly shaped; eis-chat is the proof-of-pattern reference. 5 flagship gaps +
contract-bypass + missing doctrine backstop identified. Stage C: Fable synthesis
with working positions + 5 owner forks; Stage-D fan-out = one Opus-F deep-dive.

Planning-only; PR #397 stays draft.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…nning only)

Folds the 6th topic (F-ai / NetScript AI suite) into the Road-to-0.0.1-stable
program after the A–E Plan-Gate PASS + owner ratification. Design of record
(Opus-F, 18-slice FAI-0…17 DAG) committed; plan.md carries a self-contained
Topic-F-ai section (LD-F1…F6, ai-stack epic row, DAG lane, OF-F1…F8 forks,
DRAFT supersession map, risks) that leaves the ratified A–E LD/OF numbering
untouched. Verdict: evaluate-and-harden, not rebuild.

No framework code, no GitHub mutation. F-ai has its own Plan-Gate pass pending
(F1 adversarial → F2 fixes → G OpenHands PLAN-EVAL). PR #397 stays draft.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
Stage F2 of the Topic F-ai leg. Resolves all 8 findings from the WSL-Codex
adversarial review (F1-ai-adversarial-review.md, verdict FAIL_PLAN):

- F1AI-01 (BLOCKER): remove invented beta.6 Topic-A "dashboard AI panel";
  reframe FAI-0..3 as a parity FLOOR for the OF-6 telemetry seam + stable
  DDX-19 handshake, not a Topic-A beta.6 hard-dep (proposal, plan DAG/OF-F3,
  OQ-3). No ratified Topic-A artifact edited.
- F1AI-02 (BLOCKER): FAI-17 cross-topic deps corrected T1+T6 -> T3+T6
  (T1 transitive via T3), matching Topic-B T9's declared deps.
- F1AI-03 (BLOCKER): single authoritative supersession headline
  "15 KEEP - 1 FOLD (#257->#379) - 0 close - 3 NEW (FAI-4/10/11)";
  #272 = KEEP-stable, not a fold.
- F1AI-04 (MAJOR): epic-header label-sync owner action (epic:/wave: into
  .github/labels.yml, folded into OF-F4); FAI-10/11/14 wave:defer -> wave:v1.
- F1AI-05 (MAJOR): add Archetype gate matrix (surface->archetype->gates)
  + plan epic-entry gate note.
- F1AI-06 (MAJOR): OF-F3 reclassified safe-to-defer.
- F1AI-07 (MINOR): FAI-9 renamed "beta.6 capability merge gate" + per-milestone gates.
- F1AI-08 (NIT): tighten FAI-0/FAI-6 file:line refs.

V3 harness cleanup: delete run-root commits.md; correct "append commits.md"
wording in A/B/E agent-briefs.md to the V3 commit trail.

Divergences recorded in drift.md (FAI-1, FAI-2, V3); worklog n6 rows added.
Planning artifacts only: no framework code, no GitHub mutation. PR #397 stays DRAFT.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
@rickylabs

Copy link
Copy Markdown
Owner Author

@openhands-agent model=openrouter/minimax/minimax-m3 output=pr-comment

use harness

Run a PLAN-EVAL (the harness first evaluator pass) on the Topic F-ai (AI-suite) leg of this planning run, on this PR branch plan/roadmap-expansion. You are a separate evaluator session from the Fable author and the WSL-Codex F1 reviewer — do NOT author, defend, or re-edit the plan. Judge the plan, not code. Do not run any implementation gate or comment on slices that do not yet exist.

SKILL

Activate before any work:

  • .agents/skills/netscript-harness — harness workflow + evaluator protocol (MANDATORY; read evaluator/plan-protocol.md).
  • .agents/skills/netscript-doctrine — archetype / gate / debt vocabulary for the planned surfaces.
  • .agents/skills/netscript-deno-toolchaindeno doc / jsr-audit surface inspection.
  • .agents/skills/netscript-pr — label / milestone taxonomy (for the label-sync check only; do NOT mutate GitHub).

Scope (F-ai leg only)

This is a follow-on leg. The A–E plan already PASSED PLAN-EVAL (verdict of record plan-eval.md, OpenHands run-28716441078-1, minimax M3). Evaluate ONLY the Topic F-ai addition:

  • plan.md → the "Topic F-ai — post-ratification integration" section (LD-F1…F6, OF-F1…F8, the F-ai DAG lane, supersession map, F-ai risk rows).
  • design/F-ai/{proposal,epic-and-issues,open-questions,agent-briefs}.md.
  • Confirm the F1-ai adversarial fixes landed at F2 (commit 1d3ca080). F1-ai-adversarial-review.md lists 8 findings (F1AI-01…08, verdict FAIL_PLAN); confirm each is fixed and no new gap was introduced.

Inputs (read, in this order)

  1. .llm/harness/gates/plan-gate.md — the checklist you enforce.
  2. .llm/harness/evaluator/verdict-definitions.md — verdict meanings incl. FAIL_PLAN.
  3. .llm/harness/evaluator/plan-protocol.md — your procedure.
  4. .llm/harness/gates/archetype-gate-matrix.md — the gate matrix.
  5. .llm/runs/plan-roadmap-expansion--seed/research.md, the F-ai ## Design worklog rows (n4/n5/n6) in worklog.md, and drift.md (entries FAI-1, FAI-2).
  6. .llm/runs/plan-roadmap-expansion--seed/plan.md (Topic F-ai section) + design/F-ai/*.
  7. .llm/runs/plan-roadmap-expansion--seed/F1-ai-adversarial-review.md.
  8. .github/labels.yml — for the F1AI-04 label-sync check (READ-ONLY).

Procedure

Walk the Plan-Gate checklist box by box; cite the exact plan location that satisfies each box or mark it unchecked. Run your OWN independent open-decision sweep for any F-ai decision that would force rework if deferred — if you find one the plan did not flag, that is an automatic unchecked box. Confirm the F-ai commit slices (FAI-0…17, 18 slices) are ordered, sized < 30, and each names its proving gate + files. Confirm the jsr-audit surface scan covers the F-ai public-surface deltas (plugins/ai publish:false→publishable, the @netscript/ai/otel GenAI adapter == Topic-B T9, @netscript/fresh/ai).

Specifically verify the eight F1-ai fixes: F1AI-01 (no invented beta.6 Topic-A "dashboard AI panel"; FAI-0…3 reframed as a parity floor for the OF-6 telemetry seam + stable DDX-19 handshake, not a Topic-A hard-dep), F1AI-02 (FAI-17 deps = Topic-B T3 + T6, T1 transitive — matches design/B-telemetry/epic-and-issues.md:156), F1AI-03 (single authoritative supersession headline: 15 KEEP · 1 FOLD (#257#379) · 0 close · 3 NEW (FAI-4/10/11)), F1AI-04 (label-sync owner action for epic:/wave: into .github/labels.yml; beta slices wave:v1 not wave:defer), F1AI-05 (archetype gate matrix present), F1AI-06 (OF-F3 safe-to-defer), F1AI-07 (FAI-9 = beta.6 capability merge gate + per-milestone gates), F1AI-08 (line-ref precision).

Output (this is the verdict of record)

Fill the template .llm/harness/templates/plan-eval.md and write it to a NEW file:
.llm/runs/plan-roadmap-expansion--seed/plan-eval-F-ai.md
Do NOT overwrite the existing plan-eval.md — that is the A–E verdict of record. Emit exactly one verdict: PASS or FAIL_PLAN (if FAIL_PLAN, list each unchecked box + the specific fix required). Commit it to this branch and post your summary as the PR comment.

Note on job status: if the run's agent exits 0 and the committed/summary artifact says PASS, that IS the verdict of record even if a later CI "Commit changes back to PR branch" / "Commit run trace" step reports the job as failed — that is a commit-back step failing, not a self-post decline. Do not re-run on that basis.

Do NOT create, close, or label any GitHub issue, epic, or milestone. This run is Phase-2-filing HARD-STOPPED and PR #397 stays DRAFT.

@rickylabs

rickylabs commented Jul 4, 2026

Copy link
Copy Markdown
Owner Author

OpenHands Agent — Completed

Model: openrouter/minimax/minimax-m3
Provider: OPENROUTER
Job status: failure
Trace: .llm/tmp/run/openhands/pr-397/run-28722662042-1 (committed to the PR branch)
Artifact: openhands-agent-28722662042-1

OpenHands run summary — run-28722662042-1

Run: plan(roadmap): roadmap-expansion — Fable 5 planning run
Trigger: @openhands-agent model=openrouter/minimax/minimax-m3 output=pr-comment use harness
Triggered at: PR #397 (plan/roadmap-expansion, draft) issue comment
Session kind: PLAN-EVAL (evaluator only, separate from Fable 5 author and WSL-Codex F1-ai reviewer)
Scope: Topic F-ai (AI suite) leg only — the A–E plan already PASSED PLAN-EVAL (verdict of
record plan-eval.md, OpenHands run-28716441078-1, minimax M3). This is the F-ai follow-on
evaluator pass; this session did NOT author, defend, or re-edit the plan.

Summary

Verdict: PASS for the F-ai leg of plan-roadmap-expansion--seed.

Walked the Plan-Gate checklist (.llm/harness/gates/plan-gate.md) box by box. Ran an
independent open-decision sweep for any F-ai decision the plan left out that would force rework
if deferred — found none. Verified the F2 commit 1d3ca080 (Fable 5, 2026-07-05) closed all
eight F1-ai findings (3 BLOCKER + 3 MAJOR + 1 MINOR + 1 NIT) and did not introduce a new gap.
Confirmed the 18-slice F-ai DAG (FAI-0…17, well under 30) is ordered, sized, and each slice
names its proving gate + files + dep edges. Confirmed the F-ai jsr-audit surface scan covers
all public-surface deltas (plugins/ai publish:false→publishable, the @netscript/ai/otel
GenAI-span adapter == Topic-B T9, the @netscript/fresh/ai FA3/FA4 runtime surfaces,
@netscript/fresh-ui ai registry, etc.) with the F-ai archetype gate matrix (§8.1) selected
per surface. The V3 commit-trail cleanup in agent-briefs is correct (F-ai agent-briefs has
zero commits.md references; A/B/E agent-briefs all use V3-correct "no commits.md" framing).

Open decisions (OF-F1…OF-F8): the only "must resolve now" fork is OF-F4 (milestones +
.github/labels.yml epic:/wave: sync — a pure owner action that blocks issue-filing, not
design). OF-F1 (E9/#248 co-ownership) is pre-resolved by LD-F5 with a documented single-issue
path. OF-F3 (Topic-A AI surface at beta.6) is correctly reclassified as safe-to-defer — the
F1-ai review removed the invented beta.6 Topic-A AI panel, so deferring OF-F3 changes no F-ai
slice or milestone. No FAIL_PLAN open decision in the F-ai set.

The A–E verdict of record (plan-eval.md) is unchanged. This is the F-ai leg verdict of
record (plan-eval-F-ai.md). Plan is ready for owner ratification; PR #397 stays draft. No
GitHub issues/epics/milestones/labels created, closed, or mutated (Phase-2 filing HARD-STOPPED).

Changes

  • Created: .llm/runs/plan-roadmap-expansion--seed/plan-eval-F-ai.md (138 lines, the F-ai
    leg verdict of record, filled from the .llm/harness/templates/plan-eval.md template).
    • Plan-Gate checklist table: 8/8 PASS (Research, Decisions locked, Open-decision sweep,
      Commit slices < 30, Risk register, Gate set selected, Deferred scope explicit,
      jsr-audit surface scan).
    • F1-ai fix verification table: 8/8 FIXED (F1AI-01…F1AI-08), with file:line evidence per fix.
    • Evaluator-run open-decision sweep: 11 candidate deferrals considered, 0 rework-forcing.
    • Verdict: PASS, with 6 non-blocking notes.
  • Committed: a1e66534 on plan/roadmap-expansion (1 file changed, 138 insertions; local
    branch ahead of origin/plan/roadmap-expansion by 1 commit; NOT pushed — the workflow's
    "Commit changes back to PR branch" CI step will pick it up per the trigger protocol).
  • Git config set locally (per AGENTS.md "openhands" identity default, no global config):
    user.name openhands, user.email openhands@all-hands.dev.

Validation

  • F2 commit verified: 1d3ca080 at origin/plan/roadmap-expansion HEAD (synced);
    subject "plan(F-ai): resolve F1-ai adversarial findings + V3 commits.md cleanup"; by Fable 5,
    2026-07-05. Working tree clean after this run's commit.
  • Read-only label-file check (F1AI-04): .github/labels.yml (182 lines) contains
    0 epic: label definitions and 0 wave: label definitions (only 1 comment-header
    mention). The F-ai epic-and-issues header (lines 20-30) names the owner action to add
    epic:ai-stack, epic:telemetry-revamp, and the wave:v1/wave:v1-min/wave:defer
    block. The F-ai slice labels themselves are correct (FAI-10/11/14 = wave:v1, FAI-17 =
    wave:defer for stable; no other F-ai slice carries wave:defer on a beta milestone).
  • F1AI-02 fix cross-verification: Topic-B T9's own declared deps at
    design/B-telemetry/epic-and-issues.md:156 and DAG :168 read Deps: T3, T6 /
    T3, T6 → T9. F-ai's FAI-17 dep statement now matches (T3 + T6, T1 transitive) across
    proposal, epic-and-issues, agent-briefs, OQ-1, and plan.md LD-F5.
  • F1AI-03 fix cross-verification: supersession headline 15 KEEP · 1 FOLD ([AI-stack FB4] fresh-ui: mcp-ui-widget (themed sandboxed ui:// iframe) #257[AI-stack FA4] @netscript/fresh/ai: createMcpAppCallHandler route (widget action -> tools/call, allowlist + stdio fallback + OTel) #379)
    · 0 close/supersede · 3 NEW (FAI-4/10/11)
    is single-source and mirrored verbatim across
    plan.md:177-181 and design/F-ai/epic-and-issues.md:331-334 + the table :323-333.
    [AI-stack FB6] @netscript/fresh-ui: interactive MCP-App bridge (widget-action → tools/call → re-render, deferred) #272 is KEEP stable (not a fold).
  • F1AI-05 fix cross-verification: archetype gate matrix §8.1 is present in
    design/F-ai/proposal.md:306-321 (surface → archetype → required gates); cross-referenced
    from plan.md:133-141.
  • F1AI-06 fix cross-verification: OF-F3 reclassified safe-to-defer in plan.md:161 and
    plan.md:171-173; OQ-3 rework classification in design/F-ai/open-questions.md:46-54.
  • F1AI-07 fix cross-verification: FAI-9 renamed "beta.6 capability merge gate";
    per-milestone gate table at design/F-ai/epic-and-issues.md:370-376; cross-referenced
    from plan.md:139-141.
  • F1AI-08 fix cross-verification: exact file:line citations throughout
    (stream-proxy.stub.ts:16-64, sandbox.ts:71-77, memory.ts:70, aiContractV1:377-379,
    deno.json:25, packages/fresh/deno.json:15-16).
  • V3 bookkeeping verified: design/F-ai/agent-briefs.md has zero commits.md refs;
    remaining commits.md mentions in design/{A-dashboard,B-telemetry,E-desktop}/agent-briefs.md
    are V3-correct ("no commits.md" framing). Drift V3 entry recorded at drift.md:197-204.
  • No code validation run — this is a planning-only evaluator session; per the trigger
    protocol, no implementation gate was run. Runtime, lint, e2e, and deno doc --lint are
    deferred to per-slice IMPL-EVAL.
  • No framework code changes.llm/runs/plan-roadmap-expansion--seed/plan-eval-F-ai.md
    is the only file in this run's commit.

Responses to review comments or issue comments

  • Trigger comment (@openhands-agent … use harness): this session. The PR comment
    summary is in the "Summary for PR comment" block at the bottom of
    .llm/runs/plan-roadmap-expansion--seed/plan-eval-F-ai.md. The workflow owns the actual
    PR-comment posting (the OpenHands API tool is not invoked directly per the operational
    contract).
  • No other review comments are present on PR plan(roadmap): roadmap-expansion — Fable 5 planning run (5 backlog features → Road-to-0.0.1-stable) #397 (the PR is draft and the prior
    PLAN-EVAL session's "Commit changes back to PR branch" step was pending — this run's
    verdict is the F-ai follow-on).

Remaining risks

These are non-blocking observations captured in the verdict's "Notes" section, surfaced here
for downstream IMPL-EVAL sessions and the next F-cycle (if any):

  1. Generative-UI sandbox primitive is implicitly locked to eis-chat's
    iframe sandbox="allow-scripts" + srcdoc + zero-src pattern (web-platform,
    wrap-don't-reinvent). FAI-6 IMPL-EVAL should record the exact sandbox= attribute set
    in the slice's acceptance evidence so FAI-8 (which shares the iframe primitive) inherits
    a documented posture, not an implicit one.
  2. TanStack AI exact-pin lint rule (OF-F7) should be added to the
    .llm/tools/fitness/-equivalent TanStack-touching slice's deno task deps:why evidence
    before FAI-0 lands, so the pin discipline is part of the FAI-2 e2e acceptance (drift at
    the FAI-0/FAI-1 boundary is otherwise easy to miss).
  3. Supersession map's NEW issue list (FAI-4 / FAI-10 / FAI-11) requires explicit
    owner approval before any Phase-2 filing. The map says "DRAFT — owner approves before
    any Phase-2 filing; NO closes until approved; NEVER a closing keyword on epic: NetScript AI Stack — first-class AI runtime, chat & plugin seams (anchor #219) #238." The
    next run that actually files must surface the approval back to the owner explicitly,
    not as a single "drafts are ready" line.
  4. Per-milestone gate table (design/F-ai/epic-and-issues.md:370-376) is the
    authoritative gate-set source for the F-ai epic; IMPL-EVAL must use this table, not
    the F-ai epic entry's single "merge-gate" column, to avoid the F1AI-07 conflation
    (FAI-9 gates beta.6 only, not the whole epic).
  5. No wave:defer on F-ai beta slices — verified. Only FAI-17 (stable) carries
    wave:defer; this is the canonical label per .agents/skills/netscript-pr/SKILL.md:190-210
    for stable-tier slices.
  6. The "Commit changes back to PR branch" CI step for this run may report failure if it
    has a race with the prior run-28716441078-1's commit-back step (a known CI fragility, not
    a verdict issue). Per the trigger's "Note on job status" instruction, if the agent exits
    0 and the committed verdict artifact says PASS, that IS the verdict of record — the
    later CI failure is a commit-back step failing, not a self-post decline. Do not re-run
    on that basis.
  7. The A–E plan's two "must resolve now" forks (OF-5 and OF-10) are already owner-ratified
    per the plan.md header. If the owner reverts either ratification, the F-ai plan is not
    affected directly, but the cross-topic edges (OF-F1 to Topic-B T9, OF-F2 to Topic-A DDX-19
    stable handshake) need to be re-verified for the reratified A–E plan. Not a current
    concern — the A–E ratifications stand as of 2026-07-04.

Run: https://github.com/rickylabs/netscript/actions/runs/28722662042

Stage G of the Topic F-ai leg. Transcribes the independent OpenHands PLAN-EVAL
verdict (run-28722662042-1, minimax M3, separate session) into
plan-eval-F-ai.md with a provenance header. Verdict: PASS.

- Plan-Gate 8/8 satisfied; all 8 F1-ai findings (F1AI-01..08) verified FIXED
  at F2 (1d3ca08) with file:line evidence; independent 11-candidate
  open-decision sweep found 0 rework-forcing deferrals.
- The runner committed the file as a1e66534 but the CI "commit back to PR
  branch" step failed to push it (known caveat); content recovered verbatim
  from run artifact openhands-agent-28722662042-1 and landed here.
- A-E plan-eval.md (run-28716441078-1) unchanged. Plan now fully ratified
  (A-E PASS + F-ai PASS).

Planning artifact only; no framework code, no GitHub mutation. PR #397 stays
DRAFT and Phase-2 filing remains HARD-STOPPED pending owner authorization.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
…tation)

Five reversible planning artifacts so the owner's own OF-5/OF-10 go-message
can drive the whole Phase-2 batch in one shot. ZERO GitHub mutation, PR #397
stays DRAFT, no undraft/merge/filing/closes/impl-launch this session.

- OWNER-DECISION-BRIEF.md: 4 primary unblock inputs (OF-5 opt-in SDK, OF-10
  per-capability, milestones=create beta.6/7/8, labels=3 new epic labels) each
  one-line-answerable + rework cost of each alternative; + secondary forks.
- SUPERSESSION-MAP.md: audit of 47 open issues -> 41 KEEP / 2 FOLD
  (#257->#379, #375->#E2, both via downstream PR closing-keyword) / 0
  filing-time closes; live-read corrections captured.
- filing-manifest.md: per-epic create/update index (telemetry T1-9, dashboard
  DDX-0..19, docs, desktop #E1-8, F-ai re-sequence + FAI-4/10/11) with
  taxonomy+milestone+closing-keyword+DAG per issue; empty close-list.
- proposed-labels-patch.md: unified diff for .github/labels.yml (NOT applied to
  the live file); repo create = only 3 new epic labels.
- beta5-launch-brief.md: telemetry supervisor brief (T1->T2 wave, crit path to
  beta.6, WSL Codex draft-PR-per-slice, separate-session eval).

Live-read corrections: 0.0.1-beta.5 already exists (create beta.6/7/8 only);
wave:* + epic:ai-stack/deployment already live on the repo (labels.yml file is
stale). Both contradict stale design-doc assumptions.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
@rickylabs
rickylabs marked this pull request as ready for review July 5, 2026 06:10
@rickylabs
rickylabs merged commit 029fbf3 into main Jul 5, 2026
6 checks passed
@rickylabs
rickylabs deleted the plan/roadmap-expansion branch July 5, 2026 06:11
rickylabs added a commit that referenced this pull request Jul 5, 2026
#471)

* harness(seed-run): S1 run scaffolding — supervisor.md-first activation

Run dir for codifying the plan-roadmap-expansion--seed pipeline as a
reusable v3 run shape. supervisor.md written before any other artifact,
dogfooding the identity gate the exemplar run violated (drift #2).
research.md distills the exemplar stages with citations; plan.md locks
LD-1..LD-8 (name=seed run, home=workflow/, contracts-over-template,
lane bindings by reference, dogfood acceptance).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN

* harness(seed-run): S2 profile — seed-run.md + supervisor.md template

New v3 run shape workflow/seed-run.md: planning-only runs whose
deliverable is a GitHub board. Stage contracts A-I (bootstrap ->
discovery corpus -> synthesis -> deep-dive packs -> plan lock ->
adversarial -> PLAN-EVAL -> owner ratify + one-shot filing -> handoff),
lane bindings by reference to lane-policy.md only, evidence-citation
gate, drafts-only-until-ratification boundary, GitHub-authority-after-
filing rule, scale-to-fit + when-NOT-to-use, dogfood acceptance.
Promoted from plan-roadmap-expansion--seed (PR #397; board #399-#461).

templates/supervisor.md fills the systemic gap: the identity file is
mandatory (lane-policy) but had no template, which is the plausible
root cause of the exemplar never writing one.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN

* harness(seed-run): S3 wiring — activation, README, harness skill

activation.md: bootstrap step 10 (seed-run branch) + supervisor.md
added to Mandatory Artifacts (it was mandated by lane-policy but absent
from the activation checklist - the second half of the systemic gap).
README.md: Start Here pointer + supervisor.md in the artifact list.
netscript-harness SKILL: Key Concepts row, decision-tree branch,
Reference Files row; .claude/skills mirror regenerated via
sync-claude-skills.ts (SYNCED). validate-claude-surface.ts all-ok.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN

* harness(seed-run): S4 adversarial brief — unoriented Codex review launch

Attack surfaces: contradictions with harness law, restated bindings,
fresh-agent executability, G->H ratification-boundary loopholes,
template soundness, reference integrity, mirror integrity, overclaims.
Findings-only contract; no repo edits by the reviewer.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN

* harness(seed-run): S5 adversarial fixes — all 8 findings accepted

Codex verdict: 0 blockers, 6 major, 2 minor, mirror byte-identical.
Fixes per adversarial-triage.md:
1. Mutation boundary split into two surfaces: the run draft PR is the
   always-writable commit trail; the board (issues/epics/milestones/
   repo labels) is untouchable before stage H.
2. plan/ registered in the netscript-pr branch taxonomy (seed-run only)
   instead of the harness doc contradicting the canonical skill.
3. Stage B now carries the Tier-C hard rule: workflow.js committed
   under <run-dir>/workflows/ before execution, or the corpus is not
   Stage-B proof.
4. Stage F de-hardcoded from Tier D to distinct-model invariants
   (unoriented, separate session, findings-only); tier per supervisor.md.
5. phase-registry.md scoped to multi-group runs per activation step 9.
6. drift #3: the eval exception is precisely PLAN-EVAL-skipped
   (owner-directed); OpenHands verdict retained; ships for ratification.
7. worklog/context-pack refreshed to PR #471 reality.
8. FILING-LOG.md spelling unified.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN

* harness(seed-run): S6 — transcribe IMPL-EVAL PASS verdict

OpenHands separate-session eval (minimax-m3) on PR #471: PASS,
"Recommend merge". 8/8 triage cross-check, LD-1..LD-8 conformance,
2 non-blocking observations deferred to OD-1. Commit-back step failed
(known mode) — verdict transcribed from the agent summary comment;
branch tip verified churn-free at 604e8ae. Awaiting owner ratification.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant