harness: codify the seed run — discovery-to-orchestration v3 run shape#471
Conversation
Run dir for codifying the plan-roadmap-expansion--seed pipeline as a reusable v3 run shape. supervisor.md written before any other artifact, dogfooding the identity gate the exemplar run violated (drift #2). research.md distills the exemplar stages with citations; plan.md locks LD-1..LD-8 (name=seed run, home=workflow/, contracts-over-template, lane bindings by reference, dogfood acceptance). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
New v3 run shape workflow/seed-run.md: planning-only runs whose deliverable is a GitHub board. Stage contracts A-I (bootstrap -> discovery corpus -> synthesis -> deep-dive packs -> plan lock -> adversarial -> PLAN-EVAL -> owner ratify + one-shot filing -> handoff), lane bindings by reference to lane-policy.md only, evidence-citation gate, drafts-only-until-ratification boundary, GitHub-authority-after- filing rule, scale-to-fit + when-NOT-to-use, dogfood acceptance. Promoted from plan-roadmap-expansion--seed (PR #397; board #399-#461). templates/supervisor.md fills the systemic gap: the identity file is mandatory (lane-policy) but had no template, which is the plausible root cause of the exemplar never writing one. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
activation.md: bootstrap step 10 (seed-run branch) + supervisor.md added to Mandatory Artifacts (it was mandated by lane-policy but absent from the activation checklist - the second half of the systemic gap). README.md: Start Here pointer + supervisor.md in the artifact list. netscript-harness SKILL: Key Concepts row, decision-tree branch, Reference Files row; .claude/skills mirror regenerated via sync-claude-skills.ts (SYNCED). validate-claude-surface.ts all-ok. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
Slice trail (S1–S3)
Next: WSL Codex adversarial pass → fixes → OpenHands separate-session eval → owner ratification. |
Attack surfaces: contradictions with harness law, restated bindings, fresh-agent executability, G->H ratification-boundary loopholes, template soundness, reference integrity, mirror integrity, overclaims. Findings-only contract; no repo edits by the reviewer. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
Codex verdict: 0 blockers, 6 major, 2 minor, mirror byte-identical. Fixes per adversarial-triage.md: 1. Mutation boundary split into two surfaces: the run draft PR is the always-writable commit trail; the board (issues/epics/milestones/ repo labels) is untouchable before stage H. 2. plan/ registered in the netscript-pr branch taxonomy (seed-run only) instead of the harness doc contradicting the canonical skill. 3. Stage B now carries the Tier-C hard rule: workflow.js committed under <run-dir>/workflows/ before execution, or the corpus is not Stage-B proof. 4. Stage F de-hardcoded from Tier D to distinct-model invariants (unoriented, separate session, findings-only); tier per supervisor.md. 5. phase-registry.md scoped to multi-group runs per activation step 9. 6. drift #3: the eval exception is precisely PLAN-EVAL-skipped (owner-directed); OpenHands verdict retained; ships for ratification. 7. worklog/context-pack refreshed to PR #471 reality. 8. FILING-LOG.md spelling unified. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
S4–S5: adversarial pass complete, all findings fixedS4 ( S5 (
Next: OpenHands separate-session eval, then owner ratification. This run does not self-certify or merge. |
|
@openhands-agent model=openrouter/minimax/minimax-m3 output=pr-comment You are the separate-session harness evaluator for this docs-only PR. Do NOT implement fixes; produce a verdict. Read, in order:
Evaluate:
Output a PR comment: verdict |
OpenHands Agent — CompletedModel: openrouter/minimax/minimax-m3 OpenHands separate-session harness evaluator — summary\n\nRun:
|
OpenHands separate-session eval (minimax-m3) on PR #471: PASS, "Recommend merge". 8/8 triage cross-check, LD-1..LD-8 conformance, 2 non-blocking observations deferred to OD-1. Commit-back step failed (known mode) — verdict transcribed from the agent summary comment; branch tip verified churn-free at 604e8ae. Awaiting owner ratification. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN
S6 — IMPL-EVAL PASS · ready for owner ratificationVerdict: PASS ("Recommend merge") — OpenHands separate session, minimax-m3, run 28744305384. The job's Evaluator confirmations: 5-commit slice trail matches worklog exactly · zero Two non-blocking evaluator observations deferred to OD-1 (filing-manifest template promotion): Stage B lacks a cross-ref to the Owner decision (this run does not self-certify or merge)
Slice trail: S1 |
What
Codifies the agentic pipeline that produced
plan-roadmap-expansion--seed(PR #397; board#399–#461) as a reusable Harness v3 run shape: the seed run.
.llm/harness/workflow/seed-run.md— the profile. Stage contracts A–I (bootstrap → discoverycorpus → synthesis → deep-dive packs → plan lock → adversarial → PLAN-EVAL → owner ratify +
one-shot filing → handoff), lane bindings by reference to
lane-policy.mdonly, evidence-citationgate, drafts-only-until-ratification boundary, GitHub-authority-after-filing rule, scale-to-fit +
"when NOT to use", landmine pointers, dogfood acceptance criterion.
.llm/harness/templates/supervisor.md— template for the mandatory run-identity file. Thefile is mandated by
lane-policy.mdbut had no template — the plausible root cause of theexemplar run never writing one (its supervisor identity had to be recovered by transcript
search).
workflow/activation.mdnow also listssupervisor.mdunder Mandatory Artifacts.workflow/activation.md(bootstrap step 10), harnessREADME.md,netscript-harnessSKILL (+ regenerated.claude/skills/mirror).Why
The exemplar run took a replan from zero to a fully-planned, GitHub-native, implementation-ready
board in one shot — empirical discovery (repo + docs + market), synthesized objectives, per-epic
design packs, adversarial + evaluator verdicts, owner-ratified one-shot filing. Owner directive:
make that pipeline repeatable for every big triage / deep-search / plan / orchestration job. This
PR freezes the stage contracts (not the exemplar's folder tree) so a fresh supervisor can
execute A→I from the doc alone.
Run artifacts
.llm/runs/harness-seed-run-profile--codify/—supervisor.mdwritten first (dogfooding the gatethe exemplar missed),
research.md(exemplar distillation with citations),plan.md(lockeddecisions LD-1..LD-8), worklog/drift/context-pack.
Gates
validate-claude-surface.ts— all 5 checks ok (skills mirror SYNCED, 17 skills).seed-run.mdverified to resolve..llm/harness/**/*.mdis outside the repo fmt surface (deno.jsonfmt.includeispackages/plugins ts,tsx); the 29-file raw-fmt drift there is pre-existing and untouched
(non-verdict per AGENTS.md).
Evaluation plan (this PR does not self-certify)
WSL Codex adversarial pass → fixes → OpenHands separate-session eval → owner ratification. Docs-only
change: no
packages//plugins/source touched.🤖 Generated with Claude Code
https://claude.ai/code/session_012wKHquACkXnWPDgJYhhFjN