Skip to content

Field Validation

Cory Ebert edited this page Sep 11, 2026 · 3 revisions

Field Validation

Shell examples run from the target project root. For scripts/forgeflow/ commands, use the helper path from your ForgeFlow checkout, or replace that prefix with "${CODEX_HOME:-$HOME/.codex}/forgeflow/scripts/forgeflow/" for Codex and "$HOME/.claude/forgeflow/scripts/forgeflow/" for Claude Code. Run JavaScript helpers with node and shell helpers with bash. Replace <project> with the actual project folder name before running a placeholder example.

Use this plan to validate Forgeflow on real branches across representative project types. The goal is to collect comparable local evidence, not to publish raw project data.

Trial Matrix

Run at least one branch trial in each project type:

Project Type Example Change Useful Forgeflow Signals
Frontend app form state, API integration, accessibility fix Designer findings, service path checks, accessibility classes
API service auth boundary, validation, persistence change Guardian findings, Builder data-layer findings, Verifier decisions
Monorepo package boundary, shared config, generated clients Coordinator coordination notes, scope manifest size, budget warnings
Docs/config command docs, release docs, CI config skip/thin routing, release-check output, low-noise review
Release prep version bump, changelog, installer docs release-check pass/fail, health/version guidance, public summary quality

Per-Branch Steps

For each branch:

  1. Run Branch Trial.
  2. Save a local outcome record with review.workflow set to forgeflow.
  3. If possible, record comparable no-agent and single-agent outcomes for the same change using Workflow Comparison.
  4. Generate a public-safe evaluation summary:
scripts/forgeflow/render-evaluation-report.js \
  --outcomes .forgeflow/<project>/review-outcomes.jsonl \
  --context-root .forgeflow \
  --budget-config .forgeflow-budget.json \
  --public \
  --out .forgeflow/<project>/evaluation-summary.md
  1. Review the summary using Evaluation Sharing.
  2. Store the summary using Evaluation Summary Collection.
  3. Record first-run friction separately from review quality using First-Run Friction.

Friction Log

Track friction in a local note, issue, or spreadsheet. Use the fuller template in First-Run Friction, or start with these fields:

project_type:
runtime: claude-code | codex | both
install_path: update-forgeflow | template-installer | existing-install
branch_shape:
review_mode:
context_budget_status:
time_to_first_review_minutes:
blocked_by:
fix_category: install | health | docs | template-installer | agent-routing | context-budget | other
notes:

Do not include secrets, private URLs, or source snippets.

Aggregate Evidence

For each project type, keep only aggregate values in shareable notes:

  • reviewed changes
  • confirmed findings
  • rejected findings
  • false positive rate
  • average review minutes
  • context percent saved
  • budget violations
  • first-run blockers by category

Raw review-outcomes.jsonl, context packets, and telemetry rows should stay local unless the receiving audience is allowed to see the underlying project context.

Use Evaluation Summary Collection to keep summaries organized during field validation.

Exit Criteria

Field validation is ready to turn into product fixes when repeated friction appears in the same category. Examples:

  • install friction repeats across Codex trials
  • health checks pass but users still miss restart requirements
  • context budget violations repeat in monorepos
  • public summaries need manual cleanup every time
  • routing misses a class of files in more than one project

When that happens, create a targeted fix in the relevant install, health, docs, routing, or context helper. Use Friction To Fix to classify the issue and choose the smallest fix layer.

Clone this wiki locally