-
Notifications
You must be signed in to change notification settings - Fork 3
eval_removal_plan
Goal: delete the file/run-based workflow evaluation ("eval") subsystem from
mcp-agent-builder-go, with outcome measurement continuing to come from the
producing steps' own stored outputs — no mandated route, step, or table.
Non-goal: routing evaluation (routing-evaluation.json, RoutingEvaluatedEvent,
deterministic routing, route tracing) is a separate feature and stays. Note this
does NOT include routeEvaluations.ts, which despite its name implements
eval-step route gating only (see Phase 5) — it is deleted.
Durable workflow outputs already land in the database, so Pulse needs measurement data, not a particular plan shape. Any normal workflow step or scheduled collector may produce a run-scoped, evidence-backed measurement; Pulse does not depend on which step or route produced it. Target architecture:
ordinary workflow steps (producers)
↓
run-scoped database outputs (source of truth, stays in place)
↓
Pulse / reports / Goal Advisor read producer outputs directly;
producers call record_goal_observations only when a value should
enter the generic goal-progress history
Final review rejected the earlier mandatory-migration design (an earlier draft
of contract 1.0.43:
a measurement-router → measure-outcomes route plus a workflow_metrics
table on every workflow). It imposed topology most workflows don't need,
referenced the nonexistent add_regular_step tool, assumed table-scoped DB
authorization that doesn't exist (steps get database-wide read-write), assumed
immutable run identity that producers don't guarantee (RUN_FOLDER reuses
iteration-0), and described a metrics table no runtime code implements —
Pulse and Goal Advisor consume workflow_goal_metrics /
pulse_goal_observations through get_goal_metrics. That mandatory migration
was deleted. No replacement contract migration is required: producer-owned
measurement is ordinary workflow design guidance, not an unattended pre-run
rewrite. Legacy evaluation dependencies that genuinely need redesign are
handled explicitly in Builder instead of blocking every workflow schedule.
Replacement contract (flexible placement, no mandated topology):
- Reuse an existing step's DB output when it already measures the outcome.
- Add
record_goal_observationsto that producer only when the value should enter Pulse's generic goal-progress history. - Put route-specific measurements in existing route-local steps.
- Use a shared convergence step only when multiple routes genuinely require the same calculation.
- Add a dedicated measurement step only when no existing step can safely own it.
- Make no topology change when existing measurement is already sufficient.
- Flag ambiguous evals for manual migration rather than inventing steps or routes.
- Retain old evaluation artifacts as read-only history unless a separate cleanup migration is approved.
What is still lost (no replacement): re-grading an old run after completion, grading without modifying or influencing the run, hiding learnings from the grader, keeping evaluation cost/failures separate from execution, and comparing multiple evaluator versions against the same archived output. Full removal is only correct if none of these are product requirements.
All compatibility decisions are made here, BEFORE any deletion. Deleting tools
first would strand older workspace upgrades and saved schedules that still
invoke run_full_evaluation.
- Contract migrations (
cmd/server/workflow_version_upgrades.go): per migration that invokes eval tools, decide guard (skip when eval is gone) or retire (if every supported workspace is past that contract version). - Saved schedules: rewrite schedule instructions that assume an eval pass
(e.g. Mututal-Fund "Weekly Saturday Portfolio Sync" sequences steps
"including its eval pass"); confirm no enabled schedule depends on
run_full_evaluationoutput. - Workspace data: leave
evaluation/dirs,evaluation_report.jsonfiles,costs/evaluation/ledgers, andeval_resultsrows inert on disk, or ship a one-shot cleanup migration. Either way, decide before deleting the code that reads them. - Backup scopes: drop
evaluation/evaluation-planfrom the trading and check-form backup/sync folder lists (or accept syncing stale dirs). - Deprecate at runtime: disable automatic eval by default and observe for one or two release cycles to surface hidden consumers before deleting anything.
Deprecation mechanics — DisableEval is a plain bool
(controller_types.go), and omitted tool args decode to false
(planning_exports.go), so "default disabled with explicit opt-in" cannot be
expressed with the current flag: omitted and explicit-false are
indistinguishable. Pick one before implementing:
- (a) Add
EnableEval *bool(nil= default disabled during deprecation, explicittrueopts in), or - (b) Keep
DisableEvalbut explicitly default it totrueat every entry point (run_full_workflow, single-step execute paths, scheduler, webhook).
Also hide eval mode in the frontend (workflowMode: 'eval', eval canvas,
EvaluationPopup); leave components in place until Phase 5.
Guidance only; no topology or contract migration is mandated. Must complete before Phases 2–5 delete what it replaces.
- Per-legacy-eval placement follows the Direction rules: reuse a producer's
DB output, add
record_goal_observationsonly for Pulse history, route-localize, converge only when shared, dedicate a step only when no producer can own it, otherwise change nothing, and flag ambiguity for manual migration. - Repoint Goal Advisor and report/cost guidance from eval reports /
eval_resultsto producer outputs andget_goal_metrics. - Write the measurement guidance (replacement for
evaluation-plan.md):guidance/templates/system/measurement-plan.md— placement rules, run scope and evidence, and legacy artifacts as read-only history. - Do not add a contract migration for this retirement. A universal pre-run agent turn cannot safely infer whether legacy measurements remain relevant and must not block unrelated workflow execution. Use Builder for an explicit, operator-visible redesign when a workflow still depends on a retired evaluation artifact.
Delete wholesale (no callers outside the eval subsystem except the entry points removed in Phase 3):
-
agent_go/pkg/orchestrator/agents/workflow/step_based_workflow/evaluation_execution.go(ExecuteEvaluationOnly,MaybeRunAutoEvaluation, report phase) -
agent_go/pkg/orchestrator/agents/workflow/step_based_workflow/evaluation_plan_tool.go(add_evaluation_step,update_evaluation_plan,delete_evaluation_step) -
agent_go/pkg/orchestrator/agents/workflow/step_based_workflow/evaluation_helpers.go(validate_evaluation_planregistration) agent_go/pkg/orchestrator/agents/workflow/step_based_workflow/evaluation_types.goagent_go/pkg/orchestrator/agents/workflow/step_based_workflow/eval_results_storage.goagent_go/pkg/orchestrator/agents/workflow/step_based_workflow/evaluation_score_storage.goagent_go/cmd/server/evaluation_score_storage.go- Eval-only tests:
evaluation_output_content_test.go,evaluation_plan_tool_test.go,evaluation_report_enrichment_test.go,evaluation_run_metadata_order_test.go,evaluation_skipped_sentinel_test.go,eval_results_storage_test.go,evaluation_score_storage_test.go(both dirs).
-
controller_batch_execution.go: delete the auto-eval block after group completion (keepfinalizeRunMetadata, which is needed regardless). -
controller_types.go: delete theDisableEvaloption (or theEnableEvalreplacement from Phase 0, once nothing reads it). -
webhook_step.go: delete theDisableEval = trueline (standalone steps simply never had eval). - Delete every
isEvaluationModebranch (~20 files). Highest-density files:controller_execution.go(10),interactive_workshop_manager.go(7),controller_agent_factory.go(6),controller_message_sequence.go(5),controller_orchestrator.go(4),controller_workshop.go(3),controller.go(3, incl. the field itself),step_config.go(2, collapse toplanning/),planning_exports.go(2), plus single branches incode_layout.go,controller_scripted.go,plan_snapshot.go,prompt_health.go,reflection_turn_run.go,run_provenance.go,workflow_continuation_recovery.go,workshop_retry_recovery.go, andcmd/server/workflow_phase_tools.go. -
controller_agent_factory.go: deleteTARGET_RUN_PATHread-path grants and the eval medium-tier default. -
controller_execution.go: deleteTARGET_RUN_PATHprompt injection and theIsEvaluationModetemplate var. -
controller_learning_helpers.go: delete theisEvalModeparameter fromcanReadLearnings/canWriteLearnings/resolveExecutionLearningsAccess/shouldDirectWriteLearnings(keep the routing-step exclusions). -
controller_agent_factory.go:resolveEffectiveDBAccessalready ignores its eval parameters — simplify its signature. -
agent_go/pkg/workflowtypes/types.go: deleteWorkflowStatusEvalExecutionandWorkflowStatusEvalBuilder. -
agent_go/pkg/orchestrator/types/workflow_orchestrator.go: delete theExecuteEvaluationOnlycall site. -
planning_exports.go: deleteRegisterRunFullEvaluationTooland the eval controller call site. -
planning_agent.go: deleteadd_evaluation_stepregistration. -
schedule_collision_guard.go: delete eval tool entries.
-
cmd/server/workflow_review_data.go: delete theevaluationssection (EvaluationReport*structs,CostScopeEvaluationreads,evaluation/runstiming load,loadWorkflowEvaluationReports). - Routes: delete
/workflow/pulse-eval-results(server.go, handlerhandleGetPulseEvalResultsinpulse_worklist.go) and/workflow/evaluation-reports(server.go, handlerhandleGetEvaluationReportsinworkflow.go). -
cmd/server/workspace_state.go: deleteeval_datafrom overview responses. -
cmd/server/external_tools.go: dropevaluation/*fromexternalPlanPaths. -
cmd/server/workflow_backup.go: dropevaluation/evaluation_plan.json. -
cmd/server/workflow_manifest.go,report_preview_metrics*,slack_retention.go,webhook_retention.go: delete eval references. - Cost ledger/observer: delete
CostScopeEvaluationandcosts/evaluationhandling. -
cmd/server/workflow_phase_tools.go: delete the eval-session registration. -
cmd/server/toolset_invariant_test.go: drop the five eval tools from the expected set. -
cmd/server/workflow_version_upgrades.go: execute the Phase 0 decision (guard or retire each eval-invoking migration). -
schemas/auto-improvement.schema.json: remove the deprecatedeval_stepsource, leavingmeasurementandtelemetry.
Delete:
frontend/src/components/workflow/hooks/useEvaluationPlanData.tsfrontend/src/components/workflow/hooks/useEvaluationPlanToFlow.tsfrontend/src/components/workflow/nodes/EvaluationNode.tsxfrontend/src/components/workflow/nodes/CompactEvaluationNodes.tsxfrontend/src/components/workflow/EvaluationPopup.tsxfrontend/src/components/workflow/PulseEvalSummary.tsx-
frontend/src/components/workflow/canvas/evaluationLayout.ts(+ test) -
frontend/src/components/workflow/canvas/routeEvaluations.ts(+ test) — despite the name this is eval-only: it importsEvaluationStepand implementsapplies_to_routesgating. Keeping it after deletingEvaluationStepbreaks compilation. -
frontend/src/utils/evaluationReport.ts(+ test)
Scrub (keep ordinary routing logic, delete eval branches):
-
canvas/routeTrace.ts(importsevaluationMatchesRoute; handlesisEvaluationStep/evaluation-groupnodes) -
canvas/triggerLayout.ts(importsEvaluationStep; handles eval nodes) -
canvas/WorkflowCanvas.tsx,canvas/WorkspaceViewHost.tsx,canvas/workspaceViewData.ts,hooks/usePlanToFlow.ts,nodes/RoutingStepNode.tsx,nodes/StepNode.tsx,nodes/index.ts,utils/stepConfigMatching.ts -
components/WorkflowsOverviewPage.tsx: removeevalDatastate,EvaluationPopupusage,onOpenEvalhandlers. -
components/workflow/ReportViewer.tsx: remove pulse-eval-results consumption. -
services/api.ts: removegetPulseEvalResults,getEvaluationReports,eval_datahandling. -
services/api-types.ts: removeEvaluationPlan,EvaluationStep,EvaluationReportsResponse, and related report types. -
stores/useWorkflowStore.ts: remove'eval'fromworkflowMode,evaluationPlanstate,loadEvaluationPlan. - Keep
routeTrace.ts/triggerLayout.tsroute-tracing logic androuteEvaluations-free routing display — only the eval-gating code goes. - Rebuild frontend and delete the stale
agent_go/static/assets/PulseEvalSummary-*.jsbundle.
- Delete:
agent_go/cmd/server/guidance/templates/system/evaluation-plan.md,agent_go/cmd/server/guidance/templates/improve/improve-evaluation.md. - Scrub eval references in:
report/improve-report.md,review/review-artifact-drift.md,system/file-layout.md,system/plan-drift-review.md,system/strategy-auditor.md,system/workshop-mode-flow.md,system/workspace-views.md,guidance/guidance.go,cmd/server/instructions.go(incl. Goal Advisor prompt,evaluation/runs/layout,run_retention_countwording), andinteractive_workshop_manager.goprompts referencing eval plans/reports. Repoint each to the Phase 1 measurement guidance. - Docs: delete
docs/workflow/evaluation_system.md; scrub eval mentions in the remainingdocs/workflow/*.md(about 20 files reference it). - This file (
docs/workflow/eval_removal_plan.md) should be deleted once the removal is complete.
Execute the Phase 0 decision: leave evaluation/ dirs,
evaluation_report.json files, costs/evaluation/ ledgers, and eval_results
rows inert, or run the one-shot cleanup migration. Locally affected: 7
workflows with live eval plans (trading has 18 steps), 122
evaluation_report.json files (~165 MB), eval_results rows in 6 DBs, and
eval-created domain tables (HDFC eval_behavior_audit / eval_login_download;
trading eval_profile_integrity / eval_runtime_duration /
eval_trace_integrity / eval_sc). Schedules that previously got an automatic
eval pass keep outcome tracking through their producers' existing outputs
(plus record_goal_observations where Pulse history is wanted); no new step
or route is required.
- ~46 Go test files mention evaluation. Delete the eval-only tests listed in Phase 2; update the rest (toolset invariants, workshop registration, plan snapshot/changelog, execution-only, webhook, scheduler, report metrics tests) to drop eval fixtures and assertions.
- Frontend: delete
evaluationLayout.test.ts,evaluationReport.test.ts,routeEvaluations.test.ts,RouteEvaluations.test.tsx,routeTrace.test.tseval cases,triggerLayout.test.tseval cases; update store/canvas tests.
- Module-scoped builds and tests pass (the repo root holds only
go.work; there are no root packages, so barego build ./...fails):(cd agent_go && go build ./... && go test ./...)(cd workspace && go build ./... && go test ./...)
- Frontend
tsc+ unit tests pass; production bundle contains noPulseEvalSummarychunk. - Grep gates. These strings must return zero in non-test Go and
non-test
frontend/src:-
evaluation_plan,isEvaluationMode,IsEvaluationMode,TARGET_RUN_PATH,ExecuteEvaluationOnly,MaybeRunAutoEvaluation -
run_full_evaluation,validate_evaluation_plan,add_evaluation_step,update_evaluation_plan,delete_evaluation_step -
CostScopeEvaluation,EvaluationStep,EvaluationPlan,EvaluationReportsResponse,EvalExecution -
pulse-eval-results,evaluation-reports,eval_data,EvalData,getEvaluationReports,getPulseEvalResults,applies_to_routes,routeEvaluations,evaluationMatchesRoute -
measurement-router,measure-outcomes,workflow_metrics,upgrade-measurement-route(the rejected mandatory topology) -
eval_step,costs/evaluation,evaluation/runs,evaluation_report,eval_results,eval_route_scores,improve-evaluation,evaluation-planIn guidance templates, prompts, schemas, tests, and active docs, the same strings survive only on lines that also say legacy/retired/read-only history (or in assertions pinning retired behavior). Explicitly reviewed exceptions: this plan; dated historical records underdocs/bugs/anddocs/audits/; thepulse_v2proposal docs and the retired design section inpersistent_stores_design.md;registry.test.ts(asserts the commands are retired);external_tools_test.go(asserts the retired plan path stays write-protected);declared_execution_mode_strip_test.go(asserts retired files stay byte-identical);slack_review_test.go(asserts retention leaves legacy evidence inert);ManualWorkspaceRefresh.test.ts(asserts the eval refresh is absent); the upgrade-chain guard test and the measurement-guidance render test (assert the rejected topology is absent);reflection_turn_test.go(asserts the retired table stays platform-owned);isEvaluationAgentEvent/inferTrackedExecutionKind(classify historical records only);isPlatformOwnedTable/planDriftReservedTables(keep retired eval tables out of step reach).
-
- Allowed survivors:
routing-evaluation.json,RoutingEvaluatedEvent,routeTrace.ts/triggerLayout.tsrouting logic,RoutingEvaluationCounts, cron "evaluation" wording inscheduler.go,ToolResponseEvaluation(external mcpagent context-editing type — later removed 2026-09-22 with the whole context-editing feature), pulse-reviewevaluationargument aliases (legacy compat). - End-to-end: full workflow run, scheduled run, and webhook run complete with no eval phase; review/costs/timing pages render without an eval section; Goal Advisor and reports read producer outputs and goal observations with no mandated measurement route, step, or table.
Auto-synced from docs/ on main. Edit there, not here.