-
Notifications
You must be signed in to change notification settings - Fork 3
plat 031
| Coordination | Value |
|---|---|
| Assigned agent | Codex |
| Ticket state | implemented; runtime reverify |
| Last synchronized | 2026-08-06 |
Claim this ticket in this file before implementation. During active work, update this fragment rather than the shared index; synchronize the index once at handoff, review, or completion.
-
Priority: P1
-
Owner: execution-cost persistence identity and step attribution
-
Source finding:
HARNESS-COST-LEDGER-RUN-ATTRIBUTION -
Source workflow:
Workflow/rtslatency -
Source Pulse run:
schedule-cron--42eca39a_1785886230496797000 -
Reconfirmed as:
HARNESS-RUN-ID-ALIASINGinWorkflow/hetznerssh(2026-08-06): historical cost/evaluation state underiteration-0/production-servercould be presented as today's run after rotation. -
Problem: the cost writer identifies an execution by a reusable run-folder key inside a wall-clock-date file. When one execution crosses UTC midnight, its costs are split between two files that both contain
run_folders["iteration-0/dev"]. The ledger also attributed a $14.7928535 Opus charge toexecution_only:10, although10is a schedule-message index and not a step ID inplanning/plan.json. -
Evidence:
-
costs/execution/dev/2026-08-04.jsoncontainsrun_folders["iteration-0/dev"]updated through2026-08-04T23:57:51.531744Z. -
costs/execution/dev/2026-08-05.jsoncontains the same run-folder key, created at2026-08-05T00:03:46.320904Z. - Both ledgers contain
execution_only:10; the 2026-08-04 entry records oneclaude-opus-5call costing $14.7928535. - The cron fires at 05:00 Asia/Kolkata (23:30 UTC), so crossing the ledger's UTC date boundary is normal rather than exceptional.
-
-
Impact: per-run totals cannot be reproduced reliably, adjacent executions that reuse the same run folder cannot be distinguished, and the largest charge can be assigned to a non-existent plan step. Pulse therefore cannot make trustworthy run-level cost or step-tier recommendations from this ledger shape.
-
Relationship to PLAT-009: PLAT-009 repairs how
get_cost_summaryreads and merges grouped/date shards. It cannot recover an execution identity or valid step ID that the writer never persisted. -
Required fix:
- Generate one immutable execution/run-instance ID when the schedule starts.
- Persist that ID on every phase and step cost event, independent of date or reusable run-folder path.
- Attribute a cost only to a validated plan step ID or to an explicit
non-step phase such as
schedule_message; never overload a numeric message index as a step ID. - Make the date files storage shards only. Query and aggregation identity must be the execution ID.
- Preserve the run-folder/group fields as searchable metadata, not identity.
-
Migration: old date/folder-only rows remain historical evidence but cannot always be separated safely. Mark ambiguous historical aggregates as such rather than inventing execution IDs or silently assigning them to the nearest run.
-
Acceptance: a run beginning before UTC midnight and completing afterward has one execution ID across both storage shards; querying that ID returns the exact total once; every step-attributed row references a real plan step; and two executions reusing
iteration-0/devremain distinguishable. -
Required tests: a real persistence/query round trip spanning UTC midnight, two consecutive executions with the same run folder, rejection or explicit classification of numeric schedule-message indices, and reconciliation of phase/model totals without double counting.
-
Implementation (2026-08-05, Claude Code,
mcp-agent-builder-go1bfa745d5): Writer-side fix only — this closes the identity gap the writer creates; it does not add the reader/query layer, which is PLAT-009's territory.-
ContextAwareEventBridgemints oneExecutionID(uuid) per bridge instance inNewContextAwareEventBridge. A bridge instance lives for the whole run, so this is functionally "one ID per execution," not literally the scheduler'srunID— wiring the scheduler's ownrunIDthrough would have required threading a new parameter acrossexecuteJob→executeWorkflowJob→ orchestrator construction, several layers outside the cost-persistence code path. The bridge-lifetime ID gives the same guarantee (stable across a UTC-date-shard rotation, since both shards are written by calls from the same live bridge) without that wider blast radius. Exposed viabridge.ExecutionID(). -
StepTokenData.ExecutionIDcarries it from event to persister.TokenUsageFile.ExecutionID(new,omitempty) is stamped sticky first-write instampExecutionID(base_orchestrator_tokens.go): the first call to touch a run folder's aggregate claims the ID; a later call with a different non-empty ID (a second execution reusing the same run folder) displaces it intoTokenUsageFile.PriorExecutionIDsinstead of silently overwriting it. - Numeric-index misattribution (fix item 3) is closed at two points: (a)
bridge-level —
numericOnlyIdentifierPatternreclassifies any bare-digitStepID(e.g."10") to phase"schedule_message"before persisting, regardless of which phase produced it; (b) writer-level — theby_step_and_modelkey builder (stepAggregationKey) no longer falls back to"<caller phase>:<numeric Step>"(the oldexecution_only:10shape) whenStepIDis empty; it now always uses"schedule_message:<n>". Both guards are heuristic (bare-digits-only), not a lookup against the realplanning/plan.jsonstep list — no code path was found that threads plan step IDs into the bridge/persistence layer, and building that lookup was out of scope for this pass. If a plan ever legitimately used a bare-numeral step ID, it would be misclassified; no such ID was found in this codebase's step constructors (RegularPlanStep,MessageSequencePlanStep, etc. all use descriptive slugs in every call site inspected). -
Root cause of the literal
execution_only:10in the evidence was not pinned down. EverycreateExecutionOnlyAgentcall site (controller_execution.go,controller_message_sequence.go,workflow_continuation_recovery.go) passes a realstep.GetID()/state.StepID, andcontext_aware_bridge.go's own numeric-index fallback branch requires a non-emptyStepIDbefore it can produce phase"execution_only"— so on the paths traced, the digits"10"would have to be the literal string value of some step's realGetID(), not a stray counter. The two guards above catch this defensively at the single persistence choke point regardless of which caller produced it, but the actual upstream producer of a bare-numeral step ID for this workflow type is still unidentified. -
PhaseTokenData/PhaseTokenUsageFile(costs/phase/token_usage.json) were deliberately left untouched. That file is a lifetime-cumulative aggregate across every phase-only run ever, not a per-execution record — stamping a single execution's ID onto it would misrepresent its own shape. The evidence in this ticket is entirely in the execution-scope ledger (costs/execution/...), which is what this fix addresses. - Migration: no backfill script was written. Old rows simply have an empty
ExecutionID(the field isomitempty), which is itself the "mark as ambiguous rather than invent an ID" behavior the ticket asked for — no separate migration pass was needed. -
Acceptance criteria status:
- "one execution ID across both storage shards" — met; proven by
TestStampExecutionIDGivesOneIdentityAcrossTwoIndependentDateShardFilesandTestBridgeStampsStableExecutionIDAcrossTurnsAndDiffersAcrossBridges. - "two executions reusing the same run folder remain distinguishable" —
met for detectability (
PriorExecutionIDsproves two executions touched the folder) but not for full aggregate separation — theByModel/ByStepAndModelsums for two executions sharing a run folder still merge into one running total. Splitting those would need theRunFoldersmap keyed by execution identity instead of bare run folder (fix item 4's full ask), which changes the on-disk key shape that PLAT-009's reader consumes — deliberately not done here to avoid an uncoordinated breaking change to Codex's activeruntime_reverifywork on PLAT-009. - "querying that ID returns the exact total once" — not built; no query API exists yet. PLAT-009 territory.
- "every step-attributed row references a real plan step" — the numeric
case is excluded; there is no positive validation against
planning/plan.json's actual step list, so a malformed-but-non-numeric ID would still pass through.
- "one execution ID across both storage shards" — met; proven by
-
Tests:
plat031_execution_id_test.go(pure unit tests onstampExecutionID/stepAggregationKey) andplat031_bridge_execution_id_test.go(throughContextAwareEventBridge.HandleEvent, reusing the existingrecordingTokenPersister/noopListenerfakes). Confirmed these fail to compile against the pre-fix code (stashed and re-ran), so they're actually pinned to the fix, not vacuously passing. Not tested: a genuine end-to-end round trip throughBaseOrchestrator.PersistTokenUsage's real file I/O andtime.Now()— that function has no injectable clock, so the UTC-midnight proof is at the extracted-helper level instead of a true file-system round trip.
-
-
Remaining/runtime reverify: confirm against a real
rtslatencyrun that (a) both date shards for a midnight-crossing execution carry a matchingexecution_idfield, and (b) no newexecution_only:<digits>keys appear in fresh ledger writes.
Tectonicus independently reported costs/phase/daily versus
costs/execution disagreement of 3.1× under one opaque workflow-builder
bucket. That is not evidence that the writer-side execution-ID fix regressed:
the current ticket deliberately left the lifetime-cumulative phase ledger
unchanged. It is evidence that the reader/reporting layer must not compare the
two shapes as though they were equivalent execution totals.
PLAT-009 owns the query/aggregation work. PLAT-008 now records the shared reconciliation acceptance boundary: a real execution must reconcile exactly once in every view that claims to represent that execution, or the UI must describe its scope as cumulative/non-comparable.
Platform ID: PLAT-031
Status: implemented; awaits one real run plus rotation for runtime
reverification.
The writer and reader are now genuinely execution-keyed; this supersedes the earlier sticky-field-only implementation above.
- Daily cost files retain their existing date/group sharding, but their
authoritative payload is now
executions[execution_id]with atoken_usageaggregate andrun_folder/archived_run_foldermetadata. New writes never read fromrun_folders["iteration-0/..."], so a reused active folder cannot inherit old spend. -
run_foldersremains as a legacy read-compatible projection. New readers useexecutionsand deliberately ignore that projection when a v2 record is present, preventing double-counting during migration. Old rows without an ID remain visiblylegacy:<run-folder>rather than being falsely split. - Rotation now changes only
archived_run_folderon matching execution records (for both execution and evaluation cost scopes). The immutable UUID and its totals do not move or merge wheniteration-0becomesiteration-N. - Evaluation score history now uses the same shape: every generated report has
an
evaluation_id, andscores/evaluation/...is keyed by that ID rather than a reusable target-run folder. Repeated evaluations no longer overwrite each other. - Server cost/evaluation projections expose
execution_idand archived path, so callers can show individual executions rather than attributing history to today's active run.
Focused regression coverage:
TestExecutionKeyedCostLedgerSeparatesIterationZeroReuseTestArchiveRunCostPathsUpdatesOnlyExecutionDisplayPathTestReadRunAcrossDatesUsesExecutionKeyedRecordsNotLegacyProjectionTestEvaluationLedgerSeparatesRepeatedIterationZeroEvaluations
The remaining acceptance item is runtime evidence from a production run and subsequent rotation; no migration invents IDs for ambiguous historical data.
Auto-synced from docs/ on main. Edit there, not here.