Skip to content

plat 242

github-actions[bot] edited this page Sep 20, 2026 · 1 revision

← Pulse platform index

Internal tracking resolution — 2026-09-05

Resolved locally: rotation now uses StartedAt/CreatedAt before folder-name changes. Covers social-media PUL-6403A889 and rtslatency PUL-BA1C940F. Focused scheduler regression tests pass. Timestamp-less legacy metadata retains its name-only fallback; general concurrent-run identity is not addressed.

Closed at the user's request for internal tracking. Fixes are tested local, uncommitted source changes based on 0babf193ec0efdf33511a3150f82e0b29685814e; deployment verification is pending. No live workflow or historical schedule receipt was modified. Prior investigation below is retained as history.

PLAT-242 — A schedule-run history entry's error text can name a different iteration than its own run_folder field

Coordination Value
Assigned agent Claude Code
Ticket state resolved locally — deployment verification pending
Last synchronized 2026-09-05
  • Priority: harness_issue, severity high.
  • Findings: Twitter/social-media PUL-6403A889 — the 2026-08-25 Daily Measurement & Critique schedule retained status=error, while the run's own outer run_metadata.status and route summary both read as completed (completed/completed_with_nonfatal_validation_concerns).

What was confirmed directly against real schedule-runs.json data

Read workspace-docs/Workflow/social-media/schedule-runs.json around the cited date. Chronologically:

  1. 2026-08-24T14:30 — run_folder=iteration-0, status=error, its own error: "workshop turn 1/1 (schedule-message-1) produced no response: turn ended with orchestrator_agent_error" — a real, self-contained infrastructure failure for that invocation.
  2. 2026-08-24T17:19 — run_folder=iteration-287, status=partial, error "Pulse completed partially".
  3. 2026-08-25T03:00 — run_folder=iteration-288, status=error, error: "workflow run iteration-287/default failed (its run_metadata.json records status "failed"), even though the orchestrating workshop session completed its turns without an infrastructure error."

Entry 3's own run_folder field says iteration-288, but its error text names iteration-287 — a different iteration than the one the entry itself claims to be about, and one already independently accounted for by entry 1/2 the day before. This is the exact "false terminal status" shape PUL-6403A889 describes: a run that plausibly completed on its own merits (producing iteration-288) instead recorded error, attributing someone else's already-known failure to itself.

Mechanism located, not fully pinned down

The error text is generated by reconcileWorkshopRunOutcome (agent_go/cmd/server/scheduler.go:3592), a deliberate BUG-20260729-10 fix: after a workshop chat session finishes its turns, it diffs the run-folder listing from before vs. after the invocation (plus a StartedAt/CreatedAt timestamp fallback, itself a later fix for PLAT-182's reused-iteration-0-name case) to find which folder(s) this specific invocation actually touched, then checks that folder's own run_metadata.json for status="failed".

For iteration-287 to be picked up as "touched" by the 03:00 invocation, either the name-based diff or the timestamp fallback in reconcileWorkshopRunOutcome/workshopRunProducedEvidence must have misclassified a folder that (per entry 2's own record) already existed and was already terminal before this invocation's since boundary. The exact trigger — a listing-timing race during rotation, a stale before[] snapshot, or a StartedAt/CreatedAt value on iteration-287's metadata that doesn't reflect what entry 1/2 already recorded — was not isolated with confidence in this pass. This is stated honestly rather than guessed at.

Why this is not fixed in this session

Both candidate functions live in the same critical schedule-execution path as PLAT-241's deferred fix, are already carrying multiple prior edge-case fixes (BUG-20260729-10, PLAT-182) layered onto genuinely tricky pre/post-invocation timing logic, and a wrong guess here risks reintroducing exactly the false-positive/false-negative failure detection this code was already built to prevent. Confirming the precise trigger needs either a live reproduction with instrumented timestamps around a real rotation, or a decision from whoever last touched this reconciliation logic — not a guess made against historical JSON alone.

Verification

Confirmed the anomaly and its generating code path directly against real schedule-runs.json data and the current scheduler.go source — not inferred from the finding text alone. No code changed this session.

Reverify

Not applicable — no fix shipped. Reverify by instrumenting reconcileWorkshopRunOutcome's before/after folder sets and metadata timestamps around a live rotation to isolate the exact misclassification, then deciding a fix.

Clone this wiki locally