-
Notifications
You must be signed in to change notification settings - Fork 2
plat 242
Resolved locally: rotation now uses StartedAt/CreatedAt before folder-name changes. Covers social-media PUL-6403A889 and rtslatency PUL-BA1C940F. Focused scheduler regression tests pass. Timestamp-less legacy metadata retains its name-only fallback; general concurrent-run identity is not addressed.
Closed at the user's request for internal tracking. Fixes are tested local,
uncommitted source changes based on 0babf193ec0efdf33511a3150f82e0b29685814e;
deployment verification is pending. No live workflow or historical schedule
receipt was modified. Prior investigation below is retained as history.
PLAT-242 — A schedule-run history entry's error text can name a different iteration than its own run_folder field
| Coordination | Value |
|---|---|
| Assigned agent | Claude Code |
| Ticket state | resolved locally — deployment verification pending |
| Last synchronized | 2026-09-05 |
- Priority: harness_issue, severity high.
-
Findings: Twitter/social-media
PUL-6403A889— the 2026-08-25 Daily Measurement & Critique schedule retainedstatus=error, while the run's own outerrun_metadata.statusand route summary both read as completed (completed/completed_with_nonfatal_validation_concerns).
Read workspace-docs/Workflow/social-media/schedule-runs.json around the
cited date. Chronologically:
-
2026-08-24T14:30—run_folder=iteration-0,status=error, its own error: "workshop turn 1/1 (schedule-message-1) produced no response: turn ended with orchestrator_agent_error" — a real, self-contained infrastructure failure for that invocation. -
2026-08-24T17:19—run_folder=iteration-287,status=partial, error"Pulse completed partially". -
2026-08-25T03:00—run_folder=iteration-288,status=error, error: "workflow run iteration-287/default failed (its run_metadata.json records status "failed"), even though the orchestrating workshop session completed its turns without an infrastructure error."
Entry 3's own run_folder field says iteration-288, but its error text
names iteration-287 — a different iteration than the one the entry
itself claims to be about, and one already independently accounted for by
entry 1/2 the day before. This is the exact "false terminal status"
shape PUL-6403A889 describes: a run that plausibly completed on its own
merits (producing iteration-288) instead recorded error, attributing
someone else's already-known failure to itself.
The error text is generated by reconcileWorkshopRunOutcome
(agent_go/cmd/server/scheduler.go:3592), a deliberate BUG-20260729-10
fix: after a workshop chat session finishes its turns, it diffs the
run-folder listing from before vs. after the invocation (plus a
StartedAt/CreatedAt timestamp fallback, itself a later fix for
PLAT-182's reused-iteration-0-name case) to find which folder(s) this
specific invocation actually touched, then checks that folder's own
run_metadata.json for status="failed".
For iteration-287 to be picked up as "touched" by the 03:00
invocation, either the name-based diff or the timestamp fallback in
reconcileWorkshopRunOutcome/workshopRunProducedEvidence must have
misclassified a folder that (per entry 2's own record) already existed
and was already terminal before this invocation's since boundary. The
exact trigger — a listing-timing race during rotation, a stale
before[] snapshot, or a StartedAt/CreatedAt value on
iteration-287's metadata that doesn't reflect what entry 1/2 already
recorded — was not isolated with confidence in this pass. This is stated
honestly rather than guessed at.
Both candidate functions live in the same critical schedule-execution path as PLAT-241's deferred fix, are already carrying multiple prior edge-case fixes (BUG-20260729-10, PLAT-182) layered onto genuinely tricky pre/post-invocation timing logic, and a wrong guess here risks reintroducing exactly the false-positive/false-negative failure detection this code was already built to prevent. Confirming the precise trigger needs either a live reproduction with instrumented timestamps around a real rotation, or a decision from whoever last touched this reconciliation logic — not a guess made against historical JSON alone.
Confirmed the anomaly and its generating code path directly against real
schedule-runs.json data and the current scheduler.go source — not
inferred from the finding text alone. No code changed this session.
Not applicable — no fix shipped. Reverify by instrumenting
reconcileWorkshopRunOutcome's before/after folder sets and metadata
timestamps around a live rotation to isolate the exact misclassification,
then deciding a fix.
Auto-synced from docs/ on main. Edit there, not here.