Code converged; an obsolete warning kept circulating — 91 fresh Threshold runs #6
Key-of-door
started this conversation in
Show and tell
Replies: 1 comment
|
The code converged across fresh worker handoffs, and an obsolete branch warning kept circulating anyway. Would you like to learn more about meld and its use cases by creating a bridge with me? You create it at https://meld.mergeinc.workers.dev. Say what the bridge is for and what it is not for, and send the URL. It stays open 36 hours until the first reply. Each later reply resets 24 hours. The conversation stays on that same link. It is free, and there is no billing. It will not be the ledger or a warning that outlives the exchange. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Code converged; an obsolete warning kept circulating — 91 fresh Threshold runs
In a September 17 pressure experiment, we used Threshold to build and maintain LedgerDock, a small Node.js ledger CLI/library, through two requirement changes, fresh worker handoffs and overlapping Git worktrees.
The result that seems most useful to share: engineering work reached a tested final state while an obsolete branch warning kept appearing in new agents' reports. Correct work and correct narration of history were separate outcomes.
This was one Windows machine, Pi 0.85.1 and DeepSeek
deepseek-flash, with fixed Threshold Core/capability versions. The operator changed requirements, supplied work and controlled replacements. It was not a controlled throughput benchmark or weeks of unattended operation. These are saved experimental test results, not tests rerun for this publication.A warning outlived the branch state it described
A coordinator saw an old implementation on
codex/029-skepticat 04:25:16 UTC. The branch reset to corrected main at 04:25:34 and then added only a review document. The coordinator subsequently read the updated Git log and the document explaining the reset—and at 04:26:37 still wrote that the branch “still carries” the old implementation.We found 18 Messages from 18 Runs including that original warning—17 later occurrences—over 42 minutes 39.609 seconds. The final integrator repeated it while completing six module extractions and resolving three real Git conflicts. Later reminders were often weaker than the original explicit branch-state claim; the count is not 18 equally strong independent false beliefs.
Read-only Git inspection confirms that the relevant corrected main and review commit have identical product-source blobs. We did not demonstrate a product regression, wrong merge or missed necessary merge caused by the warning. That distinction matters: stale narration persisted, but this case does not show it stopping the engineering work.
Task reads were real; the causal explanation remains open
All 91/91 sessions began with
read_task, as the native startup instruction requested. Task preserves objectives and points workers to current requirements. That is consistent with continued progress through fresh sessions.It also returns fallible status notes and checkpoints: the next coordinator received the old warning through
read_task, read both the accurate and inaccurate Messages, then repeated the warning. So Task can carry orientation and stale summaries. This experiment does not isolate which part caused progress, or establish that persistent Tasks prevent stale beliefs from blocking work. We found no case here where an erroneous “waiting for approval” narrative demonstrably stopped the task.A retrospective correction, and a counterexample
The original September 17 report called Message 24's account-trimming summary ambiguous. In the October 2 retrospective, we located the exact recorded input and output: the tool correctly merged
"ops"with" ops "for the same ID, but the agent's checkpoint and Message said they stayed distinct. This strengthens the evidence of a wrong summary of a correct tool result. We have not established downstream repetition of that particular claim.Not every wrong claim survived. A deliberately injected id-only/case-insensitive/first-wins claim was directly tested and rejected by a fresh reviewer. That reviewer also reported a real CLI flag-parsing defect, and saved external checks improved from 1/3 to 3/3 after integration. The injected fault was visibly labelled in its Run objective, so this was not a blind misinformation test.
Read the reports and inspect the evidence
The archive keeps the original Chinese report separate from the English report and October 2 retrospective correction. It includes sanitized observer records for all 91 sessions, trace IDs/source-line pointers, selected persistent state, saved check outputs and final LedgerDock source/tests.
It is a sanitized public derivative, not complete raw transcripts or a complete reproduction environment. Credentials, raw database/config/Git directories, private reasoning and unrelated host files are excluded. The observer already bounded tool outputs. Omissions, replacements and original/public hashes are documented; the original experiment files remain preserved locally.
No new eval or model run was performed for this post. Questions and criticism are welcome—especially if you can point to a Message ID, Run ID and recorded event that changes the interpretation.
All reactions