[AUTOMATED] fix(status): a dead worker printed 'running' — say DEAD - #437
Merged
Conversation
Found by round 4's captain, which wrote in its own notes that it had checked the builders with `ps`, "not with status.py, whose stale_s only ticks on a phase change and still lists b-r4-function-disasse as running". It was right, and it had to work around the tool this loop gives it for exactly this question. `stale_s` is `now - updated_at`, and `updated_at` moves on a `state update`, i.e. on a PHASE CHANGE -- not on any sign of life. Neither status module calls `os.kill`, so nothing ever asks whether the pid exists. `reap()` does ask, but only past `stale_seconds` (1800) and only when something calls it. So a worker that dies mid-phase reads `running` with a growing counter for up to half an hour, and the operator cannot tell it from one working hard on a long phase. One label, two very different states. `collect()` now records `alive` from a `kill(pid, 0)`: True, False, or None when no pid was recorded or the check itself failed -- unknowable is not the same as dead, and PermissionError means the process exists under another uid. The repipe table prints `DEAD` in the status column when a row claims `running` and the pid is gone. Verified by planting a `running` worker with pid 999999 in the inventory: the row prints DEAD, and the inventory was restored afterwards. `scripts.pipeline.status --json` still works, so the angr lane is unaffected. tools/repipe/smoke.sh 127/127. Fourth instance of one pattern this session, and the last three were found the same way -- a number checked for an unrelated reason. Provider refusals counted as kuna failures (#419); un-refuted hypotheses indistinguishable from undecided ones (#420); a builder budget cap indistinguishable from a broken build (#436); now a dead worker indistinguishable from a busy one. Each time, a status collapsed two causes into one label, and nothing failed loudly. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Found by round 4's captain, which wrote in its own notes that it had checked the
builders with
ps, "not with status.py, whose stale_s only ticks on a phase changeand still lists b-r4-function-disasse as running". It was right, and it had to work
around the tool this loop gives it for exactly this question.
stale_sisnow - updated_at, andupdated_atmoves on astate update, i.e. on aPHASE CHANGE -- not on any sign of life. Neither status module calls
os.kill, sonothing ever asks whether the pid exists.
reap()does ask, but only paststale_seconds(1800) and only when something calls it. So a worker that dies mid-phasereads
runningwith a growing counter for up to half an hour, and the operator cannottell it from one working hard on a long phase. One label, two very different states.
collect()now recordsalivefrom akill(pid, 0): True, False, or None when no pidwas recorded or the check itself failed -- unknowable is not the same as dead, and
PermissionError means the process exists under another uid. The repipe table prints
DEADin the status column when a row claimsrunningand the pid is gone.Verified by planting a
runningworker with pid 999999 in the inventory: the row printsDEAD, and the inventory was restored afterwards.
scripts.pipeline.status --jsonstillworks, so the angr lane is unaffected. tools/repipe/smoke.sh 127/127.
Fourth instance of one pattern this session, and the last three were found the same way
-- a number checked for an unrelated reason. Provider refusals counted as kuna failures
(#419); un-refuted hypotheses indistinguishable from undecided ones (#420); a builder
budget cap indistinguishable from a broken build (#436); now a dead worker
indistinguishable from a busy one. Each time, a status collapsed two causes into one
label, and nothing failed loudly.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
🤖 Generated with Claude Code
https://claude.ai/code/session_01YcFmfZndNjgfqLQVBZCdkY