[Bug] Stale supervisor pid across a reboot force-kills whatever process now holds that number #877
technicalpickles
started this conversation in
General
Replies: 1 comment
|
Implemented in #878: fix(supervisor): stop trusting a stale supervisor pid after reboot or pid reuse |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Rebooted my Mac this week (pitchfork 2.14.0, macOS 25.6), and pitchfork never came back. Every Claude Code session I had open started failing with:
launchctl print gui/501/pitchforkshowedstate = not running, so the supervisor really was dead. But every attempt to start a new one hit:Pid 1207 was not pitchfork. It was
AMPDeviceDiscoveryAgent, an unrelated Apple LaunchAgent, which had grabbed that pid fresh on boot because low pids get handed out to early system daemons.Here's the chain:
state.tomlnever gets rewritten when the machine reboots, so it still had the pre-reboot supervisor recorded with its old pid andstatus = "running".AMPDeviceDiscoveryAgent.is_running()(src/procs.rs) is a barekill(pid, 0), no check that the pid is actually pitchfork:I tried the obvious fix,
pitchfork supervisor stop, expecting it to notice the pid didn't belong to pitchfork and just clear the stale entry. Instead:It force-killed
AMPDeviceDiscoveryAgent. Looking atstop.rs,resolve_existing_supervisor(true)always passesforce: true, andkill_or_stop()only checksPROCS.is_running(existing_pid), never whether that pid is actually pitchfork. So it sent a real kill signal to whatever process now owns that number.In my case the collateral damage was a LaunchAgent that respawns on its own trigger, so no lasting harm. But this is a process-identity bug, not a pid-liveness bug, and a differently recycled pid could belong to something you really don't want to send SIGTERM to.
Repro: reboot with pitchfork having a recorded pid in
state.toml, wait for that pid to get reused by literally anything else post-boot (easy on a busy machine, low pids especially), then try to start pitchfork again.Suggested fix: store something more than the bare pid in
state.toml, like the process start time (or the full/proc/<pid>/statstarttime on Linux,KERN_PROCon macOS), and compare both before treating a pid as "the supervisor." Right now any pid match counts, forever, across reboots.Workaround if you hit this: don't run
supervisor stoporrun --force. Just confirmlaunchctl print gui/$(id -u)/pitchforkshowsnot running, thenlaunchctl kickstart -k gui/$(id -u)/pitchfork. That starts a fresh supervisor through launchd without touching the stale pid at all.All reactions