Repository navigation
v0.41.0
The first day of running 0.40.0 on a real server found homebutler stating things nobody had told it: a docker kill filed as an OOM with high confidence, a restart recorded that never happened, and a command suggested that does not exist. This release fixes those and adds nothing new — every change corrects something homebutler said without being told, because it is the build that runs on that server before 1.0, and 1.0 will be what it does. Each fix makes homebutler say only what the machine reported: oom only when Docker sent an oom event, a start time only when something started, and in one format whichever of Docker, systemd and pm2 it came from.
⚠️ Behavior changes
- A
watchincident that ended in exit 137 isoomonly when the OOM killer was involved (#295). Every 137 used to beoomwithhighconfidence, sodocker kill, adocker stopthat ran out of time andkill -9were all filed as OOM. The Docker monitor subscribed todiealone and never received theoomevent Docker sends just before an OOM death. It now subscribes tooomtoo: a die that follows its container'soomwithin ten seconds is OOM-killed, withdocker inspectas a second source, and a 137 without that evidence isunknown,low, signalSIGKILL. An agent that branched onoomnow receives it only for an OOM. Found on the 0.40.0 soak: adocker killcame back asoom/highwhiledocker inspectsaidOOMKilled=false - An incident's
prev_started_atandcurr_started_atare RFC 3339 in UTC or empty, from every monitor (#295, #297). A script that parsed the old values breaks, because each monitor wrote them in the form it received. Docker's saiddied at event time 1790772824and(post-restart), the second whether or not anything had restarted; systemd's were its local text,Wed 2026-09-30 21:46:44 KST; pm2's were sentences,restart_time was 0.prev_started_atis now the start of the run that died andcurr_started_atthe start of the run after it, or empty when there is none. systemd is asked for--timestamp=unix, and on a systemd older than 248 the local text is read only when its zone is the machine's or UTC, because Go reads an unknown abbreviation as UTC and the time would come out nine hours wrong. pm2's arepm_uptime, checked against a real pm2 around a restart; the restart count was already inrestart_count
🐛 Fixes
installfinished by suggestinghomebutler logs <app>, a command that does not exist (#296). It ishomebutler docker logs, with the container's name rather than the app's, since the two differ for pi-hole. The same kind of hint sent people to a missinghomebutler tuiin #159, so a test now finds every string that tells someone to run homebutler and resolves it against the command tree. Found on the 0.40.0 soak, which ran the hint and gotunknown command "logs"watch installandserve installtold you to enable lingering when it was already on, and to use sudo for it (#296). They now askloginctlfirst and say nothing when lingering is on. When it is off, the command they give isloginctl enable-lingerwithout sudo, which systemd allows for your own account by default, with sudo as the fallbackwatch listshowed LAST CHECKED frozen at the moment a target was added while the watch service was watching it all along (#297). The column is the lastwatch addorwatch check, and with the service running neither happens. The list now says when the service is live and since when, and what the column means. The JSON is unchanged
🧪 Tests
crash_analysis.categoryandconfidenceare in the golden file (#295). They are what an agent branches on after an incident, and they were outside the freeze, which is how a wrongoomcould have been frozen