Skip to content

v0.41.0

Choose a tag to compare

@github-actions github-actions released this 30 Sep 13:56
· 30 commits to main since this release
1e7fb0b

The first day of running 0.40.0 on a real server found homebutler stating things nobody had told it: a docker kill filed as an OOM with high confidence, a restart recorded that never happened, and a command suggested that does not exist. This release fixes those and adds nothing new — every change corrects something homebutler said without being told, because it is the build that runs on that server before 1.0, and 1.0 will be what it does. Each fix makes homebutler say only what the machine reported: oom only when Docker sent an oom event, a start time only when something started, and in one format whichever of Docker, systemd and pm2 it came from.

⚠️ Behavior changes

  • A watch incident that ended in exit 137 is oom only when the OOM killer was involved (#295). Every 137 used to be oom with high confidence, so docker kill, a docker stop that ran out of time and kill -9 were all filed as OOM. The Docker monitor subscribed to die alone and never received the oom event Docker sends just before an OOM death. It now subscribes to oom too: a die that follows its container's oom within ten seconds is OOM-killed, with docker inspect as a second source, and a 137 without that evidence is unknown, low, signal SIGKILL. An agent that branched on oom now receives it only for an OOM. Found on the 0.40.0 soak: a docker kill came back as oom/high while docker inspect said OOMKilled=false
  • An incident's prev_started_at and curr_started_at are RFC 3339 in UTC or empty, from every monitor (#295, #297). A script that parsed the old values breaks, because each monitor wrote them in the form it received. Docker's said died at event time 1790772824 and (post-restart), the second whether or not anything had restarted; systemd's were its local text, Wed 2026-09-30 21:46:44 KST; pm2's were sentences, restart_time was 0. prev_started_at is now the start of the run that died and curr_started_at the start of the run after it, or empty when there is none. systemd is asked for --timestamp=unix, and on a systemd older than 248 the local text is read only when its zone is the machine's or UTC, because Go reads an unknown abbreviation as UTC and the time would come out nine hours wrong. pm2's are pm_uptime, checked against a real pm2 around a restart; the restart count was already in restart_count

🐛 Fixes

  • install finished by suggesting homebutler logs <app>, a command that does not exist (#296). It is homebutler docker logs, with the container's name rather than the app's, since the two differ for pi-hole. The same kind of hint sent people to a missing homebutler tui in #159, so a test now finds every string that tells someone to run homebutler and resolves it against the command tree. Found on the 0.40.0 soak, which ran the hint and got unknown command "logs"
  • watch install and serve install told you to enable lingering when it was already on, and to use sudo for it (#296). They now ask loginctl first and say nothing when lingering is on. When it is off, the command they give is loginctl enable-linger without sudo, which systemd allows for your own account by default, with sudo as the fallback
  • watch list showed LAST CHECKED frozen at the moment a target was added while the watch service was watching it all along (#297). The column is the last watch add or watch check, and with the service running neither happens. The list now says when the service is live and since when, and what the column means. The JSON is unchanged

🧪 Tests

  • crash_analysis.category and confidence are in the golden file (#295). They are what an agent branches on after an incident, and they were outside the freeze, which is how a wrong oom could have been frozen