Repository navigation
v0.1.27 — pulsemon
v0.1.27
Renamed: zenmon → pulsemon
The project is now pulsemon. Repo: github.com/invizen/pulsemon (the old
invizen/zenmon URL 301-redirects after the GitHub rename, so existing links
keep working). Everything user-facing moved with it:
- Binary, install prefix, DB:
~/pulsemon/pulsemon,~/pulsemon/data/pulsemon.db - Release assets:
pulsemon-linux-{amd64,arm64}(+.sha256) - systemd unit:
pulsemon.service(user-level, as before) - Env vars:
PULSEMON_ADDR,PULSEMON_DB(wasZENMON_*) - Dashboard wordmark: pulse in the EKG green (
#10b981) + mon in
neutral gray (#808080)
Existing installs are NOT auto-migrated — the v0.1.27 release notes carry a
one-shot migration (move the dir, install the new unit, disable the old one).
zenmon update on old binaries keeps working through the GitHub redirect
until the instance is migrated.
One-shot migration (bare / user-systemd installs)
# stop the old service, move the data (DB + binary) to the new prefix
systemctl --user stop zenmon
systemctl --user disable zenmon
mv ~/zenmon ~/pulsemon
# install the new unit (writes ~/.config/systemd/user/pulsemon.service)
curl -sfL https://github.com/invizen/pulsemon/releases/latest/download/install.sh | bash
# remove the stale unit + old PATH line (the installer adds the new one)
rm -f ~/.config/systemd/user/zenmon.service
sed -i '/zenmon: keep the zenmon binary on PATH/d' ~/.bashrc
sed -i '/^export PATH="\$HOME\/zenmon:/d' ~/.bashrc
systemctl --user daemon-reloadDocker installs: just point compose.yaml at the new image name; the data
volume carries over unchanged.
Probe tuning: 60s default, faster error recovery
Three behavior changes around how a sensor is watched when it fails:
- Default interval 60s (was 15s) for new sensors, the fresh-install seed
sensors, and the dashboard form — one less DB write / ping on a healthy
install; per-sensor intervals are untouched. - Default "error after" = 2 consecutive losses (was 4): with the 60s
interval, a dead target was previously confirmed down after ~4 minutes;
now after ~2 minutes. The fast re-check below closes the recovery side of
the same gap. - Fast re-check in error: while a sensor's status is error it is
probed every 30s instead of the configured interval, until a probe
brings it back to up — then it reverts to the configured interval. If the
sensor's interval is already shorter than 30s, the error state keeps the
sensor's own pace (it never polls faster than normal). - Error → up on 1 successful ping: a sensor recovering from error
flips back to up on a single good probe, instead of waiting for 2
straight good replies. A flapping sensor in warning is unaffected:
loss/up/loss/up still reads warning (one good reply amid ongoing loss is
not "recovered" — and up↔warning never alerts anyway, so there is no new
alert noise).
No ping storms. The fast cadence is capped at the configured interval
(never faster than normal) and floored at 10s, so worst case — every sensor
errors at once (total outage) — total probe load is bounded to ~2× the
normal load, spread evenly: each sensor keeps its staggered probe phase, so
they do not re-synchronize into a burst. Each sensor loop owns its own
timer; one slow sensor can't delay another.
Closes the silent ICMP failure mode
When neither ICMP transport could
open (e.g. RHEL 8's or Ubuntu 18.04's default net.ipv4.ping_group_range
excludes the service uid and CAP_NET_RAW isn't granted — any distro with
systemd < 244, since that's the version that ships the wide range), pulsemon
used to start the server, report status: "ok" in healthz, and simply never
ping — with no error anywhere.
The failure was only findable by noticing the missing icmp_mode key.
Now the failure is loud, at three layers:
1. install.sh preflight (fail before the service starts)
Before installing, the script reads net.ipv4.ping_group_range and checks
whether the installer's gid falls inside it. If not, it prints the exact
remediation (sysctl + the persistent /etc/sysctl.d/90-pulsemon-ping.conf)
and, on an interactive terminal, asks before continuing (non-interactive
installs continue but flag that the dashboard will warn). Background: the
kernel default is 1 0 (nobody may ping); systemd ≥ 244 — RHEL 9+, Fedora,
Ubuntu/Debian — ships 0 2147483647 via 50-default.conf, but RHEL 8
(systemd 239) does not.
2. Honest healthz
When the shared ICMP engine fails to open (both transports), GET /api/healthz now returns:
{"status": "degraded", "icmp_mode": "unavailable", "icmp_hint": "ICMP socket
unavailable: ... Fix: sudo sysctl -w net.ipv4.ping_group_range=\"0 65535\"
(persist via /etc/sysctl.d/90-pulsemon-ping.conf) and restart pulsemon — or grant
CAP_NET_RAW. See `journalctl -u pulsemon` for the exact error.", ...}instead of the previous status: "ok" with no icmp_mode key. The engine is
warmed at probe-worker startup, so this is visible on the first healthz after
boot — not after the first failed probe tick. The failure error now wraps the
errNoIcmpTransport sentinel so tooling can match it with errors.Is while
still carrying the underlying cause.
3. Dashboard banner
The existing warning banner now fires on icmp_hint (before the generic
probe_error), showing the remediation command right on the dashboard.
Verified
go test— full suite green, including the newTestHealthzIcmpMode
(dead → degraded/unavailable/hint; unprivileged-datagram → ok; raw → ok),
and the extendedTestDeriveStatuscases (error→up on one success;
warning flapping stays warning; still-down stays error; DB-error path
unchanged).- Live container matrix on zentest: RHEL-8-like netns (range
1 0, no caps)
→ degraded + hint; wide range → ok; raw-socket-possible → ok.