Repository navigation
Releases: invizen/pulsemon
Release list
pulsemon v0.1.36
v0.1.36
A small security-hardening release from an external review round. Three
behavior changes: the login throttle can no longer be dodged with a rotated
X-Forwarded-For header, the login timing-equalization dummy is now
precomputed, and webhook URL validation rejects IETF-reserved and carrier-NAT
address ranges.
Security
- Login throttle ignores untrusted
X-Forwarded-Forby default. The
per-IP failure counter (7 failures in 60s → 5-minute lockout) used to key on
the firstX-Forwarded-Forentry when present, so an attacker could rotate
a spoofed header on every request and never trip the lockout — enabling
unlimited password guessing. The throttle now keys on the direct TCP
peer by default. If pulsemon sits behind a reverse proxy and you want
attempts attributed to the real client, set the new
PULSEMON_TRUSTED_PROXIESenv var (comma-separated CIDRs or bare IPs, e.g.
127.0.0.1); the header is read only when the connection's source address
is in that list. A typo'd entry is logged and skipped — it can't break
startup. - Precomputed login-timing dummy. The dummy bcrypt hash that keeps
unknown-user logins timing the same as wrong-password ones (so the
username can't be guessed from response time) is now built once at
startup from the configured bcrypt cost, instead of reformatted on every
unknown-user lookup. - Webhook URL validation rejects reserved ranges. Alert-destination
URLs now reject IETF reserved-for-testing addresses (198.18.0.0/15) and
carrier-grade NAT (100.64.0.0/10) in addition to loopback, link-local, and
unspecified. Private (LAN) ranges stay allowed — a self-hosted relay in
your own network remains a valid destination.
Verified
- Full test suite passes under
-race,go vetandgofmtclean. - XFF rotation verified live: 8 login attempts each carrying a different
spoofedX-Forwarded-For→ the 8th is rejected with 429.
v0.1.30
A focused HTTPS-sensor bugfix release: the certificate hostname check now
crosses the www boundary, so a sensor on a bare domain works when the
server presents a cert for the www. variant (and vice versa). This is the
fix behind "google.com / letsencrypt.org read not trusted" on the v0.1.29
fleet. No change to probe scheduling, status derivation, alert routing, or
the dashboard.
HTTPS sensors: accept a www-variant certificate hostname
Real CDNs and registrars commonly serve a certificate for the www
variant of a bare domain, or the other way around. Google is the canonical
case: a request to https://google.com presents a certificate whose only
subject name is www.google.com — the bare google.com is not in the cert
at all. A standard TLS client accepts that, but the strict hostname check
in v0.1.29 did not, so the sensor read error ("not trusted") even
though curl reached the site fine.
verifyHTTPCert now retries the hostname check once against the other
www variant — when the target is a bare host it also tries
www.<host>, and when the target is a www. host it also tries the bare
host — in both the system-trusted and self-signed branches. The fallback
only crosses the www. boundary: an unrelated domain, a wrong host, an
expired cert, and a rogue-CA-signed cert are all still losses, exactly as
before.
Note: this is a republish of the corrected v0.1.29 certificate work. If
your box is onv0.1.29andpulsemon updatereports "up to date," you
are on the original (pre-fix) v0.1.29 build — runpulsemon update v0.1.30explicitly, or./install.sh v0.1.30.
v0.1.29
A small HTTP-sensor hardening release: HTTPS probes now work out of the box
against self-signed endpoints, the probe's shared connection pool actually
keeps connections alive, and the dashboard's RTT labels now say "Resp" —
because an HTTP probe measures a response, not a ping. No change to probe
scheduling, status derivation, or alert routing.
HTTPS sensors: accept self-signed certificates (hostname + expiry still enforced)
Most homelab web UIs serve a self-signed or private-CA certificate, which a
strict client treats as a verification failure — so an HTTPS sensor pointed
at them read error permanently, with no way to make it monitorable short
of installing the CA into the OS trust store.
HTTPS probes now accept a certificate when it is either:
- trusted by the system roots (unchanged behavior for public and properly
chained certs), or - genuinely self-signed (issuer == subject) and its hostname and
validity period match the target.
Chain building uses the full certificate chain the server presents (leaf +
intermediates) against the system trust store, so public and CDN-fronted
certs verify the same way a standard TLS client does.
A common real-world case is handled too: many CDNs and registrars serve a
cert for the www variant of a bare domain (google.com presents a cert
whose names are www.google.com) or the reverse. When the target host
differs from the certificate's names only by the www. prefix, the check
is retried once with the other variant — so a sensor on google.com works
even though the presented cert literally names only www.google.com.
Unrelated domains still fail; the fallback only crosses the www boundary.
"Allow self-signed" deliberately does not become "allow anything": a
cert signed by an untrusted (rogue) CA, an expired cert, and a cert whose
names don't match the target are all still losses, exactly as before. The
inspector's certificate-expiry advisory keeps working on the same path.
HTTP probes now reuse keep-alive connections
The shared probe client was configured to reuse one TCP+TLS connection per
host, but the response body was closed without being read — and net/http
discards a connection whose body wasn't consumed. In practice every probe
opened a fresh connection and paid a full TCP+TLS handshake, which the
original design explicitly set out to avoid. The body is now drained up to a
1 MB cap before the connection is released: small responses (the normal
health-endpoint case) hand the connection back to the pool, and a huge page
costs one bounded read instead of a full download.
Dashboard: "Ping" → "Resp" for the response-time labels
An HTTP/HTTPS probe measures TTFB (DNS + TCP + TLS + first response
byte) — not an ICMP round trip — so the response time for a web sensor is
naturally higher than ping to the same host. The table column already
switched from "Ping" to "Resp" when HTTP sensors are visible; the three
remaining static labels now match: the inspector stat ("Resp last"), the
inspector graph ("Resp over time"), and the sort options ("Resp high→low /
Resp low→high"). The ICMP-only labels ("PING / ICMP" type option, PING
badges) are untouched.
Verified
- Full test suite green with
-race,go vetandgofmtclean, including
new tests: valid self-signed cert accepted with expiry captured; expired
self-signed rejected; self-signed with wrong hostname rejected;
CA-signed-but-untrusted rejected (presented cert still captured for the
advisory); a keep-alive regression test that three same-host probes share
one server-side connection — confirmed to fail against the pre-fix code
(3 distinct connections); and a chain-building regression test that a
public CDN-fronted cert (leaf + intermediate to a trusted root) is
accepted — confirmed to fail against the pre-fix code. - Verified against live TLS servers: self-signed/valid → accept,
self-signed/expired → reject, self-signed/wrong-host → reject,
rogue-CA/valid+right-host → reject, public leaf+intermediate → accept.
v0.1.35 — security hardening + performance
Security hardening (external review batch)
A round of independent review surfaced six issues; all are fixed here.
- Session tokens now actually expire.
ValidSessionreturned the map-hit
instead of the expiry result, so an expired token was admitted one last time
(and the UI's 5s poll kept sliding live sessions' own expiry). Expired tokens
are now rejected and deleted in one pass. - Stored XSS via sensor name closed (two layers). A crafted sensor name
could escape its string context in inline event handlers. The UI now escapes
at the JS level for everyonclicksink, and the API rejects names
containing quotes/angle brackets/backticks/backslashes/control chars or over
60 characters on create and rename (400). - Login: per-IP failure throttle + timing-oracle close. 7 failed logins
from one IP in 60s → 5-minute lockout (429 +Retry-After), enforced before
any bcrypt work; a success clears the counter. The unknown-user dummy hash's
cost is now derived from the configured bcrypt cost so it can't drift, and
every rejected login pays a constant 200ms slowdown — capping online
password guessing at ~5 tries/second/IP. No per-account lockout: a typo'd
password can never lock the operator out. - All request bodies bounded to 1 MiB. The auth gate now wraps every
request body withhttp.MaxBytesReader; oversized bodies get a clean 413
instead of being streamed into the JSON decoder. GET /api/events?n=capped at 500, matching the sensor-history endpoint
(previously unbounded).POST /api/settings/testrejects unknown/missing provider kinds with 400
naming the valid kinds, instead of answering 200{"ok":false}.
Performance
Enabled()no longer hits the database on every request. The "is auth
enabled" check ranSELECT COUNT(*) FROM userson every request, including
the dashboard's 5s poll. It's now memoized for 2s; in-app user changes
(add/remove/disable) refresh it immediately, so the dashboard is never stale
and out-of-band CLI user management is picked up within 2s.
Polish
- The sensor form's Timeout field now explains that the pre-resolve DNS
lookup has its own 5s cap, so timeouts under ~6s don't bound DNS. - Corrupt sensor rows (which shouldn't happen with the app-controlled schema)
are now logged instead of silently vanishing from the fleet view. RemoveAllUsersnaming, andtsNow()used at every writer-side timestamp
site.
Full test suite passes under -race.
v0.1.34 — Dashboard authentication (username / password)
v0.1.34 — Dashboard authentication (username / password)
The dashboard now has real accounts. Create one and it's locked down; remove them all and it's open again.
Authentication
- Username/password accounts replace the single shared token concept: each account gets its own login, stored as a bcrypt hash
- Enable it from Settings → Authentication ("Create First Account") or from the host CLI:
pulsemon auth-user add <user> <pass> - Manage accounts in Settings: add / remove, change your own password, log out
- Disable it with the "Disable Authentication" action — removes every account and reopens the dashboard
- Sessions last 30 days (sliding) via an HttpOnly cookie; a restart clears sessions (log in again)
- The last remaining account can't be removed from the account list — only the explicit disable action opens the dashboard, so a typo can't silently do it
/api/healthzstays open even when auth is on, so external monitors keep working
CLI
pulsemon auth-user add <user> <pass> # create an account (first one enables auth)
pulsemon auth-user remove <user> # delete an account (last one refused)
pulsemon auth-user list # list accounts
Settings modal
Widened (max-w-lg → max-w-3xl) so the alert routing, TLS, and authentication panels have room to breathe — especially on 1080p and up.
UI
- Login screen is now just the sign-in form (no install-specific CLI hints)
- Login gate: username + password, 30-day session
Upgrading
Existing installs: no action needed — the dashboard boots open exactly as before. When you're ready to turn auth on, create the first account from Settings or the CLI.
v0.1.33 — severity-aware sensor grouping + compact activity dates
What's new
Severity sort groups sensor types
The Severity sort now ranks by health first — error > warning > up > paused (paused stays pinned to the bottom, as before) — and keeps sensor types clustered within each band. All your HTTP servers stay grouped together; when one of them errors, it rises to the top with the other errors, and its type-neighbours follow it into that band.
Activity log dates
Entries from an earlier day now carry a compact date — 10/05/26 · 19:17:33; same-day entries stay time-only. Useful since the feed retains 24h and crosses midnight.
v0.1.32 — native TLS + activity dates + light contrast
What's new
Native TLS — the dashboard now serves HTTPS itself
No reverse proxy required. Give pulsemon a certificate and it serves the dashboard over TLS directly:
- Install-time: set
PULSEMON_CERTandPULSEMON_KEY(paths to your cert and key). When a pair is in place, the dashboard is HTTPS-only on the TLS port (443 by default) — the plain-HTTP port is not bound, so nothing is ever reachable in plaintext. - From the dashboard: Settings → TLS / HTTPS now has a certificate panel. Pick your
cert.pem+key.pem(or paste the text), choose the HTTPS port, and hit Apply. The pair is validated in memory first (cert parses, key parses, keys match) before anything is written — a mismatched pair is rejected and the running server is left untouched. Applying (or removing) the certificate restarts pulsemon automatically and the dashboard reconnects on the new scheme. Certs uploaded this way live in the data volume and survive rebuilds. - A configured-but-missing cert file logs a loud warning and keeps serving HTTP (no crash-loop); a corrupt or encrypted key fails fast at startup with a clear message.
pulsemon restart— new subcommand to re-apply a listener config manually (used by the dashboard flow under the hood).
Activity feed shows dates
Entries from an earlier day now render Tue, Oct 06 · 23:27:25; same-day entries stay time-only.
Light theme contrast
The metric-column dividers (PING | JITTER | LOSS) were near-invisible in light mode. The light-theme surface tiers are darker now, so dividers match the dark-theme gap.
Copy: pulse, not probe
User-facing wording now says "pulse"/"pulsing" to match the name (e.g. "pulsing is broken", "N pulse send errors since start").
Fixes
- The HTTP listener could default to port 80 when the listen address wasn't applied to the server object — now it always binds the configured address.
Upcoming
- Authentication for the dashboard
- Windows build
v0.1.31
What's new
Simpler alerting
- Warning is now dashboard-only in both directions — neither entering warning nor returning to up from warning sends a notification (loss flapping is noise, not a state worth paging over).
- Everything involving error still alerts, including a new
error → warning"partial recovery" card, so a flapping link improving from down-to-lossy now tells you it's getting better without waiting for full recovery. - Full recovery (
error → up) still alerts as "recovered". Sustained error re-alerts on the configured interval.
Dark / light theme
- New sun/moon toggle in the header. Dark remains the default; your choice is remembered per browser and applies before first paint (no flash).
- Both themes are fully theme-aware: cards, status pills, sparklines, the inspector graph, scrollbars, and native dropdowns.
Better response-time graph
- The inspector graph now has a left millisecond axis: 0 ms at the bottom, a dynamic top derived from the data, evenly-spaced whole-ms labels, and faint gridlines.
- The SVG renders at the panel's real size so labels stay crisp at any width.
Smoother target entry
- The sensor target field uses heavier strokes and a small letter gap so the
//in URLs can't visually merge or drop on high-DPI displays.
Upgrade
Existing installs self-update with pulsemon update (or pulsemon update v0.1.31 for a pinned install). No config or schema changes required.
Install (fresh)
curl -sfL https://raw.githubusercontent.com/invizen/pulsemon/main/install.sh | bash
pulsemon v0.1.30
v0.1.30
A focused HTTPS-sensor bugfix release: the certificate hostname check now
crosses the www boundary, so a sensor on a bare domain works when the
server presents a cert for the www. variant (and vice versa). This is the
fix behind "google.com / letsencrypt.org read not trusted" on the v0.1.29
fleet. No change to probe scheduling, status derivation, alert routing, or
the dashboard.
HTTPS sensors: accept a www-variant certificate hostname
Real CDNs and registrars commonly serve a certificate for the www
variant of a bare domain, or the other way around. Google is the canonical
case: a request to https://google.com presents a certificate whose only
subject name is www.google.com — the bare google.com is not in the cert
at all. A standard TLS client accepts that, but the strict hostname check
in v0.1.29 did not, so the sensor read error ("not trusted") even
though curl reached the site fine.
verifyHTTPCert now retries the hostname check once against the other
www variant — when the target is a bare host it also tries
www.<host>, and when the target is a www. host it also tries the bare
host — in both the system-trusted and self-signed branches. The fallback
only crosses the www. boundary: an unrelated domain, a wrong host, an
expired cert, and a rogue-CA-signed cert are all still losses, exactly as
before.
Note: this is a republish of the corrected v0.1.29 certificate work. If
your box is onv0.1.29andpulsemon updatereports "up to date," you
are on the original (pre-fix) v0.1.29 build — runpulsemon update v0.1.30explicitly, or./install.sh v0.1.30.
pulsemon v0.1.29
v0.1.29
A small HTTP-sensor hardening release: HTTPS probes now work out of the box
against self-signed endpoints, the probe's shared connection pool actually
keeps connections alive, and the dashboard's RTT labels now say "Resp" —
because an HTTP probe measures a response, not a ping. No change to probe
scheduling, status derivation, or alert routing.
HTTPS sensors: accept self-signed certificates (hostname + expiry still enforced)
Most homelab web UIs serve a self-signed or private-CA certificate, which a
strict client treats as a verification failure — so an HTTPS sensor pointed
at them read error permanently, with no way to make it monitorable short
of installing the CA into the OS trust store.
HTTPS probes now accept a certificate when it is either:
- trusted by the system roots (unchanged behavior for public and properly
chained certs), or - genuinely self-signed (issuer == subject) and its hostname and
validity period match the target.
Chain building uses the full certificate chain the server presents (leaf +
intermediates) against the system trust store, so public and CDN-fronted
certs verify the same way a standard TLS client does.
A common real-world case is handled too: many CDNs and registrars serve a
cert for the www variant of a bare domain (google.com presents a cert
whose names are www.google.com) or the reverse. When the target host
differs from the certificate's names only by the www. prefix, the check
is retried once with the other variant — so a sensor on google.com works
even though the presented cert literally names only www.google.com.
Unrelated domains still fail; the fallback only crosses the www boundary.
"Allow self-signed" deliberately does not become "allow anything": a
cert signed by an untrusted (rogue) CA, an expired cert, and a cert whose
names don't match the target are all still losses, exactly as before. The
inspector's certificate-expiry advisory keeps working on the same path.
HTTP probes now reuse keep-alive connections
The shared probe client was configured to reuse one TCP+TLS connection per
host, but the response body was closed without being read — and net/http
discards a connection whose body wasn't consumed. In practice every probe
opened a fresh connection and paid a full TCP+TLS handshake, which the
original design explicitly set out to avoid. The body is now drained up to a
1 MB cap before the connection is released: small responses (the normal
health-endpoint case) hand the connection back to the pool, and a huge page
costs one bounded read instead of a full download.
Dashboard: "Ping" → "Resp" for the response-time labels
An HTTP/HTTPS probe measures TTFB (DNS + TCP + TLS + first response
byte) — not an ICMP round trip — so the response time for a web sensor is
naturally higher than ping to the same host. The table column already
switched from "Ping" to "Resp" when HTTP sensors are visible; the three
remaining static labels now match: the inspector stat ("Resp last"), the
inspector graph ("Resp over time"), and the sort options ("Resp high→low /
Resp low→high"). The ICMP-only labels ("PING / ICMP" type option, PING
badges) are untouched.
Verified
- Full test suite green with
-race,go vetandgofmtclean, including
new tests: valid self-signed cert accepted with expiry captured; expired
self-signed rejected; self-signed with wrong hostname rejected;
CA-signed-but-untrusted rejected (presented cert still captured for the
advisory); a keep-alive regression test that three same-host probes share
one server-side connection — confirmed to fail against the pre-fix code
(3 distinct connections); and a chain-building regression test that a
public CDN-fronted cert (leaf + intermediate to a trusted root) is
accepted — confirmed to fail against the pre-fix code. - Verified against live TLS servers: self-signed/valid → accept,
self-signed/expired → reject, self-signed/wrong-host → reject,
rogue-CA/valid+right-host → reject, public leaf+intermediate → accept.
pulsemon v0.1.28
v0.1.28
A feature release: sensors can now watch web endpoints, and maintenance
mode became a real pause/resume of sensor probing instead of an alert
switch. Plus a batch of dashboard polish (Title Case labels, a global +1px
font bump, wordmark styling, card alignment). No change to how ICMP
probing, status derivation, or alert routing work.
New: HTTP / HTTPS sensors
Sensors now have a second probe type — web endpoint — next to
ping/ICMP. The type is picked in the New Sensor / Edit Sensor form
(HTTPS/HTTP listed first — hardly anything is plain HTTP anymore), and
the target must be a valid http:// or https:// URL.
- Up when the final status is 200–299; 4xx/5xx and connection errors
count as loss — the same loss/warn/down-after thresholds apply, so an
endpoint behaves exactly like a ping target from the status and alerting
point of view. - Redirects are followed (up to 5 hops). Most real endpoints sit
behind a login redirect (303 → /login/) or an http→https hop; the
probe follows to the final status, which is what decides up/loss.
Redirect loops read as loss, not success. - HTTPS targets carry a certificate-expiry advisory in the
inspector: silent while >30 days out, amber at 8–30 days, red at ≤7
days. It is advisory — an expiring cert never flips the status. - Internal single-label hosts with an explicit port are valid targets
(e.g.http://zensrv:8080). Bare labels without a port are still
rejected as before — an explicit port makes the target deliberate, and
container DNS resolves the name. - Card badges follow the target scheme: HTTPS for
https://,
HTTP only for plainhttp://, PING for ICMP.
Maintenance Mode: pause/resume sensor probing
The settings row was reworked from an "alerts on/off" toggle into a
Maintenance Mode row with a Pause / Resume icon button (the
description swaps with state: "pause all active sensors" ↔ "resume all
previously-running sensors").
- Pause silences alerts and pauses every sensor that is currently
active — they stop probing entirely. The set of active sensors is
snapshotted into the settings table first. - Resume re-activates exactly the snapshotted sensors and nothing
else. A sensor the user paused before pressing Pause stays paused. - Nothing active → refused. Pressing Pause with zero active sensors
used to be possible and left an empty snapshot, so a later resume had
nothing to restore. Now both the dashboard ("no active sensors to
pause" toast) and the API (400) refuse it.
Dashboard polish
- Global +1px font bump across the whole dashboard (arbitrary-value
and named size classes alike), sized for 1080p+ displays. - Wordmark:
pulsemonis now extrabold (800) at 17px; the version
pill and the "network sensor monitor" subtitle line up under it. - Cards without tags align with cards that have them (the tag row
keeps a minimum height instead of collapsing). - Title Case everywhere it was inconsistent: New Sensor, Edit
Sensor, Save Sensor, Clear History, Alert Destinations, Alert
Routing, All Sensors / Sensors With Selected Tags / Only Selected
Sensors, Re-Alert Interval, Maintenance Mode. - Spike (× avg) removed from the forms and inspector. The multiplier
no longer influences status, so it was dropped from the UI; the field
stays in the DB and API (server-side default 3) so nothing migrates.
Verified
- Full suite green with
-race,go vetandgofmtclean, including
new tests for: URL validation (single-label + port, schemes), redirect
following (303→200 = up with final status; redirect loop = loss), and
the maintenance cycle (enable snapshots the active set, disable
restores exactly it, user-paused sensors untouched, enable-with-nothing
active → 400). - Live container on zentest: an
http://zensrv:8080sensor followed a
303→200 chain and reported up; a full maintenance on/off cycle
re-activated exactly the three active sensors while two user-paused
sensors stayed paused.
v0.1.27
Renamed: zenmon → pulsemon
The project is now pulsemon. Repo: github.com/invizen/pulsemon (the old
invizen/zenmon URL 301-redirects after the GitHub rename, so existing links
keep working). Everything user-facing moved with it:
- Binary, install prefix, DB:
~/pulsemon/pulsemon,~/pulsemon/data/pulsemon.db - Release assets:
pulsemon-linux-{amd64,arm64}(+.sha256) - systemd unit:
pulsemon.service(user-level, as before) - Env vars:
PULSEMON_ADDR,PULSEMON_DB(wasZENMON_*) - Dashboard wordmark: pulse in the EKG green (
#10b981) + mon in
neutral gray (#808080)
Existing installs are NOT auto-migrated — the v0.1.27 release notes carry a
one-shot migration (move the dir, install the new unit, disable the old one).
zenmon update on old binaries keeps working through the GitHub redirect
until the instance is migrated.
One-shot migration (bare / user-systemd installs)
# stop the old service, move the data (DB + binary) to the new prefix
systemctl --user stop zenmon
systemctl --user disable zenmon
mv ~/zenmon ~/pulsemon
# install the new unit (writes ~/.config/systemd/user/pulsemon.service)
curl -sfL https://raw.githubusercontent.com/invizen/pulsemon/main/install.sh | bash
# remove the stale unit + old PATH line (the installer adds the new one)
rm -f ~/.config/systemd/user/zenmon.service
sed -i '/zenmon: keep the zenmon binary on PATH/d' ~/.bashrc
sed -i '/^export PATH="\$HOME\/zenmon:/d' ~/.bashrc
systemctl --user daemon-reloadDocker installs: just point compose.yaml at the new image name; the data
volume carries over unchanged.
Probe tuning: 60s default, faster error recovery
Three behavior changes around how a sensor is watched when it fails:
- Default interval 60s (was 15s) for new sensors, the fresh-install seed
sensors, and the dashboard form — one less DB write / ping on a healthy
install; per-sensor intervals are untouched. - Default "error after" = 2 consecutive losses (was 4): with the 60s
interval, a dead target was previously confirmed down after ~4 minutes;
now after ~2 minutes. The fast re-check below closes the recovery side of
the same gap. - Fast re-check in error: while a sensor's status is error it is
probed every 30s instead of the configured interval, until a probe
brings it back to up — then it reverts to the configured interval. If the
sensor's interval is already shorter than 30s, the error state keeps the
sensor's own pace (it never polls faster than normal). - Error → up on 1 successful ping: a sensor recovering from error
flips back to up on a single good probe, instead of waiting for 2
straight good replies. A flapping sensor in warning is unaffected:
loss/up/loss/up still reads warning (one good reply amid ongoing loss is
not "recovered" — and up↔warning never alerts anyway, so there is no new
alert noise).
No ping storms. The fast cadence is capped at the configured interval
(never faster than normal) and floored at 10s, so worst case — every sensor
errors at once (total outage) — total probe load is bounded to ~2× the
normal load, spread evenly: each sensor keeps its staggered probe phase, so
they do not re-synchronize into a burst. Each sensor loop owns its own
timer; one slow sensor can't delay another.
Closes the silent ICMP failure mode
When neither ICMP transport could
open (e.g. RHEL 8's or Ubuntu 18.04's default net.ipv4.ping_group_range
excludes the service uid and CAP_NET_RAW isn't granted — any distro with
systemd < 244, since that's the version that ships the wide range), pulsemon
used to start the server, report status: "ok" in healthz, and simply never
ping — with no error anywhere.
The failure was only findable by noticing the missing icmp_mode key.
Now the failure is loud, at three layers:
1. install.sh preflight (fail before the service starts)
Before installing, the script reads net.ipv4.ping_group_range and checks
whether the installer's gid falls inside it. If not, it prints the exact
remediation (sysctl + the persistent /etc/sysctl.d/90-pulsemon-ping.conf)
and, on an interactive terminal, asks before continuing (non-interactive
installs continue but flag that the dashboard will warn). Background: the
kernel default is 1 0 (nobody may ping); systemd ≥ 244 — RHEL 9+, Fedora,
Ubuntu/Debian — ships 0 2147483647 via 50-default.conf, but RHEL 8
(systemd 239) does not.
2. Honest healthz
When the shared ICMP engine fails to open (both transports), GET /api/healthz now returns:
{"status": "degraded", "icmp_mode": "unavailable", "icmp_hint": "ICMP socket
unavailable: ... Fix: sudo sysctl -w net.ipv4.ping_group_range=\"0 65535\"
(persist via /etc/sysctl.d/90-pulsemon-ping.conf) and restart pulsemon — or grant
CAP_NET_RAW. See `journalctl -u pulsemon` for the exact error.", ...}instead of the previous status: "ok" with no icmp_mode key. The engine is
warmed at probe-worker startup, so this is visible on the first healthz after
boot — not after the first failed probe tick. The failure error now wraps the
errNoIcmpTransport sentinel so tooling can match it with errors.Is while
still carrying the underlying cause.
3. Dashboard banner
The existing warning banner now fires on icmp_hint (before the generic
probe_error), showing the remediation command right on the dashboard.
Verified
go test— full suite green, including the newTestHealthzIcmpMode
(dead → degraded/unavailable/hint; unprivileged-datagram → ok; raw → ok),
and the extendedTestDeriveStatuscases (error→up on one success;
warning flapping stays warning; still-down stays error; DB-error path
unchanged).- Live container matrix on zentest: RHEL-8-like netns (range
1 0, no caps)
→ degraded + hint; wide range → ok; raw-socket-possible → ok.
v0.1.26
A review-drive...
v0.1.27 — pulsemon
v0.1.27
Renamed: zenmon → pulsemon
The project is now pulsemon. Repo: github.com/invizen/pulsemon (the old
invizen/zenmon URL 301-redirects after the GitHub rename, so existing links
keep working). Everything user-facing moved with it:
- Binary, install prefix, DB:
~/pulsemon/pulsemon,~/pulsemon/data/pulsemon.db - Release assets:
pulsemon-linux-{amd64,arm64}(+.sha256) - systemd unit:
pulsemon.service(user-level, as before) - Env vars:
PULSEMON_ADDR,PULSEMON_DB(wasZENMON_*) - Dashboard wordmark: pulse in the EKG green (
#10b981) + mon in
neutral gray (#808080)
Existing installs are NOT auto-migrated — the v0.1.27 release notes carry a
one-shot migration (move the dir, install the new unit, disable the old one).
zenmon update on old binaries keeps working through the GitHub redirect
until the instance is migrated.
One-shot migration (bare / user-systemd installs)
# stop the old service, move the data (DB + binary) to the new prefix
systemctl --user stop zenmon
systemctl --user disable zenmon
mv ~/zenmon ~/pulsemon
# install the new unit (writes ~/.config/systemd/user/pulsemon.service)
curl -sfL https://github.com/invizen/pulsemon/releases/latest/download/install.sh | bash
# remove the stale unit + old PATH line (the installer adds the new one)
rm -f ~/.config/systemd/user/zenmon.service
sed -i '/zenmon: keep the zenmon binary on PATH/d' ~/.bashrc
sed -i '/^export PATH="\$HOME\/zenmon:/d' ~/.bashrc
systemctl --user daemon-reloadDocker installs: just point compose.yaml at the new image name; the data
volume carries over unchanged.
Probe tuning: 60s default, faster error recovery
Three behavior changes around how a sensor is watched when it fails:
- Default interval 60s (was 15s) for new sensors, the fresh-install seed
sensors, and the dashboard form — one less DB write / ping on a healthy
install; per-sensor intervals are untouched. - Default "error after" = 2 consecutive losses (was 4): with the 60s
interval, a dead target was previously confirmed down after ~4 minutes;
now after ~2 minutes. The fast re-check below closes the recovery side of
the same gap. - Fast re-check in error: while a sensor's status is error it is
probed every 30s instead of the configured interval, until a probe
brings it back to up — then it reverts to the configured interval. If the
sensor's interval is already shorter than 30s, the error state keeps the
sensor's own pace (it never polls faster than normal). - Error → up on 1 successful ping: a sensor recovering from error
flips back to up on a single good probe, instead of waiting for 2
straight good replies. A flapping sensor in warning is unaffected:
loss/up/loss/up still reads warning (one good reply amid ongoing loss is
not "recovered" — and up↔warning never alerts anyway, so there is no new
alert noise).
No ping storms. The fast cadence is capped at the configured interval
(never faster than normal) and floored at 10s, so worst case — every sensor
errors at once (total outage) — total probe load is bounded to ~2× the
normal load, spread evenly: each sensor keeps its staggered probe phase, so
they do not re-synchronize into a burst. Each sensor loop owns its own
timer; one slow sensor can't delay another.
Closes the silent ICMP failure mode
When neither ICMP transport could
open (e.g. RHEL 8's or Ubuntu 18.04's default net.ipv4.ping_group_range
excludes the service uid and CAP_NET_RAW isn't granted — any distro with
systemd < 244, since that's the version that ships the wide range), pulsemon
used to start the server, report status: "ok" in healthz, and simply never
ping — with no error anywhere.
The failure was only findable by noticing the missing icmp_mode key.
Now the failure is loud, at three layers:
1. install.sh preflight (fail before the service starts)
Before installing, the script reads net.ipv4.ping_group_range and checks
whether the installer's gid falls inside it. If not, it prints the exact
remediation (sysctl + the persistent /etc/sysctl.d/90-pulsemon-ping.conf)
and, on an interactive terminal, asks before continuing (non-interactive
installs continue but flag that the dashboard will warn). Background: the
kernel default is 1 0 (nobody may ping); systemd ≥ 244 — RHEL 9+, Fedora,
Ubuntu/Debian — ships 0 2147483647 via 50-default.conf, but RHEL 8
(systemd 239) does not.
2. Honest healthz
When the shared ICMP engine fails to open (both transports), GET /api/healthz now returns:
{"status": "degraded", "icmp_mode": "unavailable", "icmp_hint": "ICMP socket
unavailable: ... Fix: sudo sysctl -w net.ipv4.ping_group_range=\"0 65535\"
(persist via /etc/sysctl.d/90-pulsemon-ping.conf) and restart pulsemon — or grant
CAP_NET_RAW. See `journalctl -u pulsemon` for the exact error.", ...}instead of the previous status: "ok" with no icmp_mode key. The engine is
warmed at probe-worker startup, so this is visible on the first healthz after
boot — not after the first failed probe tick. The failure error now wraps the
errNoIcmpTransport sentinel so tooling can match it with errors.Is while
still carrying the underlying cause.
3. Dashboard banner
The existing warning banner now fires on icmp_hint (before the generic
probe_error), showing the remediation command right on the dashboard.
Verified
go test— full suite green, including the newTestHealthzIcmpMode
(dead → degraded/unavailable/hint; unprivileged-datagram → ok; raw → ok),
and the extendedTestDeriveStatuscases (error→up on one success;
warning flapping stays warning; still-down stays error; DB-error path
unchanged).- Live container matrix on zentest: RHEL-8-like netns (range
1 0, no caps)
→ degraded + hint; wide range → ok; raw-socket-possible → ok.