Skip to content

Releases: invizen/pulsemon

pulsemon v0.1.36

Choose a tag to compare

@invizen invizen released this 08 Oct 15:44

v0.1.36

A small security-hardening release from an external review round. Three
behavior changes: the login throttle can no longer be dodged with a rotated
X-Forwarded-For header, the login timing-equalization dummy is now
precomputed, and webhook URL validation rejects IETF-reserved and carrier-NAT
address ranges.

Security

  • Login throttle ignores untrusted X-Forwarded-For by default. The
    per-IP failure counter (7 failures in 60s → 5-minute lockout) used to key on
    the first X-Forwarded-For entry when present, so an attacker could rotate
    a spoofed header on every request and never trip the lockout — enabling
    unlimited password guessing. The throttle now keys on the direct TCP
    peer
    by default. If pulsemon sits behind a reverse proxy and you want
    attempts attributed to the real client, set the new
    PULSEMON_TRUSTED_PROXIES env var (comma-separated CIDRs or bare IPs, e.g.
    127.0.0.1); the header is read only when the connection's source address
    is in that list. A typo'd entry is logged and skipped — it can't break
    startup.
  • Precomputed login-timing dummy. The dummy bcrypt hash that keeps
    unknown-user logins timing the same as wrong-password ones (so the
    username can't be guessed from response time) is now built once at
    startup from the configured bcrypt cost, instead of reformatted on every
    unknown-user lookup.
  • Webhook URL validation rejects reserved ranges. Alert-destination
    URLs now reject IETF reserved-for-testing addresses (198.18.0.0/15) and
    carrier-grade NAT (100.64.0.0/10) in addition to loopback, link-local, and
    unspecified. Private (LAN) ranges stay allowed — a self-hosted relay in
    your own network remains a valid destination.

Verified

  • Full test suite passes under -race, go vet and gofmt clean.
  • XFF rotation verified live: 8 login attempts each carrying a different
    spoofed X-Forwarded-For → the 8th is rejected with 429.

v0.1.30

A focused HTTPS-sensor bugfix release: the certificate hostname check now
crosses the www boundary, so a sensor on a bare domain works when the
server presents a cert for the www. variant (and vice versa). This is the
fix behind "google.com / letsencrypt.org read not trusted" on the v0.1.29
fleet. No change to probe scheduling, status derivation, alert routing, or
the dashboard.

HTTPS sensors: accept a www-variant certificate hostname

Real CDNs and registrars commonly serve a certificate for the www
variant of a bare domain, or the other way around. Google is the canonical
case: a request to https://google.com presents a certificate whose only
subject name is www.google.com — the bare google.com is not in the cert
at all. A standard TLS client accepts that, but the strict hostname check
in v0.1.29 did not, so the sensor read error ("not trusted") even
though curl reached the site fine.

verifyHTTPCert now retries the hostname check once against the other
www variant — when the target is a bare host it also tries
www.<host>, and when the target is a www. host it also tries the bare
host — in both the system-trusted and self-signed branches. The fallback
only crosses the www. boundary: an unrelated domain, a wrong host, an
expired cert, and a rogue-CA-signed cert are all still losses, exactly as
before.

Note: this is a republish of the corrected v0.1.29 certificate work. If
your box is on v0.1.29 and pulsemon update reports "up to date," you
are on the original (pre-fix) v0.1.29 build — run pulsemon update v0.1.30 explicitly, or ./install.sh v0.1.30.


v0.1.29

A small HTTP-sensor hardening release: HTTPS probes now work out of the box
against self-signed endpoints, the probe's shared connection pool actually
keeps connections alive, and the dashboard's RTT labels now say "Resp" —
because an HTTP probe measures a response, not a ping. No change to probe
scheduling, status derivation, or alert routing.

HTTPS sensors: accept self-signed certificates (hostname + expiry still enforced)

Most homelab web UIs serve a self-signed or private-CA certificate, which a
strict client treats as a verification failure — so an HTTPS sensor pointed
at them read error permanently, with no way to make it monitorable short
of installing the CA into the OS trust store.

HTTPS probes now accept a certificate when it is either:

  • trusted by the system roots (unchanged behavior for public and properly
    chained certs), or
  • genuinely self-signed (issuer == subject) and its hostname and
    validity period match the target.

Chain building uses the full certificate chain the server presents (leaf +
intermediates) against the system trust store, so public and CDN-fronted
certs verify the same way a standard TLS client does.

A common real-world case is handled too: many CDNs and registrars serve a
cert for the www variant of a bare domain (google.com presents a cert
whose names are www.google.com) or the reverse. When the target host
differs from the certificate's names only by the www. prefix, the check
is retried once with the other variant — so a sensor on google.com works
even though the presented cert literally names only www.google.com.
Unrelated domains still fail; the fallback only crosses the www boundary.

"Allow self-signed" deliberately does not become "allow anything": a
cert signed by an untrusted (rogue) CA, an expired cert, and a cert whose
names don't match the target are all still losses, exactly as before. The
inspector's certificate-expiry advisory keeps working on the same path.

HTTP probes now reuse keep-alive connections

The shared probe client was configured to reuse one TCP+TLS connection per
host, but the response body was closed without being read — and net/http
discards a connection whose body wasn't consumed. In practice every probe
opened a fresh connection and paid a full TCP+TLS handshake, which the
original design explicitly set out to avoid. The body is now drained up to a
1 MB cap before the connection is released: small responses (the normal
health-endpoint case) hand the connection back to the pool, and a huge page
costs one bounded read instead of a full download.

Dashboard: "Ping" → "Resp" for the response-time labels

An HTTP/HTTPS probe measures TTFB (DNS + TCP + TLS + first response
byte) — not an ICMP round trip — so the response time for a web sensor is
naturally higher than ping to the same host. The table column already
switched from "Ping" to "Resp" when HTTP sensors are visible; the three
remaining static labels now match: the inspector stat ("Resp last"), the
inspector graph ("Resp over time"), and the sort options ("Resp high→low /
Resp low→high"). The ICMP-only labels ("PING / ICMP" type option, PING
badges) are untouched.

Verified

  • Full test suite green with -race, go vet and gofmt clean, including
    new tests: valid self-signed cert accepted with expiry captured; expired
    self-signed rejected; self-signed with wrong hostname rejected;
    CA-signed-but-untrusted rejected (presented cert still captured for the
    advisory); a keep-alive regression test that three same-host probes share
    one server-side connection — confirmed to fail against the pre-fix code
    (3 distinct connections); and a chain-building regression test that a
    public CDN-fronted cert (leaf + intermediate to a trusted root) is
    accepted — confirmed to fail against the pre-fix code.
  • Verified against live TLS servers: self-signed/valid → accept,
    self-signed/expired → reject, self-signed/wrong-host → reject,
    rogue-CA/valid+right-host → reject, public leaf+intermediate → accept.

v0.1.35 — security hardening + performance

Choose a tag to compare

@invizen invizen released this 08 Oct 05:12

Security hardening (external review batch)

A round of independent review surfaced six issues; all are fixed here.

  • Session tokens now actually expire. ValidSession returned the map-hit
    instead of the expiry result, so an expired token was admitted one last time
    (and the UI's 5s poll kept sliding live sessions' own expiry). Expired tokens
    are now rejected and deleted in one pass.
  • Stored XSS via sensor name closed (two layers). A crafted sensor name
    could escape its string context in inline event handlers. The UI now escapes
    at the JS level for every onclick sink, and the API rejects names
    containing quotes/angle brackets/backticks/backslashes/control chars or over
    60 characters on create and rename (400).
  • Login: per-IP failure throttle + timing-oracle close. 7 failed logins
    from one IP in 60s → 5-minute lockout (429 + Retry-After), enforced before
    any bcrypt work; a success clears the counter. The unknown-user dummy hash's
    cost is now derived from the configured bcrypt cost so it can't drift, and
    every rejected login pays a constant 200ms slowdown — capping online
    password guessing at ~5 tries/second/IP. No per-account lockout: a typo'd
    password can never lock the operator out.
  • All request bodies bounded to 1 MiB. The auth gate now wraps every
    request body with http.MaxBytesReader; oversized bodies get a clean 413
    instead of being streamed into the JSON decoder.
  • GET /api/events?n= capped at 500, matching the sensor-history endpoint
    (previously unbounded).
  • POST /api/settings/test rejects unknown/missing provider kinds with 400
    naming the valid kinds, instead of answering 200 {"ok":false}.

Performance

  • Enabled() no longer hits the database on every request. The "is auth
    enabled" check ran SELECT COUNT(*) FROM users on every request, including
    the dashboard's 5s poll. It's now memoized for 2s; in-app user changes
    (add/remove/disable) refresh it immediately, so the dashboard is never stale
    and out-of-band CLI user management is picked up within 2s.

Polish

  • The sensor form's Timeout field now explains that the pre-resolve DNS
    lookup has its own 5s cap, so timeouts under ~6s don't bound DNS.
  • Corrupt sensor rows (which shouldn't happen with the app-controlled schema)
    are now logged instead of silently vanishing from the fleet view.
  • RemoveAllUsers naming, and tsNow() used at every writer-side timestamp
    site.

Full test suite passes under -race.

v0.1.34 — Dashboard authentication (username / password)

Choose a tag to compare

@invizen invizen released this 08 Oct 03:24

v0.1.34 — Dashboard authentication (username / password)

The dashboard now has real accounts. Create one and it's locked down; remove them all and it's open again.

Authentication

  • Username/password accounts replace the single shared token concept: each account gets its own login, stored as a bcrypt hash
  • Enable it from Settings → Authentication ("Create First Account") or from the host CLI:
    pulsemon auth-user add <user> <pass>
    
  • Manage accounts in Settings: add / remove, change your own password, log out
  • Disable it with the "Disable Authentication" action — removes every account and reopens the dashboard
  • Sessions last 30 days (sliding) via an HttpOnly cookie; a restart clears sessions (log in again)
  • The last remaining account can't be removed from the account list — only the explicit disable action opens the dashboard, so a typo can't silently do it
  • /api/healthz stays open even when auth is on, so external monitors keep working

CLI

pulsemon auth-user add <user> <pass>    # create an account (first one enables auth)
pulsemon auth-user remove <user>        # delete an account (last one refused)
pulsemon auth-user list                 # list accounts

Settings modal

Widened (max-w-lg → max-w-3xl) so the alert routing, TLS, and authentication panels have room to breathe — especially on 1080p and up.

UI

  • Login screen is now just the sign-in form (no install-specific CLI hints)
  • Login gate: username + password, 30-day session

Upgrading

Existing installs: no action needed — the dashboard boots open exactly as before. When you're ready to turn auth on, create the first account from Settings or the CLI.

v0.1.33 — severity-aware sensor grouping + compact activity dates

Choose a tag to compare

@invizen invizen released this 08 Oct 01:23

What's new

Severity sort groups sensor types

The Severity sort now ranks by health first — error > warning > up > paused (paused stays pinned to the bottom, as before) — and keeps sensor types clustered within each band. All your HTTP servers stay grouped together; when one of them errors, it rises to the top with the other errors, and its type-neighbours follow it into that band.

Activity log dates

Entries from an earlier day now carry a compact date — 10/05/26 · 19:17:33; same-day entries stay time-only. Useful since the feed retains 24h and crosses midnight.

v0.1.32 — native TLS + activity dates + light contrast

Choose a tag to compare

@invizen invizen released this 07 Oct 19:49

What's new

Native TLS — the dashboard now serves HTTPS itself

No reverse proxy required. Give pulsemon a certificate and it serves the dashboard over TLS directly:

  • Install-time: set PULSEMON_CERT and PULSEMON_KEY (paths to your cert and key). When a pair is in place, the dashboard is HTTPS-only on the TLS port (443 by default) — the plain-HTTP port is not bound, so nothing is ever reachable in plaintext.
  • From the dashboard: Settings → TLS / HTTPS now has a certificate panel. Pick your cert.pem + key.pem (or paste the text), choose the HTTPS port, and hit Apply. The pair is validated in memory first (cert parses, key parses, keys match) before anything is written — a mismatched pair is rejected and the running server is left untouched. Applying (or removing) the certificate restarts pulsemon automatically and the dashboard reconnects on the new scheme. Certs uploaded this way live in the data volume and survive rebuilds.
  • A configured-but-missing cert file logs a loud warning and keeps serving HTTP (no crash-loop); a corrupt or encrypted key fails fast at startup with a clear message.
  • pulsemon restart — new subcommand to re-apply a listener config manually (used by the dashboard flow under the hood).

Activity feed shows dates

Entries from an earlier day now render Tue, Oct 06 · 23:27:25; same-day entries stay time-only.

Light theme contrast

The metric-column dividers (PING | JITTER | LOSS) were near-invisible in light mode. The light-theme surface tiers are darker now, so dividers match the dark-theme gap.

Copy: pulse, not probe

User-facing wording now says "pulse"/"pulsing" to match the name (e.g. "pulsing is broken", "N pulse send errors since start").

Fixes

  • The HTTP listener could default to port 80 when the listen address wasn't applied to the server object — now it always binds the configured address.

Upcoming

  • Authentication for the dashboard
  • Windows build

v0.1.31

Choose a tag to compare

@invizen invizen released this 07 Oct 18:02

What's new

Simpler alerting

  • Warning is now dashboard-only in both directions — neither entering warning nor returning to up from warning sends a notification (loss flapping is noise, not a state worth paging over).
  • Everything involving error still alerts, including a new error → warning "partial recovery" card, so a flapping link improving from down-to-lossy now tells you it's getting better without waiting for full recovery.
  • Full recovery (error → up) still alerts as "recovered". Sustained error re-alerts on the configured interval.

Dark / light theme

  • New sun/moon toggle in the header. Dark remains the default; your choice is remembered per browser and applies before first paint (no flash).
  • Both themes are fully theme-aware: cards, status pills, sparklines, the inspector graph, scrollbars, and native dropdowns.

Better response-time graph

  • The inspector graph now has a left millisecond axis: 0 ms at the bottom, a dynamic top derived from the data, evenly-spaced whole-ms labels, and faint gridlines.
  • The SVG renders at the panel's real size so labels stay crisp at any width.

Smoother target entry

  • The sensor target field uses heavier strokes and a small letter gap so the // in URLs can't visually merge or drop on high-DPI displays.

Upgrade

Existing installs self-update with pulsemon update (or pulsemon update v0.1.31 for a pinned install). No config or schema changes required.

Install (fresh)

curl -sfL https://raw.githubusercontent.com/invizen/pulsemon/main/install.sh | bash

pulsemon v0.1.30

Choose a tag to compare

@invizen invizen released this 07 Oct 05:28

v0.1.30

A focused HTTPS-sensor bugfix release: the certificate hostname check now
crosses the www boundary, so a sensor on a bare domain works when the
server presents a cert for the www. variant (and vice versa). This is the
fix behind "google.com / letsencrypt.org read not trusted" on the v0.1.29
fleet. No change to probe scheduling, status derivation, alert routing, or
the dashboard.

HTTPS sensors: accept a www-variant certificate hostname

Real CDNs and registrars commonly serve a certificate for the www
variant of a bare domain, or the other way around. Google is the canonical
case: a request to https://google.com presents a certificate whose only
subject name is www.google.com — the bare google.com is not in the cert
at all. A standard TLS client accepts that, but the strict hostname check
in v0.1.29 did not, so the sensor read error ("not trusted") even
though curl reached the site fine.

verifyHTTPCert now retries the hostname check once against the other
www variant — when the target is a bare host it also tries
www.<host>, and when the target is a www. host it also tries the bare
host — in both the system-trusted and self-signed branches. The fallback
only crosses the www. boundary: an unrelated domain, a wrong host, an
expired cert, and a rogue-CA-signed cert are all still losses, exactly as
before.

Note: this is a republish of the corrected v0.1.29 certificate work. If
your box is on v0.1.29 and pulsemon update reports "up to date," you
are on the original (pre-fix) v0.1.29 build — run pulsemon update v0.1.30 explicitly, or ./install.sh v0.1.30.


pulsemon v0.1.29

Choose a tag to compare

@invizen invizen released this 07 Oct 04:42

v0.1.29

A small HTTP-sensor hardening release: HTTPS probes now work out of the box
against self-signed endpoints, the probe's shared connection pool actually
keeps connections alive, and the dashboard's RTT labels now say "Resp" —
because an HTTP probe measures a response, not a ping. No change to probe
scheduling, status derivation, or alert routing.

HTTPS sensors: accept self-signed certificates (hostname + expiry still enforced)

Most homelab web UIs serve a self-signed or private-CA certificate, which a
strict client treats as a verification failure — so an HTTPS sensor pointed
at them read error permanently, with no way to make it monitorable short
of installing the CA into the OS trust store.

HTTPS probes now accept a certificate when it is either:

  • trusted by the system roots (unchanged behavior for public and properly
    chained certs), or
  • genuinely self-signed (issuer == subject) and its hostname and
    validity period match the target.

Chain building uses the full certificate chain the server presents (leaf +
intermediates) against the system trust store, so public and CDN-fronted
certs verify the same way a standard TLS client does.

A common real-world case is handled too: many CDNs and registrars serve a
cert for the www variant of a bare domain (google.com presents a cert
whose names are www.google.com) or the reverse. When the target host
differs from the certificate's names only by the www. prefix, the check
is retried once with the other variant — so a sensor on google.com works
even though the presented cert literally names only www.google.com.
Unrelated domains still fail; the fallback only crosses the www boundary.

"Allow self-signed" deliberately does not become "allow anything": a
cert signed by an untrusted (rogue) CA, an expired cert, and a cert whose
names don't match the target are all still losses, exactly as before. The
inspector's certificate-expiry advisory keeps working on the same path.

HTTP probes now reuse keep-alive connections

The shared probe client was configured to reuse one TCP+TLS connection per
host, but the response body was closed without being read — and net/http
discards a connection whose body wasn't consumed. In practice every probe
opened a fresh connection and paid a full TCP+TLS handshake, which the
original design explicitly set out to avoid. The body is now drained up to a
1 MB cap before the connection is released: small responses (the normal
health-endpoint case) hand the connection back to the pool, and a huge page
costs one bounded read instead of a full download.

Dashboard: "Ping" → "Resp" for the response-time labels

An HTTP/HTTPS probe measures TTFB (DNS + TCP + TLS + first response
byte) — not an ICMP round trip — so the response time for a web sensor is
naturally higher than ping to the same host. The table column already
switched from "Ping" to "Resp" when HTTP sensors are visible; the three
remaining static labels now match: the inspector stat ("Resp last"), the
inspector graph ("Resp over time"), and the sort options ("Resp high→low /
Resp low→high"). The ICMP-only labels ("PING / ICMP" type option, PING
badges) are untouched.

Verified

  • Full test suite green with -race, go vet and gofmt clean, including
    new tests: valid self-signed cert accepted with expiry captured; expired
    self-signed rejected; self-signed with wrong hostname rejected;
    CA-signed-but-untrusted rejected (presented cert still captured for the
    advisory); a keep-alive regression test that three same-host probes share
    one server-side connection — confirmed to fail against the pre-fix code
    (3 distinct connections); and a chain-building regression test that a
    public CDN-fronted cert (leaf + intermediate to a trusted root) is
    accepted — confirmed to fail against the pre-fix code.
  • Verified against live TLS servers: self-signed/valid → accept,
    self-signed/expired → reject, self-signed/wrong-host → reject,
    rogue-CA/valid+right-host → reject, public leaf+intermediate → accept.

pulsemon v0.1.28

Choose a tag to compare

@invizen invizen released this 07 Oct 00:21

v0.1.28

A feature release: sensors can now watch web endpoints, and maintenance
mode became a real pause/resume of sensor probing instead of an alert
switch. Plus a batch of dashboard polish (Title Case labels, a global +1px
font bump, wordmark styling, card alignment). No change to how ICMP
probing, status derivation, or alert routing work.

New: HTTP / HTTPS sensors

Sensors now have a second probe type — web endpoint — next to
ping/ICMP. The type is picked in the New Sensor / Edit Sensor form
(HTTPS/HTTP listed first — hardly anything is plain HTTP anymore), and
the target must be a valid http:// or https:// URL.

  • Up when the final status is 200–299; 4xx/5xx and connection errors
    count as loss
    — the same loss/warn/down-after thresholds apply, so an
    endpoint behaves exactly like a ping target from the status and alerting
    point of view.
  • Redirects are followed (up to 5 hops). Most real endpoints sit
    behind a login redirect (303 → /login/) or an http→https hop; the
    probe follows to the final status, which is what decides up/loss.
    Redirect loops read as loss, not success.
  • HTTPS targets carry a certificate-expiry advisory in the
    inspector: silent while >30 days out, amber at 8–30 days, red at ≤7
    days. It is advisory — an expiring cert never flips the status.
  • Internal single-label hosts with an explicit port are valid targets
    (e.g. http://zensrv:8080). Bare labels without a port are still
    rejected as before — an explicit port makes the target deliberate, and
    container DNS resolves the name.
  • Card badges follow the target scheme: HTTPS for https://,
    HTTP only for plain http://, PING for ICMP.

Maintenance Mode: pause/resume sensor probing

The settings row was reworked from an "alerts on/off" toggle into a
Maintenance Mode row with a Pause / Resume icon button (the
description swaps with state: "pause all active sensors" ↔ "resume all
previously-running sensors").

  • Pause silences alerts and pauses every sensor that is currently
    active — they stop probing entirely. The set of active sensors is
    snapshotted into the settings table first.
  • Resume re-activates exactly the snapshotted sensors and nothing
    else. A sensor the user paused before pressing Pause stays paused.
  • Nothing active → refused. Pressing Pause with zero active sensors
    used to be possible and left an empty snapshot, so a later resume had
    nothing to restore. Now both the dashboard ("no active sensors to
    pause" toast) and the API (400) refuse it.

Dashboard polish

  • Global +1px font bump across the whole dashboard (arbitrary-value
    and named size classes alike), sized for 1080p+ displays.
  • Wordmark: pulsemon is now extrabold (800) at 17px; the version
    pill and the "network sensor monitor" subtitle line up under it.
  • Cards without tags align with cards that have them (the tag row
    keeps a minimum height instead of collapsing).
  • Title Case everywhere it was inconsistent: New Sensor, Edit
    Sensor, Save Sensor, Clear History, Alert Destinations, Alert
    Routing, All Sensors / Sensors With Selected Tags / Only Selected
    Sensors, Re-Alert Interval, Maintenance Mode.
  • Spike (× avg) removed from the forms and inspector. The multiplier
    no longer influences status, so it was dropped from the UI; the field
    stays in the DB and API (server-side default 3) so nothing migrates.

Verified

  • Full suite green with -race, go vet and gofmt clean, including
    new tests for: URL validation (single-label + port, schemes), redirect
    following (303→200 = up with final status; redirect loop = loss), and
    the maintenance cycle (enable snapshots the active set, disable
    restores exactly it, user-paused sensors untouched, enable-with-nothing
    active → 400).
  • Live container on zentest: an http://zensrv:8080 sensor followed a
    303→200 chain and reported up; a full maintenance on/off cycle
    re-activated exactly the three active sensors while two user-paused
    sensors stayed paused.

v0.1.27

Renamed: zenmon → pulsemon

The project is now pulsemon. Repo: github.com/invizen/pulsemon (the old
invizen/zenmon URL 301-redirects after the GitHub rename, so existing links
keep working). Everything user-facing moved with it:

  • Binary, install prefix, DB: ~/pulsemon/pulsemon, ~/pulsemon/data/pulsemon.db
  • Release assets: pulsemon-linux-{amd64,arm64} (+ .sha256)
  • systemd unit: pulsemon.service (user-level, as before)
  • Env vars: PULSEMON_ADDR, PULSEMON_DB (was ZENMON_*)
  • Dashboard wordmark: pulse in the EKG green (#10b981) + mon in
    neutral gray (#808080)

Existing installs are NOT auto-migrated — the v0.1.27 release notes carry a
one-shot migration (move the dir, install the new unit, disable the old one).
zenmon update on old binaries keeps working through the GitHub redirect
until the instance is migrated.

One-shot migration (bare / user-systemd installs)

# stop the old service, move the data (DB + binary) to the new prefix
systemctl --user stop zenmon
systemctl --user disable zenmon
mv ~/zenmon ~/pulsemon
# install the new unit (writes ~/.config/systemd/user/pulsemon.service)
curl -sfL https://raw.githubusercontent.com/invizen/pulsemon/main/install.sh | bash
# remove the stale unit + old PATH line (the installer adds the new one)
rm -f ~/.config/systemd/user/zenmon.service
sed -i '/zenmon: keep the zenmon binary on PATH/d' ~/.bashrc
sed -i '/^export PATH="\$HOME\/zenmon:/d' ~/.bashrc
systemctl --user daemon-reload

Docker installs: just point compose.yaml at the new image name; the data
volume carries over unchanged.

Probe tuning: 60s default, faster error recovery

Three behavior changes around how a sensor is watched when it fails:

  • Default interval 60s (was 15s) for new sensors, the fresh-install seed
    sensors, and the dashboard form — one less DB write / ping on a healthy
    install; per-sensor intervals are untouched.
  • Default "error after" = 2 consecutive losses (was 4): with the 60s
    interval, a dead target was previously confirmed down after ~4 minutes;
    now after ~2 minutes. The fast re-check below closes the recovery side of
    the same gap.
  • Fast re-check in error: while a sensor's status is error it is
    probed every 30s instead of the configured interval, until a probe
    brings it back to up — then it reverts to the configured interval. If the
    sensor's interval is already shorter than 30s, the error state keeps the
    sensor's own pace (it never polls faster than normal).
  • Error → up on 1 successful ping: a sensor recovering from error
    flips back to up on a single good probe, instead of waiting for 2
    straight good replies. A flapping sensor in warning is unaffected:
    loss/up/loss/up still reads warning (one good reply amid ongoing loss is
    not "recovered" — and up↔warning never alerts anyway, so there is no new
    alert noise).

No ping storms. The fast cadence is capped at the configured interval
(never faster than normal) and floored at 10s, so worst case — every sensor
errors at once (total outage) — total probe load is bounded to ~2× the
normal load, spread evenly: each sensor keeps its staggered probe phase, so
they do not re-synchronize into a burst. Each sensor loop owns its own
timer; one slow sensor can't delay another.

Closes the silent ICMP failure mode

When neither ICMP transport could
open (e.g. RHEL 8's or Ubuntu 18.04's default net.ipv4.ping_group_range
excludes the service uid and CAP_NET_RAW isn't granted — any distro with
systemd < 244, since that's the version that ships the wide range), pulsemon
used to start the server, report status: "ok" in healthz, and simply never
ping — with no error anywhere.
The failure was only findable by noticing the missing icmp_mode key.
Now the failure is loud, at three layers:

1. install.sh preflight (fail before the service starts)

Before installing, the script reads net.ipv4.ping_group_range and checks
whether the installer's gid falls inside it. If not, it prints the exact
remediation (sysctl + the persistent /etc/sysctl.d/90-pulsemon-ping.conf)
and, on an interactive terminal, asks before continuing (non-interactive
installs continue but flag that the dashboard will warn). Background: the
kernel default is 1 0 (nobody may ping); systemd ≥ 244 — RHEL 9+, Fedora,
Ubuntu/Debian — ships 0 2147483647 via 50-default.conf, but RHEL 8
(systemd 239) does not.

2. Honest healthz

When the shared ICMP engine fails to open (both transports), GET /api/healthz now returns:

{"status": "degraded", "icmp_mode": "unavailable", "icmp_hint": "ICMP socket
unavailable: ... Fix: sudo sysctl -w net.ipv4.ping_group_range=\"0 65535\"
(persist via /etc/sysctl.d/90-pulsemon-ping.conf) and restart pulsemon — or grant
CAP_NET_RAW. See `journalctl -u pulsemon` for the exact error.", ...}

instead of the previous status: "ok" with no icmp_mode key. The engine is
warmed at probe-worker startup, so this is visible on the first healthz after
boot — not after the first failed probe tick. The failure error now wraps the
errNoIcmpTransport sentinel so tooling can match it with errors.Is while
still carrying the underlying cause.

3. Dashboard banner

The existing warning banner now fires on icmp_hint (before the generic
probe_error), showing the remediation command right on the dashboard.

Verified

  • go test — full suite green, including the new TestHealthzIcmpMode
    (dead → degraded/unavailable/hint; unprivileged-datagram → ok; raw → ok),
    and the extended TestDeriveStatus cases (error→up on one success;
    warning flapping stays warning; still-down stays error; DB-error path
    unchanged).
  • Live container matrix on zentest: RHEL-8-like netns (range 1 0, no caps)
    → degraded + hint; wide range → ok; raw-socket-possible → ok.

v0.1.26

A review-drive...

Read more

v0.1.27 — pulsemon

Choose a tag to compare

@invizen invizen released this 06 Oct 20:30

v0.1.27

Renamed: zenmon → pulsemon

The project is now pulsemon. Repo: github.com/invizen/pulsemon (the old
invizen/zenmon URL 301-redirects after the GitHub rename, so existing links
keep working). Everything user-facing moved with it:

  • Binary, install prefix, DB: ~/pulsemon/pulsemon, ~/pulsemon/data/pulsemon.db
  • Release assets: pulsemon-linux-{amd64,arm64} (+ .sha256)
  • systemd unit: pulsemon.service (user-level, as before)
  • Env vars: PULSEMON_ADDR, PULSEMON_DB (was ZENMON_*)
  • Dashboard wordmark: pulse in the EKG green (#10b981) + mon in
    neutral gray (#808080)

Existing installs are NOT auto-migrated — the v0.1.27 release notes carry a
one-shot migration (move the dir, install the new unit, disable the old one).
zenmon update on old binaries keeps working through the GitHub redirect
until the instance is migrated.

One-shot migration (bare / user-systemd installs)

# stop the old service, move the data (DB + binary) to the new prefix
systemctl --user stop zenmon
systemctl --user disable zenmon
mv ~/zenmon ~/pulsemon
# install the new unit (writes ~/.config/systemd/user/pulsemon.service)
curl -sfL https://github.com/invizen/pulsemon/releases/latest/download/install.sh | bash
# remove the stale unit + old PATH line (the installer adds the new one)
rm -f ~/.config/systemd/user/zenmon.service
sed -i '/zenmon: keep the zenmon binary on PATH/d' ~/.bashrc
sed -i '/^export PATH="\$HOME\/zenmon:/d' ~/.bashrc
systemctl --user daemon-reload

Docker installs: just point compose.yaml at the new image name; the data
volume carries over unchanged.

Probe tuning: 60s default, faster error recovery

Three behavior changes around how a sensor is watched when it fails:

  • Default interval 60s (was 15s) for new sensors, the fresh-install seed
    sensors, and the dashboard form — one less DB write / ping on a healthy
    install; per-sensor intervals are untouched.
  • Default "error after" = 2 consecutive losses (was 4): with the 60s
    interval, a dead target was previously confirmed down after ~4 minutes;
    now after ~2 minutes. The fast re-check below closes the recovery side of
    the same gap.
  • Fast re-check in error: while a sensor's status is error it is
    probed every 30s instead of the configured interval, until a probe
    brings it back to up — then it reverts to the configured interval. If the
    sensor's interval is already shorter than 30s, the error state keeps the
    sensor's own pace (it never polls faster than normal).
  • Error → up on 1 successful ping: a sensor recovering from error
    flips back to up on a single good probe, instead of waiting for 2
    straight good replies. A flapping sensor in warning is unaffected:
    loss/up/loss/up still reads warning (one good reply amid ongoing loss is
    not "recovered" — and up↔warning never alerts anyway, so there is no new
    alert noise).

No ping storms. The fast cadence is capped at the configured interval
(never faster than normal) and floored at 10s, so worst case — every sensor
errors at once (total outage) — total probe load is bounded to ~2× the
normal load, spread evenly: each sensor keeps its staggered probe phase, so
they do not re-synchronize into a burst. Each sensor loop owns its own
timer; one slow sensor can't delay another.

Closes the silent ICMP failure mode

When neither ICMP transport could
open (e.g. RHEL 8's or Ubuntu 18.04's default net.ipv4.ping_group_range
excludes the service uid and CAP_NET_RAW isn't granted — any distro with
systemd < 244, since that's the version that ships the wide range), pulsemon
used to start the server, report status: "ok" in healthz, and simply never
ping — with no error anywhere.
The failure was only findable by noticing the missing icmp_mode key.
Now the failure is loud, at three layers:

1. install.sh preflight (fail before the service starts)

Before installing, the script reads net.ipv4.ping_group_range and checks
whether the installer's gid falls inside it. If not, it prints the exact
remediation (sysctl + the persistent /etc/sysctl.d/90-pulsemon-ping.conf)
and, on an interactive terminal, asks before continuing (non-interactive
installs continue but flag that the dashboard will warn). Background: the
kernel default is 1 0 (nobody may ping); systemd ≥ 244 — RHEL 9+, Fedora,
Ubuntu/Debian — ships 0 2147483647 via 50-default.conf, but RHEL 8
(systemd 239) does not.

2. Honest healthz

When the shared ICMP engine fails to open (both transports), GET /api/healthz now returns:

{"status": "degraded", "icmp_mode": "unavailable", "icmp_hint": "ICMP socket
unavailable: ... Fix: sudo sysctl -w net.ipv4.ping_group_range=\"0 65535\"
(persist via /etc/sysctl.d/90-pulsemon-ping.conf) and restart pulsemon — or grant
CAP_NET_RAW. See `journalctl -u pulsemon` for the exact error.", ...}

instead of the previous status: "ok" with no icmp_mode key. The engine is
warmed at probe-worker startup, so this is visible on the first healthz after
boot — not after the first failed probe tick. The failure error now wraps the
errNoIcmpTransport sentinel so tooling can match it with errors.Is while
still carrying the underlying cause.

3. Dashboard banner

The existing warning banner now fires on icmp_hint (before the generic
probe_error), showing the remediation command right on the dashboard.

Verified

  • go test — full suite green, including the new TestHealthzIcmpMode
    (dead → degraded/unavailable/hint; unprivileged-datagram → ok; raw → ok),
    and the extended TestDeriveStatus cases (error→up on one success;
    warning flapping stays warning; still-down stays error; DB-error path
    unchanged).
  • Live container matrix on zentest: RHEL-8-like netns (range 1 0, no caps)
    → degraded + hint; wide range → ok; raw-socket-possible → ok.