Releases: CTJaeger/KleverNodeHub
Release list
v0.3.87 — Stop dropping agents on overloaded hosts
Fix: agents on maxed-out hosts dropped by the dashboard
On a heavily loaded host (e.g. 12 nodes on one server, ~90ms link to the dashboard) agents were dropped every ~60-90s and agent updates failed. The dashboard log pinned the cause:
failed to handle control frame opPing: failed to write control frame opPong:
failed to acquire lock: context deadline exceeded
The agent's keepalive Ping arrives, the dashboard must answer with a Pong, but the connection's write lock is held by a conn.Write that isn't draining — the overloaded host isn't reading fast enough, so TCP backpressure builds. The Pong misses its deadline, the dashboard's read loop dies, the socket closes, and the agent sees unexpected EOF.
Two backward-compatible changes:
1. Update binary no longer travels over the WebSocket
Pushing the ~15 MB binary base64-encoded (~20 MB) through a single conn.Write was the worst offender — it held the write lock for the whole transfer and starved Pongs. Agents now advertise an http-update capability; when the dashboard sees it, the update command carries a download_path and the agent pulls the binary over a separate HTTP request (new endpoint GET /api/agent/binary/{server_id}, server_id-validated, checksum still verified before install).
Fully backward compatible, no manual step for anyone: agents that don't advertise the capability keep getting the inline base64 payload; a new agent talking to an old dashboard still accepts inline. Rollout order doesn't matter.
2. Discovery interval adapts to load (agent)
When a discovery cycle is slow (≥5s — Docker-daemon contention) the agent backs the interval off (doubling, capped at 5 min) instead of burning a constant slice of an already-starved CPU every 60s; it snaps back to 60s once discovery is fast again. Less CPU on discovery → the read loop keeps up with the socket → the backpressure eases at the root.
Both sides change (agent + dashboard). Thanks to the capability negotiation, rollout order is irrelevant and no agent needs a manual update.
Full Changelog: v0.3.86...v0.3.87
v0.3.86 — Agent no longer drops mid-update on busy hosts
Fix: agent dropped mid-update on busy hosts
On a heavily loaded host (e.g. 12 nodes on one server, Docker discovery taking ~10s instead of milliseconds), an agent update triggered from the dashboard reliably failed.
Root cause (agent side): the WebSocket keepalive closed the connection on the first missed pong. But conn.Ping is only satisfied once the read loop processes the pong frame — and during an update the read loop is busy pulling the ~15 MB agent binary. Under host contention the pong arrives late even though the link is perfectly healthy, so the connection was torn down mid-update and the dashboard reported "failed". (Discovery was already off the hot path on its own goroutine, so the earlier "discovery blocks the loop" theory no longer applied.)
Fix: the keepalive now tolerates up to 3 consecutive missed pings (counter resets on any success) before treating the connection as dead. A transient contention spike — or the update transfer itself — no longer drops the agent, while a genuinely half-open connection is still detected within ~60-75s.
⚠️ Agent-side only. The fixed binary must be rolled out once manually (as before) — after that, dashboard-driven agent updates run through cleanly even on busy hosts.
Full Changelog: v0.3.85...v0.3.86
v0.3.85 — Fast Reset (clear DB & re-bootstrap)
Fast Reset — clear a node's chain DB and re-bootstrap from the latest epoch
A new action on each node's detail page (Chain Database panel), deliberately separate from the full-history Restore DB:
- Red "Fast Reset — clear DB & re-bootstrap" button below the (yellow) restore button.
- Stops the node, deletes only the local
db/subtree (keepsconfig/, keys andvalidatorKey.pem), recreates an emptydb/owned by the container user (999:999), then recreates the container with--start-in-epochand starts it — so it fast-bootstraps a fresh epoch-aligned state. - No download. Not a full history — for operators who don't need archival history and want a quick fresh sync. Indexer/archival nodes should use Restore DB instead.
The two actions are kept strictly separate (own button, colour, confirm + warning) so a node isn't reset by accident.
New agent action node.reset-db (whitelisted), endpoint POST /api/nodes/{id}/reset-db, live progress over WebSocket.
⚠️ Both sides change — update the agent as well as the dashboard, otherwisenode.reset-dbwon't reach the node.
Full Changelog: v0.3.84...v0.3.85
v0.3.84 — Fix SQLITE_BUSY + false Agent Offline alerts
Two issues that persisted after v0.3.83 (#63).
SQLITE_BUSY every hour during decimation
The PRAGMA busy_timeout=30000 is per-connection — it only applied to the one connection that ran it; pooled connections got SQLite's 0ms default and errored SQLITE_BUSY under write contention (confirmed hourly in production logs). Fixed by SetMaxOpenConns(1): all DB access serializes on a single connection, so database/sql queues writes in Go-land — no SQLite lock contention, SQLITE_BUSY structurally impossible. The earlier revert of this setting (startup hang from a decimation backlog) no longer applies now that decimation is chunked; a 10s startup delay is added as safety.
False "Agent Offline" alerts while agents are connected
After #63's lossy heartbeat persistence, the evaluator could read a stale LastHeartbeat from the DB and fire "Agent Offline" while the agent was still alive. Fixed two ways:
- Heartbeat DB writes use
enqueueOrSpawn— never dropped (falls back to a goroutine when the worker queue is full). Node metrics stay lossy. - The evaluator takes a
hub.IsConnectedcallback and suppresses offline alerts for any agent whose WebSocket is currently live — in-memory hub state is the authoritative liveness source.
Dashboard-side only — agents unchanged.
Full Changelog: v0.3.83...v0.3.84
What's Changed
- fix: SetMaxOpenConns(1) + evaluator hub-awareness — eliminate SQLITE_BUSY and false offline alerts by @Test0rMaik in #64
Full Changelog: v0.3.83...v0.3.84
v0.3.83 — Fix agent heartbeat stalls under DB contention
Fix: agents intermittently dropped ("heartbeat stalled")
Agents across servers went Agent Offline simultaneously at irregular intervals (and reliably every hour), then reconnected within seconds — while the dashboard itself stayed healthy.
Root cause: the per-agent WebSocket read loop persisted heartbeat + node metrics to SQLite synchronously through the shared metrics lock. Whenever that lock was held a while — most reliably the hourly Decimate, which ran the whole aggregate+delete in one long transaction — every agent's read loop blocked. Since coder/websocket only emits pongs while conn.Read is running, all agents' pings timed out at the same moment.
Fix:
- Heartbeat/metrics DB writes now run on a per-connection background worker, so the read loop (and pongs) stay responsive regardless of DB contention. In-memory liveness still updates synchronously; under backpressure a metrics sample is dropped rather than blocking.
- Decimation runs in short bucket-aligned slices, each in its own transaction, releasing the lock between slices.
- Purge batches its deletes, and the SQLite
busy_timeoutis raised 5s→30s.
Dashboard-side only — agents are unchanged, so you only need to update the dashboard. New test TestDecimate_MultipleSlices.
Full Changelog: v0.3.82...v0.3.83
What's Changed
- fix: agents intermittently drop ("heartbeat stalled") under DB contention by @Test0rMaik in #63
Full Changelog: v0.3.82...v0.3.83
v0.3.82 — BLS key generation during provisioning
Generate a BLS validator key during provisioning
The Provision modal has a new "Generate a fresh BLS validator key for each node" checkbox. When ticked, provisioning runs the Klever keygenerator and places a new validatorKey.pem into each node's config dir (before the permissions step, so the key is chown'd to the container user too).
- Works for single and batch provisioning — each node gets its own key.
- The capability already existed as a standalone step (
key.generateon the node detail page); this wires it into the create flow so you can do it in one go. - Left unticked, behavior is unchanged — import or generate the key afterwards.
Full Changelog: v0.3.81...v0.3.82
Full Changelog: v0.3.81...v0.3.82
v0.3.81 — Remove validator monitoring; logs default 5s
Removed the validator monitoring page
The validator monitoring page polled an external public API (api.mainnet.klever.org) block-by-block for its block-production timeline. That contradicts NodeHub's core principle — a self-hosted node manager must not depend on a third-party API. If that API is down or rate-limits, the dashboard degrades for reasons outside the operator's control.
The whole feature is reverted: the Validators sidebar entry, the page, and the internal/dashboard/klever poller are gone. NodeHub gets its data from the operator's own nodes via the agents — never from an outside service.
Node log auto-refresh defaults to 5s
The node detail Logs panel now auto-refreshes every 5s by default instead of 10s.
Full Changelog: v0.3.80...v0.3.81
Full Changelog: v0.3.80...v0.3.81
v0.3.80 — Restore on node detail; timeline elected-only
Restore Chain DB from a single node's detail page
The chain-DB restore was only reachable from the overview's batch bar (after selecting nodes), which was easy to miss. Each node's detail page now has a Chain Database panel — network selector, one button, live progress over WebSocket — so you can restore a single node directly. The batch variant on the overview stays for multiple nodes.
Validator timeline shows only elected validators
Off-chain / non-elected validators don't produce blocks, so they no longer take up empty rows in the Block Production Timeline. The validator table above still lists them with their on-chain state.
Full Changelog: v0.3.79...v0.3.80
Full Changelog: v0.3.79...v0.3.80
v0.3.79 — Validator monitoring, Docker cleanup, restore gap fix
Validator monitoring page
New Validators entry in the sidebar tracking your managed validators (matched by BLS key) on-chain: state (elected/jailed/waiting), commission, self-stake, allowance, blocks produced/missed, plus a live block-production timeline of the last 100 blocks per validator (green = led, red = elected but missing from the signer set, grey = idle). Background poller against the Klever indexer/node API with per-IP rate-limit backoff. Endpoint GET /api/validators.
Docker image cleanup tool
New Docker Cleanup under Tools: pick a server, see every local klever-go image (tags, ID, date, size), and delete the ones you don't need. Images referenced by a container (running or stopped) are flagged in-use and protected; the agent re-verifies before deleting.
Fix: chain-DB restore no longer leaves a gap
After a restore the node was only restarted, so it kept --start-in-epoch and jumped to the latest epoch — skipping every block between the snapshot and now. The restore now recreates the container without --start-in-epoch, so the node resumes from the restored DB height and syncs forward continuously.
Fix: "initializing" badge keyed off the right signal
A fresh node whose REST API isn't up yet (no sync metrics) now shows "initializing" instead of an empty status — signal-based, not a 10-minute timer.
Full Changelog: v0.3.78...v0.3.79
What's Changed
- feat(tools): Docker image cleanup tool to reclaim disk from old node images by @Test0rMaik in #60
- feat(validators): validator monitoring page with block-production timeline by @Test0rMaik in #61
- fix: chain-DB restore recreate without --start-in-epoch; initializing badge by API readiness by @CTJaeger in #62
Full Changelog: v0.3.78...v0.3.79
v0.3.78 — Chain-DB restore, sync modes, maintenance, SW freshness
Restore chain DB from the official Klever snapshot
New Restore DB action (batch bar on the overview) replaces a node's chain database with the official Klever FullNode snapshot (kleverchain.<network>.latest.tar.gz, tens of GB) — for when a node only has the latest epoch but you need the full archival DB (e.g. for an indexer).
Per node: preflight free-disk check (refuses if it wouldn't fit), stop, rotate the old DB aside (db.old, kept for rollback if space allows), stream-download straight through gzip+tar (never staging the archive on disk), extract only the db/ subtree (a stray config/ can't clobber live config), chown 999:999, start. Multiple nodes run one at a time. Progress streams live over WebSocket; the request is fire-and-forget so an hour-long restore doesn't hold an HTTP connection.
Provisioning sync mode
The Provision modal now offers Fast bootstrap (--start-in-epoch, default), Full DB snapshot (downloads the archive before first start — an archival node in one step), or Full sync from genesis.
No more alert spam on deliberate stops
Stopping a node from the dashboard marks it as in maintenance; offline alerts are suppressed for maintenance nodes. Cleared on start/restart and self-healing — discovery clears it as soon as the node is seen running again.
Fresh nodes show "initializing"
A just-started node that legitimately reports syncing now shows an initializing badge for its first 10 minutes of uptime instead of looking like a problem.
Fix: stale assets in a long-open tab (Service Worker)
CSS/JS/pages are now network-first (cache is an offline fallback only); only icons/fonts/manifest stay cache-first. Fixes old styles being served after a deploy until a hard refresh — the broken-layout-in-an-open-tab bug.
Full Changelog: v0.3.77...v0.3.78
Full Changelog: v0.3.77...v0.3.78