Skip to content

v0.3.83 — Fix agent heartbeat stalls under DB contention

Choose a tag to compare

@CTJaeger CTJaeger released this 20 Jun 07:10
· 5 commits to main since this release

Fix: agents intermittently dropped ("heartbeat stalled")

Agents across servers went Agent Offline simultaneously at irregular intervals (and reliably every hour), then reconnected within seconds — while the dashboard itself stayed healthy.

Root cause: the per-agent WebSocket read loop persisted heartbeat + node metrics to SQLite synchronously through the shared metrics lock. Whenever that lock was held a while — most reliably the hourly Decimate, which ran the whole aggregate+delete in one long transaction — every agent's read loop blocked. Since coder/websocket only emits pongs while conn.Read is running, all agents' pings timed out at the same moment.

Fix:

  1. Heartbeat/metrics DB writes now run on a per-connection background worker, so the read loop (and pongs) stay responsive regardless of DB contention. In-memory liveness still updates synchronously; under backpressure a metrics sample is dropped rather than blocking.
  2. Decimation runs in short bucket-aligned slices, each in its own transaction, releasing the lock between slices.
  3. Purge batches its deletes, and the SQLite busy_timeout is raised 5s→30s.

Dashboard-side only — agents are unchanged, so you only need to update the dashboard. New test TestDecimate_MultipleSlices.

Full Changelog: v0.3.82...v0.3.83

What's Changed

  • fix: agents intermittently drop ("heartbeat stalled") under DB contention by @Test0rMaik in #63

Full Changelog: v0.3.82...v0.3.83