v0.3.83 — Fix agent heartbeat stalls under DB contention
Fix: agents intermittently dropped ("heartbeat stalled")
Agents across servers went Agent Offline simultaneously at irregular intervals (and reliably every hour), then reconnected within seconds — while the dashboard itself stayed healthy.
Root cause: the per-agent WebSocket read loop persisted heartbeat + node metrics to SQLite synchronously through the shared metrics lock. Whenever that lock was held a while — most reliably the hourly Decimate, which ran the whole aggregate+delete in one long transaction — every agent's read loop blocked. Since coder/websocket only emits pongs while conn.Read is running, all agents' pings timed out at the same moment.
Fix:
- Heartbeat/metrics DB writes now run on a per-connection background worker, so the read loop (and pongs) stay responsive regardless of DB contention. In-memory liveness still updates synchronously; under backpressure a metrics sample is dropped rather than blocking.
- Decimation runs in short bucket-aligned slices, each in its own transaction, releasing the lock between slices.
- Purge batches its deletes, and the SQLite
busy_timeoutis raised 5s→30s.
Dashboard-side only — agents are unchanged, so you only need to update the dashboard. New test TestDecimate_MultipleSlices.
Full Changelog: v0.3.82...v0.3.83
What's Changed
- fix: agents intermittently drop ("heartbeat stalled") under DB contention by @Test0rMaik in #63
Full Changelog: v0.3.82...v0.3.83