Skip to content

v0.3.84 — Fix SQLITE_BUSY + false Agent Offline alerts

Choose a tag to compare

@CTJaeger CTJaeger released this 23 Jun 19:51
· 3 commits to main since this release

Two issues that persisted after v0.3.83 (#63).

SQLITE_BUSY every hour during decimation

The PRAGMA busy_timeout=30000 is per-connection — it only applied to the one connection that ran it; pooled connections got SQLite's 0ms default and errored SQLITE_BUSY under write contention (confirmed hourly in production logs). Fixed by SetMaxOpenConns(1): all DB access serializes on a single connection, so database/sql queues writes in Go-land — no SQLite lock contention, SQLITE_BUSY structurally impossible. The earlier revert of this setting (startup hang from a decimation backlog) no longer applies now that decimation is chunked; a 10s startup delay is added as safety.

False "Agent Offline" alerts while agents are connected

After #63's lossy heartbeat persistence, the evaluator could read a stale LastHeartbeat from the DB and fire "Agent Offline" while the agent was still alive. Fixed two ways:

  • Heartbeat DB writes use enqueueOrSpawn — never dropped (falls back to a goroutine when the worker queue is full). Node metrics stay lossy.
  • The evaluator takes a hub.IsConnected callback and suppresses offline alerts for any agent whose WebSocket is currently live — in-memory hub state is the authoritative liveness source.

Dashboard-side only — agents unchanged.

Full Changelog: v0.3.83...v0.3.84

What's Changed

  • fix: SetMaxOpenConns(1) + evaluator hub-awareness — eliminate SQLITE_BUSY and false offline alerts by @Test0rMaik in #64

Full Changelog: v0.3.83...v0.3.84