v0.3.86 — Agent no longer drops mid-update on busy hosts
Fix: agent dropped mid-update on busy hosts
On a heavily loaded host (e.g. 12 nodes on one server, Docker discovery taking ~10s instead of milliseconds), an agent update triggered from the dashboard reliably failed.
Root cause (agent side): the WebSocket keepalive closed the connection on the first missed pong. But conn.Ping is only satisfied once the read loop processes the pong frame — and during an update the read loop is busy pulling the ~15 MB agent binary. Under host contention the pong arrives late even though the link is perfectly healthy, so the connection was torn down mid-update and the dashboard reported "failed". (Discovery was already off the hot path on its own goroutine, so the earlier "discovery blocks the loop" theory no longer applied.)
Fix: the keepalive now tolerates up to 3 consecutive missed pings (counter resets on any success) before treating the connection as dead. A transient contention spike — or the update transfer itself — no longer drops the agent, while a genuinely half-open connection is still detected within ~60-75s.
⚠️ Agent-side only. The fixed binary must be rolled out once manually (as before) — after that, dashboard-driven agent updates run through cleanly even on busy hosts.
Full Changelog: v0.3.85...v0.3.86