Repository navigation
v0.43.0 — incident quality gate: stop phantom flapping
Problem observed
On the live fleet, every brief CPU/PSI blip was producing a full IncidentStarted / Escalated / Resolved cycle — often 3–5 rows per second during noisy periods with peak=0% and confidence=0%. 20+ phantom rows in 2 minutes on one host. Over a day that's thousands of useless rows in Postgres and a dashboard buried in noise.
Two-layer fix
Agent side (v0.43.0 primary defense)
Every agent now requires before emitting IncidentStarted:
| Gate | Default | Env override |
|---|---|---|
| Min peak score | 30% | XTOP_FLEET_QUALITY_MIN_SCORE |
| Min confidence | 40% | XTOP_FLEET_QUALITY_MIN_CONF |
| Sustained ticks | 3 consecutive | XTOP_FLEET_QUALITY_MIN_TICKS |
| Gap between escalations | 15 s | XTOP_FLEET_QUALITY_ESC_GAP_SEC |
Plus invariants:
- Resolved emits only if Started did — no dangling close-outs for incidents the hub never saw.
- Signature flips rate-limited from the incident's START, not just the last escalation.
- One-off spikes produce zero emissions — by construction: one tick can't reach 3 sustained.
Hub side (defense in depth)
- Drops
peak_score==0 && confidence==0payloads (except legitimate Resolved close-outs). - Dedupes same
(agent_id, signature, update_type)within 30 s with a 2-minute TTL map. - Returns
204 No Content— suppression is silent to agents.
Impact
On the real fleet that prompted this fix:
- Before: 20+
peak=0/conf=0incidents in 2 minutes → Postgres flooded, dashboard unreadable. - After: zero. Genuine high-score incidents (e.g. the CPU contention at 77% / 93% conf) still fire, unchanged.
Upgrade
wget https://github.com/ftahirops/xtop/releases/download/v0.43.0/xtop_0.43.0-1_amd64.deb
sudo dpkg -i xtop_0.43.0-1_amd64.deb
sudo systemctl restart xtop-hub # on the hub host
sudo systemctl restart xtop-agent # on every agent hostBoth sides should upgrade, but the hub alone already kills the flood from older agents — the gate is double-enforced intentionally.
Tuning for noisy or quiet workloads
# More aggressive (capture sooner): drop thresholds on the agent
sudo systemctl edit xtop-agent
# add: Environment=XTOP_FLEET_QUALITY_MIN_SCORE=20
# Environment=XTOP_FLEET_QUALITY_MIN_TICKS=2
# More quiet (only serious events): raise them
# Environment=XTOP_FLEET_QUALITY_MIN_SCORE=50
# Environment=XTOP_FLEET_QUALITY_MIN_TICKS=5systemctl daemon-reload && systemctl restart xtop-agent after edits.
What the next release (v0.44) will add
Baseline-aware suppression: per-host, per-bottleneck p95 tracker. If the current peak is within the host's established 7-day p95 for that bottleneck, mark it as "normal variance" and skip. Requires more history than most agents have today, so shipping once hosts have accumulated data.