Severity: P3 (spurious alerts from transient network blips)
Component: Network Plane — cmd/atenet/internal/router/health.go
Audit ID: NET-8
Summary
updateComponentHealth() sets health.Healthy = healthy directly from each probe
result with no consecutive-failure threshold or smoothing. A single 500ms network blip
to Envoy's admin port (:9901) or the Kubernetes API causes /statusz to report
unhealthy for one polling interval (default 1s). While this does not affect traffic
routing (health check is observability-only), monitoring systems built on /statusz
will page on-call for a 1-second blip.
Root Cause
File: cmd/atenet/internal/router/health.go lines 147–157
func (h *HealthState) updateComponentHealth(name string, healthy bool) {
h.mu.Lock()
defer h.mu.Unlock()
h.components[name] = ComponentHealth{
Healthy: healthy, // Direct assignment — no threshold
// ...
}
}
Steps to Reproduce
# Drop Envoy admin traffic for 600ms (just over one health check interval)
kubectl exec -n ate-system <router-pod> -- \
iptables -I OUTPUT -p tcp --dport 9901 -j DROP
# Poll /statusz from another terminal
watch -n 0.5 "curl -s http://localhost:8080/statusz | jq '.components.envoy.healthy'"
# Flips to false for one cycle
kubectl exec -n ate-system <router-pod> -- \
iptables -D OUTPUT -p tcp --dport 9901 -j DROP
# Recovers after one successful probe
Suggested Fix
Track consecutive failure count per component. Only flip to unhealthy after N
consecutive failures (e.g., N=3):
type ComponentHealth struct {
Healthy bool
consecutiveFails int
// ...
}
func (h *HealthState) updateComponentHealth(name string, healthy bool) {
h.mu.Lock()
defer h.mu.Unlock()
c := h.components[name]
if healthy {
c.consecutiveFails = 0
c.Healthy = true
} else {
c.consecutiveFails++
if c.consecutiveFails >= 3 {
c.Healthy = false
}
}
h.components[name] = c
}
Severity: P3 (spurious alerts from transient network blips)
Component: Network Plane —
cmd/atenet/internal/router/health.goAudit ID: NET-8
Summary
updateComponentHealth()setshealth.Healthy = healthydirectly from each proberesult with no consecutive-failure threshold or smoothing. A single 500ms network blip
to Envoy's admin port (
:9901) or the Kubernetes API causes/statuszto reportunhealthy for one polling interval (default 1s). While this does not affect traffic
routing (health check is observability-only), monitoring systems built on
/statuszwill page on-call for a 1-second blip.
Root Cause
File:
cmd/atenet/internal/router/health.golines 147–157Steps to Reproduce
Suggested Fix
Track consecutive failure count per component. Only flip to unhealthy after N
consecutive failures (e.g., N=3):