Skip to content

v0.4.0 — Reliability Engineering & SLOs

Choose a tag to compare

@fadyy2k fadyy2k released this 20 Sep 13:24
· 7 commits to main since this release
822dcf9

Engineering intent

Treat reliability as measurable behavior: define the SLO, test the alert logic, design failure handling and make recovery executable.

Architecture delta

  • explicit 99.5% HTTP availability SLO
  • Prometheus recording rules and multi-window error-budget burn alerts
  • promtool rule-behavior tests in CI
  • project-owned /metrics endpoint and controlled HTTP failure mode
  • ServiceMonitor discovery
  • PodDisruptionBudget, topology spreading and zero-unavailable rolling updates
  • HPA + metrics-server
  • pinned kube-prometheus-stack through Argo CD
  • severity-aware Alertmanager routing topology
  • dry-run-by-default pod-failure and error-burn game days
  • executable Velero backup/restore smoke-test harness
  • operational runbooks and node-autoscaling ADR
  • signed Phase 4 image pinned by immutable GHCR digest

Evidence

Metrics code, Prometheus rules, manifests and recovery/game-day tooling are tested by Reliability CI. See Engineering Evidence.

Evidence boundary

No live-cluster recovery or external Alertmanager delivery was claimed at this phase. Those require an approved non-production environment.