Repository navigation
v0.4.0 — Reliability Engineering & SLOs
Engineering intent
Treat reliability as measurable behavior: define the SLO, test the alert logic, design failure handling and make recovery executable.
Architecture delta
- explicit 99.5% HTTP availability SLO
- Prometheus recording rules and multi-window error-budget burn alerts
promtoolrule-behavior tests in CI- project-owned
/metricsendpoint and controlled HTTP failure mode - ServiceMonitor discovery
- PodDisruptionBudget, topology spreading and zero-unavailable rolling updates
- HPA + metrics-server
- pinned kube-prometheus-stack through Argo CD
- severity-aware Alertmanager routing topology
- dry-run-by-default pod-failure and error-burn game days
- executable Velero backup/restore smoke-test harness
- operational runbooks and node-autoscaling ADR
- signed Phase 4 image pinned by immutable GHCR digest
Evidence
Metrics code, Prometheus rules, manifests and recovery/game-day tooling are tested by Reliability CI. See Engineering Evidence.
Evidence boundary
No live-cluster recovery or external Alertmanager delivery was claimed at this phase. Those require an approved non-production environment.