Skip to content

Releases: fadyy2k/platform-engineering-eks-gitops

v0.7.1 — Local Recovery & Runtime Security Evidence

Choose a tag to compare

@fadyy2k fadyy2k released this 22 Sep 14:34
c559a15

Engineering intent

Extend the v0.7 local-runtime milestone with recovery, failure-injection and deeper runtime-security evidence while preserving the boundary between local Kubernetes proof and AWS/EKS proof.

New runtime evidence

Captured on kind / Kubernetes v1.36.4:

  • signed immutable project image admitted by server-side policy evaluation
  • mutable project tag denied by the digest policy
  • syntactically valid but unknown project digest denied by ImageValidatingPolicy
  • namespace Pod Security restricted enforcement confirmed
  • Prometheus 14/14 active targets UP, including platform-demo 2/2
  • controlled error-burn injected 200 HTTP 503 responses and the Deployment recovered to 2/2
  • PDB/HPA runtime state captured after recovery
  • Velero backup Completed and namespace-mapped restore Completed, with the restored Deployment reaching 2/2 against local MinIO-backed storage
  • OpenCost returned live namespace allocation data and VPA returned a measured recommendation in Off mode
  • Falco produced a Kubernetes-attributed Critical event from a benign temporary runtime probe
  • Trivy Operator was generating vulnerability/config/RBAC/compliance reports; the demo image had no vulnerability findings at capture time, while one medium configuration finding was retained in the evidence instead of hidden

Reproducibility

Evidence boundary

This release proves local Kubernetes runtime behavior. It still does not claim:

  • GitHub→AWS OIDC role assumption;
  • S3/KMS Terraform state;
  • EKS control plane or managed nodes;
  • AWS load balancer/NAT behavior;
  • AWS-priced OpenCost data;
  • multi-region failover.

The Velero proof used local MinIO-backed storage and must not be represented as AWS backup evidence.

v0.7.0 — Local Runtime Evidence

Choose a tag to compare

@fadyy2k fadyy2k released this 22 Sep 11:39
1f002ca

Engineering intent

Add real Kubernetes runtime evidence without requiring AWS. Phase 7 proves that the GitOps, admission, observability, runtime-security, resilience and cost/right-sizing paths actually execute in a disposable Kubernetes 1.36 environment while keeping AWS-specific claims separate.

Runtime evidence

Captured on a local kind cluster running Kubernetes v1.36.4:

  • Argo CD applications for the demo, security policies, VPA and cost controls reached Healthy / Synced
  • the signed immutable project image was admitted by Kyverno server-side admission
  • nginx:1.27 was denied by the project image policy
  • Prometheus reported both platform-demo scrape targets 2/2 up and loaded the SLO alert/recording groups
  • Trivy Operator produced an in-cluster report for the pinned digest with 0 critical / 0 high / 0 medium / 0 low findings at capture time
  • a benign temporary-pod probe triggered Falco's Read sensitive file untrusted rule with Kubernetes attribution
  • a controlled pod-failure game day recovered the Deployment to 2/2 while Argo remained Healthy/Synced
  • VPA stayed non-mutating (updateMode: Off) and returned a measured recommendation
  • OpenCost returned live namespace allocation data from the in-cluster Prometheus service

Reproducibility

Evidence boundary

This release proves local Kubernetes runtime behavior, not AWS/EKS behavior. It does not claim GitHub→AWS OIDC, S3/KMS Terraform state, EKS managed nodes, AWS networking/load balancers, AWS-priced OpenCost data or multi-region failover.

OpenCost values from the local cluster use a local/default provider model and are explicitly not an AWS bill.

v0.6.0 — Live Readiness & Evidence

Choose a tag to compare

@fadyy2k fadyy2k released this 21 Sep 14:25
b7cc49d

Engineering intent

Harden the reference for real non-production activation without crossing the cloud-cost boundary prematurely. Phase 6 removes assumptions that static CI cannot prove and makes runtime evidence explicit.

Architecture delta

  • EKS example baseline moved from Kubernetes 1.31 extended support to 1.36 standard support
  • EKS upgrade policy explicitly set to STANDARD
  • monitoring-scoped resources separated from platform-demo Argo application paths
  • dedicated monitoring Argo Applications for Alertmanager/OpenCost Prometheus resources
  • read-only live-activation preflight with account/region and cost-boundary checks
  • dry-run-by-default Argo CD bootstrap pinned to chart 10.9.2 and private ClusterIP
  • live activation runbook covering bootstrap, private API access, runtime validation and teardown
  • engineering evidence ledger with real public GitHub Actions screenshots
  • CONTRIBUTING guide plus bug/design/live-validation issue forms
  • dedicated Live Readiness CI for namespace boundaries, helper tooling and EKS baseline

Reproducible evidence

Evidence boundary

This is a live-readiness release, not a claim that AWS/EKS is already running. The AWS plan/apply path remains inactive until an approved AWS account is authenticated and a reviewed dev plan is applied. Live Argo reconciliation, alert delivery, Velero restore, OpenCost allocations and VPA recommendations remain future runtime evidence.

v0.5.0 — Cost & Multi-Environment Operations

Choose a tag to compare

@fadyy2k fadyy2k released this 20 Sep 13:31
601db42

Engineering intent

Add operational economics and reusable environment structure without allowing cost tooling to mutate workloads automatically.

Architecture delta

  • OpenCost connected to the in-cluster Prometheus design
  • tested monthly node run-rate recording rule and USD 150 lab warning
  • VPA recommender-only mode (updateMode: Off)
  • read-only right-sizing report helper
  • VPC/EKS extracted into a reusable Terraform platform module
  • independent environment state/configuration retained
  • active/passive multi-region DR ADR with explicit activation criteria
  • dedicated Cost & Operations CI

Evidence

OpenCost/VPA chart rendering, cost-rule syntax and behavior tests, manifests and helper tooling are exercised by Cost & Operations CI. See Engineering Evidence.

Evidence boundary

The USD 150 alert is a lab node run-rate threshold, not a total AWS budget. This release did not claim real OpenCost allocation data, VPA recommendations or a tested regional failover.

v0.4.0 — Reliability Engineering & SLOs

Choose a tag to compare

@fadyy2k fadyy2k released this 20 Sep 13:24
822dcf9

Engineering intent

Treat reliability as measurable behavior: define the SLO, test the alert logic, design failure handling and make recovery executable.

Architecture delta

  • explicit 99.5% HTTP availability SLO
  • Prometheus recording rules and multi-window error-budget burn alerts
  • promtool rule-behavior tests in CI
  • project-owned /metrics endpoint and controlled HTTP failure mode
  • ServiceMonitor discovery
  • PodDisruptionBudget, topology spreading and zero-unavailable rolling updates
  • HPA + metrics-server
  • pinned kube-prometheus-stack through Argo CD
  • severity-aware Alertmanager routing topology
  • dry-run-by-default pod-failure and error-burn game days
  • executable Velero backup/restore smoke-test harness
  • operational runbooks and node-autoscaling ADR
  • signed Phase 4 image pinned by immutable GHCR digest

Evidence

Metrics code, Prometheus rules, manifests and recovery/game-day tooling are tested by Reliability CI. See Engineering Evidence.

Evidence boundary

No live-cluster recovery or external Alertmanager delivery was claimed at this phase. Those require an approved non-production environment.

v0.3.0 — Policy & Runtime Security

Choose a tag to compare

@fadyy2k fadyy2k released this 20 Sep 12:48
65a625b

Engineering intent

Move enforcement inside the Kubernetes trust boundary so security remains active after CI finishes.

Architecture delta

  • restricted Pod Security Admission for the demo namespace
  • default-deny NetworkPolicy with explicit DNS/application traffic
  • Kyverno ValidatingPolicy requiring project-owned images by digest
  • Kyverno ImageValidatingPolicy verifying the Cosign keyless signing identity
  • Trivy Operator for recurring vulnerability, SBOM, configuration, RBAC, infrastructure and compliance reports
  • Falco runtime detection with modern eBPF configuration
  • dedicated policy/runtime-security CI
  • positive and negative Kyverno tests, including registry signature verification

Evidence

Policy behavior and chart rendering are exercised by the repository's Policy and Runtime Security CI. See Engineering Evidence.

Evidence boundary

The release proved policy definitions and signature verification in CI. It did not claim in-cluster Trivy/Falco reports from a provisioned EKS cluster.

v0.2.0 — Identity, State & Signed Delivery

Choose a tag to compare

@fadyy2k fadyy2k released this 20 Sep 12:10
c2e00e3

Engineering intent

Replace static cloud credentials and mutable delivery assumptions with short-lived identity, protected state and signed artifacts.

Architecture delta

  • GitHub Actions → AWS OIDC federation design
  • separate scoped Terraform plan/apply roles
  • KMS-encrypted, versioned S3 Terraform state with native S3 locking
  • independent dev, staging and prod state/configuration
  • manual protected-environment apply workflow
  • project-owned Go demo service
  • GHCR build with BuildKit provenance and SBOM attestations
  • Trivy HIGH/CRITICAL image gate
  • Cosign keyless signing and workflow-identity verification
  • Kubernetes desired state pinned to a verified immutable digest

Evidence

The container supply-chain path was exercised in GitHub Actions and the signed digest was verified. Current reproducible links are collected in Engineering Evidence.

Evidence boundary

AWS bootstrap/apply remained an explicit operator action. This release did not create AWS infrastructure or claim a live OIDC-backed Terraform plan against an AWS account.

v0.1.0 — Platform Baseline

Choose a tag to compare

@fadyy2k fadyy2k released this 20 Sep 10:33

Engineering intent

Establish a reviewable platform baseline before adding identity, supply-chain or runtime-security complexity.

Architecture delta

  • Terraform-managed AWS VPC and EKS reference
  • private-first EKS API configuration
  • managed node group + EBS CSI add-on
  • Argo CD Application for the demo workload
  • hardened Kubernetes workload and encrypted gp3 StorageClass
  • Prometheus/Grafana configuration baseline
  • Terraform, Kubernetes and security CI
  • Dependabot, secret scanning, push protection and CODEOWNERS

Evidence

The baseline is reproducible through repository CI and the current Engineering Evidence ledger.

Evidence boundary

This release defined the desired infrastructure and validation path. It did not claim that an AWS/EKS environment had been provisioned.