Skip to content

Runbook Prod Cutover EventLog

Kadyapam edited this page Jul 12, 2026 · 3 revisions

Runbook — Prod Cutover: EHDB Event-Log Tier (Phase 9, Tier 1)

Status: DRAFT — awaiting user approval. NOTHING IN THIS RUNBOOK HAS BEEN EXECUTED. No prod/GKE/gcloud action has been taken, no image deployed, no flag flipped. This page is a reviewable plan for the first production EHDB cutover — serving the append-only noetl.event log from EHDB in place of the NATS JetStream + Postgres incumbent. The operator executes it later, per stage, only on the user's explicit go.

Scope: Tier 1 of five (Roadmap Phase 9). The other four tiers (projection, KV, object, vector) stay on their external incumbents and are out of scope here — each is a separate, separately-gated cutover.

Tracks: noetl/ehdb#241 (completion program) · noetl/ai-meta#166 (the WAL-index scaling pressure this tier relieves).

Audience: the operator with write access to the prod GKE cluster.


TL;DR — the three-stage plan

Stage Action Serving backend Reversible by Safe in prod today?
A Deploy worker v5.66.0, all EHDB flags off JetStream + Postgres roll image back to v5.52.0 Yes — behavior-neutral no-op
B NOETL_EHDB_ENABLED=1 + NOETL_EHDB_EVENTLOG=shadow (data-plane pool only) JetStream + Postgres (EHDB dual-writes + parity, never serves) flag → off (instant) Yes — incumbent still authoritative
C NOETL_EHDB_EVENTLOG=primary EHDB (incumbent dual-run parity-checked) flag → shadow/off (instant, zero data loss) BLOCKED — see the durability gate below

Load-bearing caveat, read before planning Stage C. The only EHDB backend the worker implements today is local_reference — a pod-local JSONL file (NOETL_EHDB_LOCAL_REFERENCE_LOG). Under shadow that file is derived and disposable, so Stage B is safe. Under primary that file becomes the authoritative event store — and a pod-local file on a multi-replica pool means (a) each replica owns a divergent "authoritative" log, and (b) the log is lost on pod restart/reschedule. Stage C is therefore not production-safe as the code ships until a durable, shared EHDB log substrate (PVC-backed volume or object-backed store) exists and the serving topology is pinned to a single writer. Stages A and B are executable now; Stage C is gated behind that durability decision (see Blast radius §C).


0. Ground facts (verify each before acting)

These were read from the repos on 2026-07-06. Prod drifts from committed manifests — re-verify live before you act.

Fact Value Source / how to re-verify
Cluster context gke_noetl-demo-19700101_us-central1_noetl-cluster kubectl config get-contexts
Namespace noetl —
Prod stack full Rust (app=noetl-server-rust); Python retired kubectl -n noetl get deploy
Current prod worker v5.52.0 (all NOETL_EHDB_* flags off) kubectl -n noetl get deploy noetl-worker-system-pool -o jsonpath='{..image}'
Target worker image v5.66.0 — carries all 5 tiers' primary-serve + Phase 10 config verb worker#166
Prod image registry us-central1-docker.pkg.dev/noetl-demo-19700101/noetl/noetl-worker-rust (pin by digest) gcloud artifacts docker images list
Rollback image v5.52.0 (current prod digest — capture it before Stage A) —
Monitoring Google Managed Prometheus (GMP), not VictoriaMetrics ci/manifests/noetl/gmp/; Cloud Monitoring
Deploy mechanism kubectl apply of ci/manifests/noetl/*-prod.yaml + kubectl set env for flag flips noetl-cqrs-publish-only-flip.md

Which worker deployment(s) carry the event-log tier

The EHDB event-log integration lives in-process in worker-rust (src/ehdb/eventlog.rs), gated by NOETL_EHDB_EVENTLOG. It runs only on data-plane roles (worker/playbook/system); the control-plane guard refuses server/gateway/api (exit 4). The natural — and likely only correct — home for the event-log tier is the system pool (noetl-worker-system-pool), because prod already makes that pool the sole event-log writer (the CQRS materializer + off-server state builder run there; see the CQRS publish-only flip).

⚠️ ASSUMPTION TO CONFIRM (A1). This runbook sets the flag on noetl-worker-system-pool only, on the premise that it is the single event writer in prod. Before Stage B, confirm no other pool (noetl-worker-rust shared pool, any subscription pool) also authors event appends. If another pool writes events, it must carry the same flag, and Stage C's single-writer requirement is not met without further work.

Env-var contract (from src/ehdb/eventlog.rs + Backend Configuration)

NOETL_EHDB_ENABLED=true                       # umbrella gate; unset ⇒ strict no-op
NOETL_EHDB_MODE=local_reference               # the only backend implemented today
NOETL_EHDB_CLIENT_ROLE=system                 # data-plane role (worker|playbook|system) — see A2
NOETL_EHDB_LOCAL_REFERENCE_LOG=/var/lib/ehdb/eventlog.jsonl   # pod-local path — see durability gate
NOETL_EHDB_EVENTLOG=off|shadow|primary        # the tier mode; default off

⚠️ ASSUMPTION TO CONFIRM (A2). The NOETL_EHDB_CLIENT_ROLE token for the system pool. Any of worker/playbook/system passes the data-plane guard; the coherence resolver only rejects control-plane roles. Confirm the canonical token the system pool should advertise. This runbook uses system.

⚠️ ASSUMPTION TO CONFIRM (A3). The Phase-A ops Helm EHDB env-rendering (noetl/ops#234) is on branch feat/ehdb-phase-a-ops and is not merged to ops main, so the prod Helm chart does not render NOETL_EHDB_* today. This runbook therefore sets the flags directly with kubectl set env (the same mechanism the CQRS flip used). If the operator prefers a declarative path, merge + pointer-bump the ops EHDB env rendering first and set the values there instead. Either way, record the final env in worker-system-pool-deployment-prod.yaml so a later kubectl apply is a no-op.

⚠️ ASSUMPTION TO CONFIRM (A4). The committed ci/manifests/noetl/worker-system-pool-deployment-prod.yaml currently pins v5.40.3 / replicas: 1, while live prod is v5.52.0 on a 2-replica affinity drive pool (noetl/ai-meta#172). The committed manifest lags live prod. Confirm the actual live deploy method (edit-manifest-then-apply vs kubectl set image vs Helm) and the actual current system-pool shape (replica count, whether it is one Deployment or shard-0/shard-1 Deployments) before Stage A. A multi-replica system pool directly bears on the Stage C durability gate.


1. Preconditions / gates (all must be true before Stage A)

  • User has approved this runbook (sign-off §7) and given an explicit go for the stage about to run.
  • worker v5.66.0 built and present in the prod Artifact Registry, pinned by digest. Capture the digest.
  • Kind dual-run is green for tier 1 — already true: ehdb-selfcheck eventlog-primary-serve passed in-cluster on kind-noetl (served_by_ehdb:true, dual_run_holds:true, reversible:true, secret-free metrics). See Roadmap Tier 1.
  • A1–A4 assumptions above confirmed with the user / by live inspection.
  • Off-hours change window agreed; on-call engineer aware and reachable.
  • Incumbent event-log retention/backup confirmed — the JetStream noetl_events stream and the Postgres noetl.event table are the rollback target; verify current retention and that a point-in-time restore path exists before touching the serving engine.
  • Rollback rehearsed — the operator has run the Stage-B and Stage-C revert commands against kind-noetl (or a staging namespace) and confirmed the incumbent resumes serving.
  • GMP scrape live for noetl_ehdb_eventlog_* — confirm the worker /metrics surface is scraped by the GMP PodMonitoring so the parity metrics are queryable in Cloud Monitoring before Stage B.
  • Current prod worker digest captured as the Stage-A rollback target.

2. Stage A — deploy v5.66.0, all EHDB flags OFF (behavior-neutral)

Goal: get the primary-serve-capable binary into prod as a no-op. With no NOETL_EHDB_* env set, the worker is byte-identical to v5.52.0 for the event path.

Steps

  1. Capture rollback target.
    CTX=gke_noetl-demo-19700101_us-central1_noetl-cluster
    kubectl --context $CTX -n noetl get deploy noetl-worker-system-pool \
      -o jsonpath='{.spec.template.spec.containers[0].image}'   # record this digest
  2. Roll the system pool to v5.66.0 (flags still absent). Use the live deploy method confirmed in A4 — either update the digest in worker-system-pool-deployment-prod.yaml and kubectl apply -f ..., or kubectl set image. This is a rolling update of the system-pool pod(s) only — low blast radius (system playbooks + materializer + off-server drive).
    # apply-manifest path (after editing the digest in the manifest):
    kubectl --context $CTX -n noetl apply -f ci/manifests/noetl/worker-system-pool-deployment-prod.yaml
    kubectl --context $CTX -n noetl rollout status deploy/noetl-worker-system-pool
  3. (If A1 finds other event-authoring pools) roll them to v5.66.0 too, same flags-off no-op.

Verification (must all hold)

  • rollout status reports complete; 0 pod restarts after ready.
  • Image digest on the running pod == the v5.66.0 digest.
  • ehdb-selfcheck config shows all-external (run inside a pod, or as a one-shot Job on the same image): bash kubectl --context $CTX -n noetl exec deploy/noetl-worker-system-pool -- ehdb-selfcheck config # expect: coherent:true, every tier backend:"external", integration_mode:"disabled", exit 0
  • noetl_ehdb_eventlog_* metrics are absent/empty (no EHDB engine opened) — confirm in Cloud Monitoring.
  • Event path unchanged — run a smoke playbook (e.g. test/simple_python); it COMPLETES; event rows land in noetl.event exactly as before; the CQRS materializer backlog stays ≈ 0.

Success criteria

v5.66.0 running, config verb reports all-external + coherent, event throughput and materializer backlog identical to the v5.52.0 baseline, zero restarts over a ≥ 15-min soak.

Abort criteria → roll back to v5.52.0

Any pod crashloop, /metrics regression, materializer backlog climb, or smoke failure. Rollback: re-apply the captured v5.52.0 digest (kubectl apply or kubectl set image) and rollout status. Stage A introduces no new state, so rollback is a pure image revert.


3. Stage B — SHADOW in prod (dual-write + parity, EHDB never serves)

Goal: exercise the EHDB event-log engine against real prod event traffic without letting it serve. In shadow, every already-authored event is mirrored into the EHDB engine and parity-checked (sequence / count / order) against the incumbent; reads and the authoritative path are untouched.

Pre-flip checklist

  • Stage A success criteria met and stable.
  • A pod-local writable path for NOETL_EHDB_LOCAL_REFERENCE_LOG exists on the system pool (an emptyDir mount is sufficient for shadow — the log is derived and disposable). Confirm the mount before setting the flag.
  • GMP scrape confirmed for noetl_ehdb_eventlog_*.

The flip (data-plane pool only; control plane untouched)

kubectl --context $CTX -n noetl set env deploy/noetl-worker-system-pool \
  NOETL_EHDB_ENABLED=true \
  NOETL_EHDB_MODE=local_reference \
  NOETL_EHDB_CLIENT_ROLE=system \
  NOETL_EHDB_LOCAL_REFERENCE_LOG=/var/lib/ehdb/eventlog.jsonl \
  NOETL_EHDB_EVENTLOG=shadow
kubectl --context $CTX -n noetl rollout status deploy/noetl-worker-system-pool

Do not set any EHDB env on noetl-server-rust / gateway — they are control-plane and would fail the coherence guard (exit 4) / never serve.

Immediate verification

  • ehdb-selfcheck config now shows eventlog: mode:"shadow", backend:"external" (shadow still serves from the incumbent), coherent:true, exit 0. Every other tier still external/off.
  • Engine opened, dual-writing: noetl_ehdb_eventlog_* series appear.

Metrics to watch (Cloud Monitoring / GMP) — the PASS bar

Signal (noetl_ehdb_eventlog_*) Expectation
parity mismatch / divergence counter stays 0 — any non-zero is a hard fail
count parity (EHDB appends vs incumbent appends) tracks 1:1
order / sequence parity monotonic, gapless, matches incumbent
shadow lag (events mirrored behind incumbent) bounded, drains to ≈ 0
append latency (EHDB mirror leg) bounded; no climbing tail; no worker CPU/mem regression
worker pod restarts 0
CQRS materializer backlog (unchanged incumbent path) still ≈ 0

Soak & PASS bar

Soak ≥ 24 h across a representative traffic cycle (must include peak). PASS = zero parity divergence over the window, shadow lag bounded and draining, append-latency and worker resource usage flat, zero restarts. Do not proceed to Stage C until the shadow window is clean and the durability gate (§C) is resolved.

Abort + rollback (instant, no redeploy)

kubectl --context $CTX -n noetl set env deploy/noetl-worker-system-pool NOETL_EHDB_EVENTLOG=off
# or drop the whole integration:
kubectl --context $CTX -n noetl set env deploy/noetl-worker-system-pool NOETL_EHDB_ENABLED-

Shadow never served, so there is nothing to un-serve — the incumbent was authoritative throughout. The pod-local EHDB log is disposable.


4. Stage C — PRIMARY flip (EHDB serves the event log authoritatively)

GATED — do not run until the durability gate (§C in Blast radius) is resolved and the user gives an explicit, separate go for the primary flip. With the shipped local_reference (pod-local JSONL) backend, primary-serve is not production-durable. This section documents the flip fully so it is ready the moment a durable/shared backend + single-writer topology are in place.

What changes: EHDB serves the event log — append, global scan, per-execution scoped read, durable tail/ack, replay — authoritatively, with the JetStream+Postgres incumbent dual-run parity-checked on every append. Event authorship is unchanged: the gateway/server remain the gatekeeper of what enters the log; only the engine underneath the append path changes. The compile-time PRIMARY_SERVE_ACTIVATED is already true in v5.66.0, so primary activates purely from the runtime flag.

Pre-flip checklist (in addition to the durability gate)

  • Stage B soaked ≥ 24 h with zero divergence.
  • Durable, shared EHDB log substrate provisioned (PVC / object-backed), or single-writer topology proven (exactly one pod ever appends), with the log on durable storage that survives pod restart.
  • Rollback command staged in a second terminal, tested against staging.
  • On-call actively watching; change window open.
  • A canary execution identified to drive end-to-end immediately after flip.

The flip

kubectl --context $CTX -n noetl set env deploy/noetl-worker-system-pool \
  NOETL_EHDB_EVENTLOG=primary
kubectl --context $CTX -n noetl rollout status deploy/noetl-worker-system-pool

Immediate verification

  • ehdb-selfcheck config shows eventlog: mode:"primary", backend:"ehdb", coherent:true, exit 0.
  • Serving proof: noetl_ehdb_eventlog_* show outcome served_primary (not primary_divergence, not primary_unavailable).
  • Ordering / sequence / scope intact — global sequence monotonic + gapless, per-execution scope resolves, durable cursor advances.
  • Dual-run parity holds — the incumbent append leg still parity-checks each append (divergence counter stays 0).
  • Canary execution drives end-to-end and reaches its terminal state; per-execution chain roots=1 / terminals=1 / dangling=0.
  • CQRS materializer / off-server drive stay healthy; 0 restarts.

Success criteria

served_primary steady, zero divergence, canary + live traffic complete normally, replay-is-truth confirmed on a spot execution, resource usage bounded.

Observation window

Hold under active watch for a full traffic cycle before declaring the tier cut over. Keep the incumbent dual-run on (it is the rollback target) — do not retire JetStream/Postgres for the event log until a later, separately-agreed step.


5. Rollback (from any stage)

Two independent levers, zero data loss — the primary path only ever appends to the EHDB KeepAll log and never mutates or deletes anything the incumbent owns, so the incumbent's store is exactly as it was.

Lever 1 — runtime flag (operational, instant, no redeploy)

# from primary → shadow (keep verifying) or → off (fully external):
kubectl --context $CTX -n noetl set env deploy/noetl-worker-system-pool \
  NOETL_EHDB_EVENTLOG=shadow
kubectl --context $CTX -n noetl rollout status deploy/noetl-worker-system-pool

The JetStream+Postgres incumbent is authoritative again the moment the new pod is ready. The EHDB log stays whole on disk for a later re-enable.

Lever 2 — compile-time kill switch (structural fallback)

Set PRIMARY_SERVE_ACTIVATED = false in worker src/ehdb/eventlog.rs, rebuild, redeploy. primary then degrades to primary_unavailable regardless of config — serving is structurally unreachable. Use this if a runtime flag flip is somehow insufficient (e.g. a config-management race re-asserting primary).

Lever 0 — image rollback (Stage A regressions)

Re-apply the captured v5.52.0 digest. Removes the primary-serve-capable binary entirely.

Confirm the incumbent is serving again

  • ehdb-selfcheck config shows eventlog: backend:"external".
  • noetl_ehdb_eventlog_* outcome is no longer served_primary.
  • A smoke execution COMPLETES and its rows are present in noetl.event via the incumbent path; materializer backlog ≈ 0.

6. Blast radius + safety

This touches only the worker/system data-plane, not the server or gateway. Event authorship, session auth, SSE/callback routing, and every control-plane surface are untouched at every stage. Only the engine underneath the append path is in scope, and only on the data-plane pool that carries the flag.

This is one tier of five. Projection, KV, object, and vector stay on their external incumbents (Postgres materializer, NATS KV, GCS/S3, Qdrant) until each is separately cut over. ehdb-selfcheck config must continue to show those four as external throughout.

Stage What could go wrong Mitigation
A v5.66.0 crashloop / metrics regression Flags-off no-op; image rollback to v5.52.0
A Committed manifest lags live prod (A4) → wrong replica count / stale env re-applied Confirm live shape first; prefer set image/set env over blind apply; reconcile the manifest after
B EHDB mirror leg adds append latency / CPU / mem to the sole-writer pool Watch append-latency + pod resources; flag → off on any regression
B Parity divergence (EHDB engine disagrees with incumbent) Divergence counter is the PASS gate; any non-zero → off + diagnose; incumbent never stopped serving
B Wrong pool / missing pool carries the flag (A1) Confirm the complete set of event-authoring pools before Stage B
C Pod-local local_reference log is non-durable / diverges across replicas Durability gate (below) — hard blocker for Stage C
C Primary serves stale/partial after pod restart Requires durable shared store + single writer before flip
C Serving divergence under load (primary_divergence) Dual-run parity on every append; Lever 1 instant revert

§C — the durability gate (hard blocker for Stage C)

✅ RESOLVED (2026-07-12). The durable_segment backend (ehdb#254 slices 1–5) + segment GC & limits-based retention (#269/#270/#271 + worker#177/#178) resolve this gate: CRC-framed fsync'd segments (restart-durable), execution-affinity single-writer + shared cold-load tier (no multi-replica divergence), and bounded disk under retention. Alesha signed off "GO to prod durable-SHADOW" on 2026-07-12 — see Sign-off — Durable Event-Log Prod Durability. Stage C (durable-primary) is still gated on a clean ≥24 h prod durable- shadow window + the retention policy set on the writer pool + R2/R3 storage-class/topology confirmation (all in the sign-off §5/§6). The durable- shadow deploy is the prod team's next operational step; no agent flips it.

The original gate analysis below is retained for context.

The worker implements exactly one EHDB backend today: local_reference, a pod-local JSONL file. That is correct and safe for shadow (derived, disposable) but not for primary, where the file is the authoritative store:

  1. Multi-replica divergence. If more than one pod on the flagged pool appends events, each writes its own pod-local log → multiple divergent "authoritative" logs. Prod's system pool has run at 1 and 2 replicas across recent history (A4) — this must be pinned to a single writer, or the backend must be shared.
  2. Restart durability. A pod-local emptyDir log is lost on pod restart/reschedule. The authoritative event log cannot live on ephemeral pod storage.

Resolution options (pick before Stage C, with the user):

  • Provision a PVC-backed volume for NOETL_EHDB_LOCAL_REFERENCE_LOG on a single-writer StatefulSet-style pod, or
  • Wait for a durable/shared EHDB log backend (object-backed or replicated) beyond local_reference, or
  • Keep the event-log tier at shadow in prod indefinitely — capture the parity evidence and the scaling relief modeling from real traffic without ever making EHDB authoritative, and revisit primary when a durable backend ships.

Stages A and B stand on their own value (a proven-in-prod shadow with zero divergence is a strong result) and do not depend on resolving this gate.

Update (2026-07-07) — the durable backend that resolves this gate has shipped + kind-soaked. The durable_segment backend (ehdb#254 slices 1–5: CRC-framed segments, fsync-per-append, crash-recovery replay, execution-affinity single-writer, shared cold-load tier) is code-complete (worker v5.70.1) and proved segment rotation + real-pod-restart crash recovery on a PVC in kind. The durability sign-off package — go/no-go checklist, evidence bundle, residual-risk register, and the extended durable-shadow → durable-primary rollout sequence (with the v5.70.1 target image + the NOETL_EHDB_EVENTLOG_BACKEND=durable_segment + PVC env deltas) — is at Sign-off — Durable Event-Log Backend, Prod Durability. Read it before planning Stage C: it supersedes the "pick an option" framing above with a concrete durable path that is still user-gated per stage.


7. Decision log / sign-off

Each stage requires an explicit user go. The operator records the outcome here (append-only) after execution.

Stage User go (date / who) Executed (date / operator) Outcome / evidence Notes
Assumptions A1–A4 confirmed ☐ — —
Preconditions (§1) met ☐ — —
A — deploy v5.66.0 flags-off ☐ — — rollback digest: ______
B — shadow ☐ — — soak window: ______
Durability gate (§C) resolved ☑ Alesha 2026-07-12 (sign-off, not a prod deploy) durable_segment + segment GC/retention — Sign-off approved chosen option: durable_segment backend; scope = durable-SHADOW only, Stage C still gated
Durable-shadow deploy (Stage A′/B′) ☐ (approved; prod-team op) — — next operational step per Sign-off §5 — not an agent action
C — primary ☐ — — re-gated: ≥24 h shadow window + retention policy set + R2/R3

Related

Clone this wiki locally