Skip to content

Runbook Durable EventLog Prod Signoff

Kadyapam edited this page Jul 15, 2026 · 4 revisions

Sign-off — Durable Event-Log Backend, Prod Durability (§C gate, slice 6)

Status: ✅ SIGNED OFF — Alesha APPROVED "GO to prod durable-SHADOW" on 2026-07-12. The durability gate (§C) is cleared for the shadow stage. The approval is for prod durable-SHADOW only — NOT durable-primary, which stays separately gated (see §6). NOTHING HERE HAS BEEN EXECUTED IN PROD — no GKE/gcloud action, no image deployed to prod, no flag flipped; enabling prod durable-shadow is the explicit next operational step for the prod team (§5 Stage A′/B′), not an agent action.

This page assembles the durability evidence for the durable_segment event-log backend behind the durability gate on the tier-1 prod cutover (Runbook §C).

Refreshed 2026-07-11: the former Stage-C blocker D11 (segment GC / R1) flipped to MET — segment GC + limits-based retention shipped + organic-soak- proven in kind; the #261 perf head-to-head was re-run on the fixed engine (deployed durable append ~6 ms, was ~740 ms). No GAP remains on the checklist. See the GO/NO-GO summary (§6).

Tracks: noetl/ehdb#254 slice 6 (durable-backend program) · noetl/ehdb#241 (completion program). Extends (does not duplicate) Runbook — Prod Cutover: Event-Log Tier and Design: Durable Event-Log Backend.


0. What we are asking Alesha to approve

The tier-1 event-log prod cutover has one hard blocker: the only serving backend shipped when the runbook was drafted was local_reference — a pod-local JSONL file that (a) is lost on pod restart and (b) diverges across replicas. The durable-backend program (ehdb#254 slices 1–5) built and kind-proved a replacement: durable_segment — CRC-framed segment files, fsync-per-append, crash-recovery replay, execution-affinity single-writer, and a shared cold-load tier.

The ask is a durability sign-off, not an execution order. Concretely, we ask Alesha to decide go / no-go on:

Accept that the durable_segment backend clears the §C durability gate in principle — durable-on-disk + crash-recoverable + single-writer are proven in unit tests and a real-pod-restart kind soak — and authorise the next gated action: a prod durable-shadow soak (durable engine dual-writing to a real prod PVC, never serving), which is what closes the handful of items only prod can verify. Stage C (durable-primary) stays separately gated on that prod-shadow soak passing plus the segment-GC decision (§4 risk R1).

A "go" does not flip any prod flag. It authorises scheduling the prod durable-shadow soak per the extended sequence in §5. Stage C requires its own, later, explicit go.


1. Go/no-go checklist — durability criteria

Each criterion the §C gate implies, with current status and the slice evidence behind it. MET = proven in code + kind. NEEDS-PROD-VERIFICATION = the mechanism is proven, but the value can only be confirmed against real GKE storage / topology / volume, which is exactly what the prod durable-shadow soak does. GAP = a real code shortfall (see risk register).

# Durability criterion Status Evidence
D1 Durable on-disk format — append-only CRC32-framed segments, magic-word framing, fsync before append returns MET Slice 1 unit tests (ehdb#253); design §on-disk format
D2 Crash recovery / replay-is-truth — torn-tail discard + truncate, complete-bad-CRC = hard error, zero-loss for every append that returned MET Slice 1 durable-eventlog-recovery verb (zero_loss, ordering_ok, scope_ok, cursor_survived); slice 5 real-pod-restart replay of 11 856 records
D3 Restart durability on real persistent storage — segments survive pod delete + reschedule MET in kind / NEEDS-PROD-VERIFICATION on GKE storage class Slice 5: sealed segment byte-identical (sha256 c2ba3d33…, 8 387 133 B) across a --force pod delete on a 2 Gi RWO PVC. Prod PVC class/CSI behaviour unverified (risk R2)
D4 Single-writer-per-shard — exactly one replica appends a shard; hash byte-identical to worker/server so the event-log shard == the drive shard MET in code + kind / NEEDS-PROD-VERIFICATION under real multi-replica affinity Slice 2 XxHash64 shard_for_i64 (ehdb#255, owner-writes/non-owner-refused invariant); slice 5 single-owner soak emitted zero routed_away. Prod topology (system pool has run at 1 and 2 replicas) unverified (risk R3)
D5 Shared / cold-load tier — a new owner hydrates a shard from a shared medium, surviving loss of the writer's pod-local disk MET in code + kind / NEEDS-PROD-VERIFICATION with a real shared medium Slice 3 SharedTierEventLog (ehdb#258); slice 5 shared tier published each segment byte-identical to the shared dir. Real RWX-PVC / object-tier medium on GKE unverified (risk R2)
D6 Segment rotation at the size cap under sustained load MET Slice 5: seg-0001 sealed at 8 387 133 B (< 8 MiB DEFAULT_SEGMENT_MAX_BYTES) → seg-0002 opened, in-cluster on real event bytes
D7 Clean parity — only mirrored, zero invalid/degraded/divergence/routed_away MET in kind / NEEDS-PROD-VERIFICATION over a ≥24 h prod-shadow window Slice 5 metrics past 11 731 with last_ok=1, last_degraded=0. Prod parity over a real traffic cycle is the shadow-soak PASS bar
D8 Fail-safe default + reversibility — unset ⇒ byte-identical JSONL; flag revert restores incumbent with no redeploy MET Slice 4 EventLogStorageBackend::from_raw (only the exact durable_segment token opts in); slice 4 selfcheck default-backend = JSONL, no durable dir
D9 Observability — ehdb-selfcheck config surfaces the backend; parity metrics scrapeable MET (code) / NEEDS-PROD-VERIFICATION (GMP scrape) Slice 4 config matrix shows eventlog_storage_backend; metrics exist. GMP PodMonitoring scrape of noetl_ehdb_eventlog_* in prod ns unverified
D10 Backend equivalence at the current worker pointer (v5.70.1) MET ehdb-reference cca0d0d → 52120a7 (v5.70.0 → v5.70.1) touches only kv.rs/vector.rs — no durable_eventlog* change; the v5.70.0 soak validates the v5.70.1 backend
D11 Segment retention / GC — bounded disk under primary at prod volume MET (2026-07-10/11) Segment GC shipped + soak-proven. Interest-based reclamation (local ehdb#269 + shared-tier watermark ehdb#270), limits-based keep-last-N retention (ehdb#271) that bounds a store with no consumer, and the periodic worker invocation (worker#177/#178) — all default-off. Base-offset + durable fsync'd reclaim.json commit point + write-forward preserve replay/cold-load/cursor. Organic in-cluster soak (merged-main image ehdb178184-merged, shadow, no consumer): the live store grew to 14 segments under load then the periodic GC self-bounded it to 2 and held, reclaimed climbing, no error, 0 new restarts, reclaim manifest committed + replay gapless-from-base. See Design → Segment GC + Limits-based retention. Residual = operational (choose the retention window / PVC sizing), see risk R1.

Verdict: D1, D2, D6, D8, D10, and now D11 are MET outright (D11 via the merged segment-GC + retention slices, drive- and organic-soak-proven in kind). D3, D4, D5, D7, D9 are mechanism-MET, value-NEEDS-PROD-VERIFICATION — none can be closed without a real prod PVC + real prod traffic, which is precisely the prod durable-shadow soak this sign-off authorises. The former Stage-C blocker (D11 / R1) is resolved — retention exists, is default-off, and lets a primary store self-bound; what remains for R1 is an operational knob (pick the retention window + PVC size), not a code gap.

Recommendation: GO to prod durable-SHADOW (unchanged — the durable / crash-recovery core is proven; the residual D3/D4/D5/D7/D9 items are prod-only and are what the shadow soak exists to close). Stage C durable-PRIMARY is now gated only on (i) a clean ≥24 h prod-shadow window and (ii) setting the retention policy for prod (NOETL_EHDB_EVENTLOG_GC=consumer_ack + NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS / min-retained sized to the chosen window) — the segment-GC decision the old verdict blocked on is no longer a code gap. Perf re-run (below) confirms the deployed durable append is ~6 ms (was ~740 ms), so R6 (perf) is materially de-risked too. The GO/NO-GO decision stays with the sign-off owner (§6).


2. Evidence bundle

Consolidated concrete proofs already gathered, with links.

Code (merged)

Slice What PR / commit
1 Segment store + CRC framing + offset index + crash-recovery replay; DurableEventLogDriver; EventLogStorageBackend selector ehdb#253 (99c4570)
2 Execution-affinity single-writer routing (affinity::shard_for_i64, AffinityRoutedEventLog); owner-writes / non-owner-refused ehdb#255 (6fbe88f)
3 Shared / object-store segment tier (SharedSegmentBackend, SharedTierEventLog); fixed-width digest-checked keys ehdb#258 (cca0d0d)
4 Worker wiring — NOETL_EHDB_EVENTLOG_BACKEND=durable_segment in worker-rust src/ehdb/eventlog_backend.rs; ehdb-selfcheck durable-eventlog verb noetl/worker#171 (9947d9b → v5.70.0)

Current worker pointer: v5.70.1 (0fb7ea5), ehdb-reference pinned 52120a7 — durable backend byte-identical to the soaked v5.70.0 (D10).

Test + selfcheck evidence

  • Slice 1: crash-recovery drive {"recovered":true, zero_loss:true, ordering_ok:true, scope_ok:true, payloads_match:true, cursor_survived:true}.
  • Slice 2: {"single_writer_holds":true, owner_appends:4, nonowner_refusals:4, single_writer_invariant:true, coldload_read_ok:true, crash_recovery_ok:true, divergence:null}; 21 new tests. Distinct exit code 6 for a non-owner refusal.
  • Slice 3: {"shared_tier_holds":true, owner_published_ok:true, nonowner_coldload_from_shared_ok:true, crash_recovery_from_shared_ok:true, shared_miss_ok:true, parity_ok:true, divergence:null}.
  • Slice 4: 13 new tests + 191 ehdb lib tests green, clippy clean; selfcheck config → eventlog_storage_backend: durable_segment; durable-eventlog → 3 events in seg-0000000000000001.eslog, no JSONL, reopened replay = 3; default backend ⇒ JSONL, no durable dir.

Kind soak (slice 5, 2026-07-07, LOCAL kind only)

Worker v5.70.0 on pool noetl-worker-rust (role worker, single-owner shard 0), NOETL_EHDB_EVENTLOG_BACKEND=durable_segment, NOETL_EHDB_EVENTLOG_DURABLE_DIR=/ehdb-durable on a 2 Gi RWO PVC (ehdb-durable-soak).

Proof Value
Accumulation thousands of real drives (automation/pft_sql_probe_v2 + tests/large_tabular_result_test) into CRC-framed shard-0000/seg-*.eslog; shared tier published each byte-identical; no JSONL written
Rotation seg-0001 sealed 8 387 133 B → seg-0002 opened (in-cluster, real bytes)
Metrics noetl_ehdb_eventlog_ops_total{outcome="mirrored"} past 11 731; last_ok=1, last_degraded=0; 0 invalid/degraded/routed_away; 0 restarts/crashloops
Crash recovery kubectl delete pod --force → fresh pod, same PVC: sealed segment byte-identical (sha256 c2ba3d33…); active tail replayed + continued; gapless sequence (…11852,11853,11855); durable-eventlog selfcheck replayed durable_replay_records: 11856 read-only on a pod that never saw the writes

Full write-up: Design: Durable Event-Log Backend — Kind soak (slice 5). ehdb#254 slice-5 evidence comment + checkbox checked.

Soak-target caveat (feeds risk R3): the soak ran on the user pool noetl-worker-rust, not the runbook's prod event-writer target noetl-worker-system-pool. The mechanism is pool-agnostic, but the prod durable-shadow soak must run on the actual event-authoring pool.


3. Alignment deltas vs the existing tier-1 cutover runbook

The tier-1 runbook was drafted against the local_reference backend. Enabling the durable backend changes three things it must pick up (recorded here so the two pages don't drift):

  1. Target image. The runbook's Stage A target is v5.66.0, which predates the durable wiring (slice 4 = worker#171 = v5.70.0). For the durable path the Stage A target must be v5.70.1 (current pointer; carries the durable backend + the KV/object/vector subject-digest fixes). Rollback target is unchanged: v5.52.0.
  2. Env contract. The durable path adds an axis the runbook's env block omits — set alongside the existing NOETL_EHDB_EVENTLOG mode flag:
    Env var Value for durable prod Meaning
    NOETL_EHDB_EVENTLOG_BACKEND durable_segment select the durable engine (fail-safe: unset ⇒ local_reference)
    NOETL_EHDB_EVENTLOG_DURABLE_DIR a PVC mount path per-shard local segment stores + derived shared/coldload roots
    NOETL_EHDB_EVENTLOG_SHARED_DIR <durable-dir>/shared (or object-tier root) shared cold-load medium
    Under durable, NOETL_EHDB_LOCAL_REFERENCE_LOG is no longer the
    authoritative store — the PVC-backed segments are.
  3. Storage substrate. The runbook's Stage B used an emptyDir (fine for disposable local_reference shadow). Durable shadow/primary needs a real PVC (or object-tier medium) so segments survive restart — the §C resolution, made concrete.

These are additive; the runbook's stage structure, rollback levers, blast-radius analysis, and A1–A4 assumptions all still apply.


4. Residual-risk register — what kind cannot prove, only prod can

# Risk (what kind can't prove) Severity Proposed prod-verification / mitigation
R1 Segment retention / GC. No compaction/GC ships; segments accumulate until the PVC fills. A code GAP (D11). RESOLVED (2026-07-10/11) — segment GC + limits-based retention shipped + soak-proven. High — Stage-C blocker → Low (operational) Code gap closed: interest-based GC (ehdb#269/#270) + limits-based keep-last-N retention (ehdb#271) + periodic worker invocation (worker#177/#178), all default-off. The organic in-cluster soak proved a shadow store (no consumer) self-bounds: 14 → 2 segments, held flat, reclaimed climbing, replay gapless-from-base, 0 restarts. Remaining is operational, not a blocker: for Stage C set NOETL_EHDB_EVENTLOG_GC=consumer_ack + NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS (and/or MIN_RETAINED_SEGMENTS) sized to the chosen retention window on the event-writer pool, + a PVC free-bytes alert as defence-in-depth. Under primary the projection/read-model tier's ack also drives interest-based reclamation. Open follow-ups (not blockers): (a) shared-medium topology — shared-object reclamation is coherent (watermark-first), but a prod shared/object-tier medium is unverified (couples with R2); (b) max-age retention deliberately deferred (mtime resets on hydrate/cold-load — keep-last-N is the robust bound); (c) sealed-segment tier-down to the object store is a future space optimisation, not required for boundedness.
R2 Real GKE storage class / CSI behaviour. Kind used standard/local-path RWO. Prod CSI (PD-SSD etc.) fsync durability semantics, RWX support (for the shared tier), volume-detach/reattach timing on node failure, and expansion are unverified. High In the prod durable-shadow soak: pick the storage class with the user; verify fsync durability (a segment written just before a node drain survives); if the shared tier uses RWX, verify the class supports it, else confine to single-writer RWO + object-tier shared medium. Kill a node during shadow and confirm segment integrity.
R3 Multi-replica shard ownership under real affinity. The soak ran single-owner (1 replica) on the user pool. Prod's event writer (noetl-worker-system-pool) has run at 1 and 2 replicas (runbook A4). A routed_away under multi-replica is the single-writer-violation signal, never exercised in prod. High Confirm the prod event-writer pool + its replica count / shard topology (runbook A1/A4). Run the durable-shadow soak on that pool with NOETL_SHARD_INDEX/NOETL_SHARD_COUNT matching the live shape; assert routed_away == 0 (or, if 2 replicas, that each shard has exactly one owner and non-owners route, not double-write).
R4 Disk-pressure eviction. A PVC filling (see R1) or node disk pressure can evict the writer pod mid-append. Kind never hit disk pressure. Medium Alert on PVC free-bytes and node DiskPressure; size headroom; confirm the pod's emptyDir/PVC quotas. Crash-recovery (D2) covers the append-interrupted case; the concern is availability, not loss.
R5 Backup / restore + DR. No backup/restore path for the durable segments is defined. The incumbent (JetStream + Postgres noetl.event) is the current DR target; once EHDB is primary, segment loss = event loss unless the incumbent dual-run is retained. Medium Keep the incumbent dual-run on through and beyond Stage C (the runbook already mandates this — do not retire JetStream/Postgres). Define a segment backup (PVC snapshot or object-tier copy) before considering incumbent retirement (a later, separate step).
R6 Performance under prod event volume. The durable fsync-per-append leg adds latency the kind soak's traffic didn't stress at peak. Materially de-risked (2026-07-11): the #264/#266/#267 O(segment) fixes are merged + deployed. Medium → Low-Medium #261 head-to-head re-run: deployed durable append ~6 ms (was ~740 ms), Layer-A authoritative ~4–16 ms flat across store size — clears an incumbent-parity p99, no size-degradation. Still watch the durable-leg append latency + worker CPU/mem vs the v5.52.0 baseline through the ≥24 h shadow window incl. peak; PASS = flat, no climbing tail; flag → off on regression. (Layer-B in-cluster throughput is podman-VM-contention-noisy — directional only; Layer A + the deployed re-measurement are the trusted numbers.)
R7 selfcheck harness artifact. ehdb-selfcheck durable-eventlog reports ok:false/parity_mismatch when pointed at a live populated shard (it asserts a fresh isolated 3-event sequence). An operator could misread this as a durability failure. Low (operational) Runbook note: on a live shard the meaningful signal is durable_replay_records (full read-only replay count) + absence of a hard CRC error, not the verb's ok field. Captured here + in the design note.

5. Recommended prod rollout sequence (durable backend)

Extends Runbook — Prod Cutover: Event-Log Tier. Only the deltas for the durable backend are given; the runbook's preconditions (§1), blast-radius (§6), and sign-off log (§7) apply unchanged. Every step is user-gated; a "go" on this sign-off authorises up to and including the durable-shadow soak, not Stage C.

Approved (2026-07-12): the prod team may execute Stage A′ → B′. Every command below is run by the prod team against the prod cluster — no agent runs kubectl/gcloud against prod. Stage C′ stays separately gated.

Stage A′ — deploy the current durable-capable image, all EHDB flags OFF (behavior-neutral)

Target the current durable-capable release — worker merged main (the #178 merge, 521fd10; kind-validated as ehdb178184-merged), which carries durable_segment plus segment GC + limits-based retention (all default-off). Use the prod release build of that main, not the stale v5.70.1 (which predates GC/retention). No NOETL_EHDB_* env ⇒ byte-identical to v5.52.0 on the event path (D8). Verify ehdb-selfcheck config = all-external, exit 0; 0 restarts.

  • Why this image: carrying the GC/retention code now means Stage C needs no further rebuild — only a flag flip — when it is later gated in.
  • Rollback: re-apply the captured v5.52.0 digest (pure image revert).

Stage B′ — prod durable-SHADOW (dual-write to PVC, never serves)

The prod analog of the kind soak; this is what closes D3/D4/D5/D7/D9 with real evidence and exercises R2/R3/R4/R6.

Declarative alternative to the imperative steps below. The same PVC + env change is staged as a reviewable, gated GitOps manifest in noetl/ops#239 (draft; do not merge until the go). Applying that PR is equivalent to steps 1–3; the kubectl set env recipe here is the manual equivalent for a non-GitOps flip. Either way, the image prerequisite (Stage A′) and the open decisions (R1/R2/R3 + who/when) must be settled first.

  1. Provision the prod PVC (storage class + access mode decided with the user per R2) and mount it on the event-writer pool at the durable dir.
  2. Enable the shadow backend on the event-writer pool confirmed in runbook A1 (system pool, or whatever authors events):
    kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> \
      NOETL_EHDB_ENABLED=true \
      NOETL_EHDB_CLIENT_ROLE=system \
      NOETL_EHDB_EVENTLOG_BACKEND=durable_segment \
      NOETL_EHDB_EVENTLOG_DURABLE_DIR=/ehdb-durable \
      NOETL_EHDB_EVENTLOG=shadow
    kubectl --context $CTX -n noetl rollout status deploy/<event-writer-pool>
  3. (Recommended) bound the shadow PVC with retention. A shadow store has no durable consumer, so interest-based GC alone reclaims nothing — enable limits-based retention so the shadow segments do not accumulate for 24 h+:
    kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> \
      NOETL_EHDB_EVENTLOG_GC=consumer_ack \
      NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS=<N> \
      NOETL_EHDB_EVENTLOG_GC_INTERVAL_SECS=<secs>
    Size <N> × 8 MiB (the default segment cap) to a comfortable fraction of the PVC; keep the default 8 MiB SEGMENT_MAX_BYTES in prod (the kind soak used a tiny 1 KiB cap only to force fast rotation). This is the same retention that was organically soak-proven (14 → 2 segments, held flat, replay intact). All default-off unless set, so it is opt-in.
  4. Do not set any EHDB env on noetl-server-rust / gateway (control-plane; fails the coherence guard, exit 4).
  • Observation window: ≥ 24 h across a full traffic cycle incl. peak (same bar as the runbook shadow stage).
  • Rollback (instant, no redeploy; disposable — shadow never served):
    kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> NOETL_EHDB_EVENTLOG=off
    # or drop the whole integration:
    kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> NOETL_EHDB_ENABLED-
    # (fail-safe: unsetting NOETL_EHDB_EVENTLOG_BACKEND alone ⇒ local_reference, byte-identical)

Stage C′ — prod durable-PRIMARY (EHDB serves authoritatively) — SEPARATELY GATED

NOT authorised by the 2026-07-12 durable-SHADOW sign-off. Only after a clean Stage B′ window and the retention policy set on the event-writer pool (the R1 segment-GC decision is now MET — retention shipped; this is the operational config step, no longer a code gap) and a separate explicit user go.

kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> NOETL_EHDB_EVENTLOG=primary
kubectl --context $CTX -n noetl rollout status deploy/<event-writer-pool>

Incumbent (JetStream + Postgres) stays dual-run parity-checked (rollback target; R5). Serving proof = outcome=served_primary, divergence counter 0, canary execution reaches terminal state.

  • Rollback (instant, zero loss — primary only appends to KeepAll): NOETL_EHDB_EVENTLOG=shadow (or off). Structural kill switch: PRIMARY_SERVE_ACTIVATED=false, rebuild, redeploy.

Metrics / alerts to watch (GMP / Cloud Monitoring)

Signal Bar
noetl_ehdb_eventlog_ops_total{outcome} — mirrored (shadow) / served_primary (primary) steady; no invalid/degraded/primary_divergence/primary_unavailable
parity divergence counter stays 0 (hard fail if non-zero)
routed_away 0 on single-writer; non-zero ⇒ single-writer violated (R3)
PVC free bytes (kubelet volume stats) alert before full — no GC ships (R1)
segment count / total bytes on the durable dir monotonic growth expected; feeds the retention decision (R1)
rotation events segments seal at ~8 MiB and roll (D6)
append latency (durable fsync leg) + worker CPU/mem bounded, flat vs v5.52.0 baseline (R6)
worker pod restarts 0
CQRS materializer backlog (incumbent path) ≈ 0 (unchanged)

6. GO / NO-GO summary for the sign-off owner (@alesha)

The decision is yours; this is the decision-ready picture as of 2026-07-11.

What changed since the first draft (2026-07-07):

  • D11 (segment GC / the R1 Stage-C blocker): GAP → MET. Segment GC + a limits-based keep-last-N retention mode shipped (merged ehdb#269/#270/#271
    • worker#177/#178, all default-off) and were organically soak-proven in kind: a shadow store with no consumer self-bounded 14 → 2 segments and held flat, reclaimed climbing, replay gapless-from-base, 0 restarts. R1 drops from High — Stage-C blocker to Low (operational): pick the retention window
    • PVC size.
  • R6 (perf): de-risked. The #261 head-to-head re-run shows the deployed durable append at ~6 ms (was ~740 ms) and Layer-A authoritative ~4–16 ms flat across store size — incumbent-parity, no size-degradation.

Checklist state: D1, D2, D6, D8, D10, D11 = MET. D3, D4, D5, D7, D9 = mechanism-MET, value-NEEDS-PROD-VERIFICATION (only a real prod PVC + traffic closes them — that IS the durable-shadow soak). No GAP remains.

Recommendation — GO to prod durable-SHADOW. The durable / crash-recovery core is proven, the former Stage-C blocker is resolved, and perf is incumbent-parity. Nothing in the checklist argues against a shadow soak (dual-write to a real PVC, never serving) — it is behaviour-neutral to the authoritative path (D8) and is exactly what closes the remaining prod-only items. Stage C (durable-primary) stays separately gated on: (i) a clean ≥24 h prod-shadow window, (ii) setting the retention policy for the event-writer pool (an operational config, no longer a code gap), and (iii) the R2/R3 prod-storage-class + event-writer-topology confirmations.

✅ DECISION — APPROVED (Alesha, 2026-07-12)

Alesha approved: GO to prod durable-SHADOW. The §C durability sign-off is complete for the shadow stage; ehdb#254 slice-6 box is checked.

  • Scope of the approval: prod durable-SHADOW only — deploy the durable-capable image (flags-off, Stage A′) then enable durable_segment as a dual-writing shadow on a real PVC (Stage B′), never serving. This is behaviour-neutral to the authoritative event log (D8).
  • NOT approved by this decision: durable-primary (Stage C′). It remains separately gated on (i) a clean ≥24 h prod-shadow window, (ii) the retention policy set on the event-writer pool, and (iii) the R2/R3 storage-class + event-writer-topology confirmations — and needs its own explicit go.
  • Next operational step (for the prod team, NOT an agent action): execute Stage A′ → B′ per §5. No agent flips a prod flag or runs kubectl/gcloud against prod.

On "GO" (now in effect): schedule Stage A′ (deploy the durable-capable image flags-off) then Stage B′ (prod durable-shadow on a real PVC) per §5, each on its own explicit go, holding the ≥24 h shadow soak. Stage C′ (durable-primary) is NOT authorised by this sign-off.

The durability sign-off (slice 6) is DONE — ehdb#254 slice-6 box is checked. The remaining prod-enablement (Stage A′ → B′) is an operational step owned by the prod team; no agent flips a prod flag or runs kubectl/gcloud against prod. Stage C′ (durable-primary) re-gates separately (see the DECISION block above).


Related

Clone this wiki locally