Repository navigation
Runbook Durable EventLog Prod Signoff
Status: ✅ SIGNED OFF — Alesha APPROVED "GO to prod durable-SHADOW" on 2026-07-12. The durability gate (§C) is cleared for the shadow stage. The approval is for prod durable-SHADOW only — NOT durable-primary, which stays separately gated (see §6). NOTHING HERE HAS BEEN EXECUTED IN PROD — no GKE/gcloud action, no image deployed to prod, no flag flipped; enabling prod durable-shadow is the explicit next operational step for the prod team (§5 Stage A′/B′), not an agent action.
This page assembles the durability evidence for the durable_segment event-log
backend behind the durability gate on the tier-1 prod cutover
(Runbook §C).
Refreshed 2026-07-11: the former Stage-C blocker D11 (segment GC / R1) flipped to MET — segment GC + limits-based retention shipped + organic-soak- proven in kind; the #261 perf head-to-head was re-run on the fixed engine (deployed durable append ~6 ms, was ~740 ms). No GAP remains on the checklist. See the GO/NO-GO summary (§6).
Tracks: noetl/ehdb#254 slice 6 (durable-backend program) · noetl/ehdb#241 (completion program). Extends (does not duplicate) Runbook — Prod Cutover: Event-Log Tier and Design: Durable Event-Log Backend.
The tier-1 event-log prod cutover has one hard blocker: the only serving
backend shipped when the runbook was drafted was local_reference — a
pod-local JSONL file that (a) is lost on pod restart and (b) diverges
across replicas. The durable-backend program (ehdb#254 slices 1–5) built and
kind-proved a replacement: durable_segment — CRC-framed segment files,
fsync-per-append, crash-recovery replay, execution-affinity single-writer,
and a shared cold-load tier.
The ask is a durability sign-off, not an execution order. Concretely, we ask Alesha to decide go / no-go on:
Accept that the
durable_segmentbackend clears the §C durability gate in principle — durable-on-disk + crash-recoverable + single-writer are proven in unit tests and a real-pod-restart kind soak — and authorise the next gated action: a prod durable-shadow soak (durable engine dual-writing to a real prod PVC, never serving), which is what closes the handful of items only prod can verify. Stage C (durable-primary) stays separately gated on that prod-shadow soak passing plus the segment-GC decision (§4 risk R1).
A "go" does not flip any prod flag. It authorises scheduling the prod durable-shadow soak per the extended sequence in §5. Stage C requires its own, later, explicit go.
Each criterion the §C gate implies, with current status and the slice evidence behind it. MET = proven in code + kind. NEEDS-PROD-VERIFICATION = the mechanism is proven, but the value can only be confirmed against real GKE storage / topology / volume, which is exactly what the prod durable-shadow soak does. GAP = a real code shortfall (see risk register).
| # | Durability criterion | Status | Evidence |
|---|---|---|---|
| D1 | Durable on-disk format — append-only CRC32-framed segments, magic-word framing, fsync before append returns |
MET | Slice 1 unit tests (ehdb#253); design §on-disk format |
| D2 | Crash recovery / replay-is-truth — torn-tail discard + truncate, complete-bad-CRC = hard error, zero-loss for every append that returned | MET | Slice 1 durable-eventlog-recovery verb (zero_loss, ordering_ok, scope_ok, cursor_survived); slice 5 real-pod-restart replay of 11 856 records |
| D3 | Restart durability on real persistent storage — segments survive pod delete + reschedule | MET in kind / NEEDS-PROD-VERIFICATION on GKE storage class | Slice 5: sealed segment byte-identical (sha256 c2ba3d33…, 8 387 133 B) across a --force pod delete on a 2 Gi RWO PVC. Prod PVC class/CSI behaviour unverified (risk R2) |
| D4 | Single-writer-per-shard — exactly one replica appends a shard; hash byte-identical to worker/server so the event-log shard == the drive shard | MET in code + kind / NEEDS-PROD-VERIFICATION under real multi-replica affinity | Slice 2 XxHash64 shard_for_i64 (ehdb#255, owner-writes/non-owner-refused invariant); slice 5 single-owner soak emitted zero routed_away. Prod topology (system pool has run at 1 and 2 replicas) unverified (risk R3) |
| D5 | Shared / cold-load tier — a new owner hydrates a shard from a shared medium, surviving loss of the writer's pod-local disk | MET in code + kind / NEEDS-PROD-VERIFICATION with a real shared medium | Slice 3 SharedTierEventLog (ehdb#258); slice 5 shared tier published each segment byte-identical to the shared dir. Real RWX-PVC / object-tier medium on GKE unverified (risk R2) |
| D6 | Segment rotation at the size cap under sustained load | MET | Slice 5: seg-0001 sealed at 8 387 133 B (< 8 MiB DEFAULT_SEGMENT_MAX_BYTES) → seg-0002 opened, in-cluster on real event bytes |
| D7 | Clean parity — only mirrored, zero invalid/degraded/divergence/routed_away
|
MET in kind / NEEDS-PROD-VERIFICATION over a ≥24 h prod-shadow window | Slice 5 metrics past 11 731 with last_ok=1, last_degraded=0. Prod parity over a real traffic cycle is the shadow-soak PASS bar |
| D8 | Fail-safe default + reversibility — unset ⇒ byte-identical JSONL; flag revert restores incumbent with no redeploy | MET | Slice 4 EventLogStorageBackend::from_raw (only the exact durable_segment token opts in); slice 4 selfcheck default-backend = JSONL, no durable dir |
| D9 | Observability — ehdb-selfcheck config surfaces the backend; parity metrics scrapeable |
MET (code) / NEEDS-PROD-VERIFICATION (GMP scrape) | Slice 4 config matrix shows eventlog_storage_backend; metrics exist. GMP PodMonitoring scrape of noetl_ehdb_eventlog_* in prod ns unverified |
| D10 | Backend equivalence at the current worker pointer (v5.70.1) | MET |
ehdb-reference cca0d0d → 52120a7 (v5.70.0 → v5.70.1) touches only kv.rs/vector.rs — no durable_eventlog* change; the v5.70.0 soak validates the v5.70.1 backend |
| D11 | Segment retention / GC — bounded disk under primary at prod volume |
MET (2026-07-10/11) |
Segment GC shipped + soak-proven. Interest-based reclamation (local ehdb#269 + shared-tier watermark ehdb#270), limits-based keep-last-N retention (ehdb#271) that bounds a store with no consumer, and the periodic worker invocation (worker#177/#178) — all default-off. Base-offset + durable fsync'd reclaim.json commit point + write-forward preserve replay/cold-load/cursor. Organic in-cluster soak (merged-main image ehdb178184-merged, shadow, no consumer): the live store grew to 14 segments under load then the periodic GC self-bounded it to 2 and held, reclaimed climbing, no error, 0 new restarts, reclaim manifest committed + replay gapless-from-base. See Design → Segment GC + Limits-based retention. Residual = operational (choose the retention window / PVC sizing), see risk R1. |
Verdict: D1, D2, D6, D8, D10, and now D11 are MET outright (D11 via
the merged segment-GC + retention slices, drive- and organic-soak-proven in
kind). D3, D4, D5, D7, D9 are mechanism-MET, value-NEEDS-PROD-VERIFICATION —
none can be closed without a real prod PVC + real prod traffic, which is
precisely the prod durable-shadow soak this sign-off authorises. The former
Stage-C blocker (D11 / R1) is resolved — retention exists, is default-off, and
lets a primary store self-bound; what remains for R1 is an operational
knob (pick the retention window + PVC size), not a code gap.
Recommendation: GO to prod durable-SHADOW (unchanged — the durable /
crash-recovery core is proven; the residual D3/D4/D5/D7/D9 items are prod-only
and are what the shadow soak exists to close). Stage C durable-PRIMARY is now
gated only on (i) a clean ≥24 h prod-shadow window and (ii) setting the
retention policy for prod (NOETL_EHDB_EVENTLOG_GC=consumer_ack +
NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS / min-retained sized to the
chosen window) — the segment-GC decision the old verdict blocked on is no
longer a code gap. Perf re-run (below) confirms the deployed durable append is
~6 ms (was ~740 ms), so R6 (perf) is materially de-risked too. The GO/NO-GO
decision stays with the sign-off owner (§6).
Consolidated concrete proofs already gathered, with links.
| Slice | What | PR / commit |
|---|---|---|
| 1 | Segment store + CRC framing + offset index + crash-recovery replay; DurableEventLogDriver; EventLogStorageBackend selector |
ehdb#253 (99c4570) |
| 2 | Execution-affinity single-writer routing (affinity::shard_for_i64, AffinityRoutedEventLog); owner-writes / non-owner-refused |
ehdb#255 (6fbe88f) |
| 3 | Shared / object-store segment tier (SharedSegmentBackend, SharedTierEventLog); fixed-width digest-checked keys |
ehdb#258 (cca0d0d) |
| 4 | Worker wiring — NOETL_EHDB_EVENTLOG_BACKEND=durable_segment in worker-rust src/ehdb/eventlog_backend.rs; ehdb-selfcheck durable-eventlog verb |
noetl/worker#171 (9947d9b → v5.70.0) |
Current worker pointer: v5.70.1 (0fb7ea5), ehdb-reference pinned
52120a7 — durable backend byte-identical to the soaked v5.70.0 (D10).
- Slice 1: crash-recovery drive
{"recovered":true, zero_loss:true, ordering_ok:true, scope_ok:true, payloads_match:true, cursor_survived:true}. - Slice 2:
{"single_writer_holds":true, owner_appends:4, nonowner_refusals:4, single_writer_invariant:true, coldload_read_ok:true, crash_recovery_ok:true, divergence:null}; 21 new tests. Distinct exit code 6 for a non-owner refusal. - Slice 3:
{"shared_tier_holds":true, owner_published_ok:true, nonowner_coldload_from_shared_ok:true, crash_recovery_from_shared_ok:true, shared_miss_ok:true, parity_ok:true, divergence:null}. - Slice 4: 13 new tests + 191 ehdb lib tests green, clippy clean; selfcheck
config→eventlog_storage_backend: durable_segment;durable-eventlog→ 3 events inseg-0000000000000001.eslog, no JSONL, reopened replay = 3; default backend ⇒ JSONL, no durable dir.
Worker v5.70.0 on pool noetl-worker-rust (role worker, single-owner
shard 0), NOETL_EHDB_EVENTLOG_BACKEND=durable_segment,
NOETL_EHDB_EVENTLOG_DURABLE_DIR=/ehdb-durable on a 2 Gi RWO PVC
(ehdb-durable-soak).
| Proof | Value |
|---|---|
| Accumulation | thousands of real drives (automation/pft_sql_probe_v2 + tests/large_tabular_result_test) into CRC-framed shard-0000/seg-*.eslog; shared tier published each byte-identical; no JSONL written |
| Rotation |
seg-0001 sealed 8 387 133 B → seg-0002 opened (in-cluster, real bytes) |
| Metrics |
noetl_ehdb_eventlog_ops_total{outcome="mirrored"} past 11 731; last_ok=1, last_degraded=0; 0 invalid/degraded/routed_away; 0 restarts/crashloops |
| Crash recovery |
kubectl delete pod --force → fresh pod, same PVC: sealed segment byte-identical (sha256 c2ba3d33…); active tail replayed + continued; gapless sequence (…11852,11853,11855); durable-eventlog selfcheck replayed durable_replay_records: 11856 read-only on a pod that never saw the writes |
Full write-up: Design: Durable Event-Log Backend — Kind soak (slice 5). ehdb#254 slice-5 evidence comment + checkbox checked.
Soak-target caveat (feeds risk R3): the soak ran on the user pool
noetl-worker-rust, not the runbook's prod event-writer target
noetl-worker-system-pool. The mechanism is pool-agnostic, but the prod
durable-shadow soak must run on the actual event-authoring pool.
The tier-1 runbook was drafted against the
local_reference backend. Enabling the durable backend changes three things it
must pick up (recorded here so the two pages don't drift):
- Target image. The runbook's Stage A target is v5.66.0, which predates the durable wiring (slice 4 = worker#171 = v5.70.0). For the durable path the Stage A target must be v5.70.1 (current pointer; carries the durable backend + the KV/object/vector subject-digest fixes). Rollback target is unchanged: v5.52.0.
-
Env contract. The durable path adds an axis the runbook's env block
omits — set alongside the existing
NOETL_EHDB_EVENTLOGmode flag:Env var Value for durable prod Meaning NOETL_EHDB_EVENTLOG_BACKENDdurable_segmentselect the durable engine (fail-safe: unset ⇒ local_reference)NOETL_EHDB_EVENTLOG_DURABLE_DIRa PVC mount path per-shard local segment stores + derived shared/coldload roots NOETL_EHDB_EVENTLOG_SHARED_DIR<durable-dir>/shared(or object-tier root)shared cold-load medium Under durable, NOETL_EHDB_LOCAL_REFERENCE_LOGis no longer theauthoritative store — the PVC-backed segments are. -
Storage substrate. The runbook's Stage B used an
emptyDir(fine for disposable local_reference shadow). Durable shadow/primary needs a real PVC (or object-tier medium) so segments survive restart — the §C resolution, made concrete.
These are additive; the runbook's stage structure, rollback levers, blast-radius analysis, and A1–A4 assumptions all still apply.
| # | Risk (what kind can't prove) | Severity | Proposed prod-verification / mitigation |
|---|---|---|---|
| R1 |
Segment retention / GC. |
|
Code gap closed: interest-based GC (ehdb#269/#270) + limits-based keep-last-N retention (ehdb#271) + periodic worker invocation (worker#177/#178), all default-off. The organic in-cluster soak proved a shadow store (no consumer) self-bounds: 14 → 2 segments, held flat, reclaimed climbing, replay gapless-from-base, 0 restarts. Remaining is operational, not a blocker: for Stage C set NOETL_EHDB_EVENTLOG_GC=consumer_ack + NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS (and/or MIN_RETAINED_SEGMENTS) sized to the chosen retention window on the event-writer pool, + a PVC free-bytes alert as defence-in-depth. Under primary the projection/read-model tier's ack also drives interest-based reclamation. Open follow-ups (not blockers): (a) shared-medium topology — shared-object reclamation is coherent (watermark-first), but a prod shared/object-tier medium is unverified (couples with R2); (b) max-age retention deliberately deferred (mtime resets on hydrate/cold-load — keep-last-N is the robust bound); (c) sealed-segment tier-down to the object store is a future space optimisation, not required for boundedness. |
| R2 |
Real GKE storage class / CSI behaviour. Kind used standard/local-path RWO. Prod CSI (PD-SSD etc.) fsync durability semantics, RWX support (for the shared tier), volume-detach/reattach timing on node failure, and expansion are unverified. |
High | In the prod durable-shadow soak: pick the storage class with the user; verify fsync durability (a segment written just before a node drain survives); if the shared tier uses RWX, verify the class supports it, else confine to single-writer RWO + object-tier shared medium. Kill a node during shadow and confirm segment integrity. |
| R3 |
Multi-replica shard ownership under real affinity. The soak ran single-owner (1 replica) on the user pool. Prod's event writer (noetl-worker-system-pool) has run at 1 and 2 replicas (runbook A4). A routed_away under multi-replica is the single-writer-violation signal, never exercised in prod. |
High | Confirm the prod event-writer pool + its replica count / shard topology (runbook A1/A4). Run the durable-shadow soak on that pool with NOETL_SHARD_INDEX/NOETL_SHARD_COUNT matching the live shape; assert routed_away == 0 (or, if 2 replicas, that each shard has exactly one owner and non-owners route, not double-write). |
| R4 | Disk-pressure eviction. A PVC filling (see R1) or node disk pressure can evict the writer pod mid-append. Kind never hit disk pressure. | Medium | Alert on PVC free-bytes and node DiskPressure; size headroom; confirm the pod's emptyDir/PVC quotas. Crash-recovery (D2) covers the append-interrupted case; the concern is availability, not loss. |
| R5 |
Backup / restore + DR. No backup/restore path for the durable segments is defined. The incumbent (JetStream + Postgres noetl.event) is the current DR target; once EHDB is primary, segment loss = event loss unless the incumbent dual-run is retained. |
Medium | Keep the incumbent dual-run on through and beyond Stage C (the runbook already mandates this — do not retire JetStream/Postgres). Define a segment backup (PVC snapshot or object-tier copy) before considering incumbent retirement (a later, separate step). |
| R6 |
Performance under prod event volume. The durable fsync-per-append leg adds latency the kind soak's traffic didn't stress at peak. Materially de-risked (2026-07-11): the #264/#266/#267 O(segment) fixes are merged + deployed. |
Medium → Low-Medium |
#261 head-to-head re-run: deployed durable append ~6 ms (was ~740 ms), Layer-A authoritative ~4–16 ms flat across store size — clears an incumbent-parity p99, no size-degradation. Still watch the durable-leg append latency + worker CPU/mem vs the v5.52.0 baseline through the ≥24 h shadow window incl. peak; PASS = flat, no climbing tail; flag → off on regression. (Layer-B in-cluster throughput is podman-VM-contention-noisy — directional only; Layer A + the deployed re-measurement are the trusted numbers.) |
| R7 |
selfcheck harness artifact. ehdb-selfcheck durable-eventlog reports ok:false/parity_mismatch when pointed at a live populated shard (it asserts a fresh isolated 3-event sequence). An operator could misread this as a durability failure. |
Low (operational) | Runbook note: on a live shard the meaningful signal is durable_replay_records (full read-only replay count) + absence of a hard CRC error, not the verb's ok field. Captured here + in the design note. |
Extends Runbook — Prod Cutover: Event-Log Tier. Only the deltas for the durable backend are given; the runbook's preconditions (§1), blast-radius (§6), and sign-off log (§7) apply unchanged. Every step is user-gated; a "go" on this sign-off authorises up to and including the durable-shadow soak, not Stage C.
Approved (2026-07-12): the prod team may execute Stage A′ → B′. Every command below is run by the prod team against the prod cluster — no agent runs
kubectl/gcloudagainst prod. Stage C′ stays separately gated.
Target the current durable-capable release — worker merged main (the
#178 merge, 521fd10; kind-validated as ehdb178184-merged), which carries
durable_segment plus segment GC + limits-based retention (all default-off).
Use the prod release build of that main, not the stale v5.70.1 (which
predates GC/retention). No NOETL_EHDB_* env ⇒ byte-identical to v5.52.0 on the
event path (D8). Verify ehdb-selfcheck config = all-external, exit 0; 0
restarts.
- Why this image: carrying the GC/retention code now means Stage C needs no further rebuild — only a flag flip — when it is later gated in.
- Rollback: re-apply the captured v5.52.0 digest (pure image revert).
The prod analog of the kind soak; this is what closes D3/D4/D5/D7/D9 with real evidence and exercises R2/R3/R4/R6.
Declarative alternative to the imperative steps below. The same PVC + env change is staged as a reviewable, gated GitOps manifest in noetl/ops#239 (draft; do not merge until the go). Applying that PR is equivalent to steps 1–3; the
kubectl set envrecipe here is the manual equivalent for a non-GitOps flip. Either way, the image prerequisite (Stage A′) and the open decisions (R1/R2/R3 + who/when) must be settled first.
- Provision the prod PVC (storage class + access mode decided with the user per R2) and mount it on the event-writer pool at the durable dir.
- Enable the shadow backend on the event-writer pool confirmed in runbook A1
(system pool, or whatever authors events):
kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> \ NOETL_EHDB_ENABLED=true \ NOETL_EHDB_CLIENT_ROLE=system \ NOETL_EHDB_EVENTLOG_BACKEND=durable_segment \ NOETL_EHDB_EVENTLOG_DURABLE_DIR=/ehdb-durable \ NOETL_EHDB_EVENTLOG=shadow kubectl --context $CTX -n noetl rollout status deploy/<event-writer-pool>
-
(Recommended) bound the shadow PVC with retention. A
shadowstore has no durable consumer, so interest-based GC alone reclaims nothing — enable limits-based retention so the shadow segments do not accumulate for 24 h+:Sizekubectl --context $CTX -n noetl set env deploy/<event-writer-pool> \ NOETL_EHDB_EVENTLOG_GC=consumer_ack \ NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS=<N> \ NOETL_EHDB_EVENTLOG_GC_INTERVAL_SECS=<secs>
<N>× 8 MiB (the default segment cap) to a comfortable fraction of the PVC; keep the default 8 MiBSEGMENT_MAX_BYTESin prod (the kind soak used a tiny 1 KiB cap only to force fast rotation). This is the same retention that was organically soak-proven (14 → 2 segments, held flat, replay intact). All default-off unless set, so it is opt-in. -
Do not set any EHDB env on
noetl-server-rust/ gateway (control-plane; fails the coherence guard, exit 4).
- Observation window: ≥ 24 h across a full traffic cycle incl. peak (same bar as the runbook shadow stage).
-
Rollback (instant, no redeploy; disposable — shadow never served):
kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> NOETL_EHDB_EVENTLOG=off # or drop the whole integration: kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> NOETL_EHDB_ENABLED- # (fail-safe: unsetting NOETL_EHDB_EVENTLOG_BACKEND alone ⇒ local_reference, byte-identical)
NOT authorised by the 2026-07-12 durable-SHADOW sign-off. Only after a clean Stage B′ window and the retention policy set on the event-writer pool (the R1 segment-GC decision is now MET — retention shipped; this is the operational config step, no longer a code gap) and a separate explicit user go.
kubectl --context $CTX -n noetl set env deploy/<event-writer-pool> NOETL_EHDB_EVENTLOG=primary
kubectl --context $CTX -n noetl rollout status deploy/<event-writer-pool>Incumbent (JetStream + Postgres) stays dual-run parity-checked (rollback target;
R5). Serving proof = outcome=served_primary, divergence counter 0, canary
execution reaches terminal state.
-
Rollback (instant, zero loss — primary only appends to
KeepAll):NOETL_EHDB_EVENTLOG=shadow(oroff). Structural kill switch:PRIMARY_SERVE_ACTIVATED=false, rebuild, redeploy.
| Signal | Bar |
|---|---|
noetl_ehdb_eventlog_ops_total{outcome} — mirrored (shadow) / served_primary (primary) |
steady; no invalid/degraded/primary_divergence/primary_unavailable
|
| parity divergence counter | stays 0 (hard fail if non-zero) |
routed_away |
0 on single-writer; non-zero ⇒ single-writer violated (R3) |
| PVC free bytes (kubelet volume stats) | alert before full — no GC ships (R1) |
| segment count / total bytes on the durable dir | monotonic growth expected; feeds the retention decision (R1) |
| rotation events | segments seal at ~8 MiB and roll (D6) |
append latency (durable fsync leg) + worker CPU/mem |
bounded, flat vs v5.52.0 baseline (R6) |
| worker pod restarts | 0 |
| CQRS materializer backlog (incumbent path) | ≈ 0 (unchanged) |
The decision is yours; this is the decision-ready picture as of 2026-07-11.
What changed since the first draft (2026-07-07):
-
D11 (segment GC / the R1 Stage-C blocker): GAP → MET. Segment GC + a
limits-based keep-last-N retention mode shipped (merged
ehdb#269/#270/#271
-
worker#177/#178,
all default-off) and were organically soak-proven in kind: a
shadowstore with no consumer self-bounded 14 → 2 segments and held flat,reclaimedclimbing, replay gapless-from-base, 0 restarts. R1 drops from High — Stage-C blocker to Low (operational): pick the retention window - PVC size.
-
worker#177/#178,
all default-off) and were organically soak-proven in kind: a
- R6 (perf): de-risked. The #261 head-to-head re-run shows the deployed durable append at ~6 ms (was ~740 ms) and Layer-A authoritative ~4–16 ms flat across store size — incumbent-parity, no size-degradation.
Checklist state: D1, D2, D6, D8, D10, D11 = MET. D3, D4, D5, D7, D9 = mechanism-MET, value-NEEDS-PROD-VERIFICATION (only a real prod PVC + traffic closes them — that IS the durable-shadow soak). No GAP remains.
Recommendation — GO to prod durable-SHADOW. The durable / crash-recovery core is proven, the former Stage-C blocker is resolved, and perf is incumbent-parity. Nothing in the checklist argues against a shadow soak (dual-write to a real PVC, never serving) — it is behaviour-neutral to the authoritative path (D8) and is exactly what closes the remaining prod-only items. Stage C (durable-primary) stays separately gated on: (i) a clean ≥24 h prod-shadow window, (ii) setting the retention policy for the event-writer pool (an operational config, no longer a code gap), and (iii) the R2/R3 prod-storage-class + event-writer-topology confirmations.
Alesha approved: GO to prod durable-SHADOW. The §C durability sign-off is complete for the shadow stage; ehdb#254 slice-6 box is checked.
-
Scope of the approval: prod durable-SHADOW only — deploy the
durable-capable image (flags-off, Stage A′) then enable
durable_segmentas a dual-writing shadow on a real PVC (Stage B′), never serving. This is behaviour-neutral to the authoritative event log (D8). - NOT approved by this decision: durable-primary (Stage C′). It remains separately gated on (i) a clean ≥24 h prod-shadow window, (ii) the retention policy set on the event-writer pool, and (iii) the R2/R3 storage-class + event-writer-topology confirmations — and needs its own explicit go.
-
Next operational step (for the prod team, NOT an agent action): execute
Stage A′ → B′ per §5.
No agent flips a prod flag or runs
kubectl/gcloudagainst prod.
On "GO" (now in effect): schedule Stage A′ (deploy the durable-capable image flags-off) then Stage B′ (prod durable-shadow on a real PVC) per §5, each on its own explicit go, holding the ≥24 h shadow soak. Stage C′ (durable-primary) is NOT authorised by this sign-off.
The durability sign-off (slice 6) is DONE — ehdb#254 slice-6 box is checked. The remaining prod-enablement (Stage A′ → B′) is an operational step owned by the prod team; no agent flips a prod flag or runs kubectl/gcloud against prod. Stage C′ (durable-primary) re-gates separately (see the DECISION block above).
- Runbook — Prod Cutover: Event-Log Tier (Phase 9, Tier 1) — the staged plan this durability sign-off unblocks (§C gate).
- Design: Durable Event-Log Backend — the slice-1..5 design + evidence this package consolidates.
-
Backend Configuration (Phase 10) — the
NOETL_EHDB_*surface +ehdb-selfcheck config. - noetl/ehdb#254 — durable-backend program + slice checklist (slice 6 = this sign-off).
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)