Skip to content

Design Durable EventLog Backend

Kadyapam edited this page Jul 12, 2026 · 14 revisions

Design: Durable Event-Log Backend (Phase 9 primary-serve prerequisite)

Status: slices 1–3 landed (2026-07-07). The durable segment store + crash recovery is merged in ehdb-reference::durable_eventlog (ehdb#253); execution-affinity single-writer routing is merged in ehdb-reference::affinity + ehdb-reference::durable_eventlog_affinity (ehdb#255); the shared / object-store segment tier is merged in ehdb-reference::durable_eventlog_shared (ehdb#258); worker wiring (noetl/worker#171) + kind soak landed; and interest-based segment GC — the R1 / D11 gap — landed in ehdb-reference::durable_eventlog (local reclamation) and ehdb-reference::durable_eventlog_shared (coherent shared-medium reclamation via a cross-replica reclaim watermark; see Segment GC + Shared-tier reclamation). All behind selectors disabled by default (local_reference stays the default; ownership defaults to single-owner; the shared tier is opt-in; NOETL_EHDB_EVENTLOG_GC=off by default). Remaining GC work is operational: worker periodic-invocation wiring + an in-cluster soak.

Tracks: noetl/ehdb#254 (durable-backend program + slice checklist) · noetl/ehdb#241 (completion program) · Design: Event-Log Core Engine (Phase 6) · Prod Cutover Runbook §C (the durability gate this resolves) · Backend Configuration (Phase 10).

Why this exists — the §C durability gate

The Phase-6 engine shipped the contract (EventLogDriver) and one implementation, LocalReferenceEventLogDriver, over a pod-local JSONL file (local_reference). That is correct and safe for shadow — a derived, disposable mirror — but the Phase-6 design note and the prod-cutover runbook both flag it as not production-durable as the authoritative store under primary:

"Production hardening (segmented append files + a sparse offset index for O(log n) sequence seeks, compaction, and cross-node replication) is deferred to the primary-serve phase; the local-reference substrate is the reference implementation of the contract, not the production disk format." — Design: Event-Log Core Engine

The runbook's §C durability gate is the hard blocker for Stage C (primary cutover) and names two failures of the pod-local file:

  1. Restart durability. A pod-local emptyDir log is lost on pod restart/reschedule. The authoritative event log cannot live on ephemeral pod storage.
  2. Multi-replica divergence. If more than one pod on the flagged pool appends, each writes its own pod-local log → multiple divergent "authoritative" logs. It must be pinned to a single writer, or the backend must be shared.

This design builds a backend that is durable (survives restart) and lays the groundwork for coherent (single-writer-per-shard). It is the runbook's "wait for a durable/shared EHDB log backend beyond local_reference" resolution option, made concrete.

The boundary this preserves (unchanged)

Same as Phase 6: event authorship is unchanged — the gateway/server remain the gatekeeper of what enters the log. This backend is the disk-and-index under an append the producer already authorized. Platform event log only — never business data. Secret-free outcomes.

Storage medium + topology

Two axes: the on-disk format (this slice) and the medium it lives on (a deployment decision — see Assumptions below).

On-disk format — segmented append files + CRC framing + offset index

<store root>/
  seg-0000000000000001.eslog     append-only segment, rolls over at a size cap
  seg-0000000000000002.eslog     ...

Each frame in a segment:

| magic u32 (0xE5DB0001) | body_len u32 | crc32 u32 | body: JSON(SegmentFrame) |
        4 bytes               4 bytes       4 bytes          body_len bytes

SegmentFrame is one of Event { global_sequence, execution_id, transaction_id, payload }, ConsumerCreate { consumer }, or Ack { consumer, sequence } — events and durable-consumer state share the one segment stream, so a single replay rebuilds the whole in-memory index (including consumer cursors) from disk.

Why this shape over the incumbent single JSONL file:

  • Rollover — a new segment starts once the active one would exceed a size cap (DEFAULT_SEGMENT_MAX_BYTES = 8 MiB); a single frame never spans two segments, so the offset index stays valid. A segment can be archived / GC'd / replicated as a unit without rewriting the whole log.
  • CRC + magic framing — distinguishes a torn tail (a crash mid-append left a truncated frame at EOF → discarded on recovery) from bit-rot (a complete frame whose CRC/magic is wrong → a hard error, matching the incumbent JSONL log's reject-corrupt-records stance). This is the zero-loss recovery contract: every append that fsync'd before returning survives; a partial write that never returned success is discarded.

In-memory offset index (bounded)

The store keeps, rebuilt on open:

  • events: Vec<(segment_id, byte offset)> — index i locates the frame for global sequence i + 1. O(1) locate by global sequence.
  • by_execution: HashMap<execution_id, Vec<global_seq>> — per-execution scope without a full scan.
  • consumer_acks: HashMap<consumer, acked_seq> + consumers_seen — durable cursors.

The index holds locations, not payloads. A read locates the frame via the index and cold-loads the payload bytes from the segment file on demand. Index memory is O(events) in small fixed-size entries, not O(total payload bytes) — the bounded-WAL-index property noetl/ai-meta#166 chases (the unbounded in-RAM WAL index that OOM'd the system pool).

O(1) open-for-append — the checkpoint sidecar (#267)

The worker rebuilds the durable stack per op (a stateless boundary), so every mirrored append pays a fresh open. Rebuilding the offset index by replaying every segment on each open is O(segment) — on a ~20 MB store that was ~0.5 s/op, the dominant deployed cost after #266 made the shared publish O(delta) (ehdb#268).

Fix: a small checkpoint.json sidecar next to the segments holding {event_count, active_segment_id, active_len, consumers_seen, consumer_acks} — exactly the O(1) state the append / ack hot path needs (next global sequence, where to write, durable cursors), not the offset index. It is rewritten after each mutating op, strictly after the frame fsync.

  • Open-for-append loads the checkpoint and returns O(1) — no replay. The offset index (events / by_execution) is materialised lazily on the first read (ensure_index_loaded, which runs the full replay = CRC + gapless integrity check). scan_global / read_execution therefore take &mut self; read-only cold-loads still replay eagerly at open (they need the full index).
  • Correctness — the checkpoint is a cache, never truth. A missing / stale / inconsistent checkpoint falls back to a full replay (replay-is-truth) and the checkpoint is rewritten. The consistency check is a strict active-segment length anchor: the highest segment id must equal active_segment_id and its on-disk length must equal active_len. Because the checkpoint is written only after the frame fsync, it can never name more durable data than the segments hold — a crash between the frame fsync and the checkpoint rewrite leaves the active segment longer than active_len, caught as inconsistent → replay recovers the extra fsync'd frame(s). No wrong global_sequence is ever assigned. The checkpoint is written without its own fsync (a lost rewrite just costs a one-time replay on the next open), so append stays a single fsync. It mirrors the shared tier's resumable-digest sidecar (#266).
  • Deployed result: per-op durable mirror dropped from ~0.509 s (ehdb266) to ~4–16 ms (ehdb267) on the same ~20 MB kind store; micro-bench per-op-open+append is flat across S=100→10 000. See Design: Performance & Load Testing.

Durability guarantees

  • fsync per append. write_frame appends the framed bytes then sync_data() before returning. The write is not acknowledged until it is on stable storage.
  • Crash recovery / replay on restart. open lists the segment files in id order and replays each frame into the index. A torn tail (truncated at EOF) is discarded and the segment file is truncated to the last intact frame (idempotent — a second reopen sees a clean file). Replay-is-truth: the in-memory state is a pure function of the durable segments. (A clean open-for-append skips this replay via the O(1) checkpoint sidecar above and replays only on a missing / stale / inconsistent checkpoint — replay stays the correctness backstop.)
  • Zero-loss bar. Every event whose append returned survives a restart, verbatim, in order, with per-execution scope intact and the durable consumer cursor preserved. Proven directly by exercise_durable_recovery.

Semantics parity with the incumbents

DurableEventLogDriver implements the same EventLogDriver trait as LocalReferenceEventLogDriver, so it is a drop-in behind the contract:

Property Guarantee
Global ordering monotonic, gapless from 1 (next = events.len() + 1)
Per-execution scope by_execution index; scoped ordered reads
Tail / subscribe durable consumer, created on first pull
Offset / ack explicit ack-after-materialize, persisted Ack frame; cursor never moves backward
Append-only + immutable frames never mutated; segments append-only
Replay is truth in-memory index rebuilt from segments on open

A dedicated test drives the same op sequence through both the durable and local-reference drivers and asserts identical observable results (sequence, created_stream, log_record_count, scan order, per-execution scope).

Single-writer-per-shard (execution-affinity — slice 2, landed)

Coherence under multiple replicas needs exactly one writer per shard. Slice 2 (ehdb#255) lands it, reusing the noetl/ai-meta#166 / #116 execution-affinity ownership so each shard is owned by exactly one replica = its sole writer.

The ownership hash is byte-identical to the worker/server. ehdb-reference::affinity::shard_for_i64 is XxHash64(seed=0) over the execution id's 8 little-endian bytes % shard_count — the same function as noetl-worker src/sharding.rs::shard_for and noetl-server sharding::shard_for (pinned twox-hash = "1.6", same major). NoETL execution ids are snowflake i64s; EHDB stores them as decimal strings, so shard_for_execution parses the string back to the i64 and routes through shard_for_i64 — ownership agrees with the worker for every real execution. ShardOwnership mirrors the worker's AffinityConfig: a replica owns exactly its shard_index bucket, so a shard_count-replica pool partitions every execution to exactly one writer.

The routing layer (durable_eventlog_affinity::AffinityRoutedEventLog, one instance = one replica) maps each execution to a per-shard DurableSegmentStore under <root>/shard-<NNNN>/:

  • append — the owner writes (Routed::Served); a non-owner is refused with no side effect (Routed::NotOwner { owner_shard }) so the caller re-routes to the owner. No bytes written, no sequence consumed, so a re-route can never double-process.
  • read (read_execution / scan_shard) — the owner serves resident; a non-owner cold-loads the durable segments read-only (ServedBy::NonOwnerColdLoad).
  • tail / ack — durable-consumer writer state, so owner-only; a non-owner is refused.

The full EventLogDriver contract holds within each shard (gapless per-shard global sequence, per-execution scope, replay-is-truth); an execution's events always land in shard-<shard_of(execution)>. The owner keeps its owned-shard stores resident behind a mutex for O(1) locate; a cold-load is a fresh DurableSegmentStore::open_read_only over the shard's directory — what a new owner or a non-owner read does, and what a pod restart replays.

open_read_only is non-mutating by construction: it refuses every write and never truncates a recovered torn tail (only the shard's single owner repairs its own tail on its own writable open), so a non-owner read can never mutate another owner's shard. That is the structural guarantee behind single-writer.

The disk format (slice 1) is unchanged — the routing layer sits on top of the per-shard segment stores without reworking the segments.

Shared / object-store segment tier (slice 3, landed)

Slice 2's non-owner cold-load reads the non-owner's own local disk — which, on a different pod with different local storage, is empty. Slice 3 (ehdb#258, ehdb-reference::durable_eventlog_shared) makes the cold-load pull from a shared durable medium so a shard survives the loss of the writer's pod-local disk — the runbook §C durability gate's "durable/shared EHDB log backend beyond local_reference" resolution, made concrete.

The pluggable medium. SharedSegmentBackend is the seam: put_segment / get_segment / list_segment_ids. FilesystemSharedBackend (a shared directory — a ReadWriteMany/ReadWriteOnce PVC on kind) is the bootstrapping impl; the same trait routes to EHDB's own durable object tier (Phase 8 object engine) later — the self-sufficiency end-state — with no change to the routing/hydrate logic.

Fixed-width, digest-integrity-checked keys. A segment is addressed by position — (shard, segment_id), both bounded integers — so the shared- store object key is a constant 40 chars (noetl.ehdb.seg.<shard:08x>.<segment_id:016x>). The key width is independent of any payload, so it can never approach a subject-length cap on whatever medium backs the trait. This deliberately avoids the trap the object tier hit (noetl/ai-meta#234: hex-encoding an arbitrary platform key into a NATS subject blew the 256-char subject cap). Each object carries a byte-length + XxHash64 content digest, so a truncated / corrupt shared object is a hard error on read, not a silently-shortened replay. (Segments are position-keyed, not content-hashed: the active segment is append-mutable, so a content-hash key would change every append and break listing — the digest is integrity metadata, not the key.)

The routing. SharedTierEventLog (one instance = one replica) wraps the slice-2 AffinityRoutedEventLog:

  • owner append — writes locally (slice-1/2 fast path, fsync'd) then publishes the shard's segments to the shared store (idempotent: a sealed segment already published at its length is skipped; the growing active segment is always re-published).
  • non-owner read (read_execution / scan_shard) — cold-loads the shard's segments from the shared store into a scratch dir and opens a slice-1 DurableSegmentStore::open_read_only over them. Replay is byte-identical + sequence-preserving — the full EventLogDriver contract holds because the materialized segments are the owner's segments.
  • new owner inheriting a shard — hydrate_owned_shard pulls the shard's segments from shared into the new owner's own empty local dir before it serves / appends, so a pod that never held the shard locally recovers it zero-loss and continues the sequence. This is the crash-recovery-from-shared path (a pod restart / reschedule onto a fresh node).
  • tail / ack — owner-only writer state; the consumer-create + ack frames are published to shared too, so a new owner inherits the durable cursor.

The durable-eventlog-shared --root <dir> [--shard-count <n>] selfcheck spins up a shard_count-replica pool with separate local disks over one shared store and proves owner-publish / non-owner-cold-load-from-shared / crash-recovery-from-shared / shared-store-miss-reads-empty / parity. Exit 0 only when SharedTierReport::holds:

{"shared_tier_holds":true,
 "report":{"shard_count":2,"executions":2,"shared_segments":2,
 "owner_published_ok":true,"nonowner_coldload_from_shared_ok":true,
 "crash_recovery_from_shared_ok":true,"shared_miss_ok":true,"parity_ok":true,
 "divergence":null}}

Bootstrapping vs end-state (the assumption to flag). The recommended bootstrapping medium is a PVC-backed shared directory (FilesystemSharedBackend) — durable-across-restart with the segment format unchanged. The self-sufficiency end-state is EHDB's own object tier holding the segments (a different SharedSegmentBackend impl behind the same trait). Open assumptions: for a PVC, a ReadWriteMany (or single-writer ReadWriteOnce) class on the target cluster; for the object-tier backend, object-store creds / workload-identity for the owning pool. Worker wiring of the selected medium is slice 4.

Backend selection (Phase-10 surface)

A new storage-backend axis, orthogonal to the Phase-10 Backend Configuration TierMode (off/shadow/primary) axis:

Axis Env var Values Meaning
Tier mode NOETL_EHDB_EVENTLOG off / shadow / primary Whether EHDB serves the event-log tier
Storage backend NOETL_EHDB_EVENTLOG_BACKEND local_reference / durable_segment Which durable engine serves when it does

EventLogStorageBackend::from_raw is fail-safe: only the exact token durable_segment selects the durable backend; unset / empty / unrecognised → local_reference (the default). So an unknown value never silently changes the authoritative store, and nothing changes until an operator opts in. Coexistence + migration from local_reference is a forward-only cutover: run durable_segment as the primary backend on a durable medium; the local_reference log stays readable for the shadow window.

Segment GC — interest-based retention (the D11 gap)

Slices 1–5 never reclaim a sealed segment: under KeepAll + primary the on-disk store grows without bound. The prod-durability sign-off recorded that as residual-risk R1 / gap D11 — the Stage-C blocker (durable-primary cutover cannot be authorised while the store grows unbounded). This section is the resolution: an interest-based segment GC that reclaims segments a durable consumer has already consumed, while preserving every crash-safety and replay-is-truth invariant slices 1–5 established.

Retention rule — the consumer-ack watermark

A sealed segment S is reclaimable when every durable consumer has acked past S's last global sequence, respecting a min_retained_segments floor:

  1. watermark = min over all durable consumers of their ack cursor. No consumer, or any consumer that has acked nothing, ⇒ watermark = 0 ⇒ nothing is reclaimed — interest-based retention never drops an event no consumer has read (JetStream-style). A lagging consumer holds the watermark down; operators alert on lag rather than losing data.
  2. Reclaim the contiguous prefix of sealed segments whose last global sequence <= watermark, keeping at least min_retained_segments most-recent segments (including the active one, which is never reclaimable).

This is safe under the platform boundary because Postgres noetl.event is the immutable full-history archive (never purged); the durable EHDB segment store is a working/streaming substrate. Reclaiming events every consumer has already projected does not lose history — it lives in noetl.event — it only bounds the derived engine's disk.

The base offset — retained log is gapless-from-reclaimed+1

Reclaiming leading segments raises the lowest retained sequence above 1. The store now tracks a base offset (reclaimed_through, 0 for an un-GC'd store): the offset index is base-shifted (events[i] locates global sequence reclaimed_through + i + 1), a scan clamps its start to the base, and a read of a reclaimed sequence reports reclaimed (absent) — distinct from corruption. event_count stays the absolute tip (base + retained), so global ordering is still monotonic-gapless from the base; the durable EventLogDriver contract holds over the retained range. When reclaimed_through == 0 every path is byte-identical to the pre-GC backend, so an un-opted-in store is unchanged.

Crash safety — the durable reclaim manifest is the commit point

Reclamation is a durable, crash-atomic sequence:

  1. Write-forward current durable-consumer state (Ack frames at each consumer's cursor) into the active segment + fsync. This is the subtle part: the ConsumerCreate / Ack frames that established a consumer's cursor may physically live in a segment about to be unlinked. Because every consumer has acked >= watermark >= reclaimed_seq, a re-asserted Ack at the current cursor re-establishes both the consumer's existence and its cursor from surviving segments alone — so replay-is-truth still reconstructs full consumer state after the old segments are gone.
  2. Commit the fsync'd reclaim.json manifest (reclaimed_through_seq + reclaimed_through_segment) — the point of no return. Unlike the checkpoint sidecar (an optimization written without fsync), the manifest is durable truth: it is fsync'd, atomically renamed, and the directory entry fsync'd.
  3. Unlink the reclaimed segment files.
  4. Replay to rebuild the index with the new base, and refresh the checkpoint sidecar.

A crash at any point recovers correctly:

  • Before the manifest commit ⇒ nothing was reclaimed; the full log replays normally (reclaimed_through = 0). Idempotent retry on the next GC.
  • After the manifest commit, before/during the unlinks ⇒ the manifest names the reclamation; on the next open a below-watermark segment still on disk is re-deleted (on both the replay path and the O(1) checkpoint-trust path, which calls the same leftover-cleanup), and the base offset is honored. A leftover is never read (it is below the base), so even a garbage leftover is harmless — it is unlinked, not replayed.

The manifest can never name reclamation that would lose data a surviving segment does not carry, because it is written strictly after the write-forward fsync. A fresh segment after an all-consumed store starts at reclaimed_segment + 1, so a new segment id can never collide with (and be re-deleted as) a reclaimed one.

Selection surface — off by default

A new retention axis, orthogonal to the tier-mode and storage-backend axes, and disabled by default exactly like EventLogStorageBackend:

Axis Env var Values Meaning
Segment GC NOETL_EHDB_EVENTLOG_GC off (default) / consumer_ack Whether consumed sealed segments are reclaimed
Min-retained floor NOETL_EHDB_EVENTLOG_GC_MIN_RETAINED_SEGMENTS integer (default 2) Segments always kept even when fully consumed

SegmentGcPolicy::from_raw is fail-safe: only consumer_ack / on / enabled turn GC on; unset / empty / unrecognised leaves it off, so an unknown value never silently starts dropping segments. GC applies only to the durable_segment backend, and only when explicitly enabled.

CLI + cost

ehdb-local-reference durable-eventlog-gc --root <dir> [--events <n>] [--segment-max-bytes <n>] [--min-retained-segments <n>] [--consumer <name>] drives a full cycle: append --events, ack a durable consumer to ~3/4 of the log, reclaim, and prove growth was bounded while the retained log stayed gapless-from-base + zero-loss and a restart recovers. Exit 0 only when SegmentGcDriveReport::holds. A representative run (40 events, 220-byte segments) reclaimed 41 → 11 segments at watermark 30 with every invariant holding. The eventlog_segment_gc criterion group measures a single reclaim_segments call at pre-warmed sizes S — GC is off the append hot path (run periodically, not per-op), and the interest watermark keeps the retained set bounded, so a reclaim replays only the surviving segments, not the whole history. Measured (4 KiB segments, ~3/4 acked): ~16.6 ms at S = 200, ~22.9 ms at S = 800 — the cost tracks the retained set (the reclaim replays surviving segments + fsyncs the manifest + write-forward), not the full history, and is comfortably affordable for a periodic maintenance pass.

Shared-tier reclamation — coherent across the shared medium (landed)

The reclamation above is local. On its own it is not coherent with the shared segment tier (slice 3): a local unlink leaves the reclaimed segments in the shared store, so cold_load / hydrate_owned_shard would re-pull them and diverge from the owner (and a new owner would un-reclaim on transfer). The shared-tier reclamation closes that.

The cross-replica reclaim watermark. A per-shard watermark object (noetl.ehdb.rw.<shard>, monotonic) on the shared medium records the highest reclaimed (seq, segment). SharedSegmentBackend gains put_reclaim_watermark / reclaim_watermark / delete_segment (default-erroring so a backend that can't persist a watermark refuses shared-tier GC rather than losing coherence; FilesystemSharedBackend implements them). A cold-load / hydrate materializes only segments above the watermark (retained_shared_ids), so it can never re-pull a reclaimed segment — and the base offset falls out of the first surviving frame on replay, exactly as locally.

SharedTierEventLog::reclaim_shard — watermark-first ordering (the crux):

  1. Reconcile — adopt any already-committed watermark into the owner's local store first (reconcile_owned_shard → reclaim_to_segment).
  2. Plan the boundary from the owner's local consumer cursors (plan_reclaim_boundary) — no mutation.
  3. Commit the shared watermark — the cross-replica point of no return. From here every reader skips <= segment.
  4. Reclaim locally to the boundary (write-forward + fsync'd manifest + unlink).
  5. Delete the shared objects <= segment.

Committing the watermark before any deletion is what removes the divergence: a non-owner / new-owner that reads mid-reclamation already skips the segments about to be removed. A crash after step 3 leaves orphan local/shared objects readers already skip — a bounded space leak the next reclaim_shard re-attempts, never a correctness bug; the owner realigns via step 1 on its next run or bring-up (reconcile_owned_shard). The only residual transient is that a crashed owner may briefly over-serve its own un-reclaimed local segments until it reconciles — benign (it shows more, never less; the data is also in noetl.event), and self-healing.

Drive + validation. SharedTierEventLog::reclaim_shard + exercise_shared_tier_gc (CLI durable-eventlog-shared-gc) prove, over a multi-replica pool on one shared store, that after the owner reclaims shard 0 both a non-owner cold-load and a fresh new-owner hydrate return exactly the owner's retained set (no re-pull, no un-reclaim on transfer), the shared objects <= watermark are pruned (bounded shared disk), and a targeted crash-window test shows reconcile_owned_shard realigns an owner whose local segments are ahead of a committed watermark. A representative 2-shard run pruned the shared store 40 → 10 objects at watermark 30 with every coherence invariant holding. The eventlog_shared_tier_gc criterion group measures reclaim_shard (local reclaim + one watermark commit + per-segment shared deletes) at ~25 ms (S=50) / ~38 ms (S=150) — off the append hot path, a periodic maintenance pass whose cost tracks the retained set plus a small constant of shared ops.

Still off by default + not auto-invoked — the shared-tier GC runs only when an operator opts in (NOETL_EHDB_EVENTLOG_GC=consumer_ack) and a caller invokes reclaim_shard; the worker periodic-invocation wiring + in-cluster soak is the next slice (as slices 1–3 landed engine-only ahead of slice-4 worker wiring).

Worker periodic invocation (the operational wiring)

The engine ships the reclamation; the worker invokes it periodically so an opted-in operator gets a self-bounding store with no manual step (noetl/worker segment-GC wiring, tracks noetl/ehdb#254).

  • Who triggers. Each worker replica runs a periodic task for the shards it owns — only a shard's single writer may reclaim it (the execution-affinity invariant), so GC is inherently per-owner and needs no central coordinator.
  • Cadence. NOETL_EHDB_EVENTLOG_GC_INTERVAL_SECS (default 0 = off). The task spawns only when every axis is opted in: durable_segment backend + NOETL_EHDB_EVENTLOG_GC=consumer_ack + a positive interval + a data-plane durable contract. All default-off ⇒ a default worker is byte-identical and the task never even spawns.
  • Back-pressure / concurrency. The worker builds the durable stack per op (stateless boundary), so the GC task and the durable append path are two writers to the same shard's files. A process-global per-shard advisory lock (acquired by both paths) serializes them per shard — GC on one shard never blocks appends on another — and closes a latent intra-replica append↔append race. GC runs on a spawn_blocking thread so its fsync/unlink I/O never stalls the async runtime. Its per-shard hold is brief (tens of ms, per the benches) and infrequent (interval-driven).
  • Observability. Each pass records to the noetl_ehdb_eventlog_gc_* metric family (outcome = reclaimed / noop / error); a reclamation logs at INFO (low frequency, only on a real reclaim). ehdb-selfcheck durable-eventlog-gc reclaims the live owned shards + reports; --drive 1 runs a self-contained shared-tier GC cycle on the real durable volume.

In-kind soak (2026-07-10, LOCAL kind, no GKE/prod). The GC-wired worker (localhost/noetl-worker:v5.72.3-ehdbgc254) rolled onto the durable_segment user pool (noetl-worker-rust, single-owner, /ehdb-durable PVC) with NOETL_EHDB_EVENTLOG_GC=consumer_ack + NOETL_EHDB_EVENTLOG_GC_INTERVAL_SECS=30. Confirmed in-cluster: the periodic task spawned (segment-GC task started … interval_secs=30); it ran healthy ticks recorded as noetl_ehdb_eventlog_gc_ops_total{outcome="noop"} with last_ok=1 — correctly reclaiming nothing on the unconsumed shadow store (the interest-based safety); ehdb-selfcheck durable-eventlog-gc --drive 1 on the real PVC (/ehdb-durable/gc-soak) reclaimed the shared store 40 → 10 objects at watermark 30 with gc_holds=true — owner replay gapless-from-base, non-owner cold-load coherent, new-owner hydrate coherent, no divergence; 0 worker restarts, no crashloop, no impact on the live path.

The shadow-consumer caveat (interest needs a consumer)

Interest-based reclamation needs a durable consumer to ack past segments (the watermark). Under the currently-deployed shadow event-log the durable store is a disposable mirror that nothing consumes (no tail/ack site), so the watermark stays 0 and the periodic task correctly reclaims nothing — a noop, not a bug. Interest-based GC becomes effective when a durable consumer acks: naturally under primary (the projection / read-model tier tails + acks the event log to build read models), or via an explicit maintenance consumer. Limits-based retention (below) bounds a shadow store without one.

Limits-based retention — keep-last-N (the shadow bound)

A retention mode reclaims independent of consumer interest so a store with no durable consumer (a shadow mirror) self-bounds. plan_reclaim folds interest + retention into one effective boundary:

interest_ceiling = consumers.exist ? Some(min ack cursor) : None   // None = no ceiling
retention_target = max_retained_segments ? Some(seq leaving N segments) : None
boundary = retention_target ? min(retention_target, interest_ceiling ?? u64::MAX)
                            : interest_ceiling ?? 0
  • Primary + consumer + retention → min(consumer-cursor, retention-target) — never reclaim ahead of a consumer and bound the store.
  • Shadow (no consumer) + retention → retention drives; the store self-bounds.
  • Shadow + no retention → 0 — reclaim nothing (the pre-retention default).
  • min_retained_segments stays the hard floor (the retention cap is clamped up to it).

The only engine change is inside plan_reclaim — the watermark-first commit, shared-tier cold-load/hydrate coherence, reclaim_to_segment, reconcile_owned_shard, and the worker per-shard lock all consume the boundary unchanged, so retention inherits every crash-safety + coherence invariant.

Keep-last-N (count-based), not max-age. Retention counts segments, which is intrinsic and survives hydrate / cold-load. Max-age is deliberately not implemented: the only per-segment timestamp is file mtime, which resets on the fresh files a new owner materializes, so max-age would mis-fire across ownership transfer / pod reschedule.

Config: NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS (unset ⇒ interest-only, unchanged) + the worker's NOETL_EHDB_EVENTLOG_SEGMENT_MAX_BYTES rollover tunable. Merged ehdb#271 + worker#178.

Organic in-kind soak (2026-07-10 → re-confirmed 2026-07-11 on merged main, LOCAL kind). The durable_segment user pool (noetl-worker-rust, shadow, no consumer — consumers_seen: []) with NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS=3 + NOETL_EHDB_EVENTLOG_SEGMENT_MAX_BYTES=1024 + interval 20 s. Run on the merged-main image localhost/noetl-worker:ehdb178184-merged (built off merged worker main — retention engine ehdb#271

  • knob worker#178 + periodic wiring worker#177), so this proves the shipped code, not a branch image.

The live shard-0 store rotated through 60+ ~1 KiB segments across event_count 25 476 → 25 656 while holding ≈ 2 segments — the periodic retention GC kept it bounded with no consumer. A controlled burst made it explicit: a +180-event burst pushed it to 14 segments (appends outpace the 20 s tick), then one tick later (no more appends) it self-bounded back to 2 segments and held flat; noetl_ehdb_eventlog_gc_ops_total{outcome="reclaimed"} climbed (1 → 3) while {outcome="error"} stayed absent, last_degraded=0, no new restarts. The reclaim manifest advanced durably (reclaimed_through_seq 25 354 → 25 464 → …) and an independent read-only reopen replayed the retained log gapless-from-base — replay + durable-recovery intact under organic reclamation with no consumer. (The earlier 2026-07-10 run on the branch image v5.72.4-ehdbgcret254 measured the same shape: 19 → 2 segments, ~22 MB → 2,147 B.)

D11 verdict

D11 (segment GC) moves from GAP → mechanism delivered + drive-tested for BOTH the single-writer local (PVC) topology AND the shared-medium topology. The reclamation engine exists, is crash-safe, preserves replay-is-truth + the gapless contract, keeps a non-owner / new-owner coherent with the reclaiming owner across the shared medium, and is unit- + drive-tested. This unblocks the Stage-C durable-primary consideration for the deployed topology (single-writer system pool on a PVC with the shared tier) — the store no longer grows unbounded under primary on either medium. The remaining work is operational, not correctness: worker periodic-invocation wiring + an in-cluster GC soak.

Worker wiring (slice 4, landed)

Slice 4 makes the durable stack selectable + functional in worker-rust (noetl/worker#171). The worker's event-log tier constructs its storage engine from the backend axis instead of hard-coding the JSONL reference driver.

Selection seam — worker-rust src/ehdb/eventlog_backend.rs. Resolves NOETL_EHDB_EVENTLOG_BACKEND (EventLogStorageBackend::from_raw, fail-safe) and, on durable_segment, builds the whole slice-1+2+3 stack: SharedTierEventLog (slice 3, shared publish / cold-load) over AffinityRoutedEventLog (slice 2, execution-affinity single-writer) over per-shard DurableEventLogDriver segment stores (slice 1). Shard ownership is built from the worker's own NOETL_SHARD_INDEX / NOETL_SHARD_COUNT (byte-identical XxHash64 to crate::sharding::shard_for), so the event-log shard a replica owns is the same shard as the drive it owns — the durable single-writer aligns with the drive-pool affinity.

Durable paths. Derived from the configured local_reference log's parent (<log-parent>/ehdb-durable/{local,shared,coldload}) unless overridden:

Env var Default Meaning
NOETL_EHDB_EVENTLOG_BACKEND local_reference Storage engine selector
NOETL_EHDB_EVENTLOG_DURABLE_DIR <log-parent>/ehdb-durable Base for per-shard local stores + derived shared/coldload roots
NOETL_EHDB_EVENTLOG_SHARED_DIR <base>/shared Shared-tier medium root (PVC / object-store)

Where it dispatches. mirror_event — the per-event shadow mirror AND the per-event primary append — routes through the selected backend; both backends yield the same EventLogAppendOutcome, so the parity path is backend-agnostic. A non-owned append under durable_segment is a neutral EventLogOutcome::RoutedAway (correct single-writer behaviour; never fires on local_reference / single-owner). PRIMARY_SERVE_ACTIVATED is unchanged and the primary path stays gated.

Config + selfcheck. The worker's Phase-10 config resolution surfaces eventlog_storage_backend in the ehdb-selfcheck config matrix. A new ehdb-selfcheck durable-eventlog verb mirrors via the selected backend and, on durable_segment, independently reopens the owning shard store read-only to prove the appends replayed from durable segments (crash-recovery), not JSONL:

{"suite":"ehdb-durable-eventlog-selfcheck","ehdb":"enabled","mode":"shadow",
 "storage_backend":"durable_segment","ok":true,"durable_replay_records":3,
 "durable_replay_ok":true}

Disabled-by-default, reversible, zero behavior change when unset. No env ⇒ byte-identical JSONL append; flip the env back to restore the incumbent with no redeploy. Still the derived EHDB fabric — event authorship unchanged.

The worker pins ehdb-reference at cca0d0d (slices 1-3 + object subject-digest fix). The in-cluster live proof landed in slice 5 (below).

Kind soak (slice 5, landed)

Slice 5 ran the durable backend under sustained real traffic in the local kind-noetl cluster and proved crash recovery on a real pod restart — 2026-07-07, LOCAL kind only (no GKE / prod; repos/noetl + repos/server untouched).

Setup. Built localhost/noetl-worker:v5.70.0 (the slice-4 durable wiring; podman save + kind load image-archive — the podman-provider-safe path) and rolled it onto the user data-plane pool noetl-worker-rust (role worker, single replica ⇒ single-owner shard 0), adding to the existing EHDB shadow env:

NOETL_EHDB_EVENTLOG_BACKEND=durable_segment
NOETL_EHDB_EVENTLOG_DURABLE_DIR=/ehdb-durable        # a 2Gi ReadWriteOnce PVC

The durable dir is a PVC (ehdb-durable-soak, standard/local-path) so segments survive a pod delete + reschedule — the substrate the runbook §C recommendation calls for. ehdb-selfcheck config confirmed eventlog_storage_backend: durable_segment, mode shadow, role worker, secret_free.

Accumulation + segment rotation (real traffic). Drove thousands of real executions (automation/pft_sql_probe_v2 + tests/large_tabular_result_test). Real drive events landed in CRC-framed segments under /ehdb-durable/local/shard-0000/, and the slice-3 shared tier published each segment byte-identical to /ehdb-durable/shared/noetl.ehdb.seg.<shard>.<seg> (+ .meta). Sustained traffic crossed the 8 MiB rollover boundary and rotation fired in-cluster: seg-…0001.eslog sealed at 8 387 133 bytes (just under DEFAULT_SEGMENT_MAX_BYTES = 8 388 608) and a fresh seg-…0002.eslog opened as the active tail — the same rollover the slice-1 unit tests prove, now confirmed on real event bytes. No JSONL log was written (the durable backend never writes /tmp/ehdb-shadow-ref.jsonl).

Metrics. noetl_ehdb_eventlog_ops_total{operation="mirror",outcome="mirrored"} advanced monotonically past 11 731 with last_ok=1, last_degraded=0, and no invalid / degraded / routed_away outcome ever emitted (single owner serves every append). Zero worker restarts / crashloops during the soak; CRC/index integrity held (a bad CRC is a hard error — none surfaced).

Crash recovery — real pod restart on the PVC. With the backlog drained, captured the sealed segment's sha256 (c2ba3d33…, 8 387 133 B) + the active tail, then force-deleted the pod (--grace-period=0 --force). A fresh pod rescheduled on the same PVC and:

  • seg-…0001.eslog was byte-identical (same sha256, same size) — the sealed segment survived the crash with zero corruption.
  • seg-…0002.eslog continued growing — the reopened store replayed the existing segments, recognised the active tail, and kept appending (no reset to seg-0001, no sequence gap).
  • ehdb-selfcheck durable-eventlog on the new pod independently reopened the shard store read-only and replayed durable_replay_records: 11856 from the segments alone — the crash-recovery replay, on a pod that never saw the original writes. (The verb's ok:false / parity_mismatch here are harness artifacts of pointing the isolated-dir selfcheck — which asserts a fresh 3-event sequence — at a live populated shard; the live path's own parity held with last_degraded=0 across all 11 731 events.)
  • New appends continued the gapless monotonic global sequence across the restart (…11852, 11853, 11855); a further 30 real drives grew the active segment (+59 KB) with mirrored the only outcome.

Backend equivalence note. The soaked image is tagged v5.70.0; the worker pointer advanced concurrently to v5.70.1 (ehdb#172 KV/vector subject-digest fix, ehdb-reference cca0d0d → 52120a7). That diff touches only kv.rs / vector.rs — no durable_eventlog* change — so the durable event-log backend is byte-identical between the two and this soak validates the backend at the current pointer.

CLI

ehdb-local-reference gains durable-eventlog-recovery --root <dir> [--consumer <name>] — the crash-recovery drive: append a deterministic event set + ack a durable-consumer cursor, reopen a fresh driver over the same root (simulated pod restart), and prove zero-loss + gapless ordering

  • per-execution scope + payload fidelity + durable-cursor survival. Exit 0 only when DurableRecoveryReport::recovered() holds. Example output:
{"driver":"ehdb-durable-segment","recovered":true,
 "report":{"appended":3,"zero_loss":true,"ordering_ok":true,"scope_ok":true,
 "payloads_match":true,"cursor_survived":true,"pending_after_restart":2,
 "divergence":null}}

Slice 2 adds three affinity verbs:

  • durable-eventlog-affinity --root <dir> [--shard-count <n>] — the full single-writer drive: spin up an n-replica pool over one root, partition a deterministic execution set, and prove owner-writes / non-owner-refused / single-writer-invariant / non-owner-cold-load-read / crash-recovery-under- ownership. Exit 0 only when AffinitySingleWriterReport::holds():

    {"single_writer_holds":true,
     "report":{"shard_count":2,"executions":4,"owner_appends":4,
     "nonowner_refusals":4,"owner_writes_ok":true,"nonowner_refused_ok":true,
     "single_writer_invariant":true,"coldload_read_ok":true,
     "crash_recovery_ok":true,"divergence":null}}
  • durable-eventlog-affinity-append --root <dir> --shard-index <n> --shard-count <n> --execution-id <id> --transaction-id <id> --payload <text> — one routed append. Exit 0 when the owner serves it, exit 6 (distinct from the 3/4/5 engine-error codes) when refused as a non-owner, so a shell / kind-soak harness can assert the routing decision directly.

  • durable-eventlog-affinity-read --root <dir> --shard-index <n> --shard-count <n> --execution-id <id> [--after <seq>] [--limit <n>] — one routed read; the JSON served_by is owner_resident or non_owner_cold_load.

Slice 3 adds the shared-tier drive:

  • durable-eventlog-shared --root <dir> [--shard-count <n>] — the full shared-tier drive: a shard_count-replica pool with separate local disks over one shared store, proving owner-publish / non-owner-cold-load-from- shared / crash-recovery-from-shared (new owner, empty local disk) / shared-store-miss-reads-empty / parity. Exit 0 only when SharedTierReport::holds.

The segment-GC slice adds the reclamation drive:

  • durable-eventlog-gc --root <dir> [--events <n>] [--segment-max-bytes <n>] [--min-retained-segments <n>] [--consumer <name>] — append --events under a small segment size, ack a durable consumer to ~3/4 of the log, reclaim consumed sealed segments, and prove growth was bounded while the retained log stayed gapless-from-base + zero-loss and a restart recovers. Exit 0 only when SegmentGcDriveReport::holds.
  • durable-eventlog-shared-gc --root <dir> [--shard-count <n>] — a shard_count-replica pool over one shared store reclaims a shard's consumed sealed segments across BOTH local + shared, and proves a non-owner cold-load AND a fresh new-owner hydrate both see exactly the owner's retained set (no re-pull, no un-reclaim on transfer, shared disk bounded). Exit 0 only when SharedTierGcReport::holds.

Assumptions where the deploy/storage substrate is unclear (flag for the user)

The on-disk format is durable; where that format lives is a deployment decision the runbook §C leaves open. The two viable media, with a recommendation:

  • PVC-backed volume (recommended for bootstrapping). Mount a PersistentVolume on a single-writer StatefulSet-style pod and point the store root at it. Simplest path to durability-that-survives-restart with a single writer; the segment format is unchanged. Open assumptions: PVC availability + a ReadWriteOnce class on the target cluster, and pinning the event-log-owning pool to a single writer (the system pool has run at 1 and 2 replicas historically — §C A4).
  • Object-store-backed segments (the self-sufficiency end-state). EHDB's own durable object tier (Phase 8 object engine) holds the segments so they are shared + cold-loadable by a new owner, aligning with the #166 Feather-state-shards-in- object-store + cold-load-on-miss machinery. Tension: EHDB aims to replace the external object store, so long-term the segments live on EHDB's own object tier — but for bootstrapping, a PVC (or the existing object store) is the pragmatic path. Open assumptions: object-store creds / workload-identity for the owning pool; the cold-load-on-miss read path is a later slice.

Recommendation: land the segment store + crash recovery now (this slice, PVC-friendly since it is just a directory of files), then add the execution-affinity single-writer routing + the shared object-store tier as the next slices, so a PVC deployment can be validated on kind before the shared-tier work.

Scope note (deliberate). This slice does not dedupe by transaction_id (the incumbent JSONL transaction log rejects duplicate transaction ids; the durable event store treats transaction_id as opaque per-event metadata). Producer-side exactly-once by transaction id, if required, is a follow-up — the real producer path assigns unique ids, and the durability/ordering/scope contract this slice delivers does not depend on it.

Remaining durable-backend slices

  1. Segment store + crash recovery — ✅ landed (ehdb#253).
  2. Execution-affinity single-writer routing — ✅ landed (ehdb#255): XxHash64 ownership (byte-identical to worker/server) so each shard has exactly one writer; non-owner writes refused (route-to-owner), non-owner reads cold-load read-only.
  3. Shared / object-store segment tier — ✅ landed (ehdb#258): SharedSegmentBackend trait (filesystem/PVC now, EHDB object tier later) + SharedTierEventLog (owner publishes segments to shared; non-owner cold-loads / new owner hydrates from shared); fixed-width digest-integrity-checked segment keys.
  4. Worker wiring — ✅ landed (noetl/worker#171): NOETL_EHDB_EVENTLOG_BACKEND=durable_segment resolved in worker-rust src/ehdb, ehdb-selfcheck durable-eventlog verb in the image. See Worker wiring below.
  5. Kind soak — ✅ landed (2026-07-07): worker v5.70.0 rolled onto the noetl-worker-rust pool with NOETL_EHDB_EVENTLOG_BACKEND=durable_segment on a PVC; thousands of real drives accumulated CRC-framed segments, rotation fired in-cluster (seg-0001 sealed at 8 387 133 B → seg-0002), metrics advanced past 11 731 with zero invalid/degraded, and crash recovery on a real pod delete + reschedule proved the sealed segment byte-identical + a read-only reopen replaying 11 856 records with a gapless continuing sequence. See Kind soak.
  6. Prod-durability sign-off — the §C gate cleared: durable + single- writer proven, then Stage C is unblocked (still user-gated, per tier). Decision package drafted 2026-07-07: Sign-off — Durable Event-Log Backend, Prod Durability (go/no-go checklist, evidence bundle, residual-risk register, extended durable-shadow → durable-primary rollout sequence). **Recommendation: GO to a prod durable-SHADOW soak; Stage C durable-PRIMARY stays gated on that soak
    • a segment-GC decision (R1).** Awaiting user go/no-go; ehdb#254 slice-6 box stays unchecked until the actual prod sign-off.
  7. Segment GC (the R1 / D11 gap) — ✅ local reclamation landed: interest-based (consumer-ack watermark) reclamation of consumed sealed segments, base-offset (retained log gapless-from-reclaimed+1), durable reclaim.json manifest as the crash-atomic commit point, write-forward of consumer state, and a min_retained_segments floor — all behind NOETL_EHDB_EVENTLOG_GC=consumer_ack (off by default). See Segment GC above.
  8. Shared-tier reclamation (coherence across the shared medium) — ✅ landed: a monotonic per-shard cross-replica reclaim watermark on the shared medium (put_reclaim_watermark / reclaim_watermark / delete_segment); SharedTierEventLog::reclaim_shard reclaims local + shared watermark-first (commit the watermark before any delete), so cold_load / hydrate_owned_shard skip <= watermark and a non-owner / new-owner stays coherent with the reclaiming owner (no re-pull, no un-reclaim on transfer); reconcile_owned_shard realigns a crashed owner. Drive exercise_shared_tier_gc (CLI durable-eventlog-shared-gc) + a crash-window reconcile test. D11 verdict: mechanism delivered + drive-tested for BOTH the single-writer local (PVC) AND the shared-medium topology, unblocking the Stage-C durable-primary consideration for the deployed shape. Remaining is operational: worker periodic-invocation wiring + an in-cluster GC soak.

Related

Clone this wiki locally