Repository navigation
Design Durable EventLog Backend
Status: slices 1–3 landed (2026-07-07). The durable segment store +
crash recovery is merged in ehdb-reference::durable_eventlog
(ehdb#253); execution-affinity
single-writer routing is merged in ehdb-reference::affinity +
ehdb-reference::durable_eventlog_affinity
(ehdb#255); the shared /
object-store segment tier is merged in
ehdb-reference::durable_eventlog_shared
(ehdb#258); worker wiring
(noetl/worker#171) + kind soak
landed; and interest-based segment GC — the R1 / D11 gap — landed in
ehdb-reference::durable_eventlog (local reclamation) and
ehdb-reference::durable_eventlog_shared (coherent shared-medium reclamation via
a cross-replica reclaim watermark; see
Segment GC +
Shared-tier reclamation).
All behind selectors disabled by default (local_reference stays the default;
ownership defaults to single-owner; the shared tier is opt-in;
NOETL_EHDB_EVENTLOG_GC=off by default). Remaining GC work is operational:
worker periodic-invocation wiring + an in-cluster soak.
Tracks: noetl/ehdb#254 (durable-backend program + slice checklist) · noetl/ehdb#241 (completion program) · Design: Event-Log Core Engine (Phase 6) · Prod Cutover Runbook §C (the durability gate this resolves) · Backend Configuration (Phase 10).
The Phase-6 engine shipped the contract (EventLogDriver) and one
implementation, LocalReferenceEventLogDriver, over a pod-local JSONL
file (local_reference). That is correct and safe for shadow — a
derived, disposable mirror — but the Phase-6 design note and the
prod-cutover runbook both flag it as not production-durable as the
authoritative store under primary:
"Production hardening (segmented append files + a sparse offset index for O(log n) sequence seeks, compaction, and cross-node replication) is deferred to the primary-serve phase; the local-reference substrate is the reference implementation of the contract, not the production disk format." — Design: Event-Log Core Engine
The runbook's §C durability gate is the hard blocker for Stage C (primary cutover) and names two failures of the pod-local file:
-
Restart durability. A pod-local
emptyDirlog is lost on pod restart/reschedule. The authoritative event log cannot live on ephemeral pod storage. - Multi-replica divergence. If more than one pod on the flagged pool appends, each writes its own pod-local log → multiple divergent "authoritative" logs. It must be pinned to a single writer, or the backend must be shared.
This design builds a backend that is durable (survives restart) and
lays the groundwork for coherent (single-writer-per-shard). It is the
runbook's "wait for a durable/shared EHDB log backend beyond
local_reference" resolution option, made concrete.
Same as Phase 6: event authorship is unchanged — the gateway/server remain the gatekeeper of what enters the log. This backend is the disk-and-index under an append the producer already authorized. Platform event log only — never business data. Secret-free outcomes.
Two axes: the on-disk format (this slice) and the medium it lives on (a deployment decision — see Assumptions below).
<store root>/
seg-0000000000000001.eslog append-only segment, rolls over at a size cap
seg-0000000000000002.eslog ...
Each frame in a segment:
| magic u32 (0xE5DB0001) | body_len u32 | crc32 u32 | body: JSON(SegmentFrame) |
4 bytes 4 bytes 4 bytes body_len bytes
SegmentFrame is one of Event { global_sequence, execution_id, transaction_id, payload }, ConsumerCreate { consumer }, or Ack { consumer, sequence } — events and durable-consumer state share the one
segment stream, so a single replay rebuilds the whole in-memory index
(including consumer cursors) from disk.
Why this shape over the incumbent single JSONL file:
-
Rollover — a new segment starts once the active one would exceed a
size cap (
DEFAULT_SEGMENT_MAX_BYTES= 8 MiB); a single frame never spans two segments, so the offset index stays valid. A segment can be archived / GC'd / replicated as a unit without rewriting the whole log. -
CRC + magic framing — distinguishes a torn tail (a crash mid-append
left a truncated frame at EOF → discarded on recovery) from bit-rot (a
complete frame whose CRC/magic is wrong → a hard error, matching the
incumbent JSONL log's reject-corrupt-records stance). This is the
zero-loss recovery contract: every append that
fsync'd before returning survives; a partial write that never returned success is discarded.
The store keeps, rebuilt on open:
-
events: Vec<(segment_id, byte offset)>— indexilocates the frame for global sequencei + 1. O(1) locate by global sequence. -
by_execution: HashMap<execution_id, Vec<global_seq>>— per-execution scope without a full scan. -
consumer_acks: HashMap<consumer, acked_seq>+consumers_seen— durable cursors.
The index holds locations, not payloads. A read locates the frame via
the index and cold-loads the payload bytes from the segment file on
demand. Index memory is O(events) in small fixed-size entries, not
O(total payload bytes) — the bounded-WAL-index property
noetl/ai-meta#166 chases (the
unbounded in-RAM WAL index that OOM'd the system pool).
O(1) open-for-append — the checkpoint sidecar (#267)
The worker rebuilds the durable stack per op (a stateless boundary), so
every mirrored append pays a fresh open. Rebuilding the offset index by
replaying every segment on each open is O(segment) — on a ~20 MB store that
was ~0.5 s/op, the dominant deployed cost after #266 made the shared publish
O(delta) (ehdb#268).
Fix: a small checkpoint.json sidecar next to the segments holding
{event_count, active_segment_id, active_len, consumers_seen, consumer_acks} —
exactly the O(1) state the append / ack hot path needs (next global sequence,
where to write, durable cursors), not the offset index. It is rewritten
after each mutating op, strictly after the frame fsync.
-
Open-for-append loads the checkpoint and returns O(1) — no replay. The
offset index (
events/by_execution) is materialised lazily on the first read (ensure_index_loaded, which runs the full replay = CRC + gapless integrity check).scan_global/read_executiontherefore take&mut self; read-only cold-loads still replay eagerly at open (they need the full index). -
Correctness — the checkpoint is a cache, never truth. A missing / stale /
inconsistent checkpoint falls back to a full replay (replay-is-truth) and the
checkpoint is rewritten. The consistency check is a strict active-segment
length anchor: the highest segment id must equal
active_segment_idand its on-disk length must equalactive_len. Because the checkpoint is written only after the framefsync, it can never name more durable data than the segments hold — a crash between the framefsyncand the checkpoint rewrite leaves the active segment longer thanactive_len, caught as inconsistent → replay recovers the extra fsync'd frame(s). No wrongglobal_sequenceis ever assigned. The checkpoint is written without its ownfsync(a lost rewrite just costs a one-time replay on the next open), so append stays a singlefsync. It mirrors the shared tier's resumable-digest sidecar (#266). - Deployed result: per-op durable mirror dropped from ~0.509 s (ehdb266) to ~4–16 ms (ehdb267) on the same ~20 MB kind store; micro-bench per-op-open+append is flat across S=100→10 000. See Design: Performance & Load Testing.
-
fsyncper append.write_frameappends the framed bytes thensync_data()before returning. The write is not acknowledged until it is on stable storage. -
Crash recovery / replay on restart.
openlists the segment files in id order and replays each frame into the index. A torn tail (truncated at EOF) is discarded and the segment file is truncated to the last intact frame (idempotent — a second reopen sees a clean file). Replay-is-truth: the in-memory state is a pure function of the durable segments. (A clean open-for-append skips this replay via the O(1) checkpoint sidecar above and replays only on a missing / stale / inconsistent checkpoint — replay stays the correctness backstop.) -
Zero-loss bar. Every event whose
appendreturned survives a restart, verbatim, in order, with per-execution scope intact and the durable consumer cursor preserved. Proven directly byexercise_durable_recovery.
DurableEventLogDriver implements the same EventLogDriver trait as
LocalReferenceEventLogDriver, so it is a drop-in behind the contract:
| Property | Guarantee |
|---|---|
| Global ordering | monotonic, gapless from 1 (next = events.len() + 1) |
| Per-execution scope |
by_execution index; scoped ordered reads |
| Tail / subscribe | durable consumer, created on first pull |
| Offset / ack | explicit ack-after-materialize, persisted Ack frame; cursor never moves backward |
| Append-only + immutable | frames never mutated; segments append-only |
| Replay is truth | in-memory index rebuilt from segments on open |
A dedicated test drives the same op sequence through both the durable
and local-reference drivers and asserts identical observable results
(sequence, created_stream, log_record_count, scan order, per-execution
scope).
Coherence under multiple replicas needs exactly one writer per shard. Slice 2 (ehdb#255) lands it, reusing the noetl/ai-meta#166 / #116 execution-affinity ownership so each shard is owned by exactly one replica = its sole writer.
The ownership hash is byte-identical to the worker/server.
ehdb-reference::affinity::shard_for_i64 is XxHash64(seed=0) over the
execution id's 8 little-endian bytes % shard_count — the same function as
noetl-worker src/sharding.rs::shard_for and noetl-server
sharding::shard_for (pinned twox-hash = "1.6", same major). NoETL
execution ids are snowflake i64s; EHDB stores them as decimal strings, so
shard_for_execution parses the string back to the i64 and routes through
shard_for_i64 — ownership agrees with the worker for every real execution.
ShardOwnership mirrors the worker's AffinityConfig: a replica owns exactly
its shard_index bucket, so a shard_count-replica pool partitions every
execution to exactly one writer.
The routing layer (durable_eventlog_affinity::AffinityRoutedEventLog,
one instance = one replica) maps each execution to a per-shard
DurableSegmentStore under <root>/shard-<NNNN>/:
-
append — the owner writes (
Routed::Served); a non-owner is refused with no side effect (Routed::NotOwner { owner_shard }) so the caller re-routes to the owner. No bytes written, no sequence consumed, so a re-route can never double-process. -
read (
read_execution/scan_shard) — the owner serves resident; a non-owner cold-loads the durable segments read-only (ServedBy::NonOwnerColdLoad). - tail / ack — durable-consumer writer state, so owner-only; a non-owner is refused.
The full EventLogDriver contract holds within each shard (gapless
per-shard global sequence, per-execution scope, replay-is-truth); an
execution's events always land in shard-<shard_of(execution)>. The owner
keeps its owned-shard stores resident behind a mutex for O(1) locate; a
cold-load is a fresh DurableSegmentStore::open_read_only over the shard's
directory — what a new owner or a non-owner read does, and what a pod restart
replays.
open_read_only is non-mutating by construction: it refuses every write
and never truncates a recovered torn tail (only the shard's single owner
repairs its own tail on its own writable open), so a non-owner read can never
mutate another owner's shard. That is the structural guarantee behind
single-writer.
The disk format (slice 1) is unchanged — the routing layer sits on top of the per-shard segment stores without reworking the segments.
Slice 2's non-owner cold-load reads the non-owner's own local disk — which,
on a different pod with different local storage, is empty. Slice 3
(ehdb#258,
ehdb-reference::durable_eventlog_shared) makes the cold-load pull from a
shared durable medium so a shard survives the loss of the writer's
pod-local disk — the runbook §C durability gate's "durable/shared EHDB log
backend beyond local_reference" resolution, made concrete.
The pluggable medium. SharedSegmentBackend is the seam: put_segment /
get_segment / list_segment_ids. FilesystemSharedBackend (a shared
directory — a ReadWriteMany/ReadWriteOnce PVC on kind) is the bootstrapping
impl; the same trait routes to EHDB's own durable object tier (Phase 8
object engine) later — the self-sufficiency end-state — with no change to the
routing/hydrate logic.
Fixed-width, digest-integrity-checked keys. A segment is addressed by
position — (shard, segment_id), both bounded integers — so the shared-
store object key is a constant 40 chars
(noetl.ehdb.seg.<shard:08x>.<segment_id:016x>). The key width is independent
of any payload, so it can never approach a subject-length cap on whatever
medium backs the trait. This deliberately avoids the trap the object tier hit
(noetl/ai-meta#234: hex-encoding
an arbitrary platform key into a NATS subject blew the 256-char subject cap).
Each object carries a byte-length + XxHash64 content digest, so a truncated /
corrupt shared object is a hard error on read, not a silently-shortened
replay. (Segments are position-keyed, not content-hashed: the active segment is
append-mutable, so a content-hash key would change every append and break
listing — the digest is integrity metadata, not the key.)
The routing. SharedTierEventLog (one instance = one replica) wraps the
slice-2 AffinityRoutedEventLog:
-
owner append — writes locally (slice-1/2 fast path,
fsync'd) then publishes the shard's segments to the shared store (idempotent: a sealed segment already published at its length is skipped; the growing active segment is always re-published). -
non-owner read (
read_execution/scan_shard) — cold-loads the shard's segments from the shared store into a scratch dir and opens a slice-1DurableSegmentStore::open_read_onlyover them. Replay is byte-identical + sequence-preserving — the fullEventLogDrivercontract holds because the materialized segments are the owner's segments. -
new owner inheriting a shard —
hydrate_owned_shardpulls the shard's segments from shared into the new owner's own empty local dir before it serves / appends, so a pod that never held the shard locally recovers it zero-loss and continues the sequence. This is the crash-recovery-from-shared path (a pod restart / reschedule onto a fresh node). - tail / ack — owner-only writer state; the consumer-create + ack frames are published to shared too, so a new owner inherits the durable cursor.
The durable-eventlog-shared --root <dir> [--shard-count <n>] selfcheck spins
up a shard_count-replica pool with separate local disks over one
shared store and proves owner-publish / non-owner-cold-load-from-shared /
crash-recovery-from-shared / shared-store-miss-reads-empty / parity. Exit 0
only when SharedTierReport::holds:
{"shared_tier_holds":true,
"report":{"shard_count":2,"executions":2,"shared_segments":2,
"owner_published_ok":true,"nonowner_coldload_from_shared_ok":true,
"crash_recovery_from_shared_ok":true,"shared_miss_ok":true,"parity_ok":true,
"divergence":null}}Bootstrapping vs end-state (the assumption to flag). The recommended
bootstrapping medium is a PVC-backed shared directory (FilesystemSharedBackend)
— durable-across-restart with the segment format unchanged. The
self-sufficiency end-state is EHDB's own object tier holding the segments
(a different SharedSegmentBackend impl behind the same trait). Open
assumptions: for a PVC, a ReadWriteMany (or single-writer ReadWriteOnce)
class on the target cluster; for the object-tier backend, object-store creds /
workload-identity for the owning pool. Worker wiring of the selected medium is
slice 4.
A new storage-backend axis, orthogonal to the Phase-10
Backend Configuration TierMode
(off/shadow/primary) axis:
| Axis | Env var | Values | Meaning |
|---|---|---|---|
| Tier mode | NOETL_EHDB_EVENTLOG |
off / shadow / primary
|
Whether EHDB serves the event-log tier |
| Storage backend | NOETL_EHDB_EVENTLOG_BACKEND |
local_reference / durable_segment
|
Which durable engine serves when it does |
EventLogStorageBackend::from_raw is fail-safe: only the exact token
durable_segment selects the durable backend; unset / empty / unrecognised
→ local_reference (the default). So an unknown value never silently
changes the authoritative store, and nothing changes until an operator
opts in. Coexistence + migration from local_reference is a
forward-only cutover: run durable_segment as the primary backend on a
durable medium; the local_reference log stays readable for the shadow
window.
Slices 1–5 never reclaim a sealed segment: under KeepAll + primary the
on-disk store grows without bound. The prod-durability sign-off recorded that
as residual-risk R1 / gap D11 — the Stage-C blocker (durable-primary
cutover cannot be authorised while the store grows unbounded). This section is
the resolution: an interest-based segment GC that reclaims segments a
durable consumer has already consumed, while preserving every crash-safety and
replay-is-truth invariant slices 1–5 established.
A sealed segment S is reclaimable when every durable consumer has acked
past S's last global sequence, respecting a min_retained_segments floor:
-
watermark = minover all durable consumers of their ack cursor. No consumer, or any consumer that has acked nothing, ⇒watermark = 0⇒ nothing is reclaimed — interest-based retention never drops an event no consumer has read (JetStream-style). A lagging consumer holds the watermark down; operators alert on lag rather than losing data. - Reclaim the contiguous prefix of sealed segments whose last global
sequence
<= watermark, keeping at leastmin_retained_segmentsmost-recent segments (including the active one, which is never reclaimable).
This is safe under the platform boundary because Postgres noetl.event is
the immutable full-history archive (never purged); the durable EHDB segment
store is a working/streaming substrate. Reclaiming events every consumer has
already projected does not lose history — it lives in noetl.event — it only
bounds the derived engine's disk.
Reclaiming leading segments raises the lowest retained sequence above 1. The
store now tracks a base offset (reclaimed_through, 0 for an un-GC'd
store): the offset index is base-shifted (events[i] locates global sequence
reclaimed_through + i + 1), a scan clamps its start to the base, and a read
of a reclaimed sequence reports reclaimed (absent) — distinct from
corruption. event_count stays the absolute tip (base + retained), so global
ordering is still monotonic-gapless from the base; the durable
EventLogDriver contract holds over the
retained range. When reclaimed_through == 0 every path is byte-identical
to the pre-GC backend, so an un-opted-in store is unchanged.
Reclamation is a durable, crash-atomic sequence:
-
Write-forward current durable-consumer state (
Ackframes at each consumer's cursor) into the active segment +fsync. This is the subtle part: theConsumerCreate/Ackframes that established a consumer's cursor may physically live in a segment about to be unlinked. Because every consumer has acked>= watermark >= reclaimed_seq, a re-assertedAckat the current cursor re-establishes both the consumer's existence and its cursor from surviving segments alone — so replay-is-truth still reconstructs full consumer state after the old segments are gone. -
Commit the
fsync'dreclaim.jsonmanifest (reclaimed_through_seq+reclaimed_through_segment) — the point of no return. Unlike the checkpoint sidecar (an optimization written withoutfsync), the manifest is durable truth: it isfsync'd, atomically renamed, and the directory entryfsync'd. - Unlink the reclaimed segment files.
- Replay to rebuild the index with the new base, and refresh the checkpoint sidecar.
A crash at any point recovers correctly:
-
Before the manifest commit ⇒ nothing was reclaimed; the full log replays
normally (
reclaimed_through = 0). Idempotent retry on the next GC. - After the manifest commit, before/during the unlinks ⇒ the manifest names the reclamation; on the next open a below-watermark segment still on disk is re-deleted (on both the replay path and the O(1) checkpoint-trust path, which calls the same leftover-cleanup), and the base offset is honored. A leftover is never read (it is below the base), so even a garbage leftover is harmless — it is unlinked, not replayed.
The manifest can never name reclamation that would lose data a surviving
segment does not carry, because it is written strictly after the
write-forward fsync. A fresh segment after an all-consumed store starts at
reclaimed_segment + 1, so a new segment id can never collide with (and be
re-deleted as) a reclaimed one.
A new retention axis, orthogonal to the tier-mode and storage-backend axes,
and disabled by default exactly like EventLogStorageBackend:
| Axis | Env var | Values | Meaning |
|---|---|---|---|
| Segment GC | NOETL_EHDB_EVENTLOG_GC |
off (default) / consumer_ack
|
Whether consumed sealed segments are reclaimed |
| Min-retained floor | NOETL_EHDB_EVENTLOG_GC_MIN_RETAINED_SEGMENTS |
integer (default 2) | Segments always kept even when fully consumed |
SegmentGcPolicy::from_raw is fail-safe: only consumer_ack / on /
enabled turn GC on; unset / empty / unrecognised leaves it off, so an
unknown value never silently starts dropping segments. GC applies only to the
durable_segment backend, and only when explicitly enabled.
ehdb-local-reference durable-eventlog-gc --root <dir> [--events <n>] [--segment-max-bytes <n>] [--min-retained-segments <n>] [--consumer <name>]
drives a full cycle: append --events, ack a durable consumer to ~3/4 of the
log, reclaim, and prove growth was bounded while the retained log stayed
gapless-from-base + zero-loss and a restart recovers. Exit 0 only when
SegmentGcDriveReport::holds. A representative run (40 events, 220-byte
segments) reclaimed 41 → 11 segments at watermark 30 with every invariant
holding. The eventlog_segment_gc criterion group measures a single
reclaim_segments call at pre-warmed sizes S — GC is off the append hot path
(run periodically, not per-op), and the interest watermark keeps the retained
set bounded, so a reclaim replays only the surviving segments, not the whole
history. Measured (4 KiB segments, ~3/4 acked): ~16.6 ms at S = 200, ~22.9 ms
at S = 800 — the cost tracks the retained set (the reclaim replays surviving
segments + fsyncs the manifest + write-forward), not the full history, and is
comfortably affordable for a periodic maintenance pass.
The reclamation above is local. On its own it is not coherent with the
shared segment tier (slice 3): a local unlink leaves the reclaimed segments
in the shared store, so cold_load / hydrate_owned_shard would re-pull them
and diverge from the owner (and a new owner would un-reclaim on transfer).
The shared-tier reclamation closes that.
The cross-replica reclaim watermark. A per-shard watermark object
(noetl.ehdb.rw.<shard>, monotonic) on the shared medium records the highest
reclaimed (seq, segment). SharedSegmentBackend gains
put_reclaim_watermark / reclaim_watermark / delete_segment
(default-erroring so a backend that can't persist a watermark refuses shared-tier
GC rather than losing coherence; FilesystemSharedBackend implements them). A
cold-load / hydrate materializes only segments above the watermark
(retained_shared_ids), so it can never re-pull a reclaimed segment — and the
base offset falls out of the first surviving frame on replay, exactly as locally.
SharedTierEventLog::reclaim_shard — watermark-first ordering (the crux):
-
Reconcile — adopt any already-committed watermark into the owner's local
store first (
reconcile_owned_shard→reclaim_to_segment). -
Plan the boundary from the owner's local consumer cursors
(
plan_reclaim_boundary) — no mutation. -
Commit the shared watermark — the cross-replica point of no return. From
here every reader skips
<= segment. -
Reclaim locally to the boundary (write-forward +
fsync'd manifest + unlink). -
Delete the shared objects
<= segment.
Committing the watermark before any deletion is what removes the divergence:
a non-owner / new-owner that reads mid-reclamation already skips the segments
about to be removed. A crash after step 3 leaves orphan local/shared objects
readers already skip — a bounded space leak the next reclaim_shard
re-attempts, never a correctness bug; the owner realigns via step 1 on its
next run or bring-up (reconcile_owned_shard). The only residual transient is
that a crashed owner may briefly over-serve its own un-reclaimed local
segments until it reconciles — benign (it shows more, never less; the data is
also in noetl.event), and self-healing.
Drive + validation. SharedTierEventLog::reclaim_shard +
exercise_shared_tier_gc (CLI durable-eventlog-shared-gc) prove, over a
multi-replica pool on one shared store, that after the owner reclaims shard 0
both a non-owner cold-load and a fresh new-owner hydrate return exactly
the owner's retained set (no re-pull, no un-reclaim on transfer), the shared
objects <= watermark are pruned (bounded shared disk), and a targeted
crash-window test shows reconcile_owned_shard realigns an owner whose local
segments are ahead of a committed watermark. A representative 2-shard run pruned
the shared store 40 → 10 objects at watermark 30 with every coherence
invariant holding. The eventlog_shared_tier_gc criterion group measures
reclaim_shard (local reclaim + one watermark commit + per-segment shared
deletes) at ~25 ms (S=50) / ~38 ms (S=150) — off the append hot path, a
periodic maintenance pass whose cost tracks the retained set plus a small
constant of shared ops.
Still off by default + not auto-invoked — the shared-tier GC runs only when
an operator opts in (NOETL_EHDB_EVENTLOG_GC=consumer_ack) and a caller invokes
reclaim_shard; the worker periodic-invocation wiring + in-cluster soak is the
next slice (as slices 1–3 landed engine-only ahead of slice-4 worker wiring).
The engine ships the reclamation; the worker invokes it periodically so an opted-in operator gets a self-bounding store with no manual step (noetl/worker segment-GC wiring, tracks noetl/ehdb#254).
- Who triggers. Each worker replica runs a periodic task for the shards it owns — only a shard's single writer may reclaim it (the execution-affinity invariant), so GC is inherently per-owner and needs no central coordinator.
-
Cadence.
NOETL_EHDB_EVENTLOG_GC_INTERVAL_SECS(default0= off). The task spawns only when every axis is opted in:durable_segmentbackend +NOETL_EHDB_EVENTLOG_GC=consumer_ack+ a positive interval + a data-plane durable contract. All default-off ⇒ a default worker is byte-identical and the task never even spawns. -
Back-pressure / concurrency. The worker builds the durable stack per op
(stateless boundary), so the GC task and the durable append path are two
writers to the same shard's files. A process-global per-shard advisory lock
(acquired by both paths) serializes them per shard — GC on one shard never
blocks appends on another — and closes a latent intra-replica append↔append
race. GC runs on a
spawn_blockingthread so itsfsync/unlink I/O never stalls the async runtime. Its per-shard hold is brief (tens of ms, per the benches) and infrequent (interval-driven). -
Observability. Each pass records to the
noetl_ehdb_eventlog_gc_*metric family (outcome=reclaimed/noop/error); a reclamation logs at INFO (low frequency, only on a real reclaim).ehdb-selfcheck durable-eventlog-gcreclaims the live owned shards + reports;--drive 1runs a self-contained shared-tier GC cycle on the real durable volume.
In-kind soak (2026-07-10, LOCAL kind, no GKE/prod). The GC-wired worker
(localhost/noetl-worker:v5.72.3-ehdbgc254) rolled onto the durable_segment
user pool (noetl-worker-rust, single-owner, /ehdb-durable PVC) with
NOETL_EHDB_EVENTLOG_GC=consumer_ack + NOETL_EHDB_EVENTLOG_GC_INTERVAL_SECS=30.
Confirmed in-cluster: the periodic task spawned (segment-GC task started … interval_secs=30); it ran healthy ticks recorded as
noetl_ehdb_eventlog_gc_ops_total{outcome="noop"} with last_ok=1 — correctly
reclaiming nothing on the unconsumed shadow store (the interest-based safety);
ehdb-selfcheck durable-eventlog-gc --drive 1 on the real PVC
(/ehdb-durable/gc-soak) reclaimed the shared store 40 → 10 objects at
watermark 30 with gc_holds=true — owner replay gapless-from-base, non-owner
cold-load coherent, new-owner hydrate coherent, no divergence; 0 worker
restarts, no crashloop, no impact on the live path.
Interest-based reclamation needs a durable consumer to ack past segments (the
watermark). Under the currently-deployed shadow event-log the durable store
is a disposable mirror that nothing consumes (no tail/ack site), so the
watermark stays 0 and the periodic task correctly reclaims nothing — a
noop, not a bug. Interest-based GC becomes effective when a durable consumer
acks: naturally under primary (the projection / read-model tier tails +
acks the event log to build read models), or via an explicit maintenance
consumer. Limits-based retention (below) bounds a shadow store without one.
A retention mode reclaims independent of consumer interest so a store with no
durable consumer (a shadow mirror) self-bounds. plan_reclaim folds interest +
retention into one effective boundary:
interest_ceiling = consumers.exist ? Some(min ack cursor) : None // None = no ceiling
retention_target = max_retained_segments ? Some(seq leaving N segments) : None
boundary = retention_target ? min(retention_target, interest_ceiling ?? u64::MAX)
: interest_ceiling ?? 0
-
Primary + consumer + retention →
min(consumer-cursor, retention-target)— never reclaim ahead of a consumer and bound the store. - Shadow (no consumer) + retention → retention drives; the store self-bounds.
-
Shadow + no retention →
0— reclaim nothing (the pre-retention default). -
min_retained_segmentsstays the hard floor (the retention cap is clamped up to it).
The only engine change is inside plan_reclaim — the watermark-first commit,
shared-tier cold-load/hydrate coherence, reclaim_to_segment, reconcile_owned_shard,
and the worker per-shard lock all consume the boundary unchanged, so retention
inherits every crash-safety + coherence invariant.
Keep-last-N (count-based), not max-age. Retention counts segments, which is intrinsic and survives hydrate / cold-load. Max-age is deliberately not implemented: the only per-segment timestamp is file mtime, which resets on the fresh files a new owner materializes, so max-age would mis-fire across ownership transfer / pod reschedule.
Config: NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS (unset ⇒ interest-only,
unchanged) + the worker's NOETL_EHDB_EVENTLOG_SEGMENT_MAX_BYTES rollover
tunable. Merged ehdb#271 +
worker#178.
Organic in-kind soak (2026-07-10 → re-confirmed 2026-07-11 on merged main,
LOCAL kind). The durable_segment user pool (noetl-worker-rust, shadow,
no consumer — consumers_seen: []) with
NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS=3 +
NOETL_EHDB_EVENTLOG_SEGMENT_MAX_BYTES=1024 + interval 20 s. Run on the
merged-main image localhost/noetl-worker:ehdb178184-merged (built off
merged worker main — retention engine ehdb#271
- knob worker#178 + periodic wiring worker#177), so this proves the shipped code, not a branch image.
The live shard-0 store rotated through 60+ ~1 KiB segments across event_count
25 476 → 25 656 while holding ≈ 2 segments — the periodic retention GC kept it
bounded with no consumer. A controlled burst made it explicit: a +180-event burst
pushed it to 14 segments (appends outpace the 20 s tick), then one tick later
(no more appends) it self-bounded back to 2 segments and held flat;
noetl_ehdb_eventlog_gc_ops_total{outcome="reclaimed"} climbed (1 → 3) while
{outcome="error"} stayed absent, last_degraded=0, no new restarts. The
reclaim manifest advanced durably (reclaimed_through_seq 25 354 → 25 464 → …)
and an independent read-only reopen replayed the retained log gapless-from-base —
replay + durable-recovery intact under organic reclamation with no consumer.
(The earlier 2026-07-10 run on the branch image v5.72.4-ehdbgcret254 measured
the same shape: 19 → 2 segments, ~22 MB → 2,147 B.)
D11 (segment GC) moves from GAP → mechanism delivered + drive-tested for
BOTH the single-writer local (PVC) topology AND the shared-medium topology.
The reclamation engine exists, is crash-safe, preserves replay-is-truth + the
gapless contract, keeps a non-owner / new-owner coherent with the reclaiming
owner across the shared medium, and is unit- + drive-tested. This unblocks the
Stage-C durable-primary consideration for the deployed topology (single-writer
system pool on a PVC with the shared tier) — the store no longer grows unbounded
under primary on either medium. The remaining work is operational, not
correctness: worker periodic-invocation wiring + an in-cluster GC soak.
Slice 4 makes the durable stack selectable + functional in worker-rust
(noetl/worker#171). The worker's
event-log tier constructs its storage engine from the backend axis instead of
hard-coding the JSONL reference driver.
Selection seam — worker-rust src/ehdb/eventlog_backend.rs. Resolves
NOETL_EHDB_EVENTLOG_BACKEND (EventLogStorageBackend::from_raw, fail-safe)
and, on durable_segment, builds the whole slice-1+2+3 stack:
SharedTierEventLog (slice 3, shared publish / cold-load) over
AffinityRoutedEventLog (slice 2, execution-affinity single-writer) over
per-shard DurableEventLogDriver segment stores (slice 1). Shard ownership is
built from the worker's own NOETL_SHARD_INDEX / NOETL_SHARD_COUNT
(byte-identical XxHash64 to crate::sharding::shard_for), so the event-log
shard a replica owns is the same shard as the drive it owns — the durable
single-writer aligns with the drive-pool affinity.
Durable paths. Derived from the configured local_reference log's parent
(<log-parent>/ehdb-durable/{local,shared,coldload}) unless overridden:
| Env var | Default | Meaning |
|---|---|---|
NOETL_EHDB_EVENTLOG_BACKEND |
local_reference |
Storage engine selector |
NOETL_EHDB_EVENTLOG_DURABLE_DIR |
<log-parent>/ehdb-durable |
Base for per-shard local stores + derived shared/coldload roots |
NOETL_EHDB_EVENTLOG_SHARED_DIR |
<base>/shared |
Shared-tier medium root (PVC / object-store) |
Where it dispatches. mirror_event — the per-event shadow mirror AND
the per-event primary append — routes through the selected backend; both
backends yield the same EventLogAppendOutcome, so the parity path is
backend-agnostic. A non-owned append under durable_segment is a neutral
EventLogOutcome::RoutedAway (correct single-writer behaviour; never fires on
local_reference / single-owner). PRIMARY_SERVE_ACTIVATED is unchanged and
the primary path stays gated.
Config + selfcheck. The worker's Phase-10 config resolution surfaces
eventlog_storage_backend in the ehdb-selfcheck config matrix. A new
ehdb-selfcheck durable-eventlog verb mirrors via the selected backend and, on
durable_segment, independently reopens the owning shard store read-only to
prove the appends replayed from durable segments (crash-recovery), not JSONL:
{"suite":"ehdb-durable-eventlog-selfcheck","ehdb":"enabled","mode":"shadow",
"storage_backend":"durable_segment","ok":true,"durable_replay_records":3,
"durable_replay_ok":true}Disabled-by-default, reversible, zero behavior change when unset. No env ⇒ byte-identical JSONL append; flip the env back to restore the incumbent with no redeploy. Still the derived EHDB fabric — event authorship unchanged.
The worker pins ehdb-reference at cca0d0d (slices 1-3 + object
subject-digest fix). The in-cluster live proof landed in slice 5 (below).
Slice 5 ran the durable backend under sustained real traffic in the local
kind-noetl cluster and proved crash recovery on a real pod restart —
2026-07-07, LOCAL kind only (no GKE / prod; repos/noetl + repos/server
untouched).
Setup. Built localhost/noetl-worker:v5.70.0 (the slice-4 durable
wiring; podman save + kind load image-archive — the podman-provider-safe
path) and rolled it onto the user data-plane pool noetl-worker-rust
(role worker, single replica ⇒ single-owner shard 0), adding to the existing
EHDB shadow env:
NOETL_EHDB_EVENTLOG_BACKEND=durable_segment
NOETL_EHDB_EVENTLOG_DURABLE_DIR=/ehdb-durable # a 2Gi ReadWriteOnce PVC
The durable dir is a PVC (ehdb-durable-soak, standard/local-path) so
segments survive a pod delete + reschedule — the substrate the runbook §C
recommendation calls for. ehdb-selfcheck config confirmed
eventlog_storage_backend: durable_segment, mode shadow, role worker,
secret_free.
Accumulation + segment rotation (real traffic). Drove thousands of real
executions (automation/pft_sql_probe_v2 + tests/large_tabular_result_test).
Real drive events landed in CRC-framed segments under
/ehdb-durable/local/shard-0000/, and the slice-3 shared tier published each
segment byte-identical to /ehdb-durable/shared/noetl.ehdb.seg.<shard>.<seg>
(+ .meta). Sustained traffic crossed the 8 MiB rollover boundary and
rotation fired in-cluster: seg-…0001.eslog sealed at 8 387 133 bytes
(just under DEFAULT_SEGMENT_MAX_BYTES = 8 388 608) and a fresh
seg-…0002.eslog opened as the active tail — the same rollover the slice-1
unit tests prove, now confirmed on real event bytes. No JSONL log was written
(the durable backend never writes /tmp/ehdb-shadow-ref.jsonl).
Metrics. noetl_ehdb_eventlog_ops_total{operation="mirror",outcome="mirrored"}
advanced monotonically past 11 731 with last_ok=1, last_degraded=0, and
no invalid / degraded / routed_away outcome ever emitted (single
owner serves every append). Zero worker restarts / crashloops during the soak;
CRC/index integrity held (a bad CRC is a hard error — none surfaced).
Crash recovery — real pod restart on the PVC. With the backlog drained,
captured the sealed segment's sha256 (c2ba3d33…, 8 387 133 B) + the active
tail, then force-deleted the pod (--grace-period=0 --force). A fresh pod
rescheduled on the same PVC and:
-
seg-…0001.eslogwas byte-identical (samesha256, same size) — the sealed segment survived the crash with zero corruption. -
seg-…0002.eslogcontinued growing — the reopened store replayed the existing segments, recognised the active tail, and kept appending (no reset toseg-0001, no sequence gap). -
ehdb-selfcheck durable-eventlogon the new pod independently reopened the shard store read-only and replayeddurable_replay_records: 11856from the segments alone — the crash-recovery replay, on a pod that never saw the original writes. (The verb'sok:false/parity_mismatchhere are harness artifacts of pointing the isolated-dir selfcheck — which asserts a fresh 3-event sequence — at a live populated shard; the live path's own parity held withlast_degraded=0across all 11 731 events.) - New appends continued the gapless monotonic global sequence across the
restart (
…11852, 11853, 11855); a further 30 real drives grew the active segment (+59 KB) withmirroredthe only outcome.
Backend equivalence note. The soaked image is tagged v5.70.0; the worker
pointer advanced concurrently to v5.70.1 (ehdb#172 KV/vector subject-digest
fix, ehdb-reference cca0d0d → 52120a7). That diff touches only
kv.rs / vector.rs — no durable_eventlog* change — so the durable
event-log backend is byte-identical between the two and this soak validates the
backend at the current pointer.
ehdb-local-reference gains durable-eventlog-recovery --root <dir> [--consumer <name>] — the crash-recovery drive: append a deterministic
event set + ack a durable-consumer cursor, reopen a fresh driver over the
same root (simulated pod restart), and prove zero-loss + gapless ordering
- per-execution scope + payload fidelity + durable-cursor survival. Exit 0
only when
DurableRecoveryReport::recovered()holds. Example output:
{"driver":"ehdb-durable-segment","recovered":true,
"report":{"appended":3,"zero_loss":true,"ordering_ok":true,"scope_ok":true,
"payloads_match":true,"cursor_survived":true,"pending_after_restart":2,
"divergence":null}}Slice 2 adds three affinity verbs:
-
durable-eventlog-affinity --root <dir> [--shard-count <n>]— the full single-writer drive: spin up ann-replica pool over one root, partition a deterministic execution set, and prove owner-writes / non-owner-refused / single-writer-invariant / non-owner-cold-load-read / crash-recovery-under- ownership. Exit 0 only whenAffinitySingleWriterReport::holds():{"single_writer_holds":true, "report":{"shard_count":2,"executions":4,"owner_appends":4, "nonowner_refusals":4,"owner_writes_ok":true,"nonowner_refused_ok":true, "single_writer_invariant":true,"coldload_read_ok":true, "crash_recovery_ok":true,"divergence":null}} -
durable-eventlog-affinity-append --root <dir> --shard-index <n> --shard-count <n> --execution-id <id> --transaction-id <id> --payload <text>— one routed append. Exit0when the owner serves it, exit6(distinct from the 3/4/5 engine-error codes) when refused as a non-owner, so a shell / kind-soak harness can assert the routing decision directly. -
durable-eventlog-affinity-read --root <dir> --shard-index <n> --shard-count <n> --execution-id <id> [--after <seq>] [--limit <n>]— one routed read; the JSONserved_byisowner_residentornon_owner_cold_load.
Slice 3 adds the shared-tier drive:
-
durable-eventlog-shared --root <dir> [--shard-count <n>]— the full shared-tier drive: ashard_count-replica pool with separate local disks over one shared store, proving owner-publish / non-owner-cold-load-from- shared / crash-recovery-from-shared (new owner, empty local disk) / shared-store-miss-reads-empty / parity. Exit 0 only whenSharedTierReport::holds.
The segment-GC slice adds the reclamation drive:
-
durable-eventlog-gc --root <dir> [--events <n>] [--segment-max-bytes <n>] [--min-retained-segments <n>] [--consumer <name>]— append--eventsunder a small segment size, ack a durable consumer to ~3/4 of the log, reclaim consumed sealed segments, and prove growth was bounded while the retained log stayed gapless-from-base + zero-loss and a restart recovers. Exit 0 only whenSegmentGcDriveReport::holds. -
durable-eventlog-shared-gc --root <dir> [--shard-count <n>]— ashard_count-replica pool over one shared store reclaims a shard's consumed sealed segments across BOTH local + shared, and proves a non-owner cold-load AND a fresh new-owner hydrate both see exactly the owner's retained set (no re-pull, no un-reclaim on transfer, shared disk bounded). Exit 0 only whenSharedTierGcReport::holds.
The on-disk format is durable; where that format lives is a deployment decision the runbook §C leaves open. The two viable media, with a recommendation:
-
PVC-backed volume (recommended for bootstrapping). Mount a
PersistentVolume on a single-writer StatefulSet-style pod and point the
store root at it. Simplest path to durability-that-survives-restart with a
single writer; the segment format is unchanged. Open assumptions: PVC
availability + a
ReadWriteOnceclass on the target cluster, and pinning the event-log-owning pool to a single writer (the system pool has run at 1 and 2 replicas historically — §C A4). - Object-store-backed segments (the self-sufficiency end-state). EHDB's own durable object tier (Phase 8 object engine) holds the segments so they are shared + cold-loadable by a new owner, aligning with the #166 Feather-state-shards-in- object-store + cold-load-on-miss machinery. Tension: EHDB aims to replace the external object store, so long-term the segments live on EHDB's own object tier — but for bootstrapping, a PVC (or the existing object store) is the pragmatic path. Open assumptions: object-store creds / workload-identity for the owning pool; the cold-load-on-miss read path is a later slice.
Recommendation: land the segment store + crash recovery now (this slice, PVC-friendly since it is just a directory of files), then add the execution-affinity single-writer routing + the shared object-store tier as the next slices, so a PVC deployment can be validated on kind before the shared-tier work.
Scope note (deliberate). This slice does not dedupe by
transaction_id (the incumbent JSONL transaction log rejects duplicate
transaction ids; the durable event store treats transaction_id as opaque
per-event metadata). Producer-side exactly-once by transaction id, if
required, is a follow-up — the real producer path assigns unique ids, and
the durability/ordering/scope contract this slice delivers does not depend
on it.
- Segment store + crash recovery — ✅ landed (ehdb#253).
- Execution-affinity single-writer routing — ✅ landed (ehdb#255): XxHash64 ownership (byte-identical to worker/server) so each shard has exactly one writer; non-owner writes refused (route-to-owner), non-owner reads cold-load read-only.
-
Shared / object-store segment tier — ✅ landed
(ehdb#258):
SharedSegmentBackendtrait (filesystem/PVC now, EHDB object tier later) +SharedTierEventLog(owner publishes segments to shared; non-owner cold-loads / new owner hydrates from shared); fixed-width digest-integrity-checked segment keys. -
Worker wiring — ✅ landed
(noetl/worker#171):
NOETL_EHDB_EVENTLOG_BACKEND=durable_segmentresolved inworker-rust src/ehdb,ehdb-selfcheck durable-eventlogverb in the image. See Worker wiring below. -
Kind soak — ✅ landed (2026-07-07): worker
v5.70.0rolled onto thenoetl-worker-rustpool withNOETL_EHDB_EVENTLOG_BACKEND=durable_segmenton a PVC; thousands of real drives accumulated CRC-framed segments, rotation fired in-cluster (seg-0001 sealed at 8 387 133 B → seg-0002), metrics advanced past 11 731 with zeroinvalid/degraded, and crash recovery on a real pod delete + reschedule proved the sealed segment byte-identical + a read-only reopen replaying 11 856 records with a gapless continuing sequence. See Kind soak. -
Prod-durability sign-off — the §C gate cleared: durable + single-
writer proven, then Stage C is unblocked (still user-gated, per tier).
Decision package drafted 2026-07-07:
Sign-off — Durable Event-Log Backend, Prod Durability
(go/no-go checklist, evidence bundle, residual-risk register, extended
durable-shadow → durable-primary rollout sequence). **Recommendation: GO to
a prod durable-SHADOW soak; Stage C durable-PRIMARY stays gated on that soak
- a segment-GC decision (R1).** Awaiting user go/no-go; ehdb#254 slice-6 box stays unchecked until the actual prod sign-off.
-
Segment GC (the R1 / D11 gap) — ✅ local reclamation landed: interest-based
(consumer-ack watermark) reclamation of consumed sealed segments, base-offset
(retained log gapless-from-
reclaimed+1), durablereclaim.jsonmanifest as the crash-atomic commit point, write-forward of consumer state, and amin_retained_segmentsfloor — all behindNOETL_EHDB_EVENTLOG_GC=consumer_ack(off by default). See Segment GC above. -
Shared-tier reclamation (coherence across the shared medium) — ✅ landed:
a monotonic per-shard cross-replica reclaim watermark on the shared medium
(
put_reclaim_watermark/reclaim_watermark/delete_segment);SharedTierEventLog::reclaim_shardreclaims local + shared watermark-first (commit the watermark before any delete), socold_load/hydrate_owned_shardskip<= watermarkand a non-owner / new-owner stays coherent with the reclaiming owner (no re-pull, no un-reclaim on transfer);reconcile_owned_shardrealigns a crashed owner. Driveexercise_shared_tier_gc(CLIdurable-eventlog-shared-gc) + a crash-window reconcile test. D11 verdict: mechanism delivered + drive-tested for BOTH the single-writer local (PVC) AND the shared-medium topology, unblocking the Stage-C durable-primary consideration for the deployed shape. Remaining is operational: worker periodic-invocation wiring + an in-cluster GC soak.
-
Sign-off — Durable Event-Log Backend, Prod Durability — the slice-6 decision package (go/no-go checklist + residual-risk register).
-
Design: Event-Log Core Engine (Phase 6) — the contract this backend implements.
-
Prod Cutover Runbook §C — the durability gate this resolves.
-
Backend Configuration (Phase 10) — the tier-mode axis this storage-backend axis is orthogonal to.
-
Roadmap — Phase 9 primary-serve trajectory.
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)