Repository navigation
Design Execution Chain Store
ehdb-l0's DurableChainStore plus ChainPopulator. Added for the
event-store redesign (RFC noetl/ai-meta#355),
whose diagnosis is that the event store is ordered by a global sequence and
partitioned by nothing, so "read this execution's chain" costs O(total store)
— and a timed-out chain read is indistinguishable from a short chain, which is
what makes the drive re-drive.
Status: v0.4.5 — the chain is ordered by FOLLOWING LINKS, never by event_id.
Shipped and on main; adopted by server v3.118.0, which is deployed to
production INERT (all three NOETL_CHAIN_* gates OFF). The duration-sized
re-ramp is the remaining gate.
⚠⚠ A CAPABILITY WAS DELIBERATELY REMOVED IN v0.4.5, and the cost is real.
v0.4.3/v0.4.4 could serve a post-restart execution carrying several NULL-prev rows
by recomputing the chain edge from log order. That worked — and it was the same
trusting of ORDER BY event_id that produced false divergence on prod. It was never
recovery; it was a guess that usually looked right. A forked execution is now
refused and the caller falls through to Postgres. Safe, and it means the store
covers FEWER executions until the write path stops producing forks: 2 of 63 recent
executions plus the whole legacy era.
⚠ Which makes "comparator green" insufficient on its own — a source that refuses everything also reports zero divergence. Read coverage alongside it.
⚠⚠ A CORE ASSUMPTION OF THIS DESIGN IS FALSE. populate_from_log recomputes edges from log order and treats that order as a stable prefix that only ever grows at the tail. An interior insert breaks it, and every position after the insert then reads as a content conflict. Fixing that is noetl/ai-meta#362 and it is a design question, not a patch: no available column is commit-ordered — created_at is stamped at MINT time, so it always agrees with event_id order and cannot detect the reordering.
Prior status, withdrawn: shipped in v0.4.4 and armed in production since 2026-09-29 17:30Z
(server v3.117.2). The prod re-ramp measured 772 comparisons, 0 failures,
100.00% agreement at up to 6 concurrent executions, with all 80 driven
executions COMPLETED and Postgres authoritative throughout.
See noetl/ai-meta#357 and
#360.
⚠ An earlier ramp on v0.4.3 was rolled back at 13:28Z on what looked like
6.3% divergence (5 in 79) and was in fact a false alarm — see Staleness is
not a conflict below. Nothing was wrong: 60 partitions, 0 persistent
mismatches, 24/24 executions COMPLETED.
⚠ Version floor: a reader built on this store needs v0.4.4 or later.
v0.4.3 has populate_from_log and coverage, but it classifies a log snapshot
behind the store as a divergence — a false alarm that fails a comparator gate
under concurrent reads, and the reason the first prod ramp was rolled back.
v0.4.2 has the store and the watermark but not populate_from_log or the
coverage record, so it cannot tell a truncated partition from a whole one;
v0.4.1 must not be adopted at all, because its populator could create authority
out of a failure. FORMAT_VERSION is 1 across v0.3.0/v0.4.2/v0.4.3/v0.4.4, so the
on-disk layout gate does not refuse an existing store when moving between them.
chain/<execution_id>/ev/<event_id> -> the serialised ChainEvent
chain/<execution_id>/seq/<020-pad exec_seq> -> that event's id
chain/<execution_id>/head -> "<seq>:<event_id>"
chain/<execution_id>/wm -> "<first_seq>:<through_seq>"
head carries its own sequence so one read serves both the I1 head check
and the next sequence number. It did not originally, and the append was O(k) in
the partition's length as a result — 308s → 45s on the scaling test once it did.
An empty partition and an execution the store has never been told about are the same bytes on disk. Both answer "no events". A reader that cannot tell them apart concludes a running execution has not started.
So wm marks an execution as populated, and chain_if_authoritative is the
function a reader must call instead of chain directly:
| state | guarded read | meaning |
|---|---|---|
no wm
|
None |
cannot answer — fall through to the authoritative log |
wm present, agrees with content |
Some(events) |
answerable |
wm present, ahead of content |
None |
the log moved on without us |
None never means "no events". Some(vec![]) does.
The append/watermark ordering is asymmetric, and the asymmetry is a fix, not an accident:
- never-seen execution → append FIRST, mark after. A crash between them leaves events with no marker: invisible, but safe (it reads as cannot answer), and nothing is lost because the authoritative log still has them.
- already-authoritative execution → mark FIRST, append after. There is already content, so a crash must not make the store under-report it.
The first draft marked before every append. On an execution the store had
never seen, an append that then failed left Authoritative{1,1} over zero
events, so the guarded read returned Some([]) and a reader concluded a
running execution had no events — the exact cliff the watermark exists to
prevent, reintroduced through the failure path.
And it is the common case, not an edge one: arming the populator mid-flight
means the first row seen for an already-running execution carries a prev the
empty store does not have, so the append fails with NotHead on essentially
every in-flight execution.
v0.4.1 carries that bug and must not be adopted. v0.4.2 is the first tag
with the asymmetric ordering.
ChainPopulator::populate_from_log takes an execution's events as the
authoritative log ordered them and recomputes the chain edge from that
order. LogEvent has no prev_event_id field at all, and that absence is the
design: the column is NULL on 643,420 of 645,677 kind rows, so a populator
trusting it would root the partition wherever the server last restarted.
It is idempotent and incremental — the stored chain must be a prefix of the
log, and only the remainder is appended. A disagreement returns
FromLog::Diverged and writes nothing; an immutable chain is not silently
rewritten to match a new reading.
chain/<execution_id>/cov -> "<root_event_id>:<total_in_log>"
The watermark answers "was this populated?". It cannot answer "was it
populated from the BEGINNING?" — its first_seq is the store's sequence,
1 for any fresh partition regardless of where the execution began. That is
exactly how a chain beginning at an execution's third event read as
authoritative.
chain_if_authoritative refuses any partition without coverage, so a partition
built by the per-event populate path is never served — deliberately, because
that path cannot know whether it saw the execution's first event.
Measured in kind 2026-09-28 with Postgres as an independent oracle (noetl/ai-meta#357).
Shared cause: the consumer stamps prev_event_id from an in-memory head
map with no DB hydration, so after a restart the next event for a running
execution carries prev = NULL — and the store cannot tell that pseudo-root
from a genuine one. In the kind database, 534 of 595 executions already carry
more than one null-prev root.
-
Mid-flight arming truncates, confidently. Postgres 3 events, store 1,
authoritative first=1 through=1, guarded readSome(1)withroot_prev=None.first=1is true of the store and false of the execution, and the answer is indistinguishable from a complete chain: contiguous sequences, the gap check passes, the root reports no predecessor. CLOSED bypopulate_from_log+ the coverage record: the partition is built from the log's first event, so it cannot be short. ⚠ Note the root signal deliberately is not an event-type name — only 62 of 595 executions start withplaybook_started; 533 start withplaybook.initialized, so a check built on the first guess would have been inert for 90% of them. The root is "position 0 in the log", which is correct across both. -
A restart freezes the chain while the watermark advances. Postgres 7,
store 6, watermark
1:7. Every later event is rejected for the same not-the-head reason, so the prefix is permanently stale and still authoritative. Fixed (ehdb#374) — the guarded read now reads that disagreement and refuses, because it is the only evidence the store has that the log moved on without it. The source of the pseudo-root is fixed on the consumer side too: a cold in-memory head map now hydrates from durable storage before stamping.
Closed in v0.4.4 (ehdb#376,
noetl/ai-meta#360) after it failed a
prod ramp.
A caller reads the log, then takes the store lock. It cannot hold a
std::sync::Mutex across a database await, so under concurrent reads of one
execution the caller that read first can reach the lock second and find the
store already ahead of its snapshot. The prefix check called that Diverged.
It is not a disagreement. populate_from_log is this store's only writer, so every
stored event arrived from some read of this same log — a store longer than the
caller's snapshot means a fresher snapshot was already applied. Hence
FromLog::StaleLog { stored_len, log_len }: benign, is_healthy(), and the
partition stays servable.
⚠⚠ But the crate must NOT decide this, and that distinction is the whole design. "The store holds more than my snapshot" is also exactly how a second writer looks — a different populator putting events in the store that are not in the log at all (noetl/ai-meta#358, a measured 13.3% real divergence). From a single snapshot the two are indistinguishable, and absorbing the second would be far worse than the bug being fixed.
So StaleLog is an unresolved report, not a verdict. This crate cannot re-read
the log; its consumer can, and does — if the log catches up it was staleness, and if
the log still ends short while the store runs ahead, those events are not in the
authoritative log and the consumer records a divergence.
⚠ The first version of this fix reasoned "populate_from_log is the sole writer,
therefore a longer store is always staleness" and shipped nothing to catch the
second case. The #358 regression test failed on the first run. The reasoning is
sound about this crate and false about the system, which is exactly why the
verdict belongs one layer up.
A content conflict — a position both sides hold, filled differently — is still
Diverged, including when the log is shorter. Classifying on length alone would
have masked a real conflict behind staleness, so the distinction is on content.
A detected content conflict now also revokes serve-trust by dropping the
coverage record. This function reports rather than repairs, so previously it wrote
nothing — and the still-matching coverage meant any reader calling
chain_if_authoritative alone was handed a chain we had just proved wrong. Events
and the watermark are left intact: this revokes trust, it does not destroy the
evidence of what disagreed.
The guarded read compares chain.len() against cov_total, so a stale snapshot
recording a smaller total makes the partition unservable just as surely as a
false divergence did. write_coverage now refuses to shrink a total under the same
root. Fixing the classification alone would only have converted a false
Diverged into a false refusal.
⭐ That guard survived its first mutation test, which is why it is worth naming: the
test meant to cover it populated 5 events then read 2, which returns StaleLog
early — before write_coverage is ever reached. It never touched the guard it was
named for. A test whose name promises one property and checks another is worse than
no test.
chain_is_complete compares the walk against the partition's span, not a
record count. That catches a hole in the middle. It does not catch a chain
whose beginning is missing — which is why the first defect above is not
detectable from inside the store and needs the coverage marker.
apply_replicated is the follower path and deliberately does not enforce
the head: async replication delivers out of order, so a follower must be able to
accept seq 5 before seq 3. append enforces the head; the two are not
interchangeable. This asymmetry was found by a mutation test that could not
construct a gap at all — because every write path refused one.
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)