Skip to content

Design Execution Chain Store

Kadyapam edited this page Sep 30, 2026 · 6 revisions

The execution-partitioned chain store

ehdb-l0's DurableChainStore plus ChainPopulator. Added for the event-store redesign (RFC noetl/ai-meta#355), whose diagnosis is that the event store is ordered by a global sequence and partitioned by nothing, so "read this execution's chain" costs O(total store) — and a timed-out chain read is indistinguishable from a short chain, which is what makes the drive re-drive.

Status: v0.4.5 — the chain is ordered by FOLLOWING LINKS, never by event_id. Shipped and on main; adopted by server v3.118.0, which is deployed to production INERT (all three NOETL_CHAIN_* gates OFF). The duration-sized re-ramp is the remaining gate.

⚠⚠ A CAPABILITY WAS DELIBERATELY REMOVED IN v0.4.5, and the cost is real. v0.4.3/v0.4.4 could serve a post-restart execution carrying several NULL-prev rows by recomputing the chain edge from log order. That worked — and it was the same trusting of ORDER BY event_id that produced false divergence on prod. It was never recovery; it was a guess that usually looked right. A forked execution is now refused and the caller falls through to Postgres. Safe, and it means the store covers FEWER executions until the write path stops producing forks: 2 of 63 recent executions plus the whole legacy era.

⚠ Which makes "comparator green" insufficient on its own — a source that refuses everything also reports zero divergence. Read coverage alongside it.

⚠⚠ A CORE ASSUMPTION OF THIS DESIGN IS FALSE. populate_from_log recomputes edges from log order and treats that order as a stable prefix that only ever grows at the tail. An interior insert breaks it, and every position after the insert then reads as a content conflict. Fixing that is noetl/ai-meta#362 and it is a design question, not a patch: no available column is commit-ordered — created_at is stamped at MINT time, so it always agrees with event_id order and cannot detect the reordering.

Prior status, withdrawn: shipped in v0.4.4 and armed in production since 2026-09-29 17:30Z (server v3.117.2). The prod re-ramp measured 772 comparisons, 0 failures, 100.00% agreement at up to 6 concurrent executions, with all 80 driven executions COMPLETED and Postgres authoritative throughout. See noetl/ai-meta#357 and #360.

⚠ An earlier ramp on v0.4.3 was rolled back at 13:28Z on what looked like 6.3% divergence (5 in 79) and was in fact a false alarm — see Staleness is not a conflict below. Nothing was wrong: 60 partitions, 0 persistent mismatches, 24/24 executions COMPLETED.

⚠ Version floor: a reader built on this store needs v0.4.4 or later. v0.4.3 has populate_from_log and coverage, but it classifies a log snapshot behind the store as a divergence — a false alarm that fails a comparator gate under concurrent reads, and the reason the first prod ramp was rolled back. v0.4.2 has the store and the watermark but not populate_from_log or the coverage record, so it cannot tell a truncated partition from a whole one; v0.4.1 must not be adopted at all, because its populator could create authority out of a failure. FORMAT_VERSION is 1 across v0.3.0/v0.4.2/v0.4.3/v0.4.4, so the on-disk layout gate does not refuse an existing store when moving between them.

Key layout

chain/<execution_id>/ev/<event_id>          -> the serialised ChainEvent
chain/<execution_id>/seq/<020-pad exec_seq> -> that event's id
chain/<execution_id>/head                   -> "<seq>:<event_id>"
chain/<execution_id>/wm                     -> "<first_seq>:<through_seq>"

head carries its own sequence so one read serves both the I1 head check and the next sequence number. It did not originally, and the append was O(k) in the partition's length as a result — 308s → 45s on the scaling test once it did.

The watermark, and the three states it exists to separate

An empty partition and an execution the store has never been told about are the same bytes on disk. Both answer "no events". A reader that cannot tell them apart concludes a running execution has not started.

So wm marks an execution as populated, and chain_if_authoritative is the function a reader must call instead of chain directly:

state guarded read meaning
no wm None cannot answer — fall through to the authoritative log
wm present, agrees with content Some(events) answerable
wm present, ahead of content None the log moved on without us

None never means "no events". Some(vec![]) does.

⚠⚠ Never create authority out of a failure

The append/watermark ordering is asymmetric, and the asymmetry is a fix, not an accident:

  • never-seen execution → append FIRST, mark after. A crash between them leaves events with no marker: invisible, but safe (it reads as cannot answer), and nothing is lost because the authoritative log still has them.
  • already-authoritative execution → mark FIRST, append after. There is already content, so a crash must not make the store under-report it.

The first draft marked before every append. On an execution the store had never seen, an append that then failed left Authoritative{1,1} over zero events, so the guarded read returned Some([]) and a reader concluded a running execution had no events — the exact cliff the watermark exists to prevent, reintroduced through the failure path.

And it is the common case, not an edge one: arming the populator mid-flight means the first row seen for an already-running execution carries a prev the empty store does not have, so the append fails with NotHead on essentially every in-flight execution.

v0.4.1 carries that bug and must not be adopted. v0.4.2 is the first tag with the asymmetric ordering.

⭐⭐ populate_from_log — the entry point a reader must use

ChainPopulator::populate_from_log takes an execution's events as the authoritative log ordered them and recomputes the chain edge from that order. LogEvent has no prev_event_id field at all, and that absence is the design: the column is NULL on 643,420 of 645,677 kind rows, so a populator trusting it would root the partition wherever the server last restarted.

It is idempotent and incremental — the stored chain must be a prefix of the log, and only the remainder is appended. A disagreement returns FromLog::Diverged and writes nothing; an immutable chain is not silently rewritten to match a new reading.

The coverage record

chain/<execution_id>/cov  ->  "<root_event_id>:<total_in_log>"

The watermark answers "was this populated?". It cannot answer "was it populated from the BEGINNING?" — its first_seq is the store's sequence, 1 for any fresh partition regardless of where the execution began. That is exactly how a chain beginning at an execution's third event read as authoritative.

chain_if_authoritative refuses any partition without coverage, so a partition built by the per-event populate path is never served — deliberately, because that path cannot know whether it saw the execution's first event.

⚠⚠ The two measured defects, and how each was closed

Measured in kind 2026-09-28 with Postgres as an independent oracle (noetl/ai-meta#357).

Shared cause: the consumer stamps prev_event_id from an in-memory head map with no DB hydration, so after a restart the next event for a running execution carries prev = NULL — and the store cannot tell that pseudo-root from a genuine one. In the kind database, 534 of 595 executions already carry more than one null-prev root.

  • Mid-flight arming truncates, confidently. Postgres 3 events, store 1, authoritative first=1 through=1, guarded read Some(1) with root_prev=None. first=1 is true of the store and false of the execution, and the answer is indistinguishable from a complete chain: contiguous sequences, the gap check passes, the root reports no predecessor. CLOSED by populate_from_log + the coverage record: the partition is built from the log's first event, so it cannot be short. ⚠ Note the root signal deliberately is not an event-type name — only 62 of 595 executions start with playbook_started; 533 start with playbook.initialized, so a check built on the first guess would have been inert for 90% of them. The root is "position 0 in the log", which is correct across both.
  • A restart freezes the chain while the watermark advances. Postgres 7, store 6, watermark 1:7. Every later event is rejected for the same not-the-head reason, so the prefix is permanently stale and still authoritative. Fixed (ehdb#374) — the guarded read now reads that disagreement and refuses, because it is the only evidence the store has that the log moved on without it. The source of the pseudo-root is fixed on the consumer side too: a cold in-memory head map now hydrates from durable storage before stamping.

⚠⚠ Staleness is not a conflict — and the same shape is a second writer

Closed in v0.4.4 (ehdb#376, noetl/ai-meta#360) after it failed a prod ramp.

A caller reads the log, then takes the store lock. It cannot hold a std::sync::Mutex across a database await, so under concurrent reads of one execution the caller that read first can reach the lock second and find the store already ahead of its snapshot. The prefix check called that Diverged.

It is not a disagreement. populate_from_log is this store's only writer, so every stored event arrived from some read of this same log — a store longer than the caller's snapshot means a fresher snapshot was already applied. Hence FromLog::StaleLog { stored_len, log_len }: benign, is_healthy(), and the partition stays servable.

⚠⚠ But the crate must NOT decide this, and that distinction is the whole design. "The store holds more than my snapshot" is also exactly how a second writer looks — a different populator putting events in the store that are not in the log at all (noetl/ai-meta#358, a measured 13.3% real divergence). From a single snapshot the two are indistinguishable, and absorbing the second would be far worse than the bug being fixed.

So StaleLog is an unresolved report, not a verdict. This crate cannot re-read the log; its consumer can, and does — if the log catches up it was staleness, and if the log still ends short while the store runs ahead, those events are not in the authoritative log and the consumer records a divergence.

⚠ The first version of this fix reasoned "populate_from_log is the sole writer, therefore a longer store is always staleness" and shipped nothing to catch the second case. The #358 regression test failed on the first run. The reasoning is sound about this crate and false about the system, which is exactly why the verdict belongs one layer up.

The alarm is NOT turned off

A content conflict — a position both sides hold, filled differently — is still Diverged, including when the log is shorter. Classifying on length alone would have masked a real conflict behind staleness, so the distinction is on content.

A detected content conflict now also revokes serve-trust by dropping the coverage record. This function reports rather than repairs, so previously it wrote nothing — and the still-matching coverage meant any reader calling chain_if_authoritative alone was handed a chain we had just proved wrong. Events and the watermark are left intact: this revokes trust, it does not destroy the evidence of what disagreed.

Coverage is monotonic — the other half of the same bug

The guarded read compares chain.len() against cov_total, so a stale snapshot recording a smaller total makes the partition unservable just as surely as a false divergence did. write_coverage now refuses to shrink a total under the same root. Fixing the classification alone would only have converted a false Diverged into a false refusal.

⭐ That guard survived its first mutation test, which is why it is worth naming: the test meant to cover it populated 5 events then read 2, which returns StaleLog early — before write_coverage is ever reached. It never touched the guard it was named for. A test whose name promises one property and checks another is worse than no test.

Completeness is about internal gaps, not truncation

chain_is_complete compares the walk against the partition's span, not a record count. That catches a hole in the middle. It does not catch a chain whose beginning is missing — which is why the first defect above is not detectable from inside the store and needs the coverage marker.

Replication

apply_replicated is the follower path and deliberately does not enforce the head: async replication delivers out of order, so a follower must be able to accept seq 5 before seq 3. append enforces the head; the two are not interchangeable. This asymmetry was found by a mutation test that could not construct a gap at all — because every write path refused one.

Related

Clone this wiki locally