Repository navigation
Program EHDB NATS Takeover
Status: L0 foundation COMPLETE (D1–D10); L1 NATS-takeover ✅ T4 COMPLETE — the EHDB feed is the PRODUCTION command bus on shastaratech prod as of 2026-07-27 (80/80, 0 loss, 0 dup). NATS remains installed as the rollback path; T5 (delete NATS) is held on the human. This page tracks EHDB's evolution into a layered platform. The NATS takeover is now L1 of that platform, not the whole program — it sits on an L0 noetl-native replicated object store that provides durability/replication.
Topology (c) per-shard-writer-as-broker: the per-shard writer owns the
durable log and owns delivery; the control plane only publishes to the
writer; workers subscribe directly (one hop = NATS parity). Built in the new
ehdb-feed crate over the L0 Watch(shard,cursor) change-feed.
T4 shipped to production 2026-07-27 — commands now flow server → writer → worker over the EHDB feed on shastaratech prod, with NATS kept installed only as the rollback path.
| Phase | What | Status |
|---|---|---|
| T0 | shadow feed: ChangeFeed primitive + networked FeedWriter/serve/FeedSubscription + latency gate
|
MERGED (#290, #291) — bus p50 57µs / p99 137µs (NATS-parity); the ~4ms first seen was posture-A fsync durability, not the bus |
| T1 | consumer groups + ack / ack_wait redelivery + shard routing (ShardConsumerGroup) |
MERGED (#292) — competing consumers, at-least-once, crash-redelivery (0 loss), committed-cursor resume |
| T2 | KEDA lag signal — per-shard backlog gauge + Prometheus /metrics
|
MERGED (#293) — ready before T4 (hard-ordering rule) |
| T3 | gateway/SPA live feed over SSE (serve_sse) |
MERGED (#294) — id:/Last-Event-ID = cursor/resume, 0-missed/0-dup reconnect |
| T4 | command-bus cutover (flip the live path off NATS) | ✅ EXECUTED ON PROD 2026-07-27 — Option-2 single command shard; shadow 0-divergence → canary 30/30 (0 system claims, 12/12 redelivery) → full flip 80/80, 0 dup, 0 loss, server NATS_pub=0, #166 state sharding intact. Server v3.58.1 / worker v5.81.1. Took three rounds — see the runbook's execution log |
| T5 | delete NATS (irreversible) |
HELD on the human — gated on ai-meta#205 (dispatch latency — root-caused; engine fix MERGED as ehdb 03e94be, adoption held to bundle with #208 so prod deploys once; the prod re-measure is still outstanding — see below) + #208 (writer-restart survival — both defects fixed + kind-validated 2026-07-30, plus a third found in the server publish path; ehdb#302 / worker#196 / server#290 held for review, bundled with #205 so prod deploys once) + #206 (out_of_order_appends on :9102) + an organic-traffic bake; couples #188
|
The bug T4 had to clear. Round 2's full flip failed on a real defect:
under concurrent publish, sparse snowflake ids could reach the single writer
out of order, land below the follower cursor, and be filtered out of the
feed forever — silently, with lag=0. Fixed by having the writer assign
the ordering key (noetl/ehdb#300 →
ehdb d4b6235), which is the only place the ascending-sort_key contract can
actually be guaranteed. Kind before/after on an identical 40-concurrent burst:
17/40 → 41/41. Tracked and closed as
noetl/ai-meta#203.
The one number that got worse — now root-caused and fixed. EHDB command
dispatch (issued→claimed) came out of the flip at p50 285–557 ms vs the
~200 ms NATS baseline — sub-second, correctness unaffected, but ~1.5–2.8×, and
the T5 go/no-go item (noetl/ai-meta#205).
Both suspects named at the time — the 250 ms poll-claim interval and the
writer-side ordering-key append — turned out not to be the cause. See
Dispatch latency after T4
below.
The latency finding (why T0's first 4ms number was not a design-fork), the full three-round prod execution log, and the one-command rollback live in Runbook — L1 Command-Bus Cutover.
The flip was correct but slower than NATS: issued → claimed p50
285–557 ms against a ~200 ms baseline. Attributed 2026-07-28 with a
per-hop harness (ehdb-feed/examples/dispatch_bench, reproduces the
deployed topology on loopback). Neither suspected cause was real:
-
The poll-claim path is not the cost. Delivery (writer ack → a member
holds the command) is 16–33 µs and flat from 3 → 24 members.
claim_nextalready parks onFeedWriter::tip_receiver();DEFAULT_POLL_INTERVAL_MS = 250caps that park soack_waitredeliveries still surface — it is not a polling floor. -
The #203 writer-assigned key is not the cost. It is
global_sequence + 1— arithmetic, no I/O.
The cost is the durability posture, exactly the dimension T0 flagged
as separately tunable. Posture A (FlushPolicy::EveryAppend) pays a
blocking fsync per command inside the engine lock, capping the bus
at ~230 cmd/s and stalling claimers (who need that lock to poll). The
control plane then held one Mutex<PublishRouter> across the whole
round-trip, making the publish path a queue of depth
concurrency × fsync — 282 ms at 64 concurrent publishers, i.e. the
prod p50.
Fixed in ehdb#301 — merged to
main as 03e94be on 2026-07-29: off-lock
fsync through a duplicated fd, FeedWriter::append_batch group commit
(never waits to fill a batch), and a pipelined publish client. Durability,
ordering and exactly-once are unchanged — a returned sort key is still
fsynced before it is handed back, and keys are still writer-assigned and
strictly ascending. Result: 281 ms → 6.7 ms p50, flat in concurrency,
225 → 6 983 cmd/s. T0's note that "group-commit amortizes" the posture-A
cost is what this implements.
Deploy posture. The engine fix is on ehdb main; the adopting pins
(server#289,
worker#195) are ready and
mergeable at 03e94be but deliberately not merged — they bundle with the
#208 writer-restart adoption so
prod rolls once and both changes are re-measured in the same rollout.
Nothing is in prod yet, and the issued → claimed re-measure against the
~200 ms NATS baseline remains the T5 go/no-go input. It cannot come from kind:
the kind worker pool peaks at 28–131 cmd/s against a ~1 000 cmd/s bus ceiling
there, so the publish queue this fix removes never forms (two identical passes
differed 4× in p50).
The bus did not survive a restart of the per-shard writer alone — a node drain,
an image bump, an OOM — in two independent ways, and both were only recoverable
by flipping NOETL_COMMAND_BUS back to NATS, which T5 removes. Fixed in
ehdb#302 (branched off 03e94be, so
the #205 group-commit work is the base), adopted by
worker#196 and
server#290.
Defect 1 — claimers never noticed. ClaimClient::claim_next blocks in a
read until a command is available, so a writer pod that goes away leaves the
socket half-open: no data, no error, and every redial path in the crate and
its callers hangs off an Err that never comes. Observed in kind as 0 of 30
commands claimed, indefinitely, with nothing in any log. Two mechanisms now
cover it, because they catch different failures:
-
TCP keepalive on every
ehdb-feedsocket (configure_stream; claim, publish, feed and SSE). A dead peer becomes an ordinaryio::Errorin ~11 s (5 s idle + 3 × 2 s probes) even when no FIN/RST arrives — the usual case when a pod's veth and conntrack entry go away with it. -
A negotiated coordinator heartbeat. A client that asks (
heartbeat_msonClaimReq::Next, default 5 s) gets one beat up front, which arms a 3-missed-beat read deadline for the connection's whole life, and another whenever a claim parks that long. This is what catches a writer that is alive but stuck; keepalive cannot, since the peer's kernel answers it.
Backward compatible in both directions: heartbeats only go to a client that asked, and a client whose peer never heartbeats disarms its deadline and falls back to keepalive-only.
Defect 2 — the restarted writer replayed its whole shard. The coordinator
was rebuilt at from_cursor = 0, so every record still in the log was
re-delivered — kind showed ehdb_feed_shard_lag{shard="0"} 2738 right after a
routine restart, draining at ~1 record/s because each stale record costs a
control-plane round-trip to learn it is already claimed. CursorStore now
persists committed_cursor() atomically beside the shard's log (the writer's
own volume) and ClaimCoordinator::resume starts there. Two details are
load-bearing:
- The stored cursor is clamped to the reopened log's tip. The engine resumes from its durable manifest, so a log that lost an unsealed active part can reopen behind the cursor — and since the writer then re-issues keys from its recovered sequence, a cursor no future key can exceed would filter out every record and take the bus permanently dark.
-
The writer seals its log on SIGTERM (worker side). Records in an unsealed
part are invisible after a reopen even though every append
fsynced them, so without the seal a rollout restart drops the tail of the command log.
ehdb_feed_shard_committed{shard="N"} is now on the writer :9102: lag reads
0 both when caught up and when a replay has just finished, so the committed
cursor is the series to watch across a restart.
Defect 3, found while validating — the server dropped one command per
restart. EhdbCommandPublisher holds live sockets; the first publish after a
writer restart discovers the dead socket by using it. Dropping the router made
the next command redial, but the one that hit the broken socket was lost —
command.issued durable in the event log, nothing on the bus, HTTP 500 to the
caller (EHDB publish failed: early eof). server#290 redials and retries (3
attempts, 250 ms apart); a retry can only duplicate, and duplicate delivery is
already what ack_wait redelivery produces.
Kind validation — direct before/after, same cluster, 2026-07-30. Option-2 single-shard topology, writer on a PVC, 3 user-pool + 2 system-pool claimers.
| Check | Released worker:5.81.1
|
Fixed |
|---|---|---|
| Writer-only restart, then fire, workers untouched | 21 issued / 0 claimed, lag frozen at 3094 for 100 s | 61 issued / 60 claimed; all 3 claimers redialed in ~2 s |
| Lag right after the restart | 3072 (the whole log) |
0 — resumed from_cursor=3246 origin="persisted"
|
| Writer restarted mid-burst (30 executions) | — | 30/30 completed, 0 dupes, lag → 0 |
| Force-kill a worker mid-burst | — | 39 issued / 39 claimed / 0 dupes |
| Pool isolation + #166 split | — | 0 cross-pool claims; 84 + 83 on the two system shards |
| Idle bus, 10 min | — | 0 spurious redials |
The no-FIN case — the one that actually wedges — is proven deterministically at
unit level (restart_recovery.rs drives it through a relay that holds the
socket open and stops forwarding); killing a listener sends a FIN, which a
blocking read does surface, so it cannot reproduce the wedge.
Deploy posture. All three PRs are held for review, stacked so prod
deploys once: ehdb#302 merges first, then the worker/server pins move to
the post-merge main SHA. The writer needs a PVC (the cursor lives beside
the log) and enough terminationGracePeriodSeconds for the seal. Nothing in
prod; the #205 dispatch-latency re-measure is still the other T5 input.
Follow-up. The SIGTERM seal races in-flight ingest: a command published in the instant the writer is terminating can be acked and still be lost with the unsealed part. In kind that cost exactly one command per restart, recovered automatically by the orchestrator re-issuing the step ~30 s later (all 30 executions completed), so it is a latency artifact rather than loss. The clean fix is to stop accepting ingest before sealing; the crash path (no seal at all) needs L0-level recovery of local unsealed parts.
New self-contained crate crates/ehdb-l0 — noetl-native object store: the
immutable parts, ClickHouse-style manifest, sparse index, and replication are
implemented in-crate (extending the #254 durable-segment store), over a
pluggable DurableSubstrate byte-sink (local filesystem now — noetl's own
store on disk / a PVC / a block device). It is not MinIO or an external S3
server; noetl never delegates the manifest/index/replication to a third-party
product. All additive + shadow, kind/local, no NATS, no prod.
| Slice | What | Status |
|---|---|---|
| L0.1 | immutable parts + meta-catalog (manifest + sparse index) + hot-local/durable-async tiering + cold-load |
MERGED (#273, 05f4d5c) |
| L0.2 | fixed per-dataset inverted index (execution-id blooms) + index-first pruning |
MERGED (#278, 3c5e394) |
| L0.3 | background small→big merge/compaction | MERGED (#278) |
| L0.5 | retention (drop-partition) + orphan reclaim/GC | MERGED (#278) |
| L0.4 | columnar-per-field codec (payload isolated) | MERGED (#278) |
| L0.6 | N-way replication of immutable parts (no consensus) |
MERGED (#279, 3eb2eff) |
| L0.7 |
dataset-generic engine (Dataset trait + L0Engine<D>) — prerequisite for D2–D10 |
MERGED (#280, bb05cf7) |
main = 0af99fc. Every slice/dataset passed the full CI gate green on
rustc/clippy 1.97.1 (cargo fmt --all --check && cargo clippy --workspace --all-targets -- -D warnings && cargo test --workspace).
Replication is noetl-native and implemented (L0.6): every immutable part
and the durable manifest is written write-once to N DurableSubstrate replicas;
PartMeta.replicas: Vec<ReplicaLocation> records every copy; reads fall back
across replicas; orphan-reclaim vacuums all of them. Immutable parts never
conflict ⇒ no consensus / no Raft (HDFS/block-replication). Proven: 24 events
→ 3 parts each replicated 3-way, then replica-0 deleted → a fresh node
cold-loads from the survivors and reproduces the exact 24 records.
The fixed dataset roster (RFC §0.1). Each dataset is its own PR with the full
CI gate green and real proof — functional + cold-load + N-way
replica-kill fallback (read_fallbacks > 0).
| Dataset | What | Access paths | PR / commit |
|---|---|---|---|
| D1 | event log | append / read-by-index / replay | (L0.1–L0.7 base) |
| D2 | command queue | enqueue / claim-by-id / unclaimed-scan | #281 |
| D3 | execution projection read-model | record-state / get-state / list | #283 |
| D4 | KV / coherence | get/put-latest · CAS · prefix-scan · delete |
#283 2362253
|
| D5 | object / blob | put (content-addressed dedup) / get / prefix-list / delete |
#284 6b40b90
|
| D6 | vector / RAG | upsert / top-k cosine / delete |
#285 eb8035e
|
| D7 | catalog | register / get (latest+pinned) / snapshot / deregister |
#286 ddf2b22
|
| D8 | runtime registration | register / heartbeat / deregister / list-live (watermark eviction) |
#287 c7b3a82
|
| D9 | system-WASM store | publish / bind (rollback) / resolve / unpublish / list |
#288 15f0077
|
| D10 | provider-facts | set-desired / set-observed (carry-forward fold) / drift-scan / forget / list |
#289 0af99fc
|
The op-log + fold shape. Every mutable-state dataset (queue, KV, read-model, catalog, runtime, vector, provider-facts) is an append-only op-log over immutable parts; current state is a fold (last-op-wins per key, tombstones dropped). Version/CAS tokens, heartbeat watermarks, and desired/observed carry-forward all ride the op record — no in-place mutation, parts stay immutable, and each dataset cold-loads + replica-fails-over for free from the L0 engine. D5 + D9 carry large opaque bytes (Arrow-IPC / result payloads; compiled WASM), so bytes are stored whole + content-addressed on the substrate with an op-log holding only the pointer registry (RFC §2.4).
Next: L0 is complete — the object-store foundation (engine + all datasets) is built and proven in kind/local. L1 (streaming / NATS takeover), L2 (KV coherence), L3 (fixed reads) are ahead, all gated behind L0. L1 is a separate build phase (touches NATS + the delivery path) and does not start without a go-ahead.
Program umbrella: noetl/ai-meta#194.
Master RFC: docs/rfc/ehdb-layered-platform.md in
noetl/ai-meta. L1 (streaming) detail:
docs/rfc/ehdb-nats-takeover-plan.md (the (c) delivery design; its HA section
is superseded by L0).
EHDB is a noetl-centric internal store. It exists solely to hold the internal information noetl requires to operate, over a FIXED set of predefined datasets (below). It is not a hosted/user-facing DB. Binding on every layer: optimize for noetl's known access patterns; no arbitrary schemas, no SQL DDL, no cost-based query planner, no secondary indexing on arbitrary columns, no multi-tenant hosted-DB surface; business/domain data is never in EHDB. This invariant shrinks each layer's design — it is the guardrail against scope creep. The external Flight-SQL surface (#178/#184) is a read-only export of noetl's own data, not a general DB.
Write-behind cache, never the system of record (refined boundary, 2026-07-15). EHDB is never the durable system of record for business data — the customer's connector-backed store is. But EHDB MAY hold business processing context transiently, as a write-behind cache. Two roles on one engine: (a) the durable control-plane log (events, commands, routing, catalog, runtime) stays lean + reference-based + permanent; (b) the transient processing cache may hold full business context but is bounded, synced/sunk to the customer's store, then evictable — not a permanent warehouse. So the earlier "EHDB never holds business payload" absolute is refined: EHDB may cache business context in-flight, but must sink it to the customer's system of record and evict it. The line: permanent log (a) stays lean; cache (b) stays bounded + sunk + evictable.
Boundary audit (code-verified 2026-07-15): 5 clean, 4 boundary-risk sharing one root cause, 1 clean-today-with-a-seam.
| Dataset | Verdict |
|---|---|
| D4 KV/coherence, D7 catalog, D8 runtime, D9 system-WASM, D10 provider-facts | ✅ in-scope — operational/control-plane metadata; D10 secrets scrubbed |
| D5 object/blob | ✅ large results: in-scope (URN + object-store byte-source — reference-only-state honored); |
| D1 events / D3 projections / D2 commands | command.rs:1493 INLINE_CONTEXT_MAX_BYTES; state_builder.rs:128 keeps context+result); command input inline+untiered (execute.rs:1218). Large payloads already reference-only. |
| D6 vector/RAG | ✅ in-scope today; latent seam — platform RAG only (system docs / catalog embeddings), control-plane-guarded, no playbook tool: wired to rag::ingest; but the ingest surface is role-permissive, so a future user-doc ingest tool would cross the line. |
The fix is four tracked issues (filed 2026-07-15; NOT built — code in
noetl/worker + orchestrate-core in noetl/server, sequenced behind the live
L1 T4 command-bus validation). Reframed to the write-behind-cache model: the
concern is permanence in the durable log, not context-in-cache.
-
noetl/ai-meta#195 — REFRAMED.
Keep the permanent
noetl.eventlean/reference-based; the transient WAL index / slim projection / state shards keep fullcontext/result(bounded 24h / GC'd) for the drive. The refinement dissolves the earlier drive-breaking risk — the drive reads the transient cache, not the permanent log, so keeping the permanent log lean doesn't touch the drive decision (state_builder.rs:124). -
noetl/ai-meta#198 — NEW, the
write-behind sink path. Sink transient business context → the customer's
system of record, and gate eviction on sink-confirmation. Builds on
existing connector tools (
postgres/http/snowflake/transfer/artifact), the result-tier byte-source (#104), and existing GC timers — the new part is sink-confirmation-gated eviction + the sink contract. -
noetl/ai-meta#196 — command
input: keep
noetl.commandlean (reference large input); transient copy is fine. - noetl/ai-meta#197 — the D6 guard test (guardrail already documented); keep the vector-upsert hook unwired for user data.
The large-payload byte-source is the transient state-vehicle (bounded, GC'd)
— role (b), fine to hold context. The sharp line is business payload
accumulating in the permanent log (noetl.event/noetl.command), kept lean
by #195/#196, with #198 letting the cache sink + evict. Boundary finding =
tracked, not unfixed.
The fixed dataset set (the whole scope): D1 event log (noetl.event) ·
D2 command queue (noetl.command/outbox) · D3 execution/projection
read-models · D4 KV/coherence (chain heads, exec descriptors, circuit, leases,
cursors) · D5 object/blob (state shards, result tier, Arrow-IPC) · D6
vector/RAG · D7 catalog (noetl.catalog) · D8 runtime registration
(noetl.runtime) · D9 system-WASM store · D10 provider facts (#189). Each has
a fixed schema, fixed sort key, and fixed access patterns. Secret values
(noetl.credential/keychain) and all business data stay out of EHDB.
"Remove NATS" kept hitting the same wall — durability, replication, HA — and the prior (c) design ended in a deferred, quarters-long per-shard Raft ("T-RF") to give the writer HA. That is the wrong layer to solve durability. The program is restructured into four layers, with all storage/durability/replication in the bottom layer:
| Layer | Role | Analog | On |
|---|---|---|---|
| L0 | Replicated object store — durable, replicated, immutable-part storage + index | VictoriaMetrics/Logs engine over S3/GCS-grade object storage | the object store |
| L1 | Streaming / subscriptions — the NATS takeover (event log, ordering, delivery) | NATS JetStream / Kafka | L0 |
| L2 | KV — coherence state, leases, cursors | Redis / NATS-KV | L0 |
| L3 | Append-log SQL — queries over event/log columns | Postgres / ClickHouse / LogsQL | L0 |
Build L0 first (explicit sequencing). L1/L2/L3 are undurable without it.
L0 resolves the HA debate. Durability + replication come from a replicated object store (multi-region / erasure-coded), so a per-shard writer's local data is a hot cache, not the source of truth. Writers become fungible (any node cold-loads sealed parts from L0 and resumes) — the per-shard Raft build (T-RF) is retired. What remains at L1 is a lightweight ordering lease (who sequences shard N now), not consensus over a replicated log.
Adopt VM's engine principles (well-documented, open source): buffered in-memory writes → flush to immutable parts → background tiered merge (small→big, per-partition) → inverted index (IndexDB) queried first to prune → columnar on-disk with high compression. VictoriaLogs is the closer reference for EHDB's event/log data: per-field columnar (each field its own column + a bloom filter pre-filter), streams grouping, per-day partitions.
Honest departure — VM keeps the hot path on LOCAL DISK. In open-source VM
and VictoriaLogs, object storage is backup-only (vmbackup); a native
object-storage tier for live parts is an unbuilt roadmap item even for
VictoriaLogs (VictoriaLogs#48),
and VM's own durability is local disk × cluster replication + backup. So EHDB
L0 = VM's engine + a live object-store durability tier VM does not itself
ship. That is the plan and it resolves our HA debate — but it is net-new
engineering beyond VM, the highest-risk L0 piece, not a copy of a proven
model.
Hot-local / durable-async composite (the low-latency key):
append → in-memory buffer (served immediately)
→ sealed immutable part on LOCAL disk (hot tier, fsync)
→ async uploader → REPLICATED OBJECT STORE (durable/replicated tier)
read → merge { buffer, local parts, object-store parts }
The hot path never blocks on object-store latency (writes/recent reads hit RAM + local parts, single-digit ms); the upload is background. Durability- window knob: VM buys speed by not fsyncing per append (loses the last few seconds on hard crash) — acceptable for metrics, not for a source-of-truth event log. EHDB's #254 already fsyncs per append; keep that (posture A) for L1's log, allow VM-style buffered flush (B) for derived/metrics tiers.
Generalize by data type — don't force blobs into the metrics layout: append/event data → VictoriaLogs columnar+bloom; KV → small parts + index; metrics → VM TSID+Gorilla+zstd; blobs → content-addressed object mapping (bytes whole in the object store, L0 holds metadata+pointer only — Phase-8's object tier already is this).
ClickHouse-style meta-catalog (table format over object storage). VM writes
the parts but doesn't give a good pointer catalog over object storage. Because
datasets are predefined, L0 adds a ClickHouse-MergeTree-style catalog (like
Iceberg/Delta manifests, but fixed-schema): per dataset, (1) a manifest —
one row per immutable part { part_id, partition, min_key, max_key, record_count, byte_size, object_store_uri } (the pointer catalog VM lacks,
versioned), and (2) a sparse primary index per part (one entry per granule
over the fixed sort key → mark/byte-offset) + per-part min/max. Lookup (e.g.
D1 "events for execution E after seq S"): manifest-prune parts by partition +
min/max (zero I/O on non-matching parts) → binary-search the sparse index to the
granule → ranged GET only that block from the part's object_store_uri →
decode. A fixed compiled path per dataset — no scan, no planner. VM's
seal/merge writes the sparse index + appends a manifest version; the async
uploader records the object-store URI. This is what makes "recent = local part,
older = object store" transparent (same catalog, different URI).
The noetl-only invariant CUTS scope vs a general DB: no query planner (fixed per-dataset resolution paths), no DDL (10 compiled-in schemas), no general inverted index (only a few fixed per-dataset indexes on known dims — this collapses the "biggest gap" below), no full-text over arbitrary fields, no SQL engine (L3 = fixed reads + read-only export), no multi-tenant hosted-DB. It shrinks every layer.
| L0 principle | EHDB today | Gap |
|---|---|---|
| Immutable parts |
Strong — #254 immutable seg-*.eslog (8 MiB rollover), archivable/replicable |
reshape segment→part, partition by day/shard |
| Buffered flush→parts | Opposite — #254 fsyncs per append | add the in-memory buffer→flush hot tier (posture knob) |
| Background merge (small→big) | Weak — has rollover + GC/retention, no merge | the merge engine (net-new) |
| Meta-catalog + index-first | Weak — offset index (seq→loc), not the manifest/sparse-index catalog | §meta-catalog (manifest + sparse index → pointer) + a few fixed per-dataset indexes (NOT a general IndexDB — invariant shrinks it); largest gap but bounded |
| Columnar per-field + blooms | Partial — Arrow columnar for result/state tiers; event log is row-framed | per-field columnar for the event tier |
| Object-store durability tier |
Partial (seam exists!) — durable_eventlog_shared.rs ships segments to a shared store + cold-load; ehdb-storage S3/GCS adapters |
the async uploader + read-merge + object-store-as-live-tier (§ high risk) |
| Blob content-addressed map |
Strong — Phase-8 ObjectBlobDriver
|
wire as L0 blob shape |
| KV / projection |
Strong — Phase-8 KvStateDriver, Phase-7 projection (primary merged) |
latency tier / L3 wiring |
| Replication | External (the plan) — from the object store, not EHDB consensus | inherit from object store |
Real net-new for L0: (1) the inverted index + index-first pruning (biggest gap), (2) the small→big merge engine, (3) the object store as a live async-durability tier (beyond-VM) + columnar-per-field for events. The immutable-part, blob, KV, projection, and shared-store-shipping pieces are substantially present (#254 + Phase-7/8). A focused build, not from scratch.
L0 (BUILD FIRST) — ~2–4 quarters (a production storage engine): L0.0 part model (reuse #254) · L0.1 async object-store durability tier + read-merge + cold-load (first slice, highest-risk) · L0.2 meta-catalog + fixed per-dataset indexes (§ — NOT a general IndexDB) · L0.3 merge engine (fixed per-dataset policy) · L0.4 columnar-per-field + blob shape · L0.5 retention-as-drop-partition.
L1 (streaming / NATS takeover) — ~1.5 quarters, gated behind L0. The (c)
one-hop delivery design minus its HA section (now L0's job): per-shard
ordering lease + change-feed + one-hop delivery + consumer-group/ack + KEDA
prometheus lag, over fungible writers. Cheaper than before — the
PVC/Raft work moved to L0. L1 cutover (delete NATS) only on a mature L0.
L2 (KV) — ~1 quarter, gated. Coherence + leases + cursors on L0. Reuse
Phase-8 KvStateDriver.
L3 (fixed reads + read-only export) — ~1 quarter, gated. NOT a SQL engine (invariant cut): the fixed control-plane read queries over the Phase-7 projection read-models + the §meta-catalog, plus the #178/#184 read-only Flight-SQL export. No planner, no DDL (down from ~2–3q — the SQL-engine ambition is cut).
Multi-year program; L0-first because every layer above is undurable until L0 exists.
Async part-uploader + read-merge over local + replicated object store, SHADOW, on #254 segments.
Proves: a writer appends to a local immutable part (fsync, hot, single-digit ms), asynchronously uploads sealed parts to a replicated object store, and serves reads by merging local + object-store parts — so (a) the hot path never blocks on object-store latency, and (b) a fresh node with no local data cold-loads the parts from the object store and reproduces the exact log (the fungible-writer property that retires T-RF). All shadow; the authoritative path is untouched.
Exit criteria (kind): hot-path isolation (append p99 with upload on ≈ off — async upload doesn't regress append latency); durability (every sealed part lands in the object store, bounded upload lag); cold-load correctness (a fresh empty-local store cold-loads from the object store and reproduces the exact records + global sequence); reversibility (flag off ⇒ byte-identical); boundary (data-plane only, secret-free); no prod/GKE.
Why this slice first: it proves the novel, beyond-VM claim (object store as a live durability tier at no hot-path latency cost) on which the HA resolution and the whole layered bet depend. Prove the risky thing first.
Locked (2026-07-15): PROGRAM INVARIANT — EHDB is noetl-internal-only, fixed predefined datasets, NOT a general DB (shrinks every layer); layered platform, L0-first; L0 = VM/VictoriaLogs write engine + ClickHouse-style meta-catalog; HA via L0 object-store replication + L1 ordering lease (per-shard-Raft "T-RF" retired); L3 is NOT a SQL engine (fixed reads + read-only export).
Open (need the user):
- L0 durability-window posture — fsync-per-append (A) for the event log vs buffered flush (B) for derived/metrics tiers (recommend A for the log).
- The replicated object store backend — GCS multi-region? S3 + erasure? in-cluster MinIO/Ceph? — now the platform's durability foundation.
- L1 latency budget — the drive-hop p99 gating the L1 cutover (still a number).
- Sequencing confirmation — L0 is a ~2–4 quarter build before L1 (NATS removal) can begin; NATS stays fully in place throughout L0.
- Master RFC
docs/rfc/ehdb-layered-platform.md; L1 detaildocs/rfc/ehdb-nats-takeover-plan.md. - Program umbrella noetl/ai-meta#194.
- EHDB tiers L0 reuses: Event-Log Core Engine, Durable Event-Log Backend, KV/Object/Vector (Phase 8), Projection Read-Model, Performance & Load Testing.
- ehdb#241, ehdb#254, ehdb#234.
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)