Skip to content
Kadyapam edited this page Oct 10, 2026 · 81 revisions

EHDB Wiki

F1–F5 remediation landed 2026-08-30 (ehdb#333–#337, all inert): the D1 durability window is now measurable from the append and live on the writer's /metrics (worker v5.125.0); an age-based seal trigger exists default-off; stale-writer fencing runs in shadow (counts, refuses nothing); per-shard Lease election issues tokens but is not authoritative; and failure domains are declared and validatable. Four switches remain and every one is owner-gated — see the four-gate plan and Consistency Invariants.

Current state — 2026-08-20. The event-log tier is primary and SERVING on prod (since 2026-08-13), and its mirror now runs asynchronously off the event-write path (since 2026-08-19) — see Runbook: the async event-log mirror for the operational surface and the one-flag rollback. Live versions: server v3.83.1, user-pool worker v5.120.0, cmdbus-writer v5.119.1. Tier modes: eventlog primary, projection / kv / object shadow; NOETL_EHDB_VECTOR is not set at all. ⚠ Only eventlog and projection have a runtime serve path (SERVE_WIRED_TIERS). Setting NOETL_EHDB_KV=primary or NOETL_EHDB_OBJECT=primary promotes nothing — the mode table is configuration surface, not capability. ⚠ Older sections below describe the pre-cutover plan and say primary is "recognised but not activated". That was true when written and is no longer true — where a page describes a plan, read it as history.

EHDB is the Event Horizon Database for the NoETL ecosystem: an Arrow-native, NoETL-domain storage system that stores operational metadata transactionally, stores analytical/historical data, carries event streams, and serves AI/RAG retrieval needs for NoETL workloads.

EHDB is not intended to be a generic database with NoETL as one user. It is a NoETL-specialized storage substrate. Over time it should absorb the platform roles currently filled by PostgreSQL, NATS JetStream, and external object stores, replacing them with a single NoETL-centric durable fabric.

⚠ Scope narrowed 2026-08-29 (#320). EHDB owns four engines — event log, projection, KV, object. The OLAP engine (ClickHouse-replacement) and the standalone ANN/vector engine (Qdrant-replacement) are removed from scope. This is a removal, not an externalisation: analytics become projections, and vector retrieval becomes a vectorized projection searched with exact cosine over bounded, execution-scoped candidate sets. See Architecture — the four engines and Consistency Invariants. Older pages below still carry the five-tier list; where they do, this supersedes them.

Mission

EHDB provides a unified foundation for:

  • NoETL operational metadata and system-of-record state.
  • First-class catalog data: tables, schemas, partitions, statistics, lineage, ACLs, tenants, and namespace bindings.
  • Historical analytical data resident in object storage.
  • Event streams, durable consumers, and replay semantics currently provided by NATS JetStream.
  • RAG-ready document, chunk, embedding, and vector-index metadata for AI-native NoETL workloads.
  • Tenant/namespace-scoped local vector similarity fixtures for early RAG lookup validation.
  • Transaction logs, MVCC snapshots, and immutable data files.
  • Arrow-native inter-service communication and serialization.
  • Distributed read/write separation across regions, clusters, and clouds.
  • Embedded role/capability policy for NoETL gateway, API, worker, playbook, and system contexts.

Design Boundary

EHDB must stay aligned with the NoETL execution model:

gateway = gatekeeper
worker = atomic compute
playbook = ephemeral blueprint
shared cache = state vehicle
event log = source of truth

No client or gateway should reach directly into EHDB internals. NoETL data-touch behavior happens inside playbook steps or EHDB-owned service boundaries. EHDB must not require persistent per-tenant AI-agent or MCP server processes to hold state between requests.

EHDB is allowed to become an embedded distributed database runtime across the NoETL ecosystem, including workers, APIs, and gateways, but embedded does not mean every role receives data-plane permissions. The core role policy treats gateways and API/admission surfaces as control-plane embedders only. Worker, playbook, and system roles receive explicit catalog, transaction, stream, object, retrieval, replication, and system library capabilities for bounded data-touch work.

Initial Architecture

EHDB
|-- Catalog
|   |-- Tables
|   |-- Schemas
|   |-- Partitions
|   |-- Statistics
|   |-- Lineage
|   `-- ACLs
|
|-- Transaction Log
|
|-- Event Streams
|   |-- Subjects
|   |-- Durable Consumers
|   |-- Replay Cursors
|   `-- Retention Policies
|
|-- Object Storage
|   |-- Parquet
|   |-- Arrow IPC
|   |-- Iceberg-compatible tables
|   `-- Blobs
|
|-- Retrieval
|   |-- Documents
|   |-- Chunks
|   |-- Embeddings
|   `-- Vector Indexes
|
|-- System Libraries
|   |-- WASM Manifests
|   |-- Environment Bindings
|   |-- Release Channels
|   `-- Host Capabilities
|
|-- Query Engine
|
`-- Replication Engine

The catalog itself lives inside EHDB. The design should avoid an external PostgreSQL dependency for EHDB metadata durability once the system reaches its self-hosting milestone.

Current Implementation Notes

  • L1 command-bus is PROD-VALIDATED and now faster than NATS (2026-07-30). The #301 latency fix and the #302 writer-restart fix deployed to shastaratech prod together (engine 86a24f9, server v3.58.3, worker v5.81.3). Dispatch issued → claimed measured p50 138.5 / p95 156.4 / p99 181.0 ms (n=60, unsaturated) against the 338 ms NATS baseline — ~2.4× faster than NATS, closing ai-meta#205. A writer pod deleted mid-load recovered in ~2.7 s with 38/38 COMPLETED, 114 = 38×3 commands, 0 dup / 0 loss, resuming from the committed cursor instead of replaying the shard — closing ai-meta#208. ⚠️ The same run found that user-playbook dispatch had been silently dead in prod for ~2.4 days from the pre-fix restart bug, masked by system-pool traffic. Two caveats worth carrying: a saturating burst measures the worker pool, not the bus (p50 2123 ms of pure queueing behind 8 slots), and kind cannot reproduce the latency measurement at all — its pool peaks at 28–131 cmd/s against a ~1000 cmd/s bus ceiling. Detail: Runbook — L1 Command-Bus Cutover, round 4.

  • L1 command-bus dispatch latency root-caused + fixed (2026-07-28). After the T4 flip the bus was correct but ran issued → claimed at p50 285–557 ms vs a ~200 ms NATS baseline. Per-hop attribution (ehdb-feed/examples/dispatch_bench) cleared both suspects — delivery is 16–33 µs and flat in member count, so the poll-claim path is not the cost, and the #203 writer-assigned key is pure arithmetic. The cost is posture-A fsync taken inside the engine lock (~4 ms, ceiling ~230 cmd/s) multiplied by the control plane holding one publish mutex across the round-trip (a queue of concurrency × fsync — 282 ms at 64 publishers). Fixed by off-lock fsync via a duplicated fd, group commit (FeedWriter::append_batch), and a pipelined publish client (ehdb#301): 281 ms → 6.7 ms p50, flat in concurrency, 225 → 6 983 cmd/s, with durability, ordering and exactly-once unchanged. Two pre-existing writer-restart defects found alongside and filed as noetl/ai-meta#208 — both block T5. See Program — EHDB NATS Takeover.

  • Performance & load testing (Phase 1) — engine micro-benchmarks landed (2026-07-08). Deterministic criterion benches over all five reference drivers + the durable segment event-log backend (ehdb-reference::benches::engine_micro, ehdb#261). Baseline on an M1 Max dev box quantifies the headline: the durable event-log backend beats the local_reference JSONL driver 2.7× at sustained append (255 vs 96 ev/s at K=1000) and stays flat (~3.9 ms/append, fsync-bound) while the reference driver degrades O(n) per op; segment rotation costs ~2%; cold replay runs ~185 K ev/s. The reference KV/object/vector tiers are O(n)-per-op (JSONL reopen) — shadow-only, not primary-serve. ⚠ Superseded for the event-log tier (2026-08-18): the per-op JSONL reopen was the dominant cost on prod and is fixed by the runtime replay cache (NOETL_EHDB_REFERENCE_RUNTIME_CACHE, worker v5.119.1) — emit_mirror 1141 ms → 110 ms per call. The O(n)-per-op statement still describes KV/object/vector, which remain shadow-only. In-cluster load + the EHDB-vs-incumbent (Postgres+JetStream) head-to-head are the next phase. See Design: Performance & Load Testing. | Measures: L0 engine | The first measurements of the production ehdb-l0 engine — the fsync is ~95% of posture-A write cost (267 vs 5 450 rec/s group-committed), and a read steps 15.6x per record at seal_max_records = 1024. Every instrument carries a planted-effect resolution control. | | North star: distributed registry | EHDB as NoETL's distributed registry and control plane — "etcd but truly distributed" plus append log, catalog relations, topology/discovery, secret references and vectors. ⚠ The inventory corrected three prior claims: D8 is in the production crate, vectors already exist, and four of seven capabilities are substantially built. The real gaps are watch (zero pub fn watch in the workspace), time-based leases, ephemeral registration, secret-ref types and tail replication (RF=1). P1 (wall-clock TTL liveness) done. |

  • EHDB completion program (Phases 6–10) — CODE-COMPLETE (2026-07-06). Every platform tier has a built + shadow-verified engine (Phases 6–8), a per-tier reversible primary cutover (Phase 9, all five kind-validated), and a consolidated backend-selection surface (Phase 10). The only remaining step is the prod/GKE cutover, gated on the user, per tier; nothing in prod has changed (prod worker v5.52.0, all NOETL_EHDB_* flags default off).

  • Live in kind (LOCAL, SHADOW, prod untouched). The event-log tier's runtime mirror is wired into the live emit_event path and proven on real drives (worker v5.67.0, worker#167 d310c7b); the KV and object tiers' runtime mirrors are now wired too (worker v5.68.0, worker#168 2ec2e2b) — KV at the SpoolRuntime::persist_circuit NATS-KV write, object at the ControlPlaneClient::object_put chokepoint every object tier funnels through (both env-armed shadow, disabled-by-default, error-isolated). The object mirror's live-drive proof landed 2026-07-07 on worker v5.69.0 in kind (both data-plane pools): a real tests/large_tabular_result_test drive externalized a >inline-budget Arrow Feather result through object_put, and the mirror metric advanced with outcome="mirrored" (system 2, user 1) — a clean flip from v5.68.0's outcome="invalid" (2), confirming the subject-length fix (ehdb#256 bbc5047, worker v5.68.1): the object subject is now noetl.obj.<sha256hex> (74 chars, well under the 256 cap), and the ref-store received the real platform-key objects. KV is now PROVEN on the live circuit path too (2026-07-07): a dedicated subscription runtime (WORKER_MODE=subscription, worker v5.69.0 + shadow env) activated the subscriptions/spool_outage_stream kind:Subscription (spool + one http-probed downstream), so SpoolRuntime::persist_circuit fires every ~2 s probe tick and on each circuit open/close → kv::mirror_live_put. The runtime's noetl_ehdb_kv_ops_total{operation="mirror",outcome="mirrored"} climbed monotonically (10 → 50+, kv_last_ok=1, degraded=0, zero invalid); an induced downstream outage tripped subscription.circuit.opened and advanced it further. The ref-store holds the real circuit record at subject noetl.kv.noetl_subscription_circuit.<hex> where the hex decodes to circuit.<subscription_id> (a short key → ~90-char subject, well under the 256 cap). And the projection tier's runtime mirror is now wired too (worker v5.69.0, worker#170 fa64e0a) — via a windowed cadence hook at the off-server state-builder drain's post-batch checkpoint (state_builder::run_drain_loop → projection::mirror_live_window), NOT a per-event hook: each drained batch is materialized into a fresh throwaway per-window store and parity-checked against an independent worker-side fold of the same window, so the batch fold produces no false key-divergence (see Design: Projection Read-Model Engine). The vector tier's mirror is code-ready + tested but deliberately not live-wired — no platform vector-upsert site exists in the worker loop today (RAG ingest is lexical, not embedding vectors), so wiring it would create a hook that never fires (documented seam: a future platform-RAG embed+upsert in executor/command.rs). Subject-length hardening — all three tiers (2026-07-08). The KV + vector registry subjects now use a fixed-width SHA-256 digest token (noetl.kv.<bucket>.<sha256hex(key)>, noetl.vec.<sha256hex(collection)>.<sha256hex(point_id)>), matching the object tier — so a long id can't overflow the 256-char Subject cap (ehdb#259 52120a7 / ehdb#260; worker pin bump worker#172). Forward-safety: the live KV proof above used short circuit.<id> keys that never overflowed, and there's no live vector site, so nothing was broken — the digest token keeps a future long-key path (KV #115 program-coherence, platform-RAG vectors) bounded. In-kind live re-proof of the new KV subject is deferred. And the read-only Query Interface serves those mirrored executions end-to-end through the noetl server v3.53.0 /api/ehdb/* API + the noetl ehdb query CLI (server#277 ec75812 + cli#61 v4.12.0). The durable event-log backend (ehdb#254) is code-complete + kind-soaked through slice 5 — segment store + crash recovery, execution-affinity single-writer, shared cold-load tier, worker wiring (NOETL_EHDB_EVENTLOG_BACKEND=durable_segment, worker v5.70.1), and a real-pod-restart kind soak (rotation + byte-identical crash recovery on a PVC). The slice-6 prod-durability sign-off package (Runbook-Durable-EventLog-Prod-Signoff) is drafted (verdict: GO to prod durable-shadow; Stage-C durable-primary gated on the soak + a segment-GC decision) — awaiting user go/no-go; no prod action taken.

  • What's NOT done yet (the program stays open — umbrellas ehdb#234 + ehdb#241). The prod/GKE per-tier cutovers (gated on the user; tier-1 runbook drafted; its Stage C durability gate is now addressed by the kind-soaked durable backend + the slice-6 sign-off package, which still gates the actual prod durable-shadow → durable-primary sequence on the user per stage). The durable-backend program is code-complete through slice 5; slice 6 (prod sign-off) awaits go/no-go. All five runtime mirrors are now code-wired: eventlog + object + projection + KV are all proven on live drives in kind (worker v5.69.0, 2026-07-07 — object flipped invalid→mirrored, projection via the windowed drain hook, KV via a live spool-subscription persist_circuit runtime); vector is code-ready + tested but documented-unreachable — no live platform vector-upsert site exists in the worker loop yet, so its hook lands with a future platform-RAG embed+upsert write site, not before. The raw-tier query seam (/api/ehdb/tiers/{tier} still returns 501 → worker query handler follow-up). Live-drive proof of the projection mirror landed 2026-07-07 on worker v5.69.0 in kind: the same tests/large_tabular_result_test drive advanced noetl_ehdb_projection_ops_total{operation="materialize",outcome="materialized"} to 4 on both data-plane pools via the windowed drain hook — parity held, no false divergence, no invalid. Event-log mirror held as a regression (mirrored 6 both pools). KV is now proven too (live persist_circuit path via a spool-subscription runtime — see the object/KV paragraph above), so all four wired tiers (event-log, object, projection, KV) fire on live drives; only vector stays deferred (no live upsert site).

  • Phase 10 (tunable-backend config surface) — IMPLEMENTED (2026-07-06). The scattered NOETL_EHDB_* per-tier flags are consolidated into one coherent schema (ehdb-reference::backends: PlatformTier / TierMode / Backend / BackendMatrix, resolved umbrella-enable → per-tier mode → derived backend primary ⇒ EHDB / else incumbent), with coherence validation + a secret-free render and an ehdb-selfcheck config verb (ehdb#252 4c0df81 + worker#166 v5.66.0). Backward-compatible, no behavior change; EHDB is the default, not a lock-in. See Backend Configuration.

  • Phase 6 (event-log core engine) — in progress. ehdb-reference now exposes the event-log core engine behind the EventLogDriver trait (append / scan_global / read_execution / tail / ack), with LocalReferenceEventLogDriver composing the append-only stream primitives over one canonical noetl_event_log stream — its sequence is the global, monotonic, gapless event-log sequence, per-execution scope rides the noetl.event.exec.<execution_id> subject, and a durable consumer serves tail/ack. compare_shadow_parity + the worker disabled-by-default shadow prove the engine tracks the authoritative log without serving it. The shadow mirror is now wired into the live ControlPlaneClient::emit_event path and PROVEN on real in-kind drives (worker v5.67.0, worker#167 d310c7b) — executions 332760742153424896 + 332760854506246144 each mirrored their 6 events into the reference tier, 0 restarts, incumbent unaffected; only the eventlog tier is wired live so far. See Design: Event-Log Core Engine. Phase 9 tier 1 — event-log primary cutover implemented + merged (2026-07-05). Under NOETL_EHDB_EVENTLOG=primary the engine now serves the log authoritatively (append + read + tail + ack + replay), dual-run parity-checked against the incumbent and reversible (ehdb#247 7f014c9 + worker#161 v5.61.0). Kind dual-run VALIDATED 2026-07-05 (in-cluster on kind-noetl node noetl-control-plane, image v5.62.0 sha256:9d222db6: served_by_ehdb:true

    • reversible:true + dual_run_holds:true + secret-free metrics; control-plane ⇒ guard_refused); prod/GKE cutover still gated on the user. See the Roadmap Phase 9 tier-1 section.
  • Phase 7 (projection / read-model engine) — shadow complete; Phase 9 tier-2 primary merged. ehdb-reference now exposes the projection engine behind the ProjectionDriver trait (apply / read_execution_state / read_event / list_executions / checkpoint), with LocalReferenceProjectionEngine consuming the Phase-6 event-log tail and materializing the read-models the Postgres materializer produces (event read-model keyed on event_id, folded execution-state, durable consumer checkpoint) into one noetl_projection_log store. Apply is idempotent / exactly-once on the global sequence; rebuild-from-log is deterministic; compare_projection_parity + the worker disabled-by-default shadow (NOETL_EHDB_PROJECTION, worker#157 eadc3a5) prove the read-models track the materializer without serving reads. Phase 9 tier 2 — projection primary cutover implemented + merged (2026-07-05). Under NOETL_EHDB_PROJECTION=primary the engine now serves the read-models authoritatively (list_executions / per-execution read_execution_state / read_event), dual-run parity-checked against the incumbent materializer and reversible (ehdb#248 d08013c + worker#162 v5.62.0 36875e3). Kind dual-run VALIDATED 2026-07-05 (same in-cluster Job on kind-noetl: served_by_ehdb:true + reversible:true + dual_run_holds:true + rows_after_revert:3 + secret-free metrics; control-plane ⇒ guard_refused); prod/GKE cutover still gated on the user. See Design: Projection / Read-Model Engine and the Roadmap Phase 9 tier-2 section.

  • The first Rust workspace is live with local reference crates for catalog, immutable object storage, stream replay, retrieval metadata, transaction replay, and reference-state rebuild from transaction records.

  • ehdb-storage now emits content-checked immutable ObjectRef values with path, byte length, and SHA-256 digest, plus verified reads and a deterministic table/snapshot object path layout. Object refs also carry geo-location and data-gravity shard pointers for future distributed placement.

  • ehdb-catalog now commits immutable table snapshot metadata over content-checked object references, tracks latest snapshots, and rejects parent-chain mismatches.

  • ehdb-reference now applies replay-complete transaction records into catalog, stream, retrieval, and system-library reference catalogs, keeping the transaction log aligned with the source-of-truth role.

  • ehdb-retrieval now provides a local exact cosine-similarity fixture over registered chunk embeddings, scoped by tenant, namespace, and embedding model.

  • ehdb-service now wraps replayed retrieval state with LocalRetrievalSearchService, returning ranked vector, text, and hybrid chunk hits plus bounded RAG context blocks and versioned local context payload codecs/execution without exposing raw embedding vectors.

  • LocalReferenceRuntime now combines the local JSONL transaction log with the replay applier, validating projected reference state before durable append and rebuilding that state from replay on reopen.

  • ehdb-local-reference summary --log <path> now exposes deterministic JSON counts from replayed local reference state across transaction, catalog, stream, retrieval, system-library, and storage domains for bounded NoETL worker/playbook diagnostics.

  • CatalogScanGrant now records durable tenant/namespace/table scan grants for principals, and CatalogMutation::GrantScan makes those grants replayable through the local reference runtime.

  • LocalArrowFlightServer can now require x-ehdb-principal metadata and enforce replayed catalog scan grants before local get_flight_info or do_get execution.

  • FlightAccessLogPolicy now keeps local Flight scan access summaries bounded: disabled mode emits nothing, while debug-only mode emits structured request summaries without tokens, principals, tenant/table identifiers, object paths, predicate values, or Arrow payloads.

  • ehdb-transaction now includes a fsynced local JSONL transaction log for crash/restart tests. This is a reference adapter; production replicated durability still belongs behind a consensus-backed log boundary.

  • ehdb-stream now includes a fsynced local JSONL stream journal for restart tests covering stream records, retention, durable consumer cursors, and ack replay. This is a reference adapter; production stream replication remains a later boundary.

  • ehdb-system now models NoETL system WASM libraries as immutable module manifests plus environment/channel bindings, so system playbook functionality can be hot-replaced by digest/revision without forcing Rust crate semantic-version churn. The local JSONL journal preserves those manifests and bindings across restart.

  • ehdb-service now defines the first service-facing latest-table scan request/result boundary over the local Arrow scanner. It returns schema, batches, and row count while keeping Arrow Flight networking, SQL planning, distributed execution, and gateway direct reads out of scope.

  • ehdb-service also now has a versioned Arrow Flight scan ticket codec that round-trips latest-table scan requests through Flight Ticket bytes and command FlightDescriptor values before any Flight server is introduced.

  • ArrowScanResult now round-trips local scan outputs through Arrow Flight FlightData messages, proving the future result-stream contract without starting a network service.

  • ArrowScanResult now builds pre-network Arrow Flight FlightInfo metadata with schema bytes, command descriptor, endpoint ticket, row count, and encoded byte count.

  • LocalArrowFlightService now provides in-process get_flight_info, get_schema, and do_get behavior over the local scan service and Flight codecs.

  • LocalArrowFlightServer now implements the generated Arrow Flight service trait for scan get_flight_info, get_schema, and do_get, enforcing the configured request metadata auth and scan scope policies without binding a port or starting a persistent runtime.

  • LocalArrowFlightServerConfig now validates the first bounded Flight lifecycle settings, header-token auth contract, tenant/namespace scan scope contract, catalog scan grant policy, and access-log policy without opening a listener.

  • Implemented scan methods now enforce the configured local max_concurrent_requests budget with fail-fast RESOURCE_EXHAUSTED responses when all request slots are occupied.

  • LocalArrowFlightListener now provides a loopback-only reference listener harness with explicit shutdown handling.

  • The loopback client smoke path now proves get_flight_info and get_schema plus do_get over real tonic/gRPC transport, including the optional header-token auth, tenant/namespace scan scope, and catalog scan grant policies.

Releases

version date what
v0.8.1 2026-10-10 The tail replicator uploads with the engine lock released (PR #405). replicate_tail held the engine across the remote put, and in noetl/server that lock is taken by the LIVE append path — p50 77 ms / p99 234 ms injected into every append colliding with a tick. Split into prepare (no I/O) → upload (free fn) → commit (no I/O).
v0.8.0 2026-10-10 B3 — the unsealed tail replicates off-box and recovers on cold load (#394, PR #404). Loss window goes from the seal interval (900 s on prod) to the replicator tick (15 s default) — bounded, not zero. See Design — Tail Replication.
v0.7.0 2026-10-10 ClaimCoordinator::inflight() — the coordinator could not report in-flight depth, so a writer's only option was a literal 0, which publishes the stalled-consumer reading rather than "unknown" (#402).
v0.6.0 2026-10-10 backfill_under_replicated — attaching a replica replicated nothing that already existed; measured on prod as 39 of 40 parts single-copy while survives_node_loss read 1 (#400). See Design — Replica Backfill.
v0.5.2 2026-10-09 active_age returns None for an empty part (#398).
v0.5.1 2026-10-09 reclaim stranded active parts (#396).
v0.5.0 2026-10-08 North-star P1–P6 (RuntimeKind + discover, watch_since, ephemeral units, SecretRef), the four state gauges + records_superseded, and Dataset::supersede_key key-level compaction. FORMAT_VERSION still 1.
v0.4.5 2026-09-29 chain ordered by following links, never by event_id

⭐ v0.5.0 is the first ehdb release cut by automation. Every v0.4.x tag before it was pushed by hand, so the version was whatever a human typed and nothing connected it to the commits. semantic-release.yml + release-ehdb now do it — see Releasing.

✅ Auto-release is ON. EHDB_RELEASE_ENABLED=true, so a merge to main carrying a feat: or fix: subject auto-cuts a version. chore: / docs: / test: and any non-conventional subject release nothing, silently — see Releasing.

⚠⚠ Auto-release is NOT auto-deploy. Cutting a tag does not roll production. A consumer only picks up a new version when both noetl/server pins (ehdb-feed and ehdb-l0) are moved together — different tags put two ehdb-l0 versions in one dependency graph — and the prod roll after that takes its own baseline, applies by digest, and records a rollback digest first. An auto-cut tag must never be read as live.

⚠ ehdb publishes no crate and no image: it is a library consumed by git tag, so a release's artifact proof is a tag whose tree a consumer can build, which release-ehdb performs (release build, the three rlibs asserted by name, the suite at the tag, and the on-disk FORMAT_VERSION recorded in a build-manifest.txt).

Consumed where

consumer pinned at status
noetl/server v0.5.0 (both pins, #d61530b2) live in prod as server v3.126.0, with the D8 runtime registry reached: the server self-registers, renews its lease, and serves /api/runtime/topology
noetl/worker v0.4.5 ⚠ not yet moved — and that ordering is deliberate: RuntimeOp carries #[serde(deny_unknown_fields)], so readers upgrade before writers

Core Pages

  • Architecture — resilient KV core: CockroachDB's layers mapped onto EHDB as implemented, minus SQL. The L4 gap is the unsealed tail, not replication in general — sealed-part N-way copy is built, the window is already measured and bounded on prod, but open_replicated and validate_replica_domains have no production callers. Phase plan + the Raft-over-the-tail recommendation. (ai-meta#339)
  • Architecture: execution model, component boundaries, durability model, protocol direction.
  • Roadmap: phased implementation plan from Rust crate skeleton to NoETL system-of-record migration.
  • Sessions Log: chronological design and implementation notes.
  • Claude Handoff: detailed continuation prompt, phase links, validation commands, and NoETL integration checklist.
  • RFC: EHDB Completion Program — Coupling + noetl Self-Sufficiency (decided, implementing): loose server↔EHDB coupling; EHDB as NoETL's self-sufficient internal storage fabric (event-log engine fixing the event-sourcing bottleneck, projections, KV/object/vector) — k8s-only end-state, tunable per tier, platform-only, never business data.
  • EHDB Query Interface: read-only query surface over NoETL platform data — server /api/ehdb/* endpoints + the noetl ehdb query CLI. Control-plane serves the projection/read-model directly and routes raw data-plane tier queries to the worker; secret-free, bounded, read-only.
  • RFC: External EHDB Driver (design — proposed, awaiting decision): an outward-facing driver so third-party apps / BI tools can open a connection and read EHDB's tiers. Recommends Arrow Flight + Flight SQL over a dedicated data-plane endpoint fronting the worker (loose coupling preserved), scoped read-only API tokens, committed/materialized-only reads. Five forks flagged for decision (wire / surface / placement / auth / read-only-MVP).
  • Program: EHDB Layered Platform (NATS takeover is L1) — L0 COMPLETE (D1–D10); L1 T0–T4 COMPLETE: the EHDB feed is the PRODUCTION command bus on shastaratech prod since 2026-07-27 (80/80, 0 loss, 0 dup; server v3.58.1 / worker v5.81.1). NATS is still installed as the rollback path; T5 (delete NATS) is held on the human. #205 (latency) and #208 (writer-restart survival) both CLEARED on the 2026-07-30 prod deploy — server v3.58.3 / worker v5.81.3, p50 138.5 ms vs the 338 ms NATS baseline. The current T5 gate is autoscaling: noetl-worker-rust still triggers on nats-jetstream lag, so deleting NATS would strand the user pool at 8 slots. The EHDB-lag scaler is built (ops#242) and applied to prod paused — and the pool has had no autoscaling since 2026-07-26 → ai-meta#210. Also open: #209 (a crash still loses the unsealed tail) and #206. Execution log: Runbook — L1 Command-Bus Cutover. PROGRAM INVARIANT: EHDB is noetl-internal-only — a fixed set of predefined datasets (event log, commands, projections, KV, object/blob, vector, catalog, runtime, system-WASM, provider-facts), NOT a general-purpose DB (no arbitrary schemas/DDL, no query planner, no hosted-DB surface; #178/#184 is a read-only export). Refined boundary (write-behind cache): EHDB is never the durable system of record for business data (the customer's connector-backed store is) — the permanent control-plane log stays lean/reference, while a transient processing cache may hold business context but is bounded, sunk to the customer's store, then evicted. EHDB becomes a layered platform, built L0-first — L0 replicated object store (durability/replication foundation, = VictoriaMetrics/VictoriaLogs write engine (buffered flush → immutable parts → background merge → columnar) + a ClickHouse-style meta-catalog (manifest + sparse index → object-store pointer; fixed-schema, no planner) → L1 streaming (the NATS takeover / (c) one-hop delivery) → L2 KV → L3 append-log SQL. L0 resolves the HA debate: durability + replication come from the object store, so writers are fungible (cold-load sealed parts from L0) and the prior per-shard-Raft "T-RF" plan is retired — only a light L1 ordering lease remains. Honest departure: VM keeps the hot path on local disk (object storage is backup-only), so L0's object-store-as- live-durability-tier is net-new beyond VM — the highest-risk piece and the first build slice (L0.1). Big reuse from #254 segments (≈ immutable parts) + Phase-8 object/KV + Phase-7 projection; real gaps = inverted index + merge engine + object-store tier. Cost: L0 ~2–4q (gate), L1 ~1.5q, L2 ~1q, L3 ~1q (SQL-engine ambition cut by the invariant — fixed reads + read-only export only). Umbrella noetl/ai-meta#194.

Repository Links

Conventions

  • Keep product code in noetl/ehdb.
  • Keep durable design pages in this wiki.
  • Track executable work in GitHub issues and the EHDB project board.
  • Keep ai-meta limited to orchestration state, submodule pointers, cross-repo coordination notes, and compact platform memory.
  • Validate containerized NoETL integration in local kind before GKE.
  • Design — Execution Chain Store — the execution-partitioned chain store and its watermark: the three states a reader must distinguish, why the append/watermark ordering is asymmetric (never create authority out of a failure; v0.4.1 must not be adopted), the two measured defects that had blocked it, and — from v0.4.4 — why a log snapshot behind the store is staleness rather than a conflict while the same shape from a second writer must still be caught, so StaleLog is an unresolved report and the verdict belongs to the consumer. ⚠⚠⚠ ROLLED BACK 2026-09-30 — 13 divergences over 839 engagements an hour after a 772/0 reading; the prefix assumption is broken by an event committing into the MIDDLE of an id-ordered read (#362) (noetl/ai-meta#357, #360)
  • Design — Replica Backfill — attaching a replica does not replicate what already exists. Uploads are enqueued on seal and nowhere else, so a substrate added to an existing store leaves every earlier part single-copy forever while survives_node_loss reads 1 (that gauge describes the declared replica set, not the parts). Measured on prod: 40 parts in the remote manifest, 39 at replica_count 1 — and the remote holds a manifest naming all 40, so it reads as a complete store and is not a recoverable one. v0.7.0 (#400)
  • Design — Tail Replication — B3: a part is local-only until it seals, so the loss window was the seal interval. The tail now replicates off-box as write-once per-batch objects and replays on cold load, taking the window from 900 s to the 15 s tick — bounded, not zero. ⚠⚠ v0.8.1 corrects the lock boundary: the remote put must not run under the engine lock the live append path takes (p50 77 ms, p99 234 ms)
  • Embedded-State Foundations — format-version gate, stored cursor, D3/D8 exercised, append-time validation, and the foca gossip crate (noetl/ai-meta#332)

Clone this wiki locally