Repository navigation
Roadmap
This roadmap turns the EHDB vision into issue-sized phases. GitHub issues are the executable source of truth for active work; this page is the durable design index.
Tracking issue: #1 Bootstrap EHDB Rust workspace and CI
Goals:
- Add
noetl/ehdbandnoetl/ehdb.wikitoai-metaas submodules. - Create the initial wiki design pages.
- Create initial tracking issues and add them to the EHDB development project.
- Scaffold a Rust workspace with stable crate boundaries.
- Add CI-ready formatting, linting, and tests.
Acceptance:
-
cargo test --workspacepasses locally. - README explains mission, boundaries, and first developer workflow.
- Wiki has Home, Architecture, Roadmap, and Sessions Log pages.
Tracking issue: #2 Design catalog-as-database metadata model
Goals:
- Define durable identifiers for tenant, namespace, table, snapshot, transaction, and object references.
- Define Arrow-native column/type metadata.
- Implement an in-memory transactional catalog reference model.
- Add MVCC-style snapshot IDs and optimistic transaction records.
Acceptance:
- Catalog table create/read/update tests pass.
- Snapshot metadata is immutable after commit.
- Catalog mutations produce typed transaction records.
Current local reference:
- Follow-up issue: #22 Add catalog table snapshot metadata
- Follow-up issue: #62 Add catalog scan grant reference model
-
InMemoryCatalogcommits immutable table snapshots with snapshot ID, optional parent snapshot, object file refs, and committing transaction ID. - Latest snapshot lookup is tracked per table.
- Snapshot commits reject missing tables, empty file sets, duplicate snapshots, and parent-chain mismatches.
-
CatalogMutation::CommitSnapshotmakes snapshot metadata replayable through the transaction log. -
CatalogScanGrantrecords tenant/namespace/table scan access for a principal, rejects missing-table and duplicate grants, and answerscan_scanlookups. -
CatalogMutation::GrantScanmakes scan grant metadata replayable throughehdb-referenceandLocalReferenceRuntime. - Catalog metadata JSON decoding rejects unknown table, snapshot, scan grant, request, table schema, and column schema fields before replay or catalog operations.
- Table schemas revalidate column identifiers before catalog state is created.
- Table schemas reject duplicate column names before catalog state is created.
- Column schema and table schema JSON decode routes through the same identifier and duplicate-column validation before metadata is accepted.
- Core identifier JSON decode routes through constructor validation while preserving the string JSON shape.
Tracking issue: #3 Define immutable object storage layer
Goals:
- Define an object-store trait for put/get/list/delete-free immutable writes.
- Implement local filesystem adapter for tests.
- Add S3-compatible, GCS, and Azure Blob designs before implementation.
- Define object naming layout for tenant/namespace/table/snapshot.
Acceptance:
- Local adapter can write/read immutable Arrow IPC and Parquet test objects.
- Object references are catalog-addressable and content-checked.
- Object references carry geo-location and data-gravity shard placement pointers for distributed storage design.
- Delete behavior is explicit and conservative.
Current local reference:
- Follow-up issue: #20 Add content-checked immutable object references
- Follow-up issue: #24 Add geo placement and data-gravity shard pointers
- Follow-up issue: #26 Add storage placement policy model
- Follow-up issue: #28 Add deterministic replication planning model
- Follow-up issue: #30 Add durable object replica registry
- Follow-up issue: #32 Add bounded local replication executor
- Follow-up issue: #34 Add Arrow IPC table write/read fixture
- Follow-up issue: #36 Add Arrow snapshot scan fixture
- Follow-up issue: #38 Add Arrow equality filter fixture
- Follow-up issue: #40 Add Arrow scan service API boundary
- Follow-up issue: #42 Add Arrow Flight scan ticket codec
- Follow-up issue: #44 Add Arrow Flight scan result stream codec
- Follow-up issue: #46 Add Arrow Flight scan info fixture
- Follow-up issue: #48 Add local Arrow Flight scan service facade
- Follow-up issue: #50 Add Arrow Flight scan service trait adapter
- Follow-up issue: #52 Add bounded Arrow Flight server lifecycle config
- Follow-up issue: #54 Add loopback Arrow Flight listener harness
- Follow-up issue: #56 Add loopback Arrow Flight client smoke test
- Follow-up issue: #58 Add Arrow Flight auth header policy contract
- Follow-up issue: #60 Add Arrow Flight scan scope metadata guard
- Follow-up issue: #62 Add catalog scan grant reference model
- Follow-up issue: #64 Enforce catalog scan grants in Flight reference path
- Follow-up issue: #66 Add bounded Flight scan access log policy
- Follow-up issue: #68 Add Arrow Flight get_schema scan adapter
- Follow-up issue: #70 Enforce bounded Flight scan request concurrency
- Follow-up issue: #72 Add local retrieval vector similarity fixture
-
ObjectRefcarries path, byte length, SHA-256 digest, geo placement, and data-gravity shard pointer. -
PlacementPolicyvalidates exactly one primary, minimum copy count, shared data-gravity shard, and no duplicate geo/shard targets. -
plan_replicationemits already-satisfied and copy-needed actions from current replicas plus placement policy. -
ObjectReplicaRegistryrecords durable replica inventory and can feed replication planning from replayed EHDB metadata. -
LocalReplicationExecutorverifies source bytes and records copy-needed replica registrations through transaction replay. - Storage metadata JSON decoding rejects unknown object ref, placement, policy target, replica, action, and plan fields before replay or planning.
-
CatalogScanGrantrecords replayable table scan grant metadata for principals before production ACL enforcement exists. -
FlightScanGrantPolicycan requirex-ehdb-principalmetadata and enforce replayedCatalogScanGrantrecords before local Flight scan execution. -
ImmutableObjectStore::get_verifiedrejects length or digest mismatches. - The local Arrow IPC fixture writes a
RecordBatch, commits a catalog snapshot over the content-checked object, and reads it back through the catalog/object boundary. - The local Arrow scan fixture resolves latest snapshots, verifies Arrow IPC objects, decodes batches, and supports named column projection.
- Direct local Arrow scan requests reject empty projection lists and duplicate projection columns before object reads.
- Direct local Arrow scan requests validate projection-column and predicate-column selector identifiers before object reads.
- The local equality filter fixture applies single-column UTF-8 and Int64 equality predicates after verified Arrow decode and before projection.
-
ehdb-serviceexposes a local service-facing latest-table scan request/result boundary over the scanner, returning schema, batches, and row count before Arrow Flight networking exists. -
ScanFlightTicketencodes that latest-table scan request into a versioned Arrow FlightTicketpayload and commandFlightDescriptorfor the future network service. - Local scan request validation rejects malformed selector identifiers, unknown ticket/request/predicate object fields, non-canonical ticket bytes, empty projection lists, and duplicate projection columns before scan execution.
-
ArrowScanResultencodes local scan outputs into Arrow FlightFlightDatamessages and decodes them back into validated results. -
ArrowScanResultbuilds pre-network Arrow FlightFlightInfometadata with schema bytes, command descriptor, endpoint ticket, row count, and encoded byte count. -
LocalArrowFlightServiceprovides in-processget_flight_info,get_schema, anddo_getbehavior over the scan ticket, info, schema, and result codecs. -
LocalArrowFlightServerimplements the generated Arrow Flight service trait for scanget_flight_info,get_schema, anddo_get, with deterministic gRPC statuses, configured request metadata auth, tenant/namespace scan scope metadata, catalog scan grant enforcement, scan command descriptor path rejection, and no bound network listener. -
FlightAccessLogPolicyprovides disabled and debug-only modes for bounded scan access summaries, excluding tokens, principals, tenant/table identifiers, object paths, predicate values, and Arrow payloads. - Implemented scan methods enforce the local
max_concurrent_requestsbudget with fail-fast gRPCRESOURCE_EXHAUSTEDresponses when all request slots are occupied. -
LocalArrowFlightServerConfigvalidates bind, message-size, concurrency, header-token auth, tenant/namespace scan scope, catalog scan grant, and access-log policy before any listener exists. -
LocalArrowFlightListenerbinds loopback-only reference listeners and shuts them down through an explicit future. - The loopback client smoke path validates
get_schema,get_flight_info, schema-aware endpoint-ticket extraction, and decodeddo_getbatches against the decoded schema, returnedFlightInfo, and expected ticket over real tonic/gRPC transport, including the optional header-token auth, tenant/namespace scan scope, and catalog scan grant policies. - Receiver paths revalidate the concrete
do_getendpoint ticket against returnedFlightInfobefore using it. - Local service and server receiver tests validate raw
SchemaResult, returnedFlightInfo, rawFlightData, and expected ticket as one response envelope before accepting decoded rows. - Table data paths follow
{tenant}/{namespace}/tables/{table}/snapshots/{snapshot}/{file}.
Tracking issue: #4 Define transaction log and MVCC snapshot boundary
Goals:
- Define transaction log trait and append/read semantics.
- Implement local file and in-memory reference logs.
- Keep consensus engine pluggable behind the trait.
- Model commit ordering and replay into catalog state.
Acceptance:
- Replay reconstructs catalog state from log records.
- Replay reconstructs reference catalog, stream, retrieval, and system library state from log records alone.
- Duplicate transaction IDs are rejected deterministically.
- Crash/restart simulation works for the local JSONL reference log.
Current local reference:
- Follow-up issue: #8 Add local durable transaction log reference
- Follow-up issue: #16 Make transaction mutations replay-complete
-
InMemoryTransactionLogprovides deterministic ordered replay for tests and benchmarks. -
LocalJsonlTransactionLogwrites one fsynced JSONLTransactionRecordper append and rebuilds replay state on open. - Local open rejects corrupt records, duplicate transaction IDs, and sequence gaps instead of silently repairing the log.
- Local open revalidates transaction envelope and mutation identifiers before accepting persisted records back into ordered replay state.
- Local open rejects unknown transaction record and mutation fields before accepting persisted records back into ordered replay state.
-
ehdb-referenceapplies replayedTransactionRecordvalues into the local reference catalogs and rejects mismatches such as unexpected stream sequence values. -
LocalReferenceRuntimevalidates projected reference state before durable append and rebuilds reference state from replay on reopen. - Follow-up issue: #18 Add local reference runtime over transaction replay
Goals:
- Model the NoETL event/command stream role currently served by NATS JetStream.
- Define typed append-only stream records, subjects, retention policies, durable consumers, replay cursors, and ack state.
- Keep any NATS bridge as a migration adapter, not the target product boundary.
Acceptance:
- A local reference stream can publish, replay, ack, and resume from a durable cursor.
- Consumer state is tenant/namespace aware.
- NoETL execution event semantics can be represented without direct NATS dependency.
Current local reference:
- Follow-up issue: #10 Add local durable stream journal reference
-
InMemoryStreamLogprovides typed stream records, retention, replay, durable consumers, and ack cursors. -
LocalJsonlStreamLogwrites fsynced JSONL entries for stream creation, consumer creation, publish, and ack operations. - Local open rebuilds retained records, next sequence, and durable consumer cursors, and rejects corrupt journal entries.
- Stream logs reject zero max-record retention before creating or journaling streams.
- Stream replay can explicitly filter retained records by subject using
exact matches, single-token
*wildcards, and terminal>tail wildcards. - Durable consumer replay can explicitly filter pending records by subject after the consumer ack cursor without moving that cursor.
- Concrete published stream subjects reject wildcard tokens; wildcard selectors live in explicit subject filters used by replay APIs.
- Concrete stream subjects and subject filters reject empty dot-delimited tokens.
- JSONL stream replay revalidates persisted record subjects before rebuilding retained stream state.
- JSONL stream replay revalidates persisted stream coordinates, consumer names, and stream record transaction IDs before applying journal entries.
- JSONL stream replay rejects zero persisted stream sequences before rebuilding retained records or durable consumer cursors.
- JSONL stream replay rejects unknown journal, stream config, stream record, and durable consumer fields before rebuilding state.
Tracking issue: #12 Add EHDB system WASM library registry
Goals:
- Store NoETL system playbook functionality as compiled WASM library manifests in EHDB.
- Mirror NoETL worker WASM dispatch with
{ path, version, digest, entry }module refs while keeping EHDB as the durable resolver. - Bind modules by tenant, namespace, environment, release channel, and logical path.
- Allow hot replacement by rebinding a channel to a new digest/revision instead of requiring crate semantic-version bumps for every fix.
- Support different implementations for different environments such as
kind,gke-prod,aws-prod, andazure-dev.
Acceptance:
- A local reference catalog can publish immutable WASM module manifests.
- Environment/channel bindings resolve deterministic active modules.
- Rebinding a stable channel replaces the implementation while retaining prior immutable module manifests.
- System library publish/bind events are represented in the transaction log and participate in cross-domain replay tests.
- A local JSONL system-library journal preserves publish/bind state and hot-replacement bindings across restart.
- WASM host execution remains a worker/system-pool concern; EHDB stores manifests, bindings, object references, capabilities, and transaction provenance.
Current local reference:
- Follow-up issue: #14 Add local durable system library registry journal
-
InMemorySystemLibraryCatalogprovides module publish, bind, and resolve behavior. -
LocalJsonlSystemLibraryCatalogwrites fsynced JSONL entries for publish and bind operations. - Local open rebuilds immutable manifests, environment/channel bindings, and hot-replacement state, and rejects corrupt journal entries.
- Local open revalidates persisted manifest and binding identifiers before rebuilding system-library state.
- Local open rejects unknown system-library journal, publish, and bind fields before rebuilding hot-replacement state.
- System-library metadata JSON decoding rejects unknown resolved manifest, plugin reference, and binding fields before worker handoff metadata is accepted.
Goals:
- Add service crate boundary for catalog and scan APIs.
- Implement Arrow Flight read path for catalog/table scan fixtures.
- Add client smoke tests.
Acceptance:
- CLI can create a table, write a small batch, and read it back as an Arrow stream.
- Flight endpoint has bounded logging and avoids high-frequency INFO noise.
Current local reference:
- Follow-up issue: #40 Add Arrow scan service API boundary
- Follow-up issue: #42 Add Arrow Flight scan ticket codec
- Follow-up issue: #44 Add Arrow Flight scan result stream codec
- Follow-up issue: #46 Add Arrow Flight scan info fixture
- Follow-up issue: #48 Add local Arrow Flight scan service facade
- Follow-up issue: #50 Add Arrow Flight scan service trait adapter
- Follow-up issue: #52 Add bounded Arrow Flight server lifecycle config
- Follow-up issue: #54 Add loopback Arrow Flight listener harness
- Follow-up issue: #56 Add loopback Arrow Flight client smoke test
- Follow-up issue: #58 Add Arrow Flight auth header policy contract
- Follow-up issue: #60 Add Arrow Flight scan scope metadata guard
- Follow-up issue: #64 Enforce catalog scan grants in Flight reference path
- Follow-up issue: #66 Add bounded Flight scan access log policy
- Follow-up issue: #68 Add Arrow Flight get_schema scan adapter
- Follow-up issue: #70 Enforce bounded Flight scan request concurrency
- Follow-up issue: #182 Validate Arrow Flight scan command descriptor path
- Follow-up issue: #184 Reject unknown Arrow Flight scan ticket fields
- Follow-up issue: #186 Require canonical Arrow Flight scan ticket encoding
-
LocalArrowScanServicewrapsLocalArrowSnapshotScannerwith a typed latest-table scan request andArrowScanResultcontaining schema, batches, and row count. -
ScanFlightTicketround-trips requests through versioned Arrow FlightTicketbytes and command descriptors, with unsupported versions and malformed payloads rejected before scan execution; decoded tenant, namespace, table, projection-column, and predicate-column identifiers are revalidated before execution/handoff, and unknown object fields plus non-canonical bytes and empty or duplicate projection selectors are rejected. - Local scan
get_flight_infoandget_schemarequest decoding rejects path descriptors and command descriptors with non-empty paths before scan execution. -
ArrowScanResultround-trips through Arrow FlightFlightDatamessages, preserving schema and row counts while rejecting empty, missing-version, unsupported-version, extra-metadata, or malformed streams; receiver-side validation can compare returned data againstFlightInfoschema, row count, and byte-count metadata. -
ArrowScanResultcan produce and validate a single-endpoint orderedFlightInfofixture from a scan ticket and result, including decodable schema IPC metadata, descriptor/ticket request consistency, positive byte-count metadata, result/ticket consistency, expected ticket validation, schema-response consistency, and a pre-network endpoint envelope. -
LocalArrowFlightServiceexecutes localget_flight_info,get_schema, anddo_getpaths without opening a network listener. -
LocalArrowFlightServeradapts the local facade to the generated Arrow Flight service trait forget_flight_info,get_schema, anddo_get; all other Flight methods return explicit unimplemented statuses. The implemented scan methods enforce the configured request metadata auth, scan scope, and catalog scan grant policies. -
FlightAccessLogPolicycontrols bounded DEBUG-only scan access summaries for decoded scan requests; disabled mode emits no summaries. Summaries exclude tokens, principals, tenant/table identifiers, object paths, predicate values, and Arrow payloads. - Implemented scan methods enforce the configured local request budget
with fail-fast gRPC
RESOURCE_EXHAUSTEDresponses. This remains a local-reference guard, not a scheduler or distributed admission controller. -
LocalArrowFlightServerConfigvalidates lifecycle guardrails and applies message limits, the reference header-token auth policy, and tenant/namespace scan scope, catalog scan grant, and access-log policies when constructing the generated service wrapper. -
LocalArrowFlightListenerbinds a loopback-only reference listener, exposes the actual bound local address, and shuts down through an explicit future. - The loopback client smoke path uses Arrow Flight client calls over
tonic/gRPC transport to validate
get_schema, expected-ticketget_flight_infovalidation, rawSchemaResultand schema-response consistency, and schema-aware validated endpoint-ticket extraction plus concrete endpoint-ticket binding beforedo_getdecoded-schema/data/FlightInfo/expected-ticket consistency checks, including the optional header-token auth, tenant/namespace scan scope, and catalog scan grant policies. - Local service and server paths use the same raw
SchemaResult/FlightInfo/FlightData/expected-ticket envelope binding before accepting decoded rows. - This remains a local-reference network boundary: no non-loopback exposure, server daemon manager, production TLS/identity, request scheduler, SQL planner, predicate pushdown, distributed execution, or gateway direct reads are introduced yet.
Goals:
- Model NoETL-native documents, chunks, embedding metadata, vector index metadata, and retrieval policies.
- Preserve tenant, execution lineage, artifact references, and embedding model identity with every retrieval object.
- Define a path to replace a permanent Qdrant dependency with EHDB-native retrieval primitives.
Acceptance:
- Documents and chunks can be registered in the catalog model.
- Embedding metadata records include model identity, dimensions, checksum/source lineage, and placement policy.
- A local reference retrieval index can answer filtered lookup fixtures.
-
VectorSearchanswers tenant/namespace/model-scoped exact cosine similarity over registered chunk embeddings with deterministic hit ordering and finite non-zero vector validation. - Retrieval metadata JSON decoding rejects unknown document, chunk, embedding, registration request, search request, and local search hit fields before replay or handoff.
-
LocalRetrievalSearchServiceprovides an in-process service-facing request/result boundary over replayed retrieval state and returns ranked chunk hits without raw embedding vectors. -
SearchTextChunksRequestadds exact local text matching with tenant/namespace scoping, match counts, deterministic ordering, and validation. -
SearchHybridChunksRequestadds weighted exact hybrid scoring over vector cosine similarity and text match counts, scoped by tenant, namespace, and embedding model. -
AssembleRetrievalContextRequestbuilds bounded local RAG context blocks from replayed hybrid search hits, preserving citation metadata, clipped text, score metadata, and total text budget accounting. -
RetrievalContextRequestPayloadandRetrievalContextResultPayloadprovide versioned local JSON byte codecs for context assembly request/result handoff, rejecting invalid decoded identifiers and unknown request/result/context/block fields plus non-canonical payload bytes before execution or handoff. -
LocalRetrievalSearchService::execute_context_payloaddecodes a request payload, assembles context from replayed local retrieval state, and returns an encoded result payload for worker/playbook tests. -
RetrievalContextPayloadExecutorConfigbounds local request/result payload bytes and rejects oversized payloads deterministically. -
RetrievalContextPayloadScopeoptionally guards local payload execution by matching decoded request tenant/namespace to the expected worker/playbook execution scope. -
RetrievalContextPayloadExecutionSummaryreports redacted local execution metadata: request/result byte counts, context block count, total text chars, truncation status, and whether scope was required, while excluding tenant IDs, namespace values, query text, chunk text, tokens, vectors, payload bytes, object paths, and principals. -
RetrievalContextPayloadExecutionReceiptPayloadprovides a versioned JSON byte codec for those redacted summaries, giving future event-log/audit plumbing a durable receipt shape without event publication or sensitive retrieval content. Receipt encode/decode validates positive payload byte counts and rejects text chars without context blocks, while rejecting unknown receipt envelope and redacted summary fields and non-canonical JSON bytes before artifact validation or event-envelope helper use. -
RetrievalContextPayloadExecution::encode_receipt_payloademits that receipt directly from a local execution result for worker/playbook tests. -
RetrievalContextPayloadExecutionArtifactsreturns result payload bytes together with bounded redacted receipt payload bytes for local worker/playbook handoff tests, and validates receipt/result consistency before callers treat the artifact pair as coherent. -
RetrievalContextPayloadExecutionReceiptEventPayloadwraps validated receipt bytes in a versioned JSON event envelope with stable subjectehdb.retrieval.context.execution.receipt, giving future stream/audit plumbing a local payload shape without automatic publication or result/context bytes. Event payload decoding rejects unknown envelope fields and non-canonical JSON bytes before replay decode, publisher helper use, or consumer handoff. -
RetrievalContextReceiptEventStreamTargetexplicitly builds and creates the receipt event stream with caller-selected retention; publish helpers do not auto-create streams. - Keep-all and positive bounded-retention setup helpers make receipt event stream retention explicit; zero bounded retention is rejected before stream creation.
-
RetrievalContextReceiptEventStreamTargetandRetrievalContextReceiptEventStreamLogexplicitly publish that event envelope to caller-supplied local stream logs, requiring caller-owned tenant, namespace, stream, mutable log, and transaction id. -
RetrievalContextReceiptEventStreamRecordandRetrievalContextReceiptEventStreamReadLogexplicitly replay and decode receipt event records from caller-supplied local stream logs, validating stable subject and payload while preserving sequence and transaction id for audit assertions. -
RetrievalContextReceiptEventDurableConsumerLogexplicitly creates local durable consumers, replays pending validated receipt events for a consumer, and acks receipt event sequences without a background subscription loop. - This remains a local reference fixture: no ANN index, full-text index, retrieval service, RPC protocol, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, gateway data path, query planner, Qdrant/search adapter, production IAM, ACL engine, scheduler, or distributed query engine.
Tracking issue: #5 Plan NoETL integration path for EHDB system store
Goals:
- Add NoETL playbook step integration for EHDB catalog operations.
- Run local kind validation before any GKE rollout.
- Keep gateway as gatekeeper and NoETL workers as atomic compute.
Acceptance:
- A NoETL playbook can write and read EHDB-backed metadata in local kind.
- No gateway direct database touch is introduced.
- Wiki and NoETL docs are updated with the public integration surface.
Current local reference:
- Follow-up issue: #228 Add NoETL runtime surface replay fixture
- Follow-up issue: #230 Add embedded NoETL role capability policy
- Follow-up issue: #232 Add local-reference summary helper
-
LocalReferenceRuntimehas a NoETL-shaped replay fixture covering catalog table/snapshot/grant metadata, stream events, retrieval document/chunk/embedding metadata, system WASM library binding, and storage replica inventory from the transaction log alone. - The fixture reopens the runtime and verifies catalog scan grants, stream consumer replay, retrieval text lookup, system library resolution, and storage replica counts without hand-mutating domain catalogs.
-
ehdb-corenow models embedded NoETL roles and capabilities so EHDB can be present inside workers, APIs, and gateways while gateway/API roles remain control-plane only and worker/playbook/system roles carry explicit data-plane permissions. -
ehdb-referencenow exposesLocalReferenceSummaryplus theehdb-local-reference summary --log <path>helper for deterministic replayed-domain counts across local transaction, catalog, stream, retrieval, system-library, and storage state.
Integration handoff phases (from the Claude handoff page):
-
Phase A — ops/runtime enablement: DONE (2026-07-04). The NoETL
ops Helm charts now render disabled-by-default, role-specific EHDB env
with the control-plane vs data-plane boundary enforced in the chart
template (
noetl/ops#234). server/api/gateway get control-plane-only env; worker/playbook/system get boundedlocal_referenceenv with a pod-local JSONL log. A kind Job (ci/manifests/noetl/ehdb/smoke-job.yaml) runs the packagedehdb-local-referencesmoke inside the NoETL image before any GKE rollout. Disabled render is byte-identical to pre-EHDB. See the Architecture page's Ops Env-Rendering Boundary section. Tracks #234. -
Phase B — worker/playbook readiness hook: DONE (2026-07-04). A
bounded, stateless readiness preflight over
read_ehdb_local_reference_summary_from_envlands innoetl/noetl(noetl.core.ehdb_readiness, PR noetl/noetl#688). It is wired for worker/playbook/system data-plane roles only; gateway/api/server never call it and a code-level guard (assert_data_plane_read_allowed) refuses a data-plane read for any control-plane role (defense-in-depth on the contract validation). The read is time-bounded (default 5s, clamped 0.1–30s) and holds no connection/state. Observability:noetl_ehdb_readiness_*metrics on the worker/metricssurface (no secret values). Disabled-by-default → strict no-op (byte-identical, no metric recorded). Exposed as a worker/playbook-local preflight command (scripts/ehdb_readiness_preflight.py) / kind smoke step — not a server endpoint. Kind-validated onkind-noetl. See the Architecture page's Worker/Playbook Readiness Hook section. Tracks #234. -
Phase C — bounded worker/playbook data-plane step: DONE (2026-07-04).
A NoETL playbook step (worker tool) that performs a bounded EHDB
local-reference operation — append and read a single domain record
through the adapter — rather than only a readiness summary. Landed in
noetl/noetl(noetl.core.ehdb_dataplaneappend_ehdb_domain_record/read_ehdb_domain_records+ adapterLocalReferenceAppendResult/LocalReferenceReadResult, PR noetl/noetl#689) over theehdb-local-reference append/readsubcommands (noetl/ehdb#235, merged3ae8950). Still disabled-by-default (strict no-op, byte-identical/metrics), still worker/playbook/system-only with a code-level control-plane guard (assert_data_plane_access_allowedrefuses gateway/api/server before any helper runs), bounded (payload byte cap + read-limit cap + short time cap), and stateless (helper opened+dropped per call). Secret-freenoetl_ehdb_dataplane_ops_total{operation,outcome}metrics. Exposed as a worker/playbook-local step CLI (scripts/ehdb_dataplane_step.py) / kind smoke — not a server endpoint. Kind-validated onkind-noetl(disabled no-op + byte-identical/metrics, worker/system/playbook append→read roundtrip, control-plane gateway guard refused, secret-free metrics). See the Architecture page's Worker/Playbook Data-Plane Step section. Tracks #234. -
Phase D — event-stream integration path: DONE (2026-07-04). A
bounded, disabled-by-default worker/playbook/system drain that
mirrors already-emitted NoETL events into a derived EHDB stream and
consumes them through a durable consumer with explicit
ack-after-materialize semantics. EHDB adds
consume/acksubcommands +consume_local_reference_event_records/ack_local_reference_event_consumer(composing with the Phase Cappendproject leg) and a read-onlyInMemoryStreamLog::consumercursor getter (noetl/ehdb#237, merged3cefba9). NoETL addsnoetl.core.ehdb_eventstream(project/consume/ack) + adapter consume/ack result types + worker/metrics+ step CLI + kind smoke (noetl/noetl#690). Still disabled-by-default (strict no-op, byte-identical/metrics), still worker/playbook/system-only with the code-level control-plane guard (assert_event_stream_access_allowed), bounded (payload + consume limit + ack sequence ≥ 1 + time caps), stateless, secret-freenoetl_ehdb_eventstream_ops_totalmetrics. Event-log-authoritative invariant: the NoETL event log stays the append-only source of truth; EHDB is a derived, auxiliary consumer that never writes back to it (no event-writer import; unit-test asserted). kind-validated onkind-noetl(disabled no-op + byte-identical/metrics, worker drain- durable cursor restart, gateway guard refused, secret-free metrics, Phase C no-regression). See the Architecture page's Worker/Playbook Event-Stream Drain section. Tracks #234.
-
Phase E — system WASM store + RAG retrieval: COMPLETE (2026-07-05,
RUST-FIRST). The EHDB side landed as
#239 (ehdb
9bb5928): boundedpublish/bind/resolvehelpers for immutable system WASM library manifests + mutable environment/channel bindings inehdb-reference. The worker side landed as noetl/worker#154 (merge3162b1d): a new in-processsrc/ehdb/systemstore.rsbridges those helpers — publish (one atomic immutable-manifest commit), bind (channel (re)bind; rebind hot-replaces the active module, prior manifests retained), resolve (read-only replay; never-bound ⇒Absentprobe) — plus anoetl_ehdb_systemstore_*metric family andehdb-selfcheckpublish-system/bind-system/resolve-system + asystem-suiteA→E driver. Every boundary the Phase C/D helpers enforce is preserved and tested: disabled-by-default strict no-op (byte-identical/metrics), control-plane guard (gateway/api/server refused before any runtime opens), bounded (module-size + capability-count caps; over-bound publishesRejected; WASM execution stays host-side/sandboxed — EHDB only catalogs the module ref), stateless (runtime opened + dropped per call), event-log-authoritative (private JSONL, nevernoetl.event). Kind-validated against the worker-rust image (localhost/local/noetl:ehdb-phase-e,kind-noetl, distinctehdb-phase-e-selfcheckpod): disabled no-op (empty metrics), enabled A–D drive, enabled Phase-E system-suite (absent→publish→bind→resolve rev1→ publish→rebind→resolve rev2), control-planeserverrefused (exit 4, no data), oversized publishRejected(exit 3), 0noetl.eventwrites, secret-free metrics, Rust stack 0 restarts. No GKE. Per the owner directive (2026-07-04) Phase E is Rust-first: the integration lives inworker-rustand invokes theehdbcrate in-process (no subprocess), Python retired. Second slice — bounded RAG retrieval — DONE (2026-07-05). The prerequisite ehdb slice landed as #240 (ehdbc2aaad5): a bounded, read-onlyretrieve_local_reference_contexthelper inehdb-referencethat runsehdb-retrieval'ssearch_textunder the hood, plus aningest_local_reference_retrieval_documentcompanion andehdb-local-referenceingest-doc/retrieveCLI verbs. Three caps are enforced inside the helper: top-k (MAX_RETRIEVAL_TOP_K= 64), per-hit result size (MAX_RETRIEVAL_MAX_CHUNK_BYTES= 64 KiB, chunk text truncated on a char boundary), and a wall-clock budget (MAX_RETRIEVAL_TIME_BUDGET_MS= 60 s, surfaced viatime_capped). Over-ceiling caps ⇒Rejected; empty query / bad tenant·namespace id ⇒Invalid— both classified without searching so a CLI maps them to distinct exit codes. The worker side landed as noetl/worker#155 (merged1ebaf2): a new in-processsrc/ehdb/rag.rsbridges the helpers —retrieve(bounded, read-only text search;NOETL_EHDB_RAG_TOP_K8/64,NOETL_EHDB_RAG_MAX_CHUNK_BYTES4 KiB/64 KiB,NOETL_EHDB_RAG_TIME_BUDGET_MS5 s/60 s) andingest— plus anoetl_ehdb_rag_*metric family andehdb-selfcheckingest-rag/retrieve-rag/rag-suite(A→E+RAG driver). Every Phase C/D/E boundary is preserved and tested: disabled-by-default no-op (byte-identical/metrics), control-plane guard (gateway/api/server refused before any runtime opens), bounded, stateless, event-log-authoritative (retrieve read-only; ingest writes only the private JSONL fabric, nevernoetl.event). Validated via the builtehdb-selfcheckA→E+RAG drive: disabled no-op (empty metrics), enabledrag-suite(ingest 3 chunks → retrieve hit with top-k truncation → empty → over-limit rejected,ok:true,metrics_secret_free), over-limitRejected(exit 3), control-planeserverrefused (exit 4, no retrieval), secret-freenoetl_ehdb_rag_*metrics, plus 8 unit tests. In-containerkind-noetlvalidation deferred — the worker-rust image (localhost/local/noetl:ehdb-rag, distinctehdb-rag-selfcheckpod) is a ~110-min cold build (theehdb-referencerev bump busts the cargo-chef dependency cache) that timed out in this environment; the builtehdb-selfcheckdrive above is the standing evidence and the same behaviours were kind-validated for the Phase C/D/E first slices. No GKE. With both slices done Phase E is complete; the broad EHDB-integration umbrella #234 stays open for the remaining roadmap (Phase 6 PostgreSQL replacement path, Phase 7 dependency collapse).
The Phase B–D worker/playbook integration glue (readiness hook, data-plane
step, event-stream drain) was originally added in the legacy Python worker
runtime (noetl/core), shelling out to the Rust ehdb-local-reference
binary as a subprocess. Prod runs worker-rust, so those disabled-by-default
Python hooks never executed in prod — a stopgap.
Re-homed (owner directive 2026-07-04): the EHDB worker/playbook integration
is now Rust-only. worker-rust owns it in process
(noetl/worker src/ehdb),
calling the ehdb-reference crate directly (summarize / append / read /
consume / ack) with no subprocess shell-out — readiness (non-fatal
bootstrap preflight), bounded data-plane append/read, and the event-stream
project/consume/ack durable-consumer drain. Every boundary the Python stopgap
enforced is preserved and tested: disabled-by-default strict no-op
(byte-identical /metrics), control-plane guard (gateway/api/server refused),
bounded + stateless, secret-free noetl_ehdb_* metrics, and the
event-log-authoritative invariant (structurally asserted — no NoETL
event-writer import).
-
worker-rust — noetl/worker#153
MERGED (
d6226a2). Shipsehdb-selfcheckfor in-image validation. -
Python retire — noetl/noetl#691
MERGED (merge commit
ff3a920f) — removes the Python EHDB modules / wiring / bundled helper binary so the Python path can never run as a parallel implementation; adds a guard test. noetlmainnow carries zeronoetl/core/ehdb_*modules. - Kind-validated against the worker-rust image (
kind-noetl): disabled no-op (byte-identical), enabled worker full drive, cross-process durable cursor, control-plane guard refused (no write), secret-free metrics, event-log-authoritative. No GKE. Rust stack undisturbed (0 restarts).
Both PRs merged; noetl/ehdb#238 is CLOSED — the re-home is complete. Phase E (system WASM store → RAG) is Rust-first from the start so it does not re-open this debt.
Decided by RFC: EHDB Completion Program — Server↔EHDB Coupling + noetl Self-Sufficiency (status: decided, implementing). Tracked as noetl/ehdb#241 (with the integration umbrella #234).
The completion trajectory makes EHDB NoETL's self-sufficient internal storage fabric — a family of core engines (event-log, projection, KV/state, object/blob, vector) under one catalog/URN namespace. The end-state: the NoETL platform runs on Kubernetes with no external infrastructure dependency for platform functionality, EHDB the default backend for every platform tier, every tier still selectable back to the incumbent engine.
Platform-only boundary (load-bearing): EHDB serves NoETL platform functionality only — event log, projections, platform KV/state, catalog, system-WASM store, platform artifacts/vector. Business data is never stored in EHDB — tenant/domain data stays in real business systems (PostgreSQL, Snowflake, Cassandra, Kafka, Elasticsearch, ClickHouse, object stores, …) reached via playbook plugins/connectors under playbook policy. "EHDB replaces PostgreSQL / Qdrant / NATS-JetStream / object store" always means NoETL's internal platform uses of those engines, never the business-facing connector targets.
The old "Phase 6: PostgreSQL Replacement Path" and "Phase 7: Dependency Collapse Path" (#6) are absorbed into this trajectory: Phase 6 (log store) + Phase 7 (projections) cover the Postgres-materializer replacement; Phases 8–9 cover dependency collapse; Phase 10 is the tunable driver surface.
Program status (2026-07-06): CODE-COMPLETE. Phases 6–8 built + shadow-verified every platform-tier engine; Phase 9 activated each tier's reversible primary cutover (all five IMPLEMENTED + MERGED + in-kind dual-run VALIDATED); Phase 10 landed the consolidated Backend Configuration surface. The whole program is code-complete and kind-validated — the only remaining step is the prod/GKE cutover, gated on the user, per tier. Nothing in prod has changed (prod worker
v5.52.0; allNOETL_EHDB_*flags default off).
Status: in progress — engine + disabled-by-default shadow slice landed
(2026-07-05); the shadow mirror is now wired into the LIVE event-emit path and
PROVEN on real in-kind drives (worker v5.67.0, 2026-07-06). Design note:
Event-Log Core Engine (Phase 6).
Cutover to serving the log from EHDB (primary) stays a later,
separately-gated step (Phase 9 tier 1, below).
⚠ Superseded 2026-08-13. That cutover has happened: the event-log tier is
primaryand serving on prod, reads resolving through the writer-fronted tier service. Its mirror went asynchronous on 2026-08-19 — see Runbook: the async event-log mirror. Read this section as the plan of record, not current state.
The headline motivation. EHDB's event-log core engine becomes the
durable persistence + ordering + serving layer for the append-only
noetl.event log, replacing the NATS JetStream + PostgreSQL
log-and-store path that is today's scaling pressure point (off-server
state builder, unbounded WAL index
noetl/ai-meta#166,
materializer soak
noetl/ai-meta#104).
Event authorship is unchanged: gateway/server remain the gatekeeper of what enters the log through the append-only producer path; non-log data-plane roles never fabricate events. What changes is the engine underneath the producer path.
Acceptance:
- EHDB event-log engine persists + orders + serves
noetl.eventwith append-only, immutable, replay-is-truth semantics preserved. - A NoETL execution flow drives end-to-end with EHDB as the log engine (behind the driver, dual-run against JetStream+Postgres for verify).
- Rollback to JetStream+Postgres documented; kind-validated before GKE.
Landed so far (design + shadow slice):
-
ehdb-reference::eventlog— the event-log core engine behind theEventLogDrivertrait (append/scan_global/read_execution/tail/ack), withLocalReferenceEventLogDrivercomposing the append-only stream primitives over one canonicalnoetl_event_logstream so its sequence is the global, monotonic, gapless event-log sequence. Per-execution scope vianoetl.event.exec.<execution_id>subject; durable-consumer tail/ack;compare_shadow_parity. CLIeventlog-*verbs. (ehdb PR #242, merged9a9b28d.) -
worker-rust src/ehdb/eventlog.rs— disabled-by-default shadow behindNOETL_EHDB_EVENTLOG=off|shadow|primary(defaultoff):shadowdual-writes each already-authored event into the engine + compares sequence/count/order parity without serving reads or touching the authoritative path;primaryrecognised but not activated (compile-time guard). ⚠ No longer true as of 2026-08-13 —primaryis activated and serving on prod; see the banner on Home. Secret-freenoetl_ehdb_eventlog_*metrics; control-plane guard;ehdb-selfcheck mirror-eventlog/eventlog-suite. -
Runtime mirror wired LIVE + proven on a real drive (worker
v5.67.0, worker#167 mergedd310c7b, releaseae4164d, 2026-07-06). The shadow mirror now fires from the real event-emit chokepointControlPlaneClient::emit_event(every worker path — EventEmitter, retry, spool, subscription, plugin — funnels through it), not only viaehdb-selfcheck/tests. Armed once at construction whenNOETL_EHDB_ENABLED+NOETL_EHDB_EVENTLOG=shadow+ data-plane role + local-reference log; disabled / off / primary / control-plane role ⇒ strict per-event no-op (byte-identical/metrics).mirror_live_eventis panic-isolated + best-effort — a mirror failure surfaces as a metered non-ok outcome and never propagates into the authoritative event path; shadow never serves; event authorship untouched. Proven live in thekind-noetlstack (SHADOW, both data-plane pools, LOCAL only): real drivesautomation/pft_sql_probe_v2→ executions332760742153424896and332760854506246144each mirrored their exactly-6 events into the reference tier (subjectnoetl.event.exec.<id>),noetl_ehdb_eventlog_ops_totalmirror/mirroredadvancing,last_ok=1/last_degraded=0, 0 restarts, while Postgresnoetl.eventpersisted all events unaffected. Only the eventlog tier is wired live so far — projection/kv/object/vector still mirror viaehdb-selfcheck/tests only. GOTCHA: the pool is not sticky onv5.67.0— a redeploy/recreate reverts it tov5.66.0and the mirror goes dark until re-rolled. Prod unchanged. See the Sessions-Log entries for the live-drive evidence.
Remaining before Phase 7: production segmented disk format + offset
index + compaction for primary-serve; sharded/multi-stream global
ordering; a JetStream+Postgres EventLogDriver for the tunable surface
(Phase 10); the projection engine (Phase 7) attaching to the tail/ack/
read-execution serving surface; primary-serve cutover (dual-run verify +
documented rollback, kind before GKE).
Status: first slice landed (2026-07-06) — the durable segment store +
crash recovery the Phase-6 note deferred as "production disk format", and
the hard blocker the
prod-cutover runbook §C durability gate
names for Stage C (the only backend today, local_reference, is a pod-local
JSONL file lost on restart + divergent across replicas → not
production-durable as the authoritative store under primary). Design note:
Durable Event-Log Backend; program +
slice checklist tracked in
noetl/ehdb#254.
Landed (ehdb#253):
ehdb-reference::durable_eventlog — DurableSegmentStore (append-only
CRC32-framed seg-*.eslog segments + size rollover, in-memory offset index,
fsync-per-append, crash-recovery replay with torn-tail discard + bit-rot
hard-error) and DurableEventLogDriver implementing the same
EventLogDriver contract, behind the EventLogStorageBackend
(local_reference | durable_segment) selector with local_reference the
default. Payloads cold-loaded via the index (bounded index memory, the
#166 property). CLI
durable-eventlog-recovery proves zero-loss + ordering + scope +
durable-cursor survival across a simulated restart. 17 tests incl. parity vs
LocalReference; clippy -D warnings + fmt + workspace test green.
In-container kind validation via the worker image is PENDING — follow-up
deploy.
Remaining slices: execution-affinity single-writer routing (XxHash64
ownership, reads route-to-owner / cold-load) → shared object-store segment
tier → worker wiring (NOETL_EHDB_EVENTLOG_BACKEND) → kind restart soak →
prod-durability sign-off (clears §C; Stage C still user-gated per tier).
Status: in progress — design + engine slice merged
(ehdb#243, e0f1c0f) +
disabled-by-default worker shadow wiring merged
(worker#157, eadc3a5)
(2026-07-05). Design note:
Projection / Read-Model Engine (Phase 7).
Read-cutover off Postgres is NOT part of this phase — it stays a later,
separately-gated step (Phase 9).
EHDB's projection engine builds + serves the materialized read-models
(execution / event / runtime state) off the event log, retiring the
PostgreSQL materializer and projected state tables. It consumes the
Phase-6 EventLogDriver tail, materializes the same read-models the Postgres
materializer produces (event read-model keyed on event_id; folded
execution-state; durable consumer checkpoint), and stays consistent under the
#103 sole-writer rule (it only reads the log; the materializer stays the sole
noetl.event writer).
Landed this slice:
-
ehdb —
ehdb-reference::projection: theProjectionDrivertrait (apply/read_execution_state/read_event/list_executions/checkpoint) +LocalReferenceProjectionEngine, composing the append-only stream primitives over onenoetl_projection_logstore. Idempotent / exactly-once apply keyed on the Phase-6 global sequence (skip<= checkpoint-
event_iddedup); deterministic rebuild-from-log;compare_projection_parity. CLIprojection-apply/projection-read-exec/projection-read-event/projection-list/projection-checkpoint/projection-from-eventlog/projection-suite.
-
-
worker (worker#157,
eadc3a5) —src/ehdb/projection.rsshadow behindNOETL_EHDB_PROJECTION=off|shadow|primary(defaultoff):shadowdual-materializes the read-models from the event-log tail alongside the Postgres materializer + compares parity without serving reads or touching the authoritative materializer;primaryrecognised but not activated (compile-timePRIMARY_SERVE_ACTIVATED = false). Pinsehdb-referenceto the merged #243 reve0f1c0f.
Acceptance (met by this slice, except the gated cutover):
- ✅ Read-models materialized from EHDB projections match the Postgres materializer output under dual-materialize (key / value / checkpoint parity, zero divergence).
- ✅ Projection rebuild-from-log is deterministic and bounded (batch-boundary independent, unit-proved).
- ✅ Read-cutover + rollback-to-Postgres documented — primary-serve IMPLEMENTED + MERGED as Phase 9 tier 2 (2026-07-05, ehdb#248 + worker#162); kind dual-run VALIDATED (2026-07-05, in-cluster); prod/GKE cutover STILL GATED on user. See Phase 9 tier 2 below.
Remaining before Phase 8: production projection store format (segmented +
indexed) + compaction for primary-serve; richer execution-state fold (full
projection_snapshot column parity); sharded/multi-store projection aligned
with server shard publish; a Postgres-materializer ProjectionDriver for the
Phase-10 tunable surface; the primary read-serving cutover off Postgres
(Phase 9, separately gated, dual-run verify + rollback, kind before GKE).
Status: in progress — all three engine slices now shadow-complete (2026-07-05): the KV/state engine slice (ehdb#244, worker KV shadow worker#158), the object/blob engine slice (ehdb#245, worker object shadow worker#159, v5.59.0), and the vector engine slice (ehdb#246, worker vector shadow worker#160, v5.60.0) — all three disabled-by-default worker shadows. Phase 8 stays in progress even though every engine is built + shadow-complete: no tier is cut over to primary — per-tier primary cutover is the separately-gated Phase 9 step. Design note: KV/State + Object/Blob + Vector Engines (Phase 8).
Runtime live-wiring (following the event-log tier, Phase 6). The Phase-8
shadow engines were exercised only by ehdb-selfcheck; wiring them into the
worker's real runtime paths is a follow-on. Status (worker
worker#168, v5.68.0, 2ec2e2b):
-
KV — wired live.
kv::mirror_live_putinvoked inSpoolRuntime::persist_circuitafter the authoritative NATS-KV circuit-stateput(bucketnoetl_subscription_circuit; the only live platform NATS-KV write in the worker). -
object — wired live.
object::mirror_live_putinvoked inControlPlaneClient::object_put, the single chokepoint every platform object tier (result-tier, state-shard, plugin intent) funnels through — mirrors ALL live object puts with digest parity. -
projection — deferred.
shadow_projectis a batch materialize that reads back the whole accumulating projection log and compares against a full authoritative execution-state fold; a per-event hook (atemit_eventor the off-server state builder'sWalEventIndex::apply) would report persistent false key-divergence. Faithful seam = a bounded windowed batch drive that also supplies the incumbent materializer's fold — larger than a call-site hook. - vector — deferred. No live platform vector-upsert exists in the worker loop today (platform RAG is in-process, read-only); a hook now would never fire. It lands with a future platform-RAG ingest/embed write site.
Each live hook arms only for NOETL_EHDB_ENABLED + NOETL_EHDB_<TIER>=shadow +
a data-plane role, is a strict no-op otherwise, and is error-isolated
(best-effort, metered, never propagated). Live-drive proof is pending the next
image redeploy (the live pool is on v5.67.0; v5.68.0 carries these hooks).
EHDB takes over NoETL's internal platform tiers currently on NATS KV (coherence state — chain heads, exec descriptors), the external object store (platform artifacts, result-tier payloads, Arrow IPC), and Qdrant (platform RAG/retrieval, catalog embeddings).
Acceptance:
- Platform KV/state, object/blob, and vector reads/writes served by EHDB
engines under dual-run verify against the incumbents.
- ✅ KV/state engine + disabled-by-default worker shadow (parity presence/value/TTL).
- ✅ Object/blob engine + disabled-by-default worker shadow (content-addressed; parity digest/length/retrievability) for state-shards + result-tier.
- ✅ Vector engine + disabled-by-default worker shadow (bounded cosine top-k;
parity id-set/rank-order/score-monotonicity) over the platform RAG
collections — the last Phase-8 engine slice, formalizing the in-process
Phase-E retrieval path behind a
VectorDriver.
- The platform-only boundary holds — no business-connector target moves.
- Per-tier rollback documented; kind-validated.
Phase 8 → Phase 9 boundary. All three engines are shadow-complete, but the box below stays unchecked — a shadow proves parity, it does not serve reads. Flipping any tier to primary (serving reads from EHDB, retiring the incumbent for that tier) is a Phase 9 step, gated and per-tier. See Phase 9 for the five per-tier cutovers this unlocks.
With Phases 6–8 landed (every platform-tier engine built + shadow-verified), the NoETL platform runs self-sufficient on Kubernetes — no external NATS/JetStream, Postgres, Qdrant, or object store required for platform functionality. Only the k8s runtime remains a hard dependency.
Phase 9 = per-tier primary cutover. Phase 8 left every tier in shadow
(dual-write + parity, incumbent still authoritative). Phase 9 flips each tier
to primary (EHDB serves reads, the incumbent is retired for that tier) as a
separately-gated step per tier — dual-run verify against the shadow parity
history, a documented per-tier rollback, kind-validated before any GKE rollout.
The five cutovers, each independent:
| # | Tier | Shadow flag (Phase 8, off→shadow) |
Phase-9 cutover (shadow→primary) |
Incumbent retired | Status |
|---|---|---|---|---|---|
| 1 | Event log | NOETL_EHDB_EVENTLOG |
serve the event log from EHDB | NATS JetStream + Postgres noetl.event
|
primary-serve IMPLEMENTED + MERGED (2026-07-05); kind dual-run VALIDATED (2026-07-05, in-cluster); prod cutover runbook drafted 2026-07-06 — awaiting user go; prod/GKE cutover STILL GATED on user |
| 2 | Projection / read-model | NOETL_EHDB_PROJECTION |
serve read-models from EHDB | Postgres materializer | primary-serve IMPLEMENTED + MERGED (2026-07-05); kind dual-run VALIDATED (2026-07-05, in-cluster); prod/GKE cutover STILL GATED on user |
| 3 | KV / state | NOETL_EHDB_KV |
serve platform KV from EHDB | NATS KV | primary-serve IMPLEMENTED + MERGED (2026-07-05); kind dual-run VALIDATED (2026-07-06, in-cluster); prod/GKE cutover STILL GATED on user |
| 4 | Object / blob | NOETL_EHDB_OBJECT |
serve state-shards + result-tier from EHDB | external object store (GCS/S3) | primary-serve IMPLEMENTED + MERGED (2026-07-05); kind dual-run VALIDATED (2026-07-06, in-cluster); prod/GKE cutover STILL GATED on user |
| 5 | Vector | NOETL_EHDB_VECTOR |
serve platform RAG retrieval from EHDB | Qdrant | primary-serve IMPLEMENTED + MERGED (2026-07-06); kind dual-run VALIDATED (2026-07-06, in-cluster); prod/GKE cutover STILL GATED on user |
Phase 9 is CODE-COMPLETE (2026-07-06): all five per-tier primary cutovers are
IMPLEMENTED + MERGED and in-kind dual-run VALIDATED in-cluster (each
activated + reversible; tiers 1–2 validated 2026-07-05, tiers 3–5 validated in a
combined in-cluster run 2026-07-06 — worker v5.65.0 image v5.65.0-p9
7d30b250d261, native-arm64, loaded into kind-noetl; Job ehdb-p9-t345 ns
ehdb-p9-validate, pod on node noetl-control-plane, container kernel
6.19.7-200.fc43.aarch64 — NOT the macOS host; pod Succeeded, 0 restarts). Every
tier stays driver-selectable back to its incumbent (Phase 10) so a cutover is
reversible. The one thing that remains is the prod/GKE cutover — still gated on
the user, per tier. Nothing in prod changed (prod runs worker v5.52.0; all
NOETL_EHDB_* flags default off).
The first per-tier primary cutover — the reason for the whole program (the
event-log bottleneck) — is built and merged, activated behind
NOETL_EHDB_EVENTLOG=primary, reversible, and dual-run-verified. The
prod/GKE cutover remains a separate later step gated on the user — nothing
in prod changed.
-
ehdb — #247 (merged
7f014c9):ehdb-reference::eventloggainsexercise_primary_serve+EventLogPrimaryEvent+EventLogPrimaryServeReport::served_by_ehdb()— one authoritative cycle through every serving leg (append → global scan → per-execution scoped read → durable tail → ack → fresh-driver replay) that asserts the JetStream+Postgres semantics are preserved (monotonic gapless global sequence, per-execution scope, durable cursor advance, replay-is-truth) and dual-run parity-checks each append against the incumbent sequence. CLI verbeventlog-primary-serve. -
worker — #161 (merged
ddf41de, releasedv5.61.07e98538):src/ehdb/eventlog.rsflipsPRIMARY_SERVE_ACTIVATEDfalse → true;primarymode now serves the append authoritatively (outcomeserved_primary, orprimary_divergenceon dual-run divergence).serve_primary_cycledrives the full authoritative cycle and then demonstrates reversibility (flip back toshadow, mirror one more event over the same log, confirm it replays whole).ehdb-selfcheck eventlog-primary-serveverb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.
Local proof (same ehdb-selfcheck binary the worker image ships; debug host
build, logic identical to release): off ⇒ byte-identical no-op (exit 0);
primary ⇒ served_by_ehdb:true + reversible:true + secret-free metrics
(exit 0); control-plane role ⇒ guard_refused (exit 4).
Kind dual-run: VALIDATED (2026-07-05). The podman VM was recovered
(stop/start reset the wedged ssh socket), the worker v5.62.0 image
(36875e3, ships both tiers + ehdb-selfcheck) was built native-arm64, loaded
into kind-noetl (image sha256:9d222db6), and ehdb-selfcheck eventlog-primary-serve ran in-cluster in a Job pod on node
noetl-control-plane (namespace ehdb-p9-validate; pod host kernel
6.19.7-...fc43.aarch64 = the kind node, not the macOS host, not GKE;
kubectl context kind-noetl, API server https://127.0.0.1:61866).
Evidence: off ⇒ ehdb:disabled + metrics_empty:true (exit 0);
primary ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true
-
dual_run_holds:true(3×count_ok/order_ok/sequence_ok,divergence:null) -
replay_matches:true+scope_ok:true+records_after_revert:4(incumbent path restored whole on flip-back toshadow, zero data loss) + secret-freenoetl_ehdb_eventlog_*metrics (exit 0); control-plane roleserver ⇒guard_refused(exit 4). PodSucceeded,0restarts. prod/GKE cutover stays GATED on the user — nothing in prod changed.
Two independent levers restore the JetStream+Postgres incumbent with zero data loss:
-
Runtime flag (operational, instant, no redeploy). Set
NOETL_EHDB_EVENTLOG=shadow(oroff) on the worker/system pool. The incumbent is authoritative again immediately. Zero data loss: the primary path only ever appends to the EHDBKeepAlllog and never mutates or deletes anything the incumbent owns, so the incumbent's store is exactly as it was and the EHDB log stays whole on disk for a later re-enable. -
Compile-time kill switch (structural). Set
PRIMARY_SERVE_ACTIVATED = falseinsrc/ehdb/eventlog.rsand redeploy —primarythen degrades toprimary_unavailableregardless of config, making serving structurally unreachable.
Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. Event authorship is never affected by either — the gateway/server stay the gatekeeper of what is appended; only the serving engine underneath changes.
The staged prod/GKE cutover procedure is drafted at
Runbook — Prod Cutover: Event-Log Tier:
Stage A (deploy worker v5.66.0 flags-off, behavior-neutral) → Stage B
(shadow in prod: dual-write + parity, incumbent still authoritative) → Stage C
(primary flip). Nothing has been executed — it is a reviewable plan the
user approves per stage. It flags one hard blocker for Stage C: the shipped
local_reference backend is a pod-local JSONL file, which is safe for shadow
but not durable/shared enough to be the authoritative store under primary
without a PVC-backed/shared substrate + single-writer topology. Stages A–B are
executable now; Stage C is gated on that durability decision.
The second per-tier primary cutover — serving the materialized read-models the
control plane queries from EHDB in place of the PostgreSQL materializer — is
built and merged, activated behind NOETL_EHDB_PROJECTION=primary,
reversible, and dual-run-verified. It mirrors the tier-1 event-log pattern
exactly. The prod/GKE cutover remains a separate later step gated on the
user — nothing in prod changed.
-
ehdb — #248 (merged
d08013c):ehdb-reference::projectiongainsexercise_primary_serve+ProjectionPrimaryInput+ProjectionPrimaryServeReport::served_by_ehdb()— one authoritative cycle through every serving leg (apply/materialize → the three read-model query contractslist_executions/ per-executionread_execution_state/read_event→ durable checkpoint → idempotent re-apply → fresh-engine replay) that asserts the PostgreSQL-materializer query contracts are preserved (identical read-models, per-execution scope, exactly-once on the global sequence, replay-is-truth) and dual-run parity-checks the served read-models against the incumbent materializer viacompare_projection_parity. CLI verbprojection-primary-serve. -
worker — #162 (merged
a56583c, releasedv5.62.036875e3):src/ehdb/projection.rsflipsPRIMARY_SERVE_ACTIVATEDfalse → true;primarymode now serves the read-models authoritatively (outcomeserved_primary, orprimary_divergenceon dual-run divergence).serve_primary_cycledrives the full authoritative cycle and then demonstrates reversibility (flip back toshadow, materialize one more execution over the same store, confirm the read-models replay whole).ehdb-selfcheck projection-primary-serveverb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.
Local proof (same ehdb-selfcheck binary the worker image ships; debug host
build, logic identical to release): off ⇒ byte-identical no-op (exit 0);
primary ⇒ served_by_ehdb:true + reversible:true + secret-free metrics
(exit 0); control-plane role ⇒ guard_refused (exit 4).
Kind dual-run: VALIDATED (2026-07-05). Batched with tier-1 on the same
v5.62.0 image (36875e3, sha256:9d222db6) in the same in-cluster Job pod on
node noetl-control-plane (namespace ehdb-p9-validate, context kind-noetl,
not the macOS host, not GKE). ehdb-selfcheck projection-primary-serve
evidence: off ⇒ ehdb:disabled + metrics_empty:true (exit 0);
primary ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true
-
dual_run_holds:true(key_ok/value_ok/checkpoint_ok,checkpoint_lag:0,divergence:null) +list_ok:true(list_count:2) +read_event_ok:true -
replay_idempotent:true+replay_matches:true+scope_ok:true -
rows_after_revert:3(Postgres-materializer read path restored whole on flip-back toshadow, zero data loss) + secret-freenoetl_ehdb_projection_*metrics (exit 0); control-plane roleserver ⇒guard_refused(exit 4). PodSucceeded,0restarts. prod/GKE cutover stays GATED on the user.
Two independent levers restore the PostgreSQL materializer read path with zero data loss:
-
Runtime flag (operational, instant, no redeploy). Set
NOETL_EHDB_PROJECTION=shadow(oroff) on the worker/system pool. The PostgreSQL materializer is the authoritative read path again immediately. Zero data loss: the primary path only ever materializes into the derived EHDBKeepAllprojection store by consuming already-authored events and never mutates or deletes anything the incumbent owns, so the incumbent read-models are exactly as they were and the EHDB store stays whole on disk for a later re-enable. -
Compile-time kill switch (structural). Set
PRIMARY_SERVE_ACTIVATED = falseinsrc/ehdb/projection.rsand redeploy —primarythen degrades toprimary_unavailableregardless of config, making serving structurally unreachable.
Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. A projection is a derived read-model built by consuming the append-only event log — it never authors an event; the cutover changes only which engine serves the read-models, never the event log (the source of truth).
The third per-tier primary cutover — serving NoETL's internal platform
KV/state tier from EHDB in place of the internal NATS-KV bucket (the worker's
noetl_subscription_circuit breaker store today, the #115 program-scale
coherence keys as they move off NATS-KV) — is built and merged, activated
behind NOETL_EHDB_KV=primary, reversible, and dual-run-verified. It mirrors
the tier-1 event-log and tier-2 projection patterns exactly. The prod/GKE cutover
remains a separate later step gated on the user — nothing in prod changed.
Business (tenant/domain) KV stays external, reached by playbook connectors — it
never flows through this tier.
-
ehdb — #249 (merged
73b1446):ehdb-reference::kvgainsexercise_primary_serve+KvPrimaryInput+KvPrimaryServeReport::served_by_ehdb()— one authoritative cycle through every serving leg (put → per-key served get → bucket scan → optimistic CAS (versioned swap + create-only conflict) → tombstone delete → absolute-TTL lease → fresh-driver replay) that asserts the NATS-KV semantics are preserved (last-writer-wins get, bucket scan, optimistic CAS, tombstone delete, absolute TTL, replay-is-truth) and dual-run parity-checks each served read against a NATS-KV mirror applied in lockstep viacompare_kv_parity. CLI verbkv-primary-serve. -
worker — #163 (merged
ba9f829, releasedv5.63.0a7925e0):src/ehdb/kv.rsflipsPRIMARY_SERVE_ACTIVATEDfalse → true;primarymode now serves the KV op authoritatively (outcomeserved_primary, orprimary_divergenceon dual-run divergence).serve_primary_cycledrives the full authoritative cycle and then demonstrates reversibility (flip back toshadow, mirror one more key over the same store, confirm the store serves the whole live set).ehdb-selfcheck kv-primary-serveverb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.
Local proof (same ehdb-selfcheck binary the worker image ships; debug host
build, logic identical to release): off ⇒ byte-identical no-op (ehdb:disabled
-
metrics_empty:true, exit 0);primary ⇒outcome:served_primary+served_by_ehdb:true+reversible:true+dual_run_holds:true(5 served-read parities,present_ok/value_ok/ttl_ok,divergence:null) +put_ok/get_ok/scan_ok/cas_ok/delete_ok/ttl_ok/replay_matchesall true +keys_after_revert:3(NATS-KV path restored whole on flip-back toshadow, zero data loss) + secret-freenoetl_ehdb_kv_*metrics (exit 0); control-plane roleserver ⇒guard_refused(exit 4).
Kind dual-run: VALIDATED 2026-07-06 (in-cluster). Worker v5.65.0 image
localhost/noetl-worker:v5.65.0-p9 (7d30b250d261, native-arm64, ships
ehdb-selfcheck) was loaded into kind-noetl and ehdb-selfcheck kv-primary-serve ran in Job ehdb-p9-t345 (ns ehdb-p9-validate) on node
noetl-control-plane — container kernel 6.19.7-200.fc43.aarch64, Alpine
3.22, NOT the macOS host; context kind-noetl, API https://127.0.0.1:61866.
Evidence: off ⇒ disabled + metrics_empty (exit 0); primary (worker
role) ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true +
dual_run_holds:true (5 served-read parities, present_ok/value_ok/ttl_ok
against a NATS-KV mirror applied in lockstep) + put_ok/get_ok/scan_ok/
cas_ok/delete_ok/ttl_ok/replay_matches all true + keys_after_revert:3
- secret-free
noetl_ehdb_kv_*metrics (exit 0);primary(server role) ⇒guard_refused(exit 4). Pod Succeeded, 0 restarts. prod/GKE cutover stays GATED on the user — nothing in prod changed.
Two independent levers restore the internal NATS-KV path with zero data loss:
-
Runtime flag (operational, instant, no redeploy). Set
NOETL_EHDB_KV=shadow(oroff) on the worker/system pool. The internal NATS-KV bucket is the authoritative KV tier again immediately. Zero data loss: the primary path only ever appends to the derived EHDBKeepAllKV stream and never mutates or deletes anything NATS-KV owns, so the NATS-KV bucket is exactly as it was and the EHDB store stays whole on disk for a later re-enable. -
Compile-time kill switch (structural). Set
PRIMARY_SERVE_ACTIVATED = falseinsrc/ehdb/kv.rsand redeploy —primarythen degrades toprimary_unavailableregardless of config, making serving structurally unreachable.
Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. A KV entry is derived platform state, not an event — this tier never authors the event log; the cutover changes only which engine serves the platform KV state, never the event log (the source of truth). Platform KV only — business KV never flows through it.
The fourth per-tier primary cutover — serving NoETL's internal platform
object/blob tier from EHDB in place of the internal external object store
(the state shards #166, state_materializer/state_reader, and the
result tier #104, result_materializer/result_resolver, both reached
today through the server's /api/internal/objects/{key} API) — is built and
merged, activated behind NOETL_EHDB_OBJECT=primary, reversible, and
dual-run-verified. It mirrors the tier-1 event-log, tier-2 projection, and
tier-3 KV patterns exactly. The prod/GKE cutover remains a separate later step
gated on the user — nothing in prod changed. Business (tenant/domain) object
buckets stay external, reached by playbook connectors — they never flow through
this tier.
-
ehdb — #250 (merged
bb42b8d):ehdb-reference::objectgainsexercise_primary_serve+ObjectPrimaryInput+ObjectPrimaryServeReport::served_by_ehdb()— one authoritative cycle through every serving leg (put → per-key digest-verified served get → prefix list → in-cluster locate → tombstone delete → fresh-driver replay) that asserts the external-store semantics are preserved (content-addressed put, digest-verified get, prefix list, in-cluster locate, tombstone delete, replay-is-truth) and dual-run digest-parity-checks each served read against an external-store mirror applied in lockstep viacompare_object_parity(the mirror digest is the independent SHA-256 of the same bytes, so the parity is exact). CLI verbobject-primary-serve. -
worker — #164 (merged
a100adf, releasedv5.64.0369e4c1):src/ehdb/object.rsflipsPRIMARY_SERVE_ACTIVATEDfalse → true;primarymode now serves the object op authoritatively (outcomeserved_primary, orprimary_divergenceon dual-run divergence).serve_primary_cycledrives the full authoritative cycle and then demonstrates reversibility (flip back toshadow, mirror one more object over the same store, confirm the store serves the whole live set).ehdb-selfcheck object-primary-serveverb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.
Local proof (same ehdb-selfcheck binary the worker image ships; debug host
build, logic identical to release): off ⇒ byte-identical no-op (ehdb:disabled
-
metrics_empty:true, exit 0);primary ⇒outcome:served_primary+served_by_ehdb:true+reversible:true+dual_run_holds:true(4 served-read digest parities,present_ok/digest_ok/length_ok/retrievable_ok,divergence:null) +put_ok/get_ok/list_ok/locate_ok/delete_ok/replay_matchesall true +keys_after_revert:3(external-store path restored whole on flip-back toshadow, zero data loss) + secret-freenoetl_ehdb_object_*metrics (exit 0); control-plane roleserver ⇒guard_refused(exit 4).
Kind dual-run: VALIDATED 2026-07-06 (in-cluster). Same worker v5.65.0
image (v5.65.0-p9 7d30b250d261) + Job ehdb-p9-t345 (ns ehdb-p9-validate,
node noetl-control-plane, container kernel 6.19.7-200.fc43.aarch64, NOT the
macOS host; context kind-noetl). Evidence for ehdb-selfcheck object-primary-serve: off ⇒ disabled no-op (exit 0); primary (worker
role) ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true +
dual_run_holds:true (4 served-read parities, digest_ok/length_ok/
present_ok/retrievable_ok — content-addressed digest integrity against the
object-store mirror) + put_ok/get_ok/list_ok/locate_ok/delete_ok +
replay_matches true + keys_after_revert:3 + secret-free noetl_ehdb_object_*
metrics (exit 0); primary (server role) ⇒ guard_refused (exit 4). Pod
Succeeded, 0 restarts. prod/GKE cutover stays GATED on the user — nothing in
prod changed.
Two independent levers restore the internal external object store path with zero data loss:
-
Runtime flag (operational, instant, no redeploy). Set
NOETL_EHDB_OBJECT=shadow(oroff) on the worker/system pool. The internal external object store is the authoritative object tier again immediately. Zero data loss: the primary path only ever appends to the derived EHDBKeepAllobject registry + content-addressed blob store and never mutates or deletes anything the external store owns, so the external store is exactly as it was and the EHDB store stays whole on disk for a later re-enable. -
Compile-time kill switch (structural). Set
PRIMARY_SERVE_ACTIVATED = falseinsrc/ehdb/object.rsand redeploy —primarythen degrades toprimary_unavailableregardless of config, making serving structurally unreachable.
Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. An object is a derived platform artifact (content-derivable from the WAL), not an event — this tier never authors the event log; the cutover changes only which engine serves the platform object bytes, never the event log (the source of truth). Platform object tier only (state shards + result tier) — business object buckets never flow through it.
The fifth and final per-tier primary cutover — serving NoETL's internal
platform vector tier from EHDB in place of the internal Qdrant retrieval path
(the platform RAG / catalog embeddings reached in-process via the Phase-E
retrieval path) — is built and merged, activated behind
NOETL_EHDB_VECTOR=primary, reversible, and dual-run-verified. It mirrors the
tier-1 event-log, tier-2 projection, tier-3 KV, and tier-4 object patterns
exactly. With this tier all five Phase-9 primary-serve activations are
implemented. The prod/GKE cutover remains a separate later step gated on the
user — nothing in prod changed. Business (tenant/domain) vector collections stay
external, reached by playbook connectors — they never flow through this tier.
-
ehdb — #251 (merged
0f47fe3):ehdb-reference::vectorgainsexercise_primary_serve+VectorPrimaryInput+VectorPrimaryServeReport::served_by_ehdb()— one authoritative cycle through every serving leg (upsert → served cosine top-k query → tombstone delete → fresh-driver replay) that asserts the Qdrant retrieval semantics are preserved (bounded cosine top-k ranking, tombstone delete, replay-is-truth) and dual-run parity-checks each served query — id-set + rank-order + score-monotonicity (compare_vector_parity) — against a Qdrant mirror ranked in lockstep with the identical cosine scoring (an independent computation, not a copy of the engine's output, so the parity is exact). CLI verbvector-primary-serve. -
worker — #165 (merged
681782c, releasedv5.65.0ceedbba):src/ehdb/vector.rsflipsPRIMARY_SERVE_ACTIVATEDfalse → true;primarymode now serves the retrieval op authoritatively (outcomeserved_primary, orprimary_divergenceon dual-run divergence).serve_primary_cycledrives the full authoritative cycle and then demonstrates reversibility (flip back toshadow, mirror one more point over the same index, confirm the collection serves the whole live set).ehdb-selfcheck vector-primary-serveverb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.
Local proof (same ehdb-selfcheck binary the worker image ships; debug host
build, logic identical to release): off ⇒ byte-identical no-op (ehdb:disabled
-
metrics_empty:true, exit 0);primary ⇒outcome:served_primary+served_by_ehdb:true+reversible:true+dual_run_holds:true(3 served-query parities,ids_ok/order_ok/monotonic_ok,divergence:null) +upsert_ok/query_ok/delete_ok/replay_matchesall true +candidates_after_revert:3(Qdrant retrieval path restored whole on flip-back toshadow, zero data loss) + secret-freenoetl_ehdb_vector_*metrics (exit 0); control-plane roleserver ⇒guard_refused(exit 4); over-limit dimensionality⇒rejected(exit 3).
Kind dual-run: VALIDATED 2026-07-06 (in-cluster). Same worker v5.65.0
image (v5.65.0-p9 7d30b250d261, the LATEST release — contains all five tiers
from sequential merges) + Job ehdb-p9-t345 (ns ehdb-p9-validate, node
noetl-control-plane, container kernel 6.19.7-200.fc43.aarch64, NOT the macOS
host; context kind-noetl). Evidence for ehdb-selfcheck vector-primary-serve:
off ⇒ disabled no-op (exit 0); primary (worker role) ⇒
outcome:served_primary + served_by_ehdb:true + reversible:true +
dual_run_holds:true (3 top-k parities, ids_ok (id-set) / order_ok
(rank-order) / monotonic_ok (score-monotonicity) against the Qdrant mirror) +
upsert_ok/query_ok/delete_ok + query_returned:3 (bounded top-k cap) +
replay_matches true + candidates_after_revert:3 + secret-free
noetl_ehdb_vector_* metrics (exit 0); primary (server role) ⇒ guard_refused
(exit 4). Pod Succeeded, 0 restarts. This is the last per-tier in-kind
validation — Phase 9 is now CODE-COMPLETE. prod/GKE cutover stays GATED on
the user — nothing in prod changed.
Two independent levers restore the internal Qdrant retrieval path with zero data loss:
-
Runtime flag (operational, instant, no redeploy). Set
NOETL_EHDB_VECTOR=shadow(oroff) on the worker/system pool. The internal Qdrant retrieval path is the authoritative vector tier again immediately. Zero data loss: the primary path only ever appends to the derived EHDBKeepAllvector index and never mutates or deletes anything Qdrant owns, so Qdrant is exactly as it was and the EHDB index stays whole on disk for a later re-enable. -
Compile-time kill switch (structural). Set
PRIMARY_SERVE_ACTIVATED = falseinsrc/ehdb/vector.rsand redeploy —primarythen degrades toprimary_unavailableregardless of config, making serving structurally unreachable.
Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. A vector point is a derived platform index entry (content-derivable from the platform RAG source), not an event — this tier never authors the event log; the cutover changes only which engine serves platform retrieval, never the event log (the source of truth). Platform vectors only (RAG / catalog embeddings) — business vector collections never flow through it.
Acceptance:
- A reference deployment runs the full platform on k8s with EHDB as the only storage fabric (no external infra pods/services for platform tiers).
- Migration/import/export adapters retained for onboarding + cloud durability, but not required at run time.
- Explicit cutover gates per dependency, each dual-run verified; kind-validated before every GKE rollout.
The per-tier backend-selection surface that makes EHDB the default while keeping every platform tier selectable back to the incumbent engine — so Phase 9's self-sufficiency is a default, not a lock-in. Full reference: Backend Configuration.
| Platform tier | Default | Selectable back to |
|---|---|---|
| Event-log driver | EHDB | NATS JetStream + Postgres |
| Projection driver | EHDB | Postgres materializer |
| KV / state driver | EHDB | NATS KV |
| Object / blob driver | EHDB | External object store / Postgres |
| Vector driver | EHDB | Qdrant |
Acceptance:
- Per-tier driver interface with ≥2 selectable backends per tier (EHDB + incumbent).
- Default deployment runs EHDB across all platform tiers; an operator overlay pins any tier back to its incumbent.
- Each tier selection is dual-run verifiable with a documented rollback; kind-validated; no GKE rollout without it.
Status: IMPLEMENTED (2026-07-06). The config surface landed — a single
coherent schema resolved from the existing NOETL_EHDB_* env, no breaking
rename:
-
ehdb — #252 (merged
4c0df81):ehdb-reference::backends—PlatformTier(the five tiers, each with its env var + incumbent),TierMode {off|shadow|primary},Backend {ehdb|external}+backend_for_mode(primary⇒ EHDB, else the incumbent), andBackendMatrixwith coherencevalidate()(rejectsshadow/primarywithoutNOETL_EHDB_ENABLED, or a data-plane tier on a control-plane role) + a secret-freeto_json(). Pure data — reads no env, opens no engine. -
worker — #166 (released
v5.66.0):src/ehdb/backends.rsresolve(&EnvMap)maps the process env into the matrix by reading each tier's mode through the same<Tier>Mode::from_envparser the runtime dispatch uses (backward-compatible by construction — the consolidated view can never drift). Newehdb-selfcheck configverb prints the resolved 5-tier backend+mode matrix (secret-free; exit 0 coherent, 4 incoherent).
Selfcheck evidence (ehdb-selfcheck config, batched with the Phase-9 kind
runs): all-external default ⇒ 5×backend:external, exit 0; all-EHDB (enabled
worker, 5×primary) ⇒ 5×backend:ehdb, exit 0; mixed (log+vector primary,
projection shadow) ⇒ per-tier ehdb/external, exit 0; primary without
enable ⇒ coherent:false, exit 4; data-plane tier on gateway role ⇒
coherent:false, exit 4; sensitive-keyed env present ⇒ secret_free:true,
0 leaks. No behavior change; disabled-by-default strict no-op intact.
With Phase 10 the whole EHDB completion program (Phases 6–10) is
code-complete and kind-validated. Every platform tier has a built +
shadow-verified engine, a per-tier reversible primary cutover (Phase 9), and
a consolidated backend-selection surface (Phase 10). The only remaining step
is the prod/GKE cutover — gated on the user, per tier; nothing in prod
has changed (prod runs worker v5.52.0, all NOETL_EHDB_* flags default off).
Tracking issue: #261 Performance & load testing for EHDB engine tiers. Design: Design: Performance & Load Testing.
Two layers:
-
Phase 1 — engine micro-benchmarks (Rust, deterministic) — LANDED
(2026-07-08). criterion benches over all five reference drivers + the
durable segment event-log backend
(
crates/ehdb-reference/benches/engine_micro.rs), with a committed baseline. Headline: the durable event-log backend beats thelocal_referenceJSONL driver 2.7× at sustained append (255 vs 96 ev/s at K=1000) and stays flat (~3.9 ms/append,fsync-bound) while the reference driver degradesO(n)per op; segment rotation ~2%; cold replay ~185 K ev/s. KV/object/vector reference drivers areO(n)-per-op shadow fixtures (no durable backend yet). -
Phase 2 — in-cluster end-to-end load (kind) — DESIGN ONLY, pending
scoping. Drive real traffic through the worker with
NOETL_EHDB_*flags armed, readnoetl_ehdb_*metrics + latency, and run the EHDB-vs-incumbent (Postgres+NATS-JetStream) head-to-head. Directional only (podman-VM constraints); Layer-A micro-benches stay the reliable signal. Proposed SLO strawman on the design page awaits platform-owner confirmation.
- Rust-first implementation.
- NoETL-domain-specific implementation; avoid generic database scope creep unless it directly serves the NoETL platform.
- Product code stays in
noetl/ehdb, notai-meta. - Design is maintained in this wiki as public surfaces evolve.
- Issues track work;
ai-metamemory tracks cross-repo platform decisions only. - Logging changes include flood-check rationale.
- Container image work validates in local kind before GKE.
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)