Skip to content

Roadmap

Kadyapam edited this page Aug 20, 2026 · 147 revisions

Roadmap

This roadmap turns the EHDB vision into issue-sized phases. GitHub issues are the executable source of truth for active work; this page is the durable design index.

Phase 0: Project Bootstrap

Tracking issue: #1 Bootstrap EHDB Rust workspace and CI

Goals:

  • Add noetl/ehdb and noetl/ehdb.wiki to ai-meta as submodules.
  • Create the initial wiki design pages.
  • Create initial tracking issues and add them to the EHDB development project.
  • Scaffold a Rust workspace with stable crate boundaries.
  • Add CI-ready formatting, linting, and tests.

Acceptance:

  • cargo test --workspace passes locally.
  • README explains mission, boundaries, and first developer workflow.
  • Wiki has Home, Architecture, Roadmap, and Sessions Log pages.

Phase 1: Typed Core And Catalog Model

Tracking issue: #2 Design catalog-as-database metadata model

Goals:

  • Define durable identifiers for tenant, namespace, table, snapshot, transaction, and object references.
  • Define Arrow-native column/type metadata.
  • Implement an in-memory transactional catalog reference model.
  • Add MVCC-style snapshot IDs and optimistic transaction records.

Acceptance:

  • Catalog table create/read/update tests pass.
  • Snapshot metadata is immutable after commit.
  • Catalog mutations produce typed transaction records.

Current local reference:

  • Follow-up issue: #22 Add catalog table snapshot metadata
  • Follow-up issue: #62 Add catalog scan grant reference model
  • InMemoryCatalog commits immutable table snapshots with snapshot ID, optional parent snapshot, object file refs, and committing transaction ID.
  • Latest snapshot lookup is tracked per table.
  • Snapshot commits reject missing tables, empty file sets, duplicate snapshots, and parent-chain mismatches.
  • CatalogMutation::CommitSnapshot makes snapshot metadata replayable through the transaction log.
  • CatalogScanGrant records tenant/namespace/table scan access for a principal, rejects missing-table and duplicate grants, and answers can_scan lookups.
  • CatalogMutation::GrantScan makes scan grant metadata replayable through ehdb-reference and LocalReferenceRuntime.
  • Catalog metadata JSON decoding rejects unknown table, snapshot, scan grant, request, table schema, and column schema fields before replay or catalog operations.
  • Table schemas revalidate column identifiers before catalog state is created.
  • Table schemas reject duplicate column names before catalog state is created.
  • Column schema and table schema JSON decode routes through the same identifier and duplicate-column validation before metadata is accepted.
  • Core identifier JSON decode routes through constructor validation while preserving the string JSON shape.

Phase 2: Object Storage Layer

Tracking issue: #3 Define immutable object storage layer

Goals:

  • Define an object-store trait for put/get/list/delete-free immutable writes.
  • Implement local filesystem adapter for tests.
  • Add S3-compatible, GCS, and Azure Blob designs before implementation.
  • Define object naming layout for tenant/namespace/table/snapshot.

Acceptance:

  • Local adapter can write/read immutable Arrow IPC and Parquet test objects.
  • Object references are catalog-addressable and content-checked.
  • Object references carry geo-location and data-gravity shard placement pointers for distributed storage design.
  • Delete behavior is explicit and conservative.

Current local reference:

  • Follow-up issue: #20 Add content-checked immutable object references
  • Follow-up issue: #24 Add geo placement and data-gravity shard pointers
  • Follow-up issue: #26 Add storage placement policy model
  • Follow-up issue: #28 Add deterministic replication planning model
  • Follow-up issue: #30 Add durable object replica registry
  • Follow-up issue: #32 Add bounded local replication executor
  • Follow-up issue: #34 Add Arrow IPC table write/read fixture
  • Follow-up issue: #36 Add Arrow snapshot scan fixture
  • Follow-up issue: #38 Add Arrow equality filter fixture
  • Follow-up issue: #40 Add Arrow scan service API boundary
  • Follow-up issue: #42 Add Arrow Flight scan ticket codec
  • Follow-up issue: #44 Add Arrow Flight scan result stream codec
  • Follow-up issue: #46 Add Arrow Flight scan info fixture
  • Follow-up issue: #48 Add local Arrow Flight scan service facade
  • Follow-up issue: #50 Add Arrow Flight scan service trait adapter
  • Follow-up issue: #52 Add bounded Arrow Flight server lifecycle config
  • Follow-up issue: #54 Add loopback Arrow Flight listener harness
  • Follow-up issue: #56 Add loopback Arrow Flight client smoke test
  • Follow-up issue: #58 Add Arrow Flight auth header policy contract
  • Follow-up issue: #60 Add Arrow Flight scan scope metadata guard
  • Follow-up issue: #62 Add catalog scan grant reference model
  • Follow-up issue: #64 Enforce catalog scan grants in Flight reference path
  • Follow-up issue: #66 Add bounded Flight scan access log policy
  • Follow-up issue: #68 Add Arrow Flight get_schema scan adapter
  • Follow-up issue: #70 Enforce bounded Flight scan request concurrency
  • Follow-up issue: #72 Add local retrieval vector similarity fixture
  • ObjectRef carries path, byte length, SHA-256 digest, geo placement, and data-gravity shard pointer.
  • PlacementPolicy validates exactly one primary, minimum copy count, shared data-gravity shard, and no duplicate geo/shard targets.
  • plan_replication emits already-satisfied and copy-needed actions from current replicas plus placement policy.
  • ObjectReplicaRegistry records durable replica inventory and can feed replication planning from replayed EHDB metadata.
  • LocalReplicationExecutor verifies source bytes and records copy-needed replica registrations through transaction replay.
  • Storage metadata JSON decoding rejects unknown object ref, placement, policy target, replica, action, and plan fields before replay or planning.
  • CatalogScanGrant records replayable table scan grant metadata for principals before production ACL enforcement exists.
  • FlightScanGrantPolicy can require x-ehdb-principal metadata and enforce replayed CatalogScanGrant records before local Flight scan execution.
  • ImmutableObjectStore::get_verified rejects length or digest mismatches.
  • The local Arrow IPC fixture writes a RecordBatch, commits a catalog snapshot over the content-checked object, and reads it back through the catalog/object boundary.
  • The local Arrow scan fixture resolves latest snapshots, verifies Arrow IPC objects, decodes batches, and supports named column projection.
  • Direct local Arrow scan requests reject empty projection lists and duplicate projection columns before object reads.
  • Direct local Arrow scan requests validate projection-column and predicate-column selector identifiers before object reads.
  • The local equality filter fixture applies single-column UTF-8 and Int64 equality predicates after verified Arrow decode and before projection.
  • ehdb-service exposes a local service-facing latest-table scan request/result boundary over the scanner, returning schema, batches, and row count before Arrow Flight networking exists.
  • ScanFlightTicket encodes that latest-table scan request into a versioned Arrow Flight Ticket payload and command FlightDescriptor for the future network service.
  • Local scan request validation rejects malformed selector identifiers, unknown ticket/request/predicate object fields, non-canonical ticket bytes, empty projection lists, and duplicate projection columns before scan execution.
  • ArrowScanResult encodes local scan outputs into Arrow Flight FlightData messages and decodes them back into validated results.
  • ArrowScanResult builds pre-network Arrow Flight FlightInfo metadata with schema bytes, command descriptor, endpoint ticket, row count, and encoded byte count.
  • LocalArrowFlightService provides in-process get_flight_info, get_schema, and do_get behavior over the scan ticket, info, schema, and result codecs.
  • LocalArrowFlightServer implements the generated Arrow Flight service trait for scan get_flight_info, get_schema, and do_get, with deterministic gRPC statuses, configured request metadata auth, tenant/namespace scan scope metadata, catalog scan grant enforcement, scan command descriptor path rejection, and no bound network listener.
  • FlightAccessLogPolicy provides disabled and debug-only modes for bounded scan access summaries, excluding tokens, principals, tenant/table identifiers, object paths, predicate values, and Arrow payloads.
  • Implemented scan methods enforce the local max_concurrent_requests budget with fail-fast gRPC RESOURCE_EXHAUSTED responses when all request slots are occupied.
  • LocalArrowFlightServerConfig validates bind, message-size, concurrency, header-token auth, tenant/namespace scan scope, catalog scan grant, and access-log policy before any listener exists.
  • LocalArrowFlightListener binds loopback-only reference listeners and shuts them down through an explicit future.
  • The loopback client smoke path validates get_schema, get_flight_info, schema-aware endpoint-ticket extraction, and decoded do_get batches against the decoded schema, returned FlightInfo, and expected ticket over real tonic/gRPC transport, including the optional header-token auth, tenant/namespace scan scope, and catalog scan grant policies.
  • Receiver paths revalidate the concrete do_get endpoint ticket against returned FlightInfo before using it.
  • Local service and server receiver tests validate raw SchemaResult, returned FlightInfo, raw FlightData, and expected ticket as one response envelope before accepting decoded rows.
  • Table data paths follow {tenant}/{namespace}/tables/{table}/snapshots/{snapshot}/{file}.

Phase 3: Transaction Log Boundary

Tracking issue: #4 Define transaction log and MVCC snapshot boundary

Goals:

  • Define transaction log trait and append/read semantics.
  • Implement local file and in-memory reference logs.
  • Keep consensus engine pluggable behind the trait.
  • Model commit ordering and replay into catalog state.

Acceptance:

  • Replay reconstructs catalog state from log records.
  • Replay reconstructs reference catalog, stream, retrieval, and system library state from log records alone.
  • Duplicate transaction IDs are rejected deterministically.
  • Crash/restart simulation works for the local JSONL reference log.

Current local reference:

  • Follow-up issue: #8 Add local durable transaction log reference
  • Follow-up issue: #16 Make transaction mutations replay-complete
  • InMemoryTransactionLog provides deterministic ordered replay for tests and benchmarks.
  • LocalJsonlTransactionLog writes one fsynced JSONL TransactionRecord per append and rebuilds replay state on open.
  • Local open rejects corrupt records, duplicate transaction IDs, and sequence gaps instead of silently repairing the log.
  • Local open revalidates transaction envelope and mutation identifiers before accepting persisted records back into ordered replay state.
  • Local open rejects unknown transaction record and mutation fields before accepting persisted records back into ordered replay state.
  • ehdb-reference applies replayed TransactionRecord values into the local reference catalogs and rejects mismatches such as unexpected stream sequence values.
  • LocalReferenceRuntime validates projected reference state before durable append and rebuilds reference state from replay on reopen.
  • Follow-up issue: #18 Add local reference runtime over transaction replay

Phase 3.5: NoETL Stream Log Boundary

Goals:

  • Model the NoETL event/command stream role currently served by NATS JetStream.
  • Define typed append-only stream records, subjects, retention policies, durable consumers, replay cursors, and ack state.
  • Keep any NATS bridge as a migration adapter, not the target product boundary.

Acceptance:

  • A local reference stream can publish, replay, ack, and resume from a durable cursor.
  • Consumer state is tenant/namespace aware.
  • NoETL execution event semantics can be represented without direct NATS dependency.

Current local reference:

  • Follow-up issue: #10 Add local durable stream journal reference
  • InMemoryStreamLog provides typed stream records, retention, replay, durable consumers, and ack cursors.
  • LocalJsonlStreamLog writes fsynced JSONL entries for stream creation, consumer creation, publish, and ack operations.
  • Local open rebuilds retained records, next sequence, and durable consumer cursors, and rejects corrupt journal entries.
  • Stream logs reject zero max-record retention before creating or journaling streams.
  • Stream replay can explicitly filter retained records by subject using exact matches, single-token * wildcards, and terminal > tail wildcards.
  • Durable consumer replay can explicitly filter pending records by subject after the consumer ack cursor without moving that cursor.
  • Concrete published stream subjects reject wildcard tokens; wildcard selectors live in explicit subject filters used by replay APIs.
  • Concrete stream subjects and subject filters reject empty dot-delimited tokens.
  • JSONL stream replay revalidates persisted record subjects before rebuilding retained stream state.
  • JSONL stream replay revalidates persisted stream coordinates, consumer names, and stream record transaction IDs before applying journal entries.
  • JSONL stream replay rejects zero persisted stream sequences before rebuilding retained records or durable consumer cursors.
  • JSONL stream replay rejects unknown journal, stream config, stream record, and durable consumer fields before rebuilding state.

Phase 3.75: System WASM Library Registry

Tracking issue: #12 Add EHDB system WASM library registry

Goals:

  • Store NoETL system playbook functionality as compiled WASM library manifests in EHDB.
  • Mirror NoETL worker WASM dispatch with { path, version, digest, entry } module refs while keeping EHDB as the durable resolver.
  • Bind modules by tenant, namespace, environment, release channel, and logical path.
  • Allow hot replacement by rebinding a channel to a new digest/revision instead of requiring crate semantic-version bumps for every fix.
  • Support different implementations for different environments such as kind, gke-prod, aws-prod, and azure-dev.

Acceptance:

  • A local reference catalog can publish immutable WASM module manifests.
  • Environment/channel bindings resolve deterministic active modules.
  • Rebinding a stable channel replaces the implementation while retaining prior immutable module manifests.
  • System library publish/bind events are represented in the transaction log and participate in cross-domain replay tests.
  • A local JSONL system-library journal preserves publish/bind state and hot-replacement bindings across restart.
  • WASM host execution remains a worker/system-pool concern; EHDB stores manifests, bindings, object references, capabilities, and transaction provenance.

Current local reference:

  • Follow-up issue: #14 Add local durable system library registry journal
  • InMemorySystemLibraryCatalog provides module publish, bind, and resolve behavior.
  • LocalJsonlSystemLibraryCatalog writes fsynced JSONL entries for publish and bind operations.
  • Local open rebuilds immutable manifests, environment/channel bindings, and hot-replacement state, and rejects corrupt journal entries.
  • Local open revalidates persisted manifest and binding identifiers before rebuilding system-library state.
  • Local open rejects unknown system-library journal, publish, and bind fields before rebuilding hot-replacement state.
  • System-library metadata JSON decoding rejects unknown resolved manifest, plugin reference, and binding fields before worker handoff metadata is accepted.

Phase 4: Arrow Flight Read Path

Goals:

  • Add service crate boundary for catalog and scan APIs.
  • Implement Arrow Flight read path for catalog/table scan fixtures.
  • Add client smoke tests.

Acceptance:

  • CLI can create a table, write a small batch, and read it back as an Arrow stream.
  • Flight endpoint has bounded logging and avoids high-frequency INFO noise.

Current local reference:

  • Follow-up issue: #40 Add Arrow scan service API boundary
  • Follow-up issue: #42 Add Arrow Flight scan ticket codec
  • Follow-up issue: #44 Add Arrow Flight scan result stream codec
  • Follow-up issue: #46 Add Arrow Flight scan info fixture
  • Follow-up issue: #48 Add local Arrow Flight scan service facade
  • Follow-up issue: #50 Add Arrow Flight scan service trait adapter
  • Follow-up issue: #52 Add bounded Arrow Flight server lifecycle config
  • Follow-up issue: #54 Add loopback Arrow Flight listener harness
  • Follow-up issue: #56 Add loopback Arrow Flight client smoke test
  • Follow-up issue: #58 Add Arrow Flight auth header policy contract
  • Follow-up issue: #60 Add Arrow Flight scan scope metadata guard
  • Follow-up issue: #64 Enforce catalog scan grants in Flight reference path
  • Follow-up issue: #66 Add bounded Flight scan access log policy
  • Follow-up issue: #68 Add Arrow Flight get_schema scan adapter
  • Follow-up issue: #70 Enforce bounded Flight scan request concurrency
  • Follow-up issue: #182 Validate Arrow Flight scan command descriptor path
  • Follow-up issue: #184 Reject unknown Arrow Flight scan ticket fields
  • Follow-up issue: #186 Require canonical Arrow Flight scan ticket encoding
  • LocalArrowScanService wraps LocalArrowSnapshotScanner with a typed latest-table scan request and ArrowScanResult containing schema, batches, and row count.
  • ScanFlightTicket round-trips requests through versioned Arrow Flight Ticket bytes and command descriptors, with unsupported versions and malformed payloads rejected before scan execution; decoded tenant, namespace, table, projection-column, and predicate-column identifiers are revalidated before execution/handoff, and unknown object fields plus non-canonical bytes and empty or duplicate projection selectors are rejected.
  • Local scan get_flight_info and get_schema request decoding rejects path descriptors and command descriptors with non-empty paths before scan execution.
  • ArrowScanResult round-trips through Arrow Flight FlightData messages, preserving schema and row counts while rejecting empty, missing-version, unsupported-version, extra-metadata, or malformed streams; receiver-side validation can compare returned data against FlightInfo schema, row count, and byte-count metadata.
  • ArrowScanResult can produce and validate a single-endpoint ordered FlightInfo fixture from a scan ticket and result, including decodable schema IPC metadata, descriptor/ticket request consistency, positive byte-count metadata, result/ticket consistency, expected ticket validation, schema-response consistency, and a pre-network endpoint envelope.
  • LocalArrowFlightService executes local get_flight_info, get_schema, and do_get paths without opening a network listener.
  • LocalArrowFlightServer adapts the local facade to the generated Arrow Flight service trait for get_flight_info, get_schema, and do_get; all other Flight methods return explicit unimplemented statuses. The implemented scan methods enforce the configured request metadata auth, scan scope, and catalog scan grant policies.
  • FlightAccessLogPolicy controls bounded DEBUG-only scan access summaries for decoded scan requests; disabled mode emits no summaries. Summaries exclude tokens, principals, tenant/table identifiers, object paths, predicate values, and Arrow payloads.
  • Implemented scan methods enforce the configured local request budget with fail-fast gRPC RESOURCE_EXHAUSTED responses. This remains a local-reference guard, not a scheduler or distributed admission controller.
  • LocalArrowFlightServerConfig validates lifecycle guardrails and applies message limits, the reference header-token auth policy, and tenant/namespace scan scope, catalog scan grant, and access-log policies when constructing the generated service wrapper.
  • LocalArrowFlightListener binds a loopback-only reference listener, exposes the actual bound local address, and shuts down through an explicit future.
  • The loopback client smoke path uses Arrow Flight client calls over tonic/gRPC transport to validate get_schema, expected-ticket get_flight_info validation, raw SchemaResult and schema-response consistency, and schema-aware validated endpoint-ticket extraction plus concrete endpoint-ticket binding before do_get decoded-schema/data/FlightInfo/expected-ticket consistency checks, including the optional header-token auth, tenant/namespace scan scope, and catalog scan grant policies.
  • Local service and server paths use the same raw SchemaResult/ FlightInfo/FlightData/expected-ticket envelope binding before accepting decoded rows.
  • This remains a local-reference network boundary: no non-loopback exposure, server daemon manager, production TLS/identity, request scheduler, SQL planner, predicate pushdown, distributed execution, or gateway direct reads are introduced yet.

Phase 4.5: RAG And Retrieval Primitives

Goals:

  • Model NoETL-native documents, chunks, embedding metadata, vector index metadata, and retrieval policies.
  • Preserve tenant, execution lineage, artifact references, and embedding model identity with every retrieval object.
  • Define a path to replace a permanent Qdrant dependency with EHDB-native retrieval primitives.

Acceptance:

  • Documents and chunks can be registered in the catalog model.
  • Embedding metadata records include model identity, dimensions, checksum/source lineage, and placement policy.
  • A local reference retrieval index can answer filtered lookup fixtures.
  • VectorSearch answers tenant/namespace/model-scoped exact cosine similarity over registered chunk embeddings with deterministic hit ordering and finite non-zero vector validation.
  • Retrieval metadata JSON decoding rejects unknown document, chunk, embedding, registration request, search request, and local search hit fields before replay or handoff.
  • LocalRetrievalSearchService provides an in-process service-facing request/result boundary over replayed retrieval state and returns ranked chunk hits without raw embedding vectors.
  • SearchTextChunksRequest adds exact local text matching with tenant/namespace scoping, match counts, deterministic ordering, and validation.
  • SearchHybridChunksRequest adds weighted exact hybrid scoring over vector cosine similarity and text match counts, scoped by tenant, namespace, and embedding model.
  • AssembleRetrievalContextRequest builds bounded local RAG context blocks from replayed hybrid search hits, preserving citation metadata, clipped text, score metadata, and total text budget accounting.
  • RetrievalContextRequestPayload and RetrievalContextResultPayload provide versioned local JSON byte codecs for context assembly request/result handoff, rejecting invalid decoded identifiers and unknown request/result/context/block fields plus non-canonical payload bytes before execution or handoff.
  • LocalRetrievalSearchService::execute_context_payload decodes a request payload, assembles context from replayed local retrieval state, and returns an encoded result payload for worker/playbook tests.
  • RetrievalContextPayloadExecutorConfig bounds local request/result payload bytes and rejects oversized payloads deterministically.
  • RetrievalContextPayloadScope optionally guards local payload execution by matching decoded request tenant/namespace to the expected worker/playbook execution scope.
  • RetrievalContextPayloadExecutionSummary reports redacted local execution metadata: request/result byte counts, context block count, total text chars, truncation status, and whether scope was required, while excluding tenant IDs, namespace values, query text, chunk text, tokens, vectors, payload bytes, object paths, and principals.
  • RetrievalContextPayloadExecutionReceiptPayload provides a versioned JSON byte codec for those redacted summaries, giving future event-log/audit plumbing a durable receipt shape without event publication or sensitive retrieval content. Receipt encode/decode validates positive payload byte counts and rejects text chars without context blocks, while rejecting unknown receipt envelope and redacted summary fields and non-canonical JSON bytes before artifact validation or event-envelope helper use.
  • RetrievalContextPayloadExecution::encode_receipt_payload emits that receipt directly from a local execution result for worker/playbook tests.
  • RetrievalContextPayloadExecutionArtifacts returns result payload bytes together with bounded redacted receipt payload bytes for local worker/playbook handoff tests, and validates receipt/result consistency before callers treat the artifact pair as coherent.
  • RetrievalContextPayloadExecutionReceiptEventPayload wraps validated receipt bytes in a versioned JSON event envelope with stable subject ehdb.retrieval.context.execution.receipt, giving future stream/audit plumbing a local payload shape without automatic publication or result/context bytes. Event payload decoding rejects unknown envelope fields and non-canonical JSON bytes before replay decode, publisher helper use, or consumer handoff.
  • RetrievalContextReceiptEventStreamTarget explicitly builds and creates the receipt event stream with caller-selected retention; publish helpers do not auto-create streams.
  • Keep-all and positive bounded-retention setup helpers make receipt event stream retention explicit; zero bounded retention is rejected before stream creation.
  • RetrievalContextReceiptEventStreamTarget and RetrievalContextReceiptEventStreamLog explicitly publish that event envelope to caller-supplied local stream logs, requiring caller-owned tenant, namespace, stream, mutable log, and transaction id.
  • RetrievalContextReceiptEventStreamRecord and RetrievalContextReceiptEventStreamReadLog explicitly replay and decode receipt event records from caller-supplied local stream logs, validating stable subject and payload while preserving sequence and transaction id for audit assertions.
  • RetrievalContextReceiptEventDurableConsumerLog explicitly creates local durable consumers, replays pending validated receipt events for a consumer, and acks receipt event sequences without a background subscription loop.
  • This remains a local reference fixture: no ANN index, full-text index, retrieval service, RPC protocol, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, gateway data path, query planner, Qdrant/search adapter, production IAM, ACL engine, scheduler, or distributed query engine.

Phase 5: NoETL Integration Prototype

Tracking issue: #5 Plan NoETL integration path for EHDB system store

Goals:

  • Add NoETL playbook step integration for EHDB catalog operations.
  • Run local kind validation before any GKE rollout.
  • Keep gateway as gatekeeper and NoETL workers as atomic compute.

Acceptance:

  • A NoETL playbook can write and read EHDB-backed metadata in local kind.
  • No gateway direct database touch is introduced.
  • Wiki and NoETL docs are updated with the public integration surface.

Current local reference:

  • Follow-up issue: #228 Add NoETL runtime surface replay fixture
  • Follow-up issue: #230 Add embedded NoETL role capability policy
  • Follow-up issue: #232 Add local-reference summary helper
  • LocalReferenceRuntime has a NoETL-shaped replay fixture covering catalog table/snapshot/grant metadata, stream events, retrieval document/chunk/embedding metadata, system WASM library binding, and storage replica inventory from the transaction log alone.
  • The fixture reopens the runtime and verifies catalog scan grants, stream consumer replay, retrieval text lookup, system library resolution, and storage replica counts without hand-mutating domain catalogs.
  • ehdb-core now models embedded NoETL roles and capabilities so EHDB can be present inside workers, APIs, and gateways while gateway/API roles remain control-plane only and worker/playbook/system roles carry explicit data-plane permissions.
  • ehdb-reference now exposes LocalReferenceSummary plus the ehdb-local-reference summary --log <path> helper for deterministic replayed-domain counts across local transaction, catalog, stream, retrieval, system-library, and storage state.

Integration handoff phases (from the Claude handoff page):

  • Phase A — ops/runtime enablement: DONE (2026-07-04). The NoETL ops Helm charts now render disabled-by-default, role-specific EHDB env with the control-plane vs data-plane boundary enforced in the chart template (noetl/ops#234). server/api/gateway get control-plane-only env; worker/playbook/system get bounded local_reference env with a pod-local JSONL log. A kind Job (ci/manifests/noetl/ehdb/smoke-job.yaml) runs the packaged ehdb-local-reference smoke inside the NoETL image before any GKE rollout. Disabled render is byte-identical to pre-EHDB. See the Architecture page's Ops Env-Rendering Boundary section. Tracks #234.
  • Phase B — worker/playbook readiness hook: DONE (2026-07-04). A bounded, stateless readiness preflight over read_ehdb_local_reference_summary_from_env lands in noetl/noetl (noetl.core.ehdb_readiness, PR noetl/noetl#688). It is wired for worker/playbook/system data-plane roles only; gateway/api/server never call it and a code-level guard (assert_data_plane_read_allowed) refuses a data-plane read for any control-plane role (defense-in-depth on the contract validation). The read is time-bounded (default 5s, clamped 0.1–30s) and holds no connection/state. Observability: noetl_ehdb_readiness_* metrics on the worker /metrics surface (no secret values). Disabled-by-default → strict no-op (byte-identical, no metric recorded). Exposed as a worker/playbook-local preflight command (scripts/ehdb_readiness_preflight.py) / kind smoke step — not a server endpoint. Kind-validated on kind-noetl. See the Architecture page's Worker/Playbook Readiness Hook section. Tracks #234.
  • Phase C — bounded worker/playbook data-plane step: DONE (2026-07-04). A NoETL playbook step (worker tool) that performs a bounded EHDB local-reference operation — append and read a single domain record through the adapter — rather than only a readiness summary. Landed in noetl/noetl (noetl.core.ehdb_dataplane append_ehdb_domain_record / read_ehdb_domain_records + adapter LocalReferenceAppendResult / LocalReferenceReadResult, PR noetl/noetl#689) over the ehdb-local-reference append/read subcommands (noetl/ehdb#235, merged 3ae8950). Still disabled-by-default (strict no-op, byte-identical /metrics), still worker/playbook/system-only with a code-level control-plane guard (assert_data_plane_access_allowed refuses gateway/api/server before any helper runs), bounded (payload byte cap + read-limit cap + short time cap), and stateless (helper opened+dropped per call). Secret-free noetl_ehdb_dataplane_ops_total{operation,outcome} metrics. Exposed as a worker/playbook-local step CLI (scripts/ehdb_dataplane_step.py) / kind smoke — not a server endpoint. Kind-validated on kind-noetl (disabled no-op + byte-identical /metrics, worker/system/playbook append→read roundtrip, control-plane gateway guard refused, secret-free metrics). See the Architecture page's Worker/Playbook Data-Plane Step section. Tracks #234.
  • Phase D — event-stream integration path: DONE (2026-07-04). A bounded, disabled-by-default worker/playbook/system drain that mirrors already-emitted NoETL events into a derived EHDB stream and consumes them through a durable consumer with explicit ack-after-materialize semantics. EHDB adds consume / ack subcommands + consume_local_reference_event_records / ack_local_reference_event_consumer (composing with the Phase C append project leg) and a read-only InMemoryStreamLog::consumer cursor getter (noetl/ehdb#237, merged 3cefba9). NoETL adds noetl.core.ehdb_eventstream (project / consume / ack) + adapter consume/ack result types + worker /metrics + step CLI + kind smoke (noetl/noetl#690). Still disabled-by-default (strict no-op, byte-identical /metrics), still worker/playbook/system-only with the code-level control-plane guard (assert_event_stream_access_allowed), bounded (payload + consume limit + ack sequence ≥ 1 + time caps), stateless, secret-free noetl_ehdb_eventstream_ops_total metrics. Event-log-authoritative invariant: the NoETL event log stays the append-only source of truth; EHDB is a derived, auxiliary consumer that never writes back to it (no event-writer import; unit-test asserted). kind-validated on kind-noetl (disabled no-op + byte-identical /metrics, worker drain
    • durable cursor restart, gateway guard refused, secret-free metrics, Phase C no-regression). See the Architecture page's Worker/Playbook Event-Stream Drain section. Tracks #234.
  • Phase E — system WASM store + RAG retrieval: COMPLETE (2026-07-05, RUST-FIRST). The EHDB side landed as #239 (ehdb 9bb5928): bounded publish / bind / resolve helpers for immutable system WASM library manifests + mutable environment/channel bindings in ehdb-reference. The worker side landed as noetl/worker#154 (merge 3162b1d): a new in-process src/ehdb/systemstore.rs bridges those helpers — publish (one atomic immutable-manifest commit), bind (channel (re)bind; rebind hot-replaces the active module, prior manifests retained), resolve (read-only replay; never-bound ⇒ Absent probe) — plus a noetl_ehdb_systemstore_* metric family and ehdb-selfcheck publish-system/bind-system/resolve-system + a system-suite A→E driver. Every boundary the Phase C/D helpers enforce is preserved and tested: disabled-by-default strict no-op (byte-identical /metrics), control-plane guard (gateway/api/server refused before any runtime opens), bounded (module-size + capability-count caps; over-bound publishes Rejected; WASM execution stays host-side/sandboxed — EHDB only catalogs the module ref), stateless (runtime opened + dropped per call), event-log-authoritative (private JSONL, never noetl.event). Kind-validated against the worker-rust image (localhost/local/noetl:ehdb-phase-e, kind-noetl, distinct ehdb-phase-e-selfcheck pod): disabled no-op (empty metrics), enabled A–D drive, enabled Phase-E system-suite (absent→publish→bind→resolve rev1→ publish→rebind→resolve rev2), control-plane server refused (exit 4, no data), oversized publish Rejected (exit 3), 0 noetl.event writes, secret-free metrics, Rust stack 0 restarts. No GKE. Per the owner directive (2026-07-04) Phase E is Rust-first: the integration lives in worker-rust and invokes the ehdb crate in-process (no subprocess), Python retired. Second slice — bounded RAG retrieval — DONE (2026-07-05). The prerequisite ehdb slice landed as #240 (ehdb c2aaad5): a bounded, read-only retrieve_local_reference_context helper in ehdb-reference that runs ehdb-retrieval's search_text under the hood, plus an ingest_local_reference_retrieval_document companion and ehdb-local-reference ingest-doc / retrieve CLI verbs. Three caps are enforced inside the helper: top-k (MAX_RETRIEVAL_TOP_K = 64), per-hit result size (MAX_RETRIEVAL_MAX_CHUNK_BYTES = 64 KiB, chunk text truncated on a char boundary), and a wall-clock budget (MAX_RETRIEVAL_TIME_BUDGET_MS = 60 s, surfaced via time_capped). Over-ceiling caps ⇒ Rejected; empty query / bad tenant·namespace id ⇒ Invalid — both classified without searching so a CLI maps them to distinct exit codes. The worker side landed as noetl/worker#155 (merge d1ebaf2): a new in-process src/ehdb/rag.rs bridges the helpers — retrieve (bounded, read-only text search; NOETL_EHDB_RAG_TOP_K 8/64, NOETL_EHDB_RAG_MAX_CHUNK_BYTES 4 KiB/64 KiB, NOETL_EHDB_RAG_TIME_BUDGET_MS 5 s/60 s) and ingest — plus a noetl_ehdb_rag_* metric family and ehdb-selfcheck ingest-rag / retrieve-rag / rag-suite (A→E+RAG driver). Every Phase C/D/E boundary is preserved and tested: disabled-by-default no-op (byte-identical /metrics), control-plane guard (gateway/api/server refused before any runtime opens), bounded, stateless, event-log-authoritative (retrieve read-only; ingest writes only the private JSONL fabric, never noetl.event). Validated via the built ehdb-selfcheck A→E+RAG drive: disabled no-op (empty metrics), enabled rag-suite (ingest 3 chunks → retrieve hit with top-k truncation → empty → over-limit rejected, ok:true, metrics_secret_free), over-limit Rejected (exit 3), control-plane server refused (exit 4, no retrieval), secret-free noetl_ehdb_rag_* metrics, plus 8 unit tests. In-container kind-noetl validation deferred — the worker-rust image (localhost/local/noetl:ehdb-rag, distinct ehdb-rag-selfcheck pod) is a ~110-min cold build (the ehdb-reference rev bump busts the cargo-chef dependency cache) that timed out in this environment; the built ehdb-selfcheck drive above is the standing evidence and the same behaviours were kind-validated for the Phase C/D/E first slices. No GKE. With both slices done Phase E is complete; the broad EHDB-integration umbrella #234 stays open for the remaining roadmap (Phase 6 PostgreSQL replacement path, Phase 7 dependency collapse).

Integration debt: re-home Phase B–D worker integration into worker-rust — DONE (2026-07-04)

The Phase B–D worker/playbook integration glue (readiness hook, data-plane step, event-stream drain) was originally added in the legacy Python worker runtime (noetl/core), shelling out to the Rust ehdb-local-reference binary as a subprocess. Prod runs worker-rust, so those disabled-by-default Python hooks never executed in prod — a stopgap.

Re-homed (owner directive 2026-07-04): the EHDB worker/playbook integration is now Rust-only. worker-rust owns it in process (noetl/worker src/ehdb), calling the ehdb-reference crate directly (summarize / append / read / consume / ack) with no subprocess shell-out — readiness (non-fatal bootstrap preflight), bounded data-plane append/read, and the event-stream project/consume/ack durable-consumer drain. Every boundary the Python stopgap enforced is preserved and tested: disabled-by-default strict no-op (byte-identical /metrics), control-plane guard (gateway/api/server refused), bounded + stateless, secret-free noetl_ehdb_* metrics, and the event-log-authoritative invariant (structurally asserted — no NoETL event-writer import).

  • worker-rust — noetl/worker#153 MERGED (d6226a2). Ships ehdb-selfcheck for in-image validation.
  • Python retire — noetl/noetl#691 MERGED (merge commit ff3a920f) — removes the Python EHDB modules / wiring / bundled helper binary so the Python path can never run as a parallel implementation; adds a guard test. noetl main now carries zero noetl/core/ehdb_* modules.
  • Kind-validated against the worker-rust image (kind-noetl): disabled no-op (byte-identical), enabled worker full drive, cross-process durable cursor, control-plane guard refused (no write), secret-free metrics, event-log-authoritative. No GKE. Rust stack undisturbed (0 restarts).

Both PRs merged; noetl/ehdb#238 is CLOSED — the re-home is complete. Phase E (system WASM store → RAG) is Rust-first from the start so it does not re-open this debt.

EHDB completion program (Phases 6–10)

Decided by RFC: EHDB Completion Program — Server↔EHDB Coupling + noetl Self-Sufficiency (status: decided, implementing). Tracked as noetl/ehdb#241 (with the integration umbrella #234).

The completion trajectory makes EHDB NoETL's self-sufficient internal storage fabric — a family of core engines (event-log, projection, KV/state, object/blob, vector) under one catalog/URN namespace. The end-state: the NoETL platform runs on Kubernetes with no external infrastructure dependency for platform functionality, EHDB the default backend for every platform tier, every tier still selectable back to the incumbent engine.

Platform-only boundary (load-bearing): EHDB serves NoETL platform functionality only — event log, projections, platform KV/state, catalog, system-WASM store, platform artifacts/vector. Business data is never stored in EHDB — tenant/domain data stays in real business systems (PostgreSQL, Snowflake, Cassandra, Kafka, Elasticsearch, ClickHouse, object stores, …) reached via playbook plugins/connectors under playbook policy. "EHDB replaces PostgreSQL / Qdrant / NATS-JetStream / object store" always means NoETL's internal platform uses of those engines, never the business-facing connector targets.

The old "Phase 6: PostgreSQL Replacement Path" and "Phase 7: Dependency Collapse Path" (#6) are absorbed into this trajectory: Phase 6 (log store) + Phase 7 (projections) cover the Postgres-materializer replacement; Phases 8–9 cover dependency collapse; Phase 10 is the tunable driver surface.

Program status (2026-07-06): CODE-COMPLETE. Phases 6–8 built + shadow-verified every platform-tier engine; Phase 9 activated each tier's reversible primary cutover (all five IMPLEMENTED + MERGED + in-kind dual-run VALIDATED); Phase 10 landed the consolidated Backend Configuration surface. The whole program is code-complete and kind-validated — the only remaining step is the prod/GKE cutover, gated on the user, per tier. Nothing in prod has changed (prod worker v5.52.0; all NOETL_EHDB_* flags default off).

Phase 6: Event-Log Core Engine (event-sourcing bottleneck fix)

Status: in progress — engine + disabled-by-default shadow slice landed (2026-07-05); the shadow mirror is now wired into the LIVE event-emit path and PROVEN on real in-kind drives (worker v5.67.0, 2026-07-06). Design note: Event-Log Core Engine (Phase 6). Cutover to serving the log from EHDB (primary) stays a later, separately-gated step (Phase 9 tier 1, below).

⚠ Superseded 2026-08-13. That cutover has happened: the event-log tier is primary and serving on prod, reads resolving through the writer-fronted tier service. Its mirror went asynchronous on 2026-08-19 — see Runbook: the async event-log mirror. Read this section as the plan of record, not current state.

The headline motivation. EHDB's event-log core engine becomes the durable persistence + ordering + serving layer for the append-only noetl.event log, replacing the NATS JetStream + PostgreSQL log-and-store path that is today's scaling pressure point (off-server state builder, unbounded WAL index noetl/ai-meta#166, materializer soak noetl/ai-meta#104).

Event authorship is unchanged: gateway/server remain the gatekeeper of what enters the log through the append-only producer path; non-log data-plane roles never fabricate events. What changes is the engine underneath the producer path.

Acceptance:

  • EHDB event-log engine persists + orders + serves noetl.event with append-only, immutable, replay-is-truth semantics preserved.
  • A NoETL execution flow drives end-to-end with EHDB as the log engine (behind the driver, dual-run against JetStream+Postgres for verify).
  • Rollback to JetStream+Postgres documented; kind-validated before GKE.

Landed so far (design + shadow slice):

  • ehdb-reference::eventlog — the event-log core engine behind the EventLogDriver trait (append / scan_global / read_execution / tail / ack), with LocalReferenceEventLogDriver composing the append-only stream primitives over one canonical noetl_event_log stream so its sequence is the global, monotonic, gapless event-log sequence. Per-execution scope via noetl.event.exec.<execution_id> subject; durable-consumer tail/ack; compare_shadow_parity. CLI eventlog-* verbs. (ehdb PR #242, merged 9a9b28d.)
  • worker-rust src/ehdb/eventlog.rs — disabled-by-default shadow behind NOETL_EHDB_EVENTLOG=off|shadow|primary (default off): shadow dual-writes each already-authored event into the engine + compares sequence/count/order parity without serving reads or touching the authoritative path; primary recognised but not activated (compile-time guard). ⚠ No longer true as of 2026-08-13 — primary is activated and serving on prod; see the banner on Home. Secret-free noetl_ehdb_eventlog_* metrics; control-plane guard; ehdb-selfcheck mirror-eventlog / eventlog-suite.
  • Runtime mirror wired LIVE + proven on a real drive (worker v5.67.0, worker#167 merged d310c7b, release ae4164d, 2026-07-06). The shadow mirror now fires from the real event-emit chokepoint ControlPlaneClient::emit_event (every worker path — EventEmitter, retry, spool, subscription, plugin — funnels through it), not only via ehdb-selfcheck/tests. Armed once at construction when NOETL_EHDB_ENABLED + NOETL_EHDB_EVENTLOG=shadow + data-plane role + local-reference log; disabled / off / primary / control-plane role ⇒ strict per-event no-op (byte-identical /metrics). mirror_live_event is panic-isolated + best-effort — a mirror failure surfaces as a metered non-ok outcome and never propagates into the authoritative event path; shadow never serves; event authorship untouched. Proven live in the kind-noetl stack (SHADOW, both data-plane pools, LOCAL only): real drives automation/pft_sql_probe_v2 → executions 332760742153424896 and 332760854506246144 each mirrored their exactly-6 events into the reference tier (subject noetl.event.exec.<id>), noetl_ehdb_eventlog_ops_total mirror/mirrored advancing, last_ok=1/last_degraded=0, 0 restarts, while Postgres noetl.event persisted all events unaffected. Only the eventlog tier is wired live so far — projection/kv/object/vector still mirror via ehdb-selfcheck/tests only. GOTCHA: the pool is not sticky on v5.67.0 — a redeploy/recreate reverts it to v5.66.0 and the mirror goes dark until re-rolled. Prod unchanged. See the Sessions-Log entries for the live-drive evidence.

Remaining before Phase 7: production segmented disk format + offset index + compaction for primary-serve; sharded/multi-stream global ordering; a JetStream+Postgres EventLogDriver for the tunable surface (Phase 10); the projection engine (Phase 7) attaching to the tail/ack/ read-execution serving surface; primary-serve cutover (dual-run verify + documented rollback, kind before GKE).

Durable event-log backend — prereq for prod-primary

Status: first slice landed (2026-07-06) — the durable segment store + crash recovery the Phase-6 note deferred as "production disk format", and the hard blocker the prod-cutover runbook §C durability gate names for Stage C (the only backend today, local_reference, is a pod-local JSONL file lost on restart + divergent across replicas → not production-durable as the authoritative store under primary). Design note: Durable Event-Log Backend; program + slice checklist tracked in noetl/ehdb#254.

Landed (ehdb#253): ehdb-reference::durable_eventlog — DurableSegmentStore (append-only CRC32-framed seg-*.eslog segments + size rollover, in-memory offset index, fsync-per-append, crash-recovery replay with torn-tail discard + bit-rot hard-error) and DurableEventLogDriver implementing the same EventLogDriver contract, behind the EventLogStorageBackend (local_reference | durable_segment) selector with local_reference the default. Payloads cold-loaded via the index (bounded index memory, the #166 property). CLI durable-eventlog-recovery proves zero-loss + ordering + scope + durable-cursor survival across a simulated restart. 17 tests incl. parity vs LocalReference; clippy -D warnings + fmt + workspace test green. In-container kind validation via the worker image is PENDING — follow-up deploy.

Remaining slices: execution-affinity single-writer routing (XxHash64 ownership, reads route-to-owner / cold-load) → shared object-store segment tier → worker wiring (NOETL_EHDB_EVENTLOG_BACKEND) → kind restart soak → prod-durability sign-off (clears §C; Stage C still user-gated per tier).

Phase 7: Projection / Read-Model Engine

Status: in progress — design + engine slice merged (ehdb#243, e0f1c0f) + disabled-by-default worker shadow wiring merged (worker#157, eadc3a5) (2026-07-05). Design note: Projection / Read-Model Engine (Phase 7). Read-cutover off Postgres is NOT part of this phase — it stays a later, separately-gated step (Phase 9).

EHDB's projection engine builds + serves the materialized read-models (execution / event / runtime state) off the event log, retiring the PostgreSQL materializer and projected state tables. It consumes the Phase-6 EventLogDriver tail, materializes the same read-models the Postgres materializer produces (event read-model keyed on event_id; folded execution-state; durable consumer checkpoint), and stays consistent under the #103 sole-writer rule (it only reads the log; the materializer stays the sole noetl.event writer).

Landed this slice:

  • ehdb — ehdb-reference::projection: the ProjectionDriver trait (apply / read_execution_state / read_event / list_executions / checkpoint) + LocalReferenceProjectionEngine, composing the append-only stream primitives over one noetl_projection_log store. Idempotent / exactly-once apply keyed on the Phase-6 global sequence (skip <= checkpoint
    • event_id dedup); deterministic rebuild-from-log; compare_projection_parity. CLI projection-apply / projection-read-exec / projection-read-event / projection-list / projection-checkpoint / projection-from-eventlog / projection-suite.
  • worker (worker#157, eadc3a5) — src/ehdb/projection.rs shadow behind NOETL_EHDB_PROJECTION=off|shadow|primary (default off): shadow dual-materializes the read-models from the event-log tail alongside the Postgres materializer + compares parity without serving reads or touching the authoritative materializer; primary recognised but not activated (compile-time PRIMARY_SERVE_ACTIVATED = false). Pins ehdb-reference to the merged #243 rev e0f1c0f.

Acceptance (met by this slice, except the gated cutover):

  • ✅ Read-models materialized from EHDB projections match the Postgres materializer output under dual-materialize (key / value / checkpoint parity, zero divergence).
  • ✅ Projection rebuild-from-log is deterministic and bounded (batch-boundary independent, unit-proved).
  • ✅ Read-cutover + rollback-to-Postgres documented — primary-serve IMPLEMENTED + MERGED as Phase 9 tier 2 (2026-07-05, ehdb#248 + worker#162); kind dual-run VALIDATED (2026-07-05, in-cluster); prod/GKE cutover STILL GATED on user. See Phase 9 tier 2 below.

Remaining before Phase 8: production projection store format (segmented + indexed) + compaction for primary-serve; richer execution-state fold (full projection_snapshot column parity); sharded/multi-store projection aligned with server shard publish; a Postgres-materializer ProjectionDriver for the Phase-10 tunable surface; the primary read-serving cutover off Postgres (Phase 9, separately gated, dual-run verify + rollback, kind before GKE).

Phase 8: KV/State + Object/Blob + Vector Engines

Status: in progress — all three engine slices now shadow-complete (2026-07-05): the KV/state engine slice (ehdb#244, worker KV shadow worker#158), the object/blob engine slice (ehdb#245, worker object shadow worker#159, v5.59.0), and the vector engine slice (ehdb#246, worker vector shadow worker#160, v5.60.0) — all three disabled-by-default worker shadows. Phase 8 stays in progress even though every engine is built + shadow-complete: no tier is cut over to primary — per-tier primary cutover is the separately-gated Phase 9 step. Design note: KV/State + Object/Blob + Vector Engines (Phase 8).

Runtime live-wiring (following the event-log tier, Phase 6). The Phase-8 shadow engines were exercised only by ehdb-selfcheck; wiring them into the worker's real runtime paths is a follow-on. Status (worker worker#168, v5.68.0, 2ec2e2b):

  • KV — wired live. kv::mirror_live_put invoked in SpoolRuntime::persist_circuit after the authoritative NATS-KV circuit-state put (bucket noetl_subscription_circuit; the only live platform NATS-KV write in the worker).
  • object — wired live. object::mirror_live_put invoked in ControlPlaneClient::object_put, the single chokepoint every platform object tier (result-tier, state-shard, plugin intent) funnels through — mirrors ALL live object puts with digest parity.
  • projection — deferred. shadow_project is a batch materialize that reads back the whole accumulating projection log and compares against a full authoritative execution-state fold; a per-event hook (at emit_event or the off-server state builder's WalEventIndex::apply) would report persistent false key-divergence. Faithful seam = a bounded windowed batch drive that also supplies the incumbent materializer's fold — larger than a call-site hook.
  • vector — deferred. No live platform vector-upsert exists in the worker loop today (platform RAG is in-process, read-only); a hook now would never fire. It lands with a future platform-RAG ingest/embed write site.

Each live hook arms only for NOETL_EHDB_ENABLED + NOETL_EHDB_<TIER>=shadow + a data-plane role, is a strict no-op otherwise, and is error-isolated (best-effort, metered, never propagated). Live-drive proof is pending the next image redeploy (the live pool is on v5.67.0; v5.68.0 carries these hooks).

EHDB takes over NoETL's internal platform tiers currently on NATS KV (coherence state — chain heads, exec descriptors), the external object store (platform artifacts, result-tier payloads, Arrow IPC), and Qdrant (platform RAG/retrieval, catalog embeddings).

Acceptance:

  • Platform KV/state, object/blob, and vector reads/writes served by EHDB engines under dual-run verify against the incumbents.
    • ✅ KV/state engine + disabled-by-default worker shadow (parity presence/value/TTL).
    • ✅ Object/blob engine + disabled-by-default worker shadow (content-addressed; parity digest/length/retrievability) for state-shards + result-tier.
    • ✅ Vector engine + disabled-by-default worker shadow (bounded cosine top-k; parity id-set/rank-order/score-monotonicity) over the platform RAG collections — the last Phase-8 engine slice, formalizing the in-process Phase-E retrieval path behind a VectorDriver.
  • The platform-only boundary holds — no business-connector target moves.
  • Per-tier rollback documented; kind-validated.

Phase 8 → Phase 9 boundary. All three engines are shadow-complete, but the box below stays unchecked — a shadow proves parity, it does not serve reads. Flipping any tier to primary (serving reads from EHDB, retiring the incumbent for that tier) is a Phase 9 step, gated and per-tier. See Phase 9 for the five per-tier cutovers this unlocks.

Phase 9: External-Dependency Retirement (Kubernetes-only)

With Phases 6–8 landed (every platform-tier engine built + shadow-verified), the NoETL platform runs self-sufficient on Kubernetes — no external NATS/JetStream, Postgres, Qdrant, or object store required for platform functionality. Only the k8s runtime remains a hard dependency.

Phase 9 = per-tier primary cutover. Phase 8 left every tier in shadow (dual-write + parity, incumbent still authoritative). Phase 9 flips each tier to primary (EHDB serves reads, the incumbent is retired for that tier) as a separately-gated step per tier — dual-run verify against the shadow parity history, a documented per-tier rollback, kind-validated before any GKE rollout. The five cutovers, each independent:

# Tier Shadow flag (Phase 8, off→shadow) Phase-9 cutover (shadow→primary) Incumbent retired Status
1 Event log NOETL_EHDB_EVENTLOG serve the event log from EHDB NATS JetStream + Postgres noetl.event primary-serve IMPLEMENTED + MERGED (2026-07-05); kind dual-run VALIDATED (2026-07-05, in-cluster); prod cutover runbook drafted 2026-07-06 — awaiting user go; prod/GKE cutover STILL GATED on user
2 Projection / read-model NOETL_EHDB_PROJECTION serve read-models from EHDB Postgres materializer primary-serve IMPLEMENTED + MERGED (2026-07-05); kind dual-run VALIDATED (2026-07-05, in-cluster); prod/GKE cutover STILL GATED on user
3 KV / state NOETL_EHDB_KV serve platform KV from EHDB NATS KV primary-serve IMPLEMENTED + MERGED (2026-07-05); kind dual-run VALIDATED (2026-07-06, in-cluster); prod/GKE cutover STILL GATED on user
4 Object / blob NOETL_EHDB_OBJECT serve state-shards + result-tier from EHDB external object store (GCS/S3) primary-serve IMPLEMENTED + MERGED (2026-07-05); kind dual-run VALIDATED (2026-07-06, in-cluster); prod/GKE cutover STILL GATED on user
5 Vector NOETL_EHDB_VECTOR serve platform RAG retrieval from EHDB Qdrant primary-serve IMPLEMENTED + MERGED (2026-07-06); kind dual-run VALIDATED (2026-07-06, in-cluster); prod/GKE cutover STILL GATED on user

Phase 9 is CODE-COMPLETE (2026-07-06): all five per-tier primary cutovers are IMPLEMENTED + MERGED and in-kind dual-run VALIDATED in-cluster (each activated + reversible; tiers 1–2 validated 2026-07-05, tiers 3–5 validated in a combined in-cluster run 2026-07-06 — worker v5.65.0 image v5.65.0-p9 7d30b250d261, native-arm64, loaded into kind-noetl; Job ehdb-p9-t345 ns ehdb-p9-validate, pod on node noetl-control-plane, container kernel 6.19.7-200.fc43.aarch64 — NOT the macOS host; pod Succeeded, 0 restarts). Every tier stays driver-selectable back to its incumbent (Phase 10) so a cutover is reversible. The one thing that remains is the prod/GKE cutover — still gated on the user, per tier. Nothing in prod changed (prod runs worker v5.52.0; all NOETL_EHDB_* flags default off).

Tier 1 (event log) — primary-serve implemented + merged (2026-07-05)

The first per-tier primary cutover — the reason for the whole program (the event-log bottleneck) — is built and merged, activated behind NOETL_EHDB_EVENTLOG=primary, reversible, and dual-run-verified. The prod/GKE cutover remains a separate later step gated on the user — nothing in prod changed.

  • ehdb — #247 (merged 7f014c9): ehdb-reference::eventlog gains exercise_primary_serve + EventLogPrimaryEvent + EventLogPrimaryServeReport::served_by_ehdb() — one authoritative cycle through every serving leg (append → global scan → per-execution scoped read → durable tail → ack → fresh-driver replay) that asserts the JetStream+Postgres semantics are preserved (monotonic gapless global sequence, per-execution scope, durable cursor advance, replay-is-truth) and dual-run parity-checks each append against the incumbent sequence. CLI verb eventlog-primary-serve.
  • worker — #161 (merged ddf41de, released v5.61.0 7e98538): src/ehdb/eventlog.rs flips PRIMARY_SERVE_ACTIVATED false → true; primary mode now serves the append authoritatively (outcome served_primary, or primary_divergence on dual-run divergence). serve_primary_cycle drives the full authoritative cycle and then demonstrates reversibility (flip back to shadow, mirror one more event over the same log, confirm it replays whole). ehdb-selfcheck eventlog-primary-serve verb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.

Local proof (same ehdb-selfcheck binary the worker image ships; debug host build, logic identical to release): off ⇒ byte-identical no-op (exit 0); primary ⇒ served_by_ehdb:true + reversible:true + secret-free metrics (exit 0); control-plane role ⇒ guard_refused (exit 4).

Kind dual-run: VALIDATED (2026-07-05). The podman VM was recovered (stop/start reset the wedged ssh socket), the worker v5.62.0 image (36875e3, ships both tiers + ehdb-selfcheck) was built native-arm64, loaded into kind-noetl (image sha256:9d222db6), and ehdb-selfcheck eventlog-primary-serve ran in-cluster in a Job pod on node noetl-control-plane (namespace ehdb-p9-validate; pod host kernel 6.19.7-...fc43.aarch64 = the kind node, not the macOS host, not GKE; kubectl context kind-noetl, API server https://127.0.0.1:61866). Evidence: off ⇒ ehdb:disabled + metrics_empty:true (exit 0); primary ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true

  • dual_run_holds:true (3× count_ok/order_ok/sequence_ok, divergence:null)
  • replay_matches:true + scope_ok:true + records_after_revert:4 (incumbent path restored whole on flip-back to shadow, zero data loss) + secret-free noetl_ehdb_eventlog_* metrics (exit 0); control-plane role server ⇒ guard_refused (exit 4). Pod Succeeded, 0 restarts. prod/GKE cutover stays GATED on the user — nothing in prod changed.
Tier-1 rollback procedure (event log)

Two independent levers restore the JetStream+Postgres incumbent with zero data loss:

  1. Runtime flag (operational, instant, no redeploy). Set NOETL_EHDB_EVENTLOG=shadow (or off) on the worker/system pool. The incumbent is authoritative again immediately. Zero data loss: the primary path only ever appends to the EHDB KeepAll log and never mutates or deletes anything the incumbent owns, so the incumbent's store is exactly as it was and the EHDB log stays whole on disk for a later re-enable.
  2. Compile-time kill switch (structural). Set PRIMARY_SERVE_ACTIVATED = false in src/ehdb/eventlog.rs and redeploy — primary then degrades to primary_unavailable regardless of config, making serving structurally unreachable.

Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. Event authorship is never affected by either — the gateway/server stay the gatekeeper of what is appended; only the serving engine underneath changes.

Tier-1 prod cutover runbook (drafted 2026-07-06 — awaiting user go)

The staged prod/GKE cutover procedure is drafted at Runbook — Prod Cutover: Event-Log Tier: Stage A (deploy worker v5.66.0 flags-off, behavior-neutral) → Stage B (shadow in prod: dual-write + parity, incumbent still authoritative) → Stage C (primary flip). Nothing has been executed — it is a reviewable plan the user approves per stage. It flags one hard blocker for Stage C: the shipped local_reference backend is a pod-local JSONL file, which is safe for shadow but not durable/shared enough to be the authoritative store under primary without a PVC-backed/shared substrate + single-writer topology. Stages A–B are executable now; Stage C is gated on that durability decision.

Tier 2 (projection / read-model) — primary-serve implemented + merged (2026-07-05)

The second per-tier primary cutover — serving the materialized read-models the control plane queries from EHDB in place of the PostgreSQL materializer — is built and merged, activated behind NOETL_EHDB_PROJECTION=primary, reversible, and dual-run-verified. It mirrors the tier-1 event-log pattern exactly. The prod/GKE cutover remains a separate later step gated on the user — nothing in prod changed.

  • ehdb — #248 (merged d08013c): ehdb-reference::projection gains exercise_primary_serve + ProjectionPrimaryInput + ProjectionPrimaryServeReport::served_by_ehdb() — one authoritative cycle through every serving leg (apply/materialize → the three read-model query contracts list_executions / per-execution read_execution_state / read_event → durable checkpoint → idempotent re-apply → fresh-engine replay) that asserts the PostgreSQL-materializer query contracts are preserved (identical read-models, per-execution scope, exactly-once on the global sequence, replay-is-truth) and dual-run parity-checks the served read-models against the incumbent materializer via compare_projection_parity. CLI verb projection-primary-serve.
  • worker — #162 (merged a56583c, released v5.62.0 36875e3): src/ehdb/projection.rs flips PRIMARY_SERVE_ACTIVATED false → true; primary mode now serves the read-models authoritatively (outcome served_primary, or primary_divergence on dual-run divergence). serve_primary_cycle drives the full authoritative cycle and then demonstrates reversibility (flip back to shadow, materialize one more execution over the same store, confirm the read-models replay whole). ehdb-selfcheck projection-primary-serve verb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.

Local proof (same ehdb-selfcheck binary the worker image ships; debug host build, logic identical to release): off ⇒ byte-identical no-op (exit 0); primary ⇒ served_by_ehdb:true + reversible:true + secret-free metrics (exit 0); control-plane role ⇒ guard_refused (exit 4).

Kind dual-run: VALIDATED (2026-07-05). Batched with tier-1 on the same v5.62.0 image (36875e3, sha256:9d222db6) in the same in-cluster Job pod on node noetl-control-plane (namespace ehdb-p9-validate, context kind-noetl, not the macOS host, not GKE). ehdb-selfcheck projection-primary-serve evidence: off ⇒ ehdb:disabled + metrics_empty:true (exit 0); primary ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true

  • dual_run_holds:true (key_ok/value_ok/checkpoint_ok, checkpoint_lag:0, divergence:null) + list_ok:true (list_count:2) + read_event_ok:true
  • replay_idempotent:true + replay_matches:true + scope_ok:true
  • rows_after_revert:3 (Postgres-materializer read path restored whole on flip-back to shadow, zero data loss) + secret-free noetl_ehdb_projection_* metrics (exit 0); control-plane role server ⇒ guard_refused (exit 4). Pod Succeeded, 0 restarts. prod/GKE cutover stays GATED on the user.
Tier-2 rollback procedure (projection / read-model)

Two independent levers restore the PostgreSQL materializer read path with zero data loss:

  1. Runtime flag (operational, instant, no redeploy). Set NOETL_EHDB_PROJECTION=shadow (or off) on the worker/system pool. The PostgreSQL materializer is the authoritative read path again immediately. Zero data loss: the primary path only ever materializes into the derived EHDB KeepAll projection store by consuming already-authored events and never mutates or deletes anything the incumbent owns, so the incumbent read-models are exactly as they were and the EHDB store stays whole on disk for a later re-enable.
  2. Compile-time kill switch (structural). Set PRIMARY_SERVE_ACTIVATED = false in src/ehdb/projection.rs and redeploy — primary then degrades to primary_unavailable regardless of config, making serving structurally unreachable.

Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. A projection is a derived read-model built by consuming the append-only event log — it never authors an event; the cutover changes only which engine serves the read-models, never the event log (the source of truth).

Tier 3 (KV / state) — primary-serve implemented + merged (2026-07-05)

The third per-tier primary cutover — serving NoETL's internal platform KV/state tier from EHDB in place of the internal NATS-KV bucket (the worker's noetl_subscription_circuit breaker store today, the #115 program-scale coherence keys as they move off NATS-KV) — is built and merged, activated behind NOETL_EHDB_KV=primary, reversible, and dual-run-verified. It mirrors the tier-1 event-log and tier-2 projection patterns exactly. The prod/GKE cutover remains a separate later step gated on the user — nothing in prod changed. Business (tenant/domain) KV stays external, reached by playbook connectors — it never flows through this tier.

  • ehdb — #249 (merged 73b1446): ehdb-reference::kv gains exercise_primary_serve + KvPrimaryInput + KvPrimaryServeReport::served_by_ehdb() — one authoritative cycle through every serving leg (put → per-key served get → bucket scan → optimistic CAS (versioned swap + create-only conflict) → tombstone delete → absolute-TTL lease → fresh-driver replay) that asserts the NATS-KV semantics are preserved (last-writer-wins get, bucket scan, optimistic CAS, tombstone delete, absolute TTL, replay-is-truth) and dual-run parity-checks each served read against a NATS-KV mirror applied in lockstep via compare_kv_parity. CLI verb kv-primary-serve.
  • worker — #163 (merged ba9f829, released v5.63.0 a7925e0): src/ehdb/kv.rs flips PRIMARY_SERVE_ACTIVATED false → true; primary mode now serves the KV op authoritatively (outcome served_primary, or primary_divergence on dual-run divergence). serve_primary_cycle drives the full authoritative cycle and then demonstrates reversibility (flip back to shadow, mirror one more key over the same store, confirm the store serves the whole live set). ehdb-selfcheck kv-primary-serve verb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.

Local proof (same ehdb-selfcheck binary the worker image ships; debug host build, logic identical to release): off ⇒ byte-identical no-op (ehdb:disabled

  • metrics_empty:true, exit 0); primary ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (5 served-read parities, present_ok/value_ok/ttl_ok, divergence:null) + put_ok/get_ok/scan_ok/cas_ok/delete_ok/ttl_ok/replay_matches all true + keys_after_revert:3 (NATS-KV path restored whole on flip-back to shadow, zero data loss) + secret-free noetl_ehdb_kv_* metrics (exit 0); control-plane role server ⇒ guard_refused (exit 4).

Kind dual-run: VALIDATED 2026-07-06 (in-cluster). Worker v5.65.0 image localhost/noetl-worker:v5.65.0-p9 (7d30b250d261, native-arm64, ships ehdb-selfcheck) was loaded into kind-noetl and ehdb-selfcheck kv-primary-serve ran in Job ehdb-p9-t345 (ns ehdb-p9-validate) on node noetl-control-plane — container kernel 6.19.7-200.fc43.aarch64, Alpine 3.22, NOT the macOS host; context kind-noetl, API https://127.0.0.1:61866. Evidence: off ⇒ disabled + metrics_empty (exit 0); primary (worker role) ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (5 served-read parities, present_ok/value_ok/ttl_ok against a NATS-KV mirror applied in lockstep) + put_ok/get_ok/scan_ok/ cas_ok/delete_ok/ttl_ok/replay_matches all true + keys_after_revert:3

  • secret-free noetl_ehdb_kv_* metrics (exit 0); primary (server role) ⇒ guard_refused (exit 4). Pod Succeeded, 0 restarts. prod/GKE cutover stays GATED on the user — nothing in prod changed.
Tier-3 rollback procedure (KV / state)

Two independent levers restore the internal NATS-KV path with zero data loss:

  1. Runtime flag (operational, instant, no redeploy). Set NOETL_EHDB_KV=shadow (or off) on the worker/system pool. The internal NATS-KV bucket is the authoritative KV tier again immediately. Zero data loss: the primary path only ever appends to the derived EHDB KeepAll KV stream and never mutates or deletes anything NATS-KV owns, so the NATS-KV bucket is exactly as it was and the EHDB store stays whole on disk for a later re-enable.
  2. Compile-time kill switch (structural). Set PRIMARY_SERVE_ACTIVATED = false in src/ehdb/kv.rs and redeploy — primary then degrades to primary_unavailable regardless of config, making serving structurally unreachable.

Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. A KV entry is derived platform state, not an event — this tier never authors the event log; the cutover changes only which engine serves the platform KV state, never the event log (the source of truth). Platform KV only — business KV never flows through it.

Tier 4 (object / blob) — primary-serve implemented + merged (2026-07-05)

The fourth per-tier primary cutover — serving NoETL's internal platform object/blob tier from EHDB in place of the internal external object store (the state shards #166, state_materializer/state_reader, and the result tier #104, result_materializer/result_resolver, both reached today through the server's /api/internal/objects/{key} API) — is built and merged, activated behind NOETL_EHDB_OBJECT=primary, reversible, and dual-run-verified. It mirrors the tier-1 event-log, tier-2 projection, and tier-3 KV patterns exactly. The prod/GKE cutover remains a separate later step gated on the user — nothing in prod changed. Business (tenant/domain) object buckets stay external, reached by playbook connectors — they never flow through this tier.

  • ehdb — #250 (merged bb42b8d): ehdb-reference::object gains exercise_primary_serve + ObjectPrimaryInput + ObjectPrimaryServeReport::served_by_ehdb() — one authoritative cycle through every serving leg (put → per-key digest-verified served get → prefix list → in-cluster locate → tombstone delete → fresh-driver replay) that asserts the external-store semantics are preserved (content-addressed put, digest-verified get, prefix list, in-cluster locate, tombstone delete, replay-is-truth) and dual-run digest-parity-checks each served read against an external-store mirror applied in lockstep via compare_object_parity (the mirror digest is the independent SHA-256 of the same bytes, so the parity is exact). CLI verb object-primary-serve.
  • worker — #164 (merged a100adf, released v5.64.0 369e4c1): src/ehdb/object.rs flips PRIMARY_SERVE_ACTIVATED false → true; primary mode now serves the object op authoritatively (outcome served_primary, or primary_divergence on dual-run divergence). serve_primary_cycle drives the full authoritative cycle and then demonstrates reversibility (flip back to shadow, mirror one more object over the same store, confirm the store serves the whole live set). ehdb-selfcheck object-primary-serve verb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.

Local proof (same ehdb-selfcheck binary the worker image ships; debug host build, logic identical to release): off ⇒ byte-identical no-op (ehdb:disabled

  • metrics_empty:true, exit 0); primary ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (4 served-read digest parities, present_ok/digest_ok/length_ok/retrievable_ok, divergence:null) + put_ok/get_ok/list_ok/locate_ok/delete_ok/replay_matches all true + keys_after_revert:3 (external-store path restored whole on flip-back to shadow, zero data loss) + secret-free noetl_ehdb_object_* metrics (exit 0); control-plane role server ⇒ guard_refused (exit 4).

Kind dual-run: VALIDATED 2026-07-06 (in-cluster). Same worker v5.65.0 image (v5.65.0-p9 7d30b250d261) + Job ehdb-p9-t345 (ns ehdb-p9-validate, node noetl-control-plane, container kernel 6.19.7-200.fc43.aarch64, NOT the macOS host; context kind-noetl). Evidence for ehdb-selfcheck object-primary-serve: off ⇒ disabled no-op (exit 0); primary (worker role) ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (4 served-read parities, digest_ok/length_ok/ present_ok/retrievable_ok — content-addressed digest integrity against the object-store mirror) + put_ok/get_ok/list_ok/locate_ok/delete_ok + replay_matches true + keys_after_revert:3 + secret-free noetl_ehdb_object_* metrics (exit 0); primary (server role) ⇒ guard_refused (exit 4). Pod Succeeded, 0 restarts. prod/GKE cutover stays GATED on the user — nothing in prod changed.

Tier-4 rollback procedure (object / blob)

Two independent levers restore the internal external object store path with zero data loss:

  1. Runtime flag (operational, instant, no redeploy). Set NOETL_EHDB_OBJECT=shadow (or off) on the worker/system pool. The internal external object store is the authoritative object tier again immediately. Zero data loss: the primary path only ever appends to the derived EHDB KeepAll object registry + content-addressed blob store and never mutates or deletes anything the external store owns, so the external store is exactly as it was and the EHDB store stays whole on disk for a later re-enable.
  2. Compile-time kill switch (structural). Set PRIMARY_SERVE_ACTIVATED = false in src/ehdb/object.rs and redeploy — primary then degrades to primary_unavailable regardless of config, making serving structurally unreachable.

Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. An object is a derived platform artifact (content-derivable from the WAL), not an event — this tier never authors the event log; the cutover changes only which engine serves the platform object bytes, never the event log (the source of truth). Platform object tier only (state shards + result tier) — business object buckets never flow through it.

Tier 5 (vector) — primary-serve implemented + merged (2026-07-06)

The fifth and final per-tier primary cutover — serving NoETL's internal platform vector tier from EHDB in place of the internal Qdrant retrieval path (the platform RAG / catalog embeddings reached in-process via the Phase-E retrieval path) — is built and merged, activated behind NOETL_EHDB_VECTOR=primary, reversible, and dual-run-verified. It mirrors the tier-1 event-log, tier-2 projection, tier-3 KV, and tier-4 object patterns exactly. With this tier all five Phase-9 primary-serve activations are implemented. The prod/GKE cutover remains a separate later step gated on the user — nothing in prod changed. Business (tenant/domain) vector collections stay external, reached by playbook connectors — they never flow through this tier.

  • ehdb — #251 (merged 0f47fe3): ehdb-reference::vector gains exercise_primary_serve + VectorPrimaryInput + VectorPrimaryServeReport::served_by_ehdb() — one authoritative cycle through every serving leg (upsert → served cosine top-k query → tombstone delete → fresh-driver replay) that asserts the Qdrant retrieval semantics are preserved (bounded cosine top-k ranking, tombstone delete, replay-is-truth) and dual-run parity-checks each served query — id-set + rank-order + score-monotonicity (compare_vector_parity) — against a Qdrant mirror ranked in lockstep with the identical cosine scoring (an independent computation, not a copy of the engine's output, so the parity is exact). CLI verb vector-primary-serve.
  • worker — #165 (merged 681782c, released v5.65.0 ceedbba): src/ehdb/vector.rs flips PRIMARY_SERVE_ACTIVATED false → true; primary mode now serves the retrieval op authoritatively (outcome served_primary, or primary_divergence on dual-run divergence). serve_primary_cycle drives the full authoritative cycle and then demonstrates reversibility (flip back to shadow, mirror one more point over the same index, confirm the collection serves the whole live set). ehdb-selfcheck vector-primary-serve verb emits the served-by-EHDB proof + reversibility + secret-free metrics; off/disabled stays a byte-identical no-op.

Local proof (same ehdb-selfcheck binary the worker image ships; debug host build, logic identical to release): off ⇒ byte-identical no-op (ehdb:disabled

  • metrics_empty:true, exit 0); primary ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (3 served-query parities, ids_ok/order_ok/monotonic_ok, divergence:null) + upsert_ok/query_ok/delete_ok/replay_matches all true + candidates_after_revert:3 (Qdrant retrieval path restored whole on flip-back to shadow, zero data loss) + secret-free noetl_ehdb_vector_* metrics (exit 0); control-plane role server ⇒ guard_refused (exit 4); over-limit dimensionality ⇒ rejected (exit 3).

Kind dual-run: VALIDATED 2026-07-06 (in-cluster). Same worker v5.65.0 image (v5.65.0-p9 7d30b250d261, the LATEST release — contains all five tiers from sequential merges) + Job ehdb-p9-t345 (ns ehdb-p9-validate, node noetl-control-plane, container kernel 6.19.7-200.fc43.aarch64, NOT the macOS host; context kind-noetl). Evidence for ehdb-selfcheck vector-primary-serve: off ⇒ disabled no-op (exit 0); primary (worker role) ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (3 top-k parities, ids_ok (id-set) / order_ok (rank-order) / monotonic_ok (score-monotonicity) against the Qdrant mirror) + upsert_ok/query_ok/delete_ok + query_returned:3 (bounded top-k cap) + replay_matches true + candidates_after_revert:3 + secret-free noetl_ehdb_vector_* metrics (exit 0); primary (server role) ⇒ guard_refused (exit 4). Pod Succeeded, 0 restarts. This is the last per-tier in-kind validation — Phase 9 is now CODE-COMPLETE. prod/GKE cutover stays GATED on the user — nothing in prod changed.

Tier-5 rollback procedure (vector)

Two independent levers restore the internal Qdrant retrieval path with zero data loss:

  1. Runtime flag (operational, instant, no redeploy). Set NOETL_EHDB_VECTOR=shadow (or off) on the worker/system pool. The internal Qdrant retrieval path is the authoritative vector tier again immediately. Zero data loss: the primary path only ever appends to the derived EHDB KeepAll vector index and never mutates or deletes anything Qdrant owns, so Qdrant is exactly as it was and the EHDB index stays whole on disk for a later re-enable.
  2. Compile-time kill switch (structural). Set PRIMARY_SERVE_ACTIVATED = false in src/ehdb/vector.rs and redeploy — primary then degrades to primary_unavailable regardless of config, making serving structurally unreachable.

Lever 1 is the operational rollback used during a cutover window; lever 2 is the belt-and-suspenders structural revert. A vector point is a derived platform index entry (content-derivable from the platform RAG source), not an event — this tier never authors the event log; the cutover changes only which engine serves platform retrieval, never the event log (the source of truth). Platform vectors only (RAG / catalog embeddings) — business vector collections never flow through it.

Acceptance:

  • A reference deployment runs the full platform on k8s with EHDB as the only storage fabric (no external infra pods/services for platform tiers).
  • Migration/import/export adapters retained for onboarding + cloud durability, but not required at run time.
  • Explicit cutover gates per dependency, each dual-run verified; kind-validated before every GKE rollout.

Phase 10: Tunable-Backend Config Surface

The per-tier backend-selection surface that makes EHDB the default while keeping every platform tier selectable back to the incumbent engine — so Phase 9's self-sufficiency is a default, not a lock-in. Full reference: Backend Configuration.

Platform tier Default Selectable back to
Event-log driver EHDB NATS JetStream + Postgres
Projection driver EHDB Postgres materializer
KV / state driver EHDB NATS KV
Object / blob driver EHDB External object store / Postgres
Vector driver EHDB Qdrant

Acceptance:

  • Per-tier driver interface with ≥2 selectable backends per tier (EHDB + incumbent).
  • Default deployment runs EHDB across all platform tiers; an operator overlay pins any tier back to its incumbent.
  • Each tier selection is dual-run verifiable with a documented rollback; kind-validated; no GKE rollout without it.

Status: IMPLEMENTED (2026-07-06). The config surface landed — a single coherent schema resolved from the existing NOETL_EHDB_* env, no breaking rename:

  • ehdb — #252 (merged 4c0df81): ehdb-reference::backends — PlatformTier (the five tiers, each with its env var + incumbent), TierMode {off|shadow|primary}, Backend {ehdb|external} + backend_for_mode (primary ⇒ EHDB, else the incumbent), and BackendMatrix with coherence validate() (rejects shadow/primary without NOETL_EHDB_ENABLED, or a data-plane tier on a control-plane role) + a secret-free to_json(). Pure data — reads no env, opens no engine.
  • worker — #166 (released v5.66.0): src/ehdb/backends.rs resolve(&EnvMap) maps the process env into the matrix by reading each tier's mode through the same <Tier>Mode::from_env parser the runtime dispatch uses (backward-compatible by construction — the consolidated view can never drift). New ehdb-selfcheck config verb prints the resolved 5-tier backend+mode matrix (secret-free; exit 0 coherent, 4 incoherent).

Selfcheck evidence (ehdb-selfcheck config, batched with the Phase-9 kind runs): all-external default ⇒ 5×backend:external, exit 0; all-EHDB (enabled worker, 5×primary) ⇒ 5×backend:ehdb, exit 0; mixed (log+vector primary, projection shadow) ⇒ per-tier ehdb/external, exit 0; primary without enable ⇒ coherent:false, exit 4; data-plane tier on gateway role ⇒ coherent:false, exit 4; sensitive-keyed env present ⇒ secret_free:true, 0 leaks. No behavior change; disabled-by-default strict no-op intact.

With Phase 10 the whole EHDB completion program (Phases 6–10) is code-complete and kind-validated. Every platform tier has a built + shadow-verified engine, a per-tier reversible primary cutover (Phase 9), and a consolidated backend-selection surface (Phase 10). The only remaining step is the prod/GKE cutover — gated on the user, per tier; nothing in prod has changed (prod runs worker v5.52.0, all NOETL_EHDB_* flags default off).

Performance & Load Testing

Tracking issue: #261 Performance & load testing for EHDB engine tiers. Design: Design: Performance & Load Testing.

Two layers:

  • Phase 1 — engine micro-benchmarks (Rust, deterministic) — LANDED (2026-07-08). criterion benches over all five reference drivers + the durable segment event-log backend (crates/ehdb-reference/benches/engine_micro.rs), with a committed baseline. Headline: the durable event-log backend beats the local_reference JSONL driver 2.7× at sustained append (255 vs 96 ev/s at K=1000) and stays flat (~3.9 ms/append, fsync-bound) while the reference driver degrades O(n) per op; segment rotation ~2%; cold replay ~185 K ev/s. KV/object/vector reference drivers are O(n)-per-op shadow fixtures (no durable backend yet).
  • Phase 2 — in-cluster end-to-end load (kind) — DESIGN ONLY, pending scoping. Drive real traffic through the worker with NOETL_EHDB_* flags armed, read noetl_ehdb_* metrics + latency, and run the EHDB-vs-incumbent (Postgres+NATS-JetStream) head-to-head. Directional only (podman-VM constraints); Layer-A micro-benches stay the reliable signal. Proposed SLO strawman on the design page awaits platform-owner confirmation.

Standing Engineering Rules

  • Rust-first implementation.
  • NoETL-domain-specific implementation; avoid generic database scope creep unless it directly serves the NoETL platform.
  • Product code stays in noetl/ehdb, not ai-meta.
  • Design is maintained in this wiki as public surfaces evolve.
  • Issues track work; ai-meta memory tracks cross-repo platform decisions only.
  • Logging changes include flood-check rationale.
  • Container image work validates in local kind before GKE.

Clone this wiki locally