Skip to content

Architecture

Kadyapam edited this page Jul 5, 2026 · 124 revisions

Architecture

EHDB is an Arrow-native NoETL-domain storage system. It is designed to store operational metadata, catalog state, event streams, AI/RAG retrieval state, and historical analytical data while keeping NoETL's execution model as the architectural boundary.

Product Shape

EHDB combines NoETL-specific versions of capabilities that are currently spread across external dependencies:

  • FoundationDB-style small transactional metadata.
  • NATS JetStream-style append-only streams, durable consumers, replay cursors, and retention policy.
  • Iceberg-style table metadata and snapshot semantics.
  • Object-store-native immutable file storage.
  • Qdrant-style embedding/vector retrieval metadata for RAG workloads.
  • NoETL system WASM library registry for hot-replaceable system playbook functionality.
  • ClickHouse-style analytical read paths over columnar data.
  • Arrow Flight and Arrow IPC as native data movement formats.
  • Multi-cloud replication and placement policy.

EHDB should remain domain specific to NoETL. It may be open source, but the product is not a generic database first. Its priority is replacing NoETL's platform dependencies with one NoETL-centric storage fabric: PostgreSQL, NATS JetStream, external object stores, Qdrant, and ClickHouse become EHDB capabilities over time.

NoETL Execution Model Fit

EHDB must honor the NoETL boundary:

gateway = gatekeeper
worker = atomic compute
playbook = ephemeral blueprint
shared cache = state vehicle
event log = source of truth

Implications:

  • Gateway routes and authorizes requests; it does not embed data-touch logic.
  • Gateway and API/admission components may embed EHDB libraries for control-plane planning and validation, but the default embedded role policy denies catalog, transaction, stream, object, retrieval, replication, and system-library data-plane capabilities to those roles.
  • Workers perform bounded compute units and do not hold slots while waiting for external callbacks.
  • Worker, playbook, and system roles receive explicit EHDB data-plane capabilities so the same embedded database substrate can serve NoETL compute without opening gateway shortcuts.
  • EHDB clients should use explicit APIs/protocols, not direct access to internal object layouts.
  • Event logs, stream logs, and transaction logs are authoritative; caches are accelerators only.
  • Tenant state must be explicit in catalog metadata, placement policy, and ACLs.

The server↔EHDB coupling (loose: control-plane stateless server, data-plane worker; the seam is the event log + catalog/URN namespace, not a per-instance binding) and the completion program (EHDB as NoETL's self-sufficient internal storage fabric — an event-log core engine that fixes the event-sourcing bottleneck, plus projection, KV/state, object/blob, and vector engines; k8s-only end-state, tunable per tier, platform-only — never business data) are decided in RFC: EHDB Completion Program — Server↔EHDB Coupling + noetl Self-Sufficiency. Event authorship rules are unchanged (gateway/server gatekeep what is appended; data-plane roles never fabricate events); the storage + ordering + projection engine underneath the append-only producer path becomes EHDB.

Embedded NoETL Roles

ehdb-core defines the first reusable embedded role/capability policy:

Role Default scope
gateway Control-plane only
api Control-plane only
worker Data-plane capabilities for bounded compute
playbook Data-plane capabilities for step-scoped work
system Control-plane plus data-plane maintenance capabilities

Capabilities distinguish admission/planning from storage work:

  • control_plane
  • catalog_read, catalog_write
  • transaction_append
  • stream_append, stream_consume
  • object_read, object_write
  • retrieval_read, retrieval_write
  • replication_plan
  • system_library_resolve

This is the foundation for EHDB as NoETL's embedded distributed database: the runtime can be present in workers, APIs, and gateways, while capability checks prevent gateway/API roles from becoming direct database readers or writers.

Ops Env-Rendering Boundary (Integration Phase A)

The role/capability policy above is enforced a second time at the deployment layer, so a mis-set Helm value cannot hand a control-plane role a data-plane storage handle. The NoETL ops Helm charts (noetl/ops automation/helm/noetl + automation/helm/gateway) render EHDB env per role, and the control-plane vs data-plane split lives in the chart template (templates/_ehdb.tpl), not in values.yaml:

Role Rendered EHDB env Storage handle
server, api, gateway NOETL_EHDB_MODE=control_plane, NOETL_EHDB_CAPABILITIES=control_plane never NOETL_EHDB_LOCAL_REFERENCE_LOG
worker, playbook, system NOETL_EHDB_MODE=local_reference + NOETL_EHDB_LOCAL_REFERENCE_LOG pod-local JSONL log under NOETL_DATA_DIR

Two invariants make this safe:

  • Disabled by default. ehdb.enabled: false in both charts. When off, the render is byte-identical to pre-EHDB — no NOETL_EHDB_* env, no attach, no behavior change. Enabling is an explicit values overlay, optionally scoped per role via ehdb.roles.<role>.enabled.
  • Boundary is not values-configurable. The role→plane mapping is hardcoded in the template keyed on the client role. An operator editing values.yaml cannot leak a local_reference log into a gateway/api/server role — the template only ever emits the control-plane env for those roles. This mirrors the runtime contract noetl.core.ehdb_contract.validate_ehdb_integration_contract, which rejects a control-plane role carrying a local-reference log and a data-plane role in a gateway/api/server role.

The packaged ehdb-local-reference helper is exercised in the worker local-reference (data-plane) path by a kind Job at noetl/ops ci/manifests/noetl/ehdb/smoke-job.yaml, run inside the NoETL image with imagePullPolicy: Never before any GKE rollout. This is ops/runtime enablement only — it wires the disabled-by-default env and the smoke gate; it does not change EHDB's default behavior.

Worker/Playbook Readiness Hook (Integration Phase B)

Phase A rendered the env; Phase B turns it into a bounded, stateless readiness preflight that a worker/playbook/system process runs at bootstrap (or as a standalone command / kind smoke step). It lives in noetl/noetl at noetl.core.ehdb_readiness.evaluate_ehdb_readiness, built over noetl.core.ehdb_adapter.read_ehdb_local_reference_summary_from_env.

The hook reads the Phase-A local-reference summary (deterministic domain counts) and returns a secret-free, structured result — outcome + role + counts + duration — classified as one of:

Outcome Meaning Ready?
disabled EHDB off (the default) — no read, no metric yes (no-op)
control_plane control-plane role — no data-plane read performed yes
ready / empty bounded summary read (non-zero / all-zero counts) yes
truncated bounded time cap tripped — degraded yes (degraded)
unavailable helper missing / errored — degraded yes (degraded)
guard_refused control-plane role handed a data-plane env no
invalid misconfigured EHDB env no

Boundary properties, all enforced in code:

  • Control-plane roles never read data. gateway/api/server return control_plane without touching the helper. A dedicated guard, assert_data_plane_read_allowed(role), raises before any read for a control-plane role — defense-in-depth on top of validate_ehdb_integration_contract. A control-plane role handed a local_reference env classifies as guard_refused (not ready).
  • Bounded + stateless. The read shells out to the bounded ehdb-local-reference summary helper under a short time cap (default 5s, clamped 0.1–30s via NOETL_EHDB_READINESS_TIMEOUT_SECONDS) and holds no connection or per-request state.
  • Disabled = byte-identical. When EHDB is off the evaluation is a strict no-op: no read, and no metric recorded, so the worker /metrics output is unchanged.
  • Not a server endpoint. Readiness is a worker/playbook-local command (scripts/ehdb_readiness_preflight.py) / kind smoke step, and a non-fatal preflight call in the worker bootstrap — never a gateway or server HTTP route, so the gatekeeper boundary is preserved.

Observability: process-local noetl_ehdb_readiness_checks_total (by outcome), noetl_ehdb_readiness_ready / _degraded gauges, and noetl_ehdb_readiness_last_duration_seconds, surfaced through the worker /metrics render. Metric text carries only the outcome label and aggregate counters — no log path, count values, or helper stderr.

Landed in PR noetl/noetl#688; tracks #234. Kind-validated on kind-noetl (disabled no-op, worker local-reference bounded read, gateway control-plane guard) with the real ehdb-local-reference helper.

Worker/Playbook Data-Plane Step (Integration Phase C)

Phase B exposed a readiness summary preflight. Phase C adds the first bounded data-plane operation: append and read a single domain record through the local-reference adapter — still disabled-by-default, still worker/playbook/system-only, with no gateway/API/server data access. It lives in noetl/noetl at noetl.core.ehdb_dataplane (append_ehdb_domain_record / read_ehdb_domain_records), built over the ehdb-local-reference append/read subcommands (noetl/ehdb#235, merged 3ae8950).

New surface:

  • noetl/core/ehdb_dataplane.py — append_ehdb_domain_record / read_ehdb_domain_records, with app-side snowflake transaction ids.
  • ehdb_adapter.py — append/read invocation builders + typed LocalReferenceAppendResult / LocalReferenceReadResult; the payload reaches the helper verbatim so the appended record is byte-exact.
  • worker/metrics.py — renders the data-plane metrics on the worker /metrics surface (renders nothing when disabled).
  • scripts/ehdb_dataplane_step.py — worker/playbook-local step CLI (append/read), not a server endpoint; scripts/smoke_ehdb_dataplane.py — kind smoke.

Boundary properties, all enforced in code:

  • Disabled by default → strict no-op. No append/read, no file, no metric — byte-identical to a build without EHDB.
  • Control-plane guard. assert_data_plane_access_allowed refuses gateway/api/server before any helper runs (defense-in-depth on top of the contract). A control-plane role handed a data-plane env classifies as guard_refused and performs no write.
  • Bounded. Payload byte cap (NOETL_EHDB_DATAPLANE_MAX_PAYLOAD_BYTES, default 65536, clamped) + read-limit cap (NOETL_EHDB_DATAPLANE_MAX_READ_LIMIT, default 1000, clamped) + short helper time cap. Over-bound requests are rejected before the helper runs.
  • Stateless. The bounded helper is opened and dropped per call.
  • Secret-free metrics. noetl_ehdb_dataplane_ops_total{operation,outcome} + last-op gauges; no stream/subject/payload/log-path values leak.

Landed in PR noetl/noetl#689; tracks #234. Validation: 23 targeted tests pass (incl. a real-binary append/read roundtrip), 84 existing EHDB tests still green; kind-validated on kind-noetl (disabled no-op + byte-identical /metrics, worker/system/playbook append→read roundtrip, control-plane gateway guard refused with no write, secret-free metrics). Nothing on GKE/prod.

Worker/Playbook Event-Stream Drain (Integration Phase D)

Phase C added a bounded domain-record append/read. Phase D adds the event-stream integration path: a bounded, disabled-by-default worker/playbook/system drain that mirrors already-emitted NoETL events into a derived EHDB stream and consumes them through a durable consumer with explicit ack-after-materialize semantics — still worker/playbook/system-only, with no gateway/API/server data access.

Event-log-authoritative invariant. The NoETL event log (noetl.event in Postgres / NATS JetStream) stays the authoritative, append-only source of truth. EHDB is a derived, auxiliary consumer of already-emitted NoETL events — the drain never writes back to the NoETL event log. It only ever runs the ehdb-local-reference helper against the separate EHDB JSONL fabric named by NOETL_EHDB_LOCAL_REFERENCE_LOG. "Project" mirrors an event that was already committed to the authoritative log; it does not emit one. There is deliberately no NoETL event-writer import in the drain module, and a unit test asserts that invariant structurally.

EHDB surface (noetl/ehdb#237, merged 3cefba9), composing with the existing Phase C append (the project leg):

  • consume subcommand + consume_local_reference_event_records — pull up to --limit records for a durable consumer after its ack cursor, creating the consumer on first pull, without moving the cursor. A never-written stream reports exists:false without creating state.
  • ack subcommand + ack_local_reference_event_consumer — advance the consumer cursor after materialize; rejects backwards moves and unknown or zero sequences atomically (replay-preview before any write).
  • InMemoryStreamLog::consumer — read-only durable-consumer lookup (including the ack cursor) for reporting the cursor on a bounded pull.

NoETL surface (noetl/noetl#690):

  • noetl/core/ehdb_eventstream.py — project / consume / ack over the local-reference adapter, with app-side snowflake transaction ids. The drain is: project a NoETL event into the derived stream → durable-consumer consume the pending batch → ack after materialize → re-consume returns only the unacked tail (durable cursor restart).
  • ehdb_adapter.py — consume/ack invocation builders + typed LocalReferenceConsumeResult / LocalReferenceAckResult.
  • worker/metrics.py — renders noetl_ehdb_eventstream_* on the worker /metrics surface (renders nothing when disabled).
  • scripts/ehdb_eventstream_step.py — worker/playbook-local step CLI (project/consume/ack), not a server endpoint; scripts/smoke_ehdb_eventstream.py — kind smoke.

Boundary properties, all enforced in code:

  • Disabled by default → strict no-op. No project/consume/ack, no file, no metric — byte-identical to a build without EHDB.
  • Control-plane guard. assert_event_stream_access_allowed refuses gateway/api/server before any helper runs. A control-plane role handed a data-plane env classifies as guard_refused and performs no write.
  • Bounded. Payload byte cap (NOETL_EHDB_EVENTSTREAM_MAX_PAYLOAD_BYTES, default 65536, clamped) + consume-batch cap (NOETL_EHDB_EVENTSTREAM_MAX_CONSUME_LIMIT, default 1000, clamped) + ack sequence ≥ 1 + short helper time cap. Over-bound requests are rejected before the helper runs.
  • Stateless. The bounded helper is opened and dropped per call. No long-lived subscription or per-tenant state between requests.
  • Secret-free metrics. noetl_ehdb_eventstream_ops_total{operation,outcome} + last-op gauges; no stream/subject/consumer/payload/log-path values leak.

Validation: 30 ehdb-reference tests pass (publish→consume→ack→reopen drain proving durable cursor restart, missing-stream absent probe, limit without cursor move, backwards/unknown/zero-sequence ack rejection); 27 NoETL Phase D tests pass (incl. a real-binary drain), 131 EHDB tests green with no regression; kind-validated on kind-noetl (disabled no-op

  • byte-identical /metrics, enabled worker drain + cursor restart, control-plane gateway guard refused with no write, secret-free metrics, Phase C append/read no-regression). Nothing on GKE/prod.

Architecture Decision: Rust-first EHDB worker integration (Python as thin wrapper)

Decision (owner directive, 2026-07-04): the EHDB worker/playbook integration is Rust-first. It must live in worker-rust and invoke the ehdb crate in-process; the Python code in noetl/core becomes a thin wrapper/binding over the Rust core, not a parallel implementation.

Why this is a decision and not just the current state:

  • The EHDB storage engine is already Rust. The ehdb crate/binary provides append / read / project / consume / ack — the actual storage-and-stream engine is Rust today.
  • The Phase B–D integration glue is Python, in the legacy runtime. The worker/playbook readiness hook (Phase B), data-plane step (Phase C), and event-stream drain (Phase D) were added in the legacy Python worker runtime (noetl/core), as a thin adapter that shells out to the Rust ehdb-local-reference binary as a subprocess.
  • Prod runs worker-rust, so those Python hooks do not execute in prod. The Phase B–D Python hooks only exercise against the Python worker. They are a disabled-by-default stopgap — useful for validating the integration contract and boundary in Python, but not the runtime that ships to production.

Target design:

  • worker-rust owns the integration in-process. It calls the ehdb crate directly, with no subprocess shell-out to the ehdb-local-reference binary.
  • Python reduces to a thin binding over the Rust core. This is the "Polars model": Python bindings over a Rust core, not a second implementation kept in sync by hand.

Implemented (2026-07-04). The re-home shipped: worker-rust now owns the Phase B–D integration in process (noetl/worker src/ehdb, noetl/worker#153 MERGED d6226a2) — readiness preflight, bounded data-plane append/read, and the event-stream project/consume/ack durable-consumer drain, calling the ehdb-reference crate directly with no subprocess. The Python EHDB path is retired (noetl/noetl#691 MERGED ff3a920f) so it is no longer a parallel implementation. Every boundary carries over and is tested (disabled-by-default byte-identical no-op, control-plane guard, bounded + stateless, secret-free metrics, event-log-authoritative). See the Roadmap Integration debt section for the kind-validation evidence. Phase E (system WASM store → RAG) is Rust-first from the start.

First-Class Catalog

The catalog is not an external service bolted onto the database. It is stored inside EHDB as native transactional tables.

Catalog domains:

  • Tenants and namespaces.
  • Databases, schemas, tables, and views.
  • Columns and Arrow data types.
  • Partitions, manifests, snapshots, and file references.
  • Statistics, indexes, and data skipping metadata.
  • Lineage and execution provenance.
  • ACLs, grants, and policy bindings.
  • RAG collections, document/chunk references, embedding metadata, and retrieval policy bindings.
  • System WASM library manifests, host capability grants, environment bindings, and release channels.

The catalog must be small, transactional, and strongly consistent for metadata mutations. Analytical data files can be immutable and eventually replicated under explicit placement rules.

The current local reference includes a narrow scan-grant model. CatalogScanGrant records tenant, namespace, table ID, principal, and granting transaction ID. InMemoryCatalog::can_scan answers the simple grant lookup, and CatalogMutation::GrantScan makes the metadata replayable through the transaction log and LocalReferenceRuntime. This is durable ACL metadata that the local Arrow Flight reference service can enforce for scans. Production IAM, policy composition, revocation, and non-loopback service exposure remain future surfaces.

The local catalog reference now models immutable table snapshots. Snapshots attach content-checked ObjectRef file sets to a table, record the committing transaction ID, and maintain a parent pointer to the previous latest snapshot. Parent-chain validation keeps snapshot history linear until branching, rollback, or compaction semantics are explicitly designed.

LocalReferenceSummary opens the local JSONL transaction log, replays it through LocalReferenceRuntime, and reports deterministic JSON counts for transactions, catalog tables/snapshots/scan grants, streams, retrieval documents/chunks/embeddings, system-library manifests/bindings, and storage replicas. The ehdb-local-reference summary --log <path> binary is a bounded local helper for NoETL worker/playbook diagnostics and integration tests. It is not a daemon, SQL/API endpoint, gateway read path, distributed executor, or production storage mutation surface.

Catalog metadata JSON decoding is strict for tables, snapshots, scan grants, create-table requests, snapshot commits, scan grant requests, table schemas, and column schemas. Unknown persisted fields are rejected before replay or catalog operations treat the metadata as valid catalog state. Table schemas revalidate column identifiers and reject duplicate column names before catalog state is created, keeping Arrow projection and predicate selectors unambiguous. Schema JSON decode follows the same constructor validation path. Core identifier JSON decode also routes through constructor validation, preserving the string JSON shape while rejecting malformed tenant, namespace, table, transaction, stream, retrieval, and related identifiers.

Event Stream Model

EHDB should subsume the NATS JetStream role for NoETL, not embed a NATS server permanently.

Stream domains:

  • Execution events and command logs.
  • Playbook callback and continuation streams.
  • Durable consumer state, replay cursors, ack policy, and retention.
  • Materializer/projector offsets.
  • Tenant-scoped subjects and routing policy.

The initial design should model streams as typed, append-only EHDB logs with explicit consumer cursors. NATS compatibility can exist during migration, but the target state is EHDB-owned stream durability.

RAG And Retrieval Model

EHDB should support AI-native NoETL workloads directly.

Retrieval domains:

  • Documents and artifacts produced by playbooks.
  • Chunk metadata, text spans, checksums, and source lineage.
  • Embedding model identity and vector dimensions.
  • Vector index metadata and placement policy.
  • Hybrid retrieval over metadata filters, text, vectors, and execution lineage.

The local reference retrieval catalog includes an exact cosine-similarity fixture over registered chunk embeddings. VectorSearch is scoped by tenant, namespace, and embedding model, validates finite non-zero query and embedding vectors, applies dimension compatibility, and returns deterministically ordered hits. This is a local correctness fixture for RAG semantics, not a production ANN index, retrieval daemon, gateway data path, Qdrant adapter, or distributed query engine.

Retrieval metadata JSON decoding is strict for documents, chunks, embeddings, registration requests, vector/text/hybrid search requests, and local search hits. Unknown fields are rejected before persisted or handed-off RAG metadata is treated as valid retrieval state.

LocalRetrievalSearchService adds the in-process service-facing boundary over that replayed retrieval state. It accepts SearchSimilarChunksRequest values and returns ranked chunk hits with document identity, text, checksum, model, dimensions, and score while excluding raw embedding vectors from the response. This is the shape a future worker/playbook step can call; it is not a gateway route, networked retrieval API, persistent service, or tenant-side process. The same service now accepts SearchTextChunksRequest values for exact case-insensitive local substring matching, returning chunk hits with match counts and deterministic ordering. This remains a correctness fixture, not a full-text index, BM25 engine, external search adapter, or distributed search service. SearchHybridChunksRequest combines exact cosine similarity and exact text match counts with caller-provided non-negative weights, producing a deterministic local score over model-scoped embeddings and replayed chunk text. This proves hybrid RAG semantics without adding a query planner, vector ANN index, full-text engine, gateway route, or distributed retrieval service. AssembleRetrievalContextRequest builds on the hybrid search boundary to assemble bounded RetrievalContextBlock values with chunk identity, document identity, ordinal, checksum, score metadata, clipped text, and total text budget accounting. This gives NoETL workers/playbooks a deterministic local context shape for RAG tests without adding prompt template rendering, LLM invocation, a network service, gateway route, retrieval daemon, or persistent per-tenant process. RetrievalContextRequestPayload and RetrievalContextResultPayload wrap context assembly requests/results in explicit versioned JSON byte payloads. Malformed JSON, unsupported versions, and invalid decoded identifiers are rejected deterministically before execution or handoff. Unknown object fields in the request envelope, assembly request, result envelope, context object, or context blocks are rejected on the same boundary, keeping local RAG handoff payloads exact. Decoded bytes must also match the canonical EHDB encoding produced by each payload's encode method. This is a local worker/playbook payload boundary, not an RPC protocol, Arrow Flight endpoint, gateway route, prompt engine, or production retrieval API. LocalRetrievalSearchService::execute_context_payload decodes a request payload, assembles context from replayed local retrieval state, and encodes the result payload. This gives worker/playbook tests a complete in-process handoff loop without introducing a daemon, network endpoint, gateway data path, or persistent tenant process. RetrievalContextPayloadExecutorConfig adds positive request/result byte limits to that local handoff. Oversized request payloads fail before decode, and oversized encoded result payloads fail before returning bytes. These are local reference guardrails, not a scheduler, quota system, or production admission controller. RetrievalContextPayloadScope adds an optional local tenant/namespace guard before context assembly. The scope-aware executor rejects decoded request payloads whose embedded tenant or namespace differs from the expected worker/playbook execution scope. This is a correctness fixture, not production IAM, ACL integration, gateway authorization, or a policy engine. RetrievalContextPayloadExecutionSummary adds redacted metrics/audit metadata for the same local worker/playbook handoff. Summary-returning execution APIs report request/result byte counts, context block count, total text chars, truncation status, and whether an execution scope was required. They deliberately exclude tenant IDs, namespace values, query text, chunk text, tokens, embedding vectors, payload bytes, object paths, and principals. This is not a logging sink, production policy surface, network API, retrieval daemon, or gateway data path. RetrievalContextPayloadExecutionReceiptPayload wraps that redacted summary in a versioned JSON byte codec. The receipt is the local durable shape future event-log/audit plumbing can store or route while still excluding retrieval-sensitive content. It does not publish an event, write a stream record, emit logs, open a network API, or create a retrieval service. Receipt encode/decode validates positive request/result byte counts and rejects text chars when there are no context blocks. Receipt decoding also rejects unknown envelope and redacted summary fields, plus non-canonical JSON bytes, before artifact validation, receipt decode handoff, or event-envelope helper use. RetrievalContextPayloadExecution::encode_receipt_payload keeps that receipt emission attached to the local execution result, so worker/playbook tests can return result bytes and receipt bytes without duplicating wrapper construction or widening the data boundary. RetrievalContextPayloadExecutionArtifacts adds the bounded local artifact shape for that handoff: result payload bytes plus redacted receipt payload bytes. The executor config now includes a receipt payload byte limit that artifact helpers enforce before returning. Artifact validation decodes and validates the receipt, rejects empty payload arrays, and verifies the receipt summary result byte count matches the actual result payload length. RetrievalContextPayloadExecutionReceiptEventPayload wraps validated receipt bytes in a versioned JSON event envelope with the stable subject ehdb.retrieval.context.execution.receipt. This gives future EHDB stream/audit plumbing a local stream-ready payload shape without publishing to streams, logging, opening a network API, creating a retrieval service, or carrying result payload/context bytes. Event payload decoding rejects unknown envelope fields and non-canonical JSON bytes before replay decode, publisher helper use, or consumer handoff. RetrievalContextReceiptEventStreamTarget can build the local stream configuration and explicitly create the receipt event stream with caller-selected retention. Setup is caller-driven, and publishing does not auto-create streams. Convenience helpers create keep-all streams or positive bounded-retention streams; zero bounded retention is rejected before touching the stream log. RetrievalContextReceiptEventStreamTarget and RetrievalContextReceiptEventStreamLog define the explicit local publisher contract for worker/playbook code that chooses to append that event envelope to an EHDB stream log. The caller provides tenant, namespace, stream name, mutable stream log, and transaction id; there is still no automatic publication, background task, gateway data path, or persistent service process. RetrievalContextReceiptEventStreamRecord and RetrievalContextReceiptEventStreamReadLog add the matching explicit local read path. Worker/playbook audit tests can replay receipt event records from caller-supplied local logs, validating the stable subject and event payload while preserving stream sequence and transaction id. This is not a background consumer, subscription loop, network API, gateway data path, or persistent service process. RetrievalContextReceiptEventDurableConsumerLog adds caller-controlled durable-consumer helpers for receipt event streams: create a consumer, replay pending validated receipt events for that consumer, and ack a receipt event sequence. The cursor belongs to the local stream log and is only advanced by explicit caller action; there is no scheduler, subscription loop, gateway data path, or persistent service process.

The goal is to replace a permanent Qdrant dependency with EHDB-native retrieval primitives that preserve NoETL tenant, lineage, and execution context.

System WASM Library Model

NoETL already has a worker-side WASM plug-in host for system-pool playbooks: commands with tool_kind: "wasm" resolve a catalog-addressed module reference { path, version, entry }, load the module by digest, and invoke it through an explicit host capability ring.

EHDB owns the durable catalog side of that model. System playbook functionality should be stored as compiled WASM libraries, not as hard-coded server branches or always-on per-tenant agent processes.

System library domains:

  • Immutable module manifests: path, revision, digest, entry export, target, object path, byte length, host capabilities, and transaction provenance.
  • Mutable bindings: tenant, namespace, environment, release channel, path, revision, and digest.
  • Environment-specific implementations: for example kind, gke-prod, aws-prod, or azure-dev can bind the same logical system library to different compiled modules.
  • Hot replacement: a stable channel can be rebound to a new digest and revision without changing the Rust crate version or forcing callers to chase semantic version bumps.
  • Local restartability: the LocalJsonlSystemLibraryCatalog journal persists publish and bind operations, then rebuilds manifests and environment/channel bindings on open. Replay revalidates persisted manifest and binding identifiers before rebuilding system-library state, and rejects unknown journal, publish, or bind fields instead of silently ignoring them.
  • Strict metadata decoding: resolved module manifests, NoETL WASM plugin references, and environment/channel bindings reject unknown JSON fields before metadata is accepted for a worker handoff.

This model keeps the execution boundary intact. Workers remain atomic compute, gateway remains a gatekeeper, and EHDB stores the authoritative library/binding state. The WASM host receives only explicit capability imports such as event publish, object put, result put, catalog mutation, stream publish, or retrieval write.

Storage Model

EHDB separates metadata, logs, and data:

Layer Shape Durability
Transaction log Ordered metadata mutations Replicated consensus log
Stream logs Ordered NoETL event/command records EHDB append log + retention
Catalog tables Typed metadata state MVCC over transactional metadata
System libraries WASM manifests and environment bindings EHDB catalog + object layer
Retrieval indexes Chunk, embedding, vector index metadata EHDB catalog + index files
Data files Arrow IPC, Parquet, blobs EHDB object layer
Snapshots Manifest and file sets Catalog + object references

Initial external object stores are migration adapters, not the final product boundary:

  • S3-compatible storage.
  • Google Cloud Storage.
  • Azure Blob Storage.
  • Local filesystem adapter for tests.

The target state is that NoETL sees EHDB as its storage system. EHDB may use cloud object APIs internally, but NoETL should not depend on a separate object-store product abstraction for ordinary platform state.

The local storage reference uses immutable object writes. ObjectRef contains the relative EHDB object path, byte length, SHA-256 digest, geo location, and data-gravity shard pointer. Verified reads recompute length and digest and fail on mismatch, giving the developer loop a concrete corruption/tamper check before distributed replication exists. Table data files use a deterministic tenant/namespace/table/snapshot layout:

{tenant}/{namespace}/tables/{table}/snapshots/{snapshot}/{file}

The local Arrow IPC table fixture is the first analytical-data-path proof over this layout. It writes an Arrow RecordBatch as an immutable IPC object, commits a catalog snapshot that points at the content-checked object reference, and reads the latest snapshot back through verified object reads before decoding Arrow. This remains a local fixture; Arrow Flight, distributed query execution, and gateway read paths are separate future service surfaces.

The local Arrow scan fixture layers a minimal read operation on top of that proof. It resolves the latest catalog snapshot, verifies each Arrow IPC object, decodes RecordBatch values, and can project named columns in caller-specified order. Empty projection lists and duplicate projection columns are rejected before object reads so direct local scans match the service/Flight projection contract. Projection-column identifiers are validated on the direct scanner boundary. Predicate pushdown, SQL planning, distributed execution, and Arrow Flight service APIs remain future surfaces.

The local equality filter fixture adds the first predicate step after decode. It applies single-column equality filters to Arrow batches after verified reads and before projection, initially for UTF-8 and Int64 columns. Predicate-column identifiers are validated before object reads. This is not predicate pushdown: no object statistics, partition pruning, SQL planner, or gateway read path is introduced.

ehdb-service introduces the first service-facing read boundary without starting a network service. LocalArrowScanService accepts a typed latest-table scan request, delegates to the local scanner, and returns an ArrowScanResult containing the Arrow schema, batches, and row count. This is the API seam for the future Arrow Flight read path; it is not a gateway shortcut, SQL planner, distributed executor, or predicate pushdown layer.

The Arrow Flight scan ticket codec is the next contract layer. A ScanFlightTicket wraps a latest-table scan request in a versioned payload, round-trips through Arrow Flight Ticket bytes, and can produce a command FlightDescriptor. The version check makes incompatible request payloads fail before execution, and identifier validation keeps decoded tenant, namespace, and table names inside the same typed boundary before ticket bytes are produced or accepted. Projection and equality-predicate column selectors are validated on the same boundary, so malformed selector identifiers, empty projection lists, and duplicate projection columns fail before local scan execution. Unknown object fields in the ticket, embedded request, or equality predicate also fail before execution, keeping the v1 JSON payload exact instead of silently narrowing future-shaped input. Decoded bytes must also match the canonical EHDB encoding produced by ScanFlightTicket::encode, so pretty-printed or reordered JSON does not cross the local Flight boundary as an accepted ticket. The codec does not start a Flight server, allocate persistent tenant processes, or authorize the gateway to read storage internals directly.

The result stream codec completes the local request/response contract for this phase. ArrowScanResult encodes schema and batches into Arrow Flight FlightData messages and decodes them back into validated Arrow batches with row counts. Produced streams carry an EHDB scan-result stream version marker on the first message and keep later message metadata empty. Decode rejects empty streams, missing markers, unsupported markers, unexpected later metadata, or malformed Arrow streams deterministically. This is the shape a future do_get implementation can return, but no network service, SQL planner, predicate pushdown, or distributed executor is introduced by the codec. Receiver-side validation can also check a returned FlightData stream against FlightInfo schema, row count, and encoded byte-count metadata before accepting decoded batches.

The pre-network FlightInfo fixture adds the discovery metadata shape. Given a ScanFlightTicket and an ArrowScanResult, EHDB can produce schema IPC bytes, a command descriptor, one endpoint ticket, ordered result metadata, total record count, and encoded byte count. A local validator accepts that fixture shape and rejects unsupported fixture metadata, missing or malformed schema IPC bytes, unordered results, negative record counts, non-positive byte counts, missing endpoint tickets, multiple endpoints, or descriptor/ticket request mismatches. Produced fixtures are also checked against the result schema, row count, encoded byte count, and expected scan ticket. This is the shape a future get_flight_info endpoint can return, but the local endpoint envelope stays pre-network: no locations, expiration, or endpoint app metadata. It still does not start a network server, allocate persistent tenant processes, or let the gateway read data directly.

LocalArrowFlightService is the in-process facade over those contracts. It provides local get_flight_info, get_schema, and do_get methods by decoding Flight scan descriptors/tickets, executing bounded local scans, encoding schema-only responses as SchemaResult, and encoding record batches as FlightData. This validates service behavior before introducing broader server runtime, ports, request metadata auth, concurrency, or logging policy. Gateway remains a gatekeeper and does not touch EHDB storage internals.

LocalArrowFlightServer is the first generated Arrow Flight service trait adapter. It accepts tonic FlightDescriptor and Ticket requests, delegates to LocalArrowFlightService, returns schema-only SchemaResult values, streams FlightData, and maps EHDB errors to deterministic gRPC statuses. Unsupported Flight methods return explicit UNIMPLEMENTED statuses. Scan descriptor requests must be EHDB command descriptors with empty paths; path descriptors and command descriptors that carry path entries fail before local scan execution. This is still not a bound network daemon: port binding, runtime lifecycle, TLS/external identity, request concurrency, access-log policy, SQL planning, predicate pushdown, and distributed execution remain future work. Gateway remains the gatekeeper, not a storage reader.

LocalArrowFlightServerConfig adds the lifecycle guardrail before a listener exists. The config records the intended bind address, maximum decode/encode message sizes, request-concurrency budget, auth policy, tenant/namespace scope policy, catalog scan grant policy, and access-log policy. It rejects zero-sized bounds and rejects unauthenticated non-loopback binds. Building a service from this config applies gRPC message limits to the generated server wrapper and passes the configured metadata policies into implemented scan calls, but still does not bind a socket, start a daemon, implement TLS/external identity, schedule requests, or emit INFO-level/high-volume access logs. Gateway remains the gatekeeper.

The local max_concurrent_requests budget is enforced with a fail-fast semaphore across implemented scan calls: get_flight_info, get_schema, and do_get. If all local request slots are occupied, the generated service returns gRPC RESOURCE_EXHAUSTED. This is a reference lifecycle guard only; it does not introduce a request queue, scheduler, distributed admission controller, or persistent tenant process.

FlightAccessLogPolicy keeps local scan access logging explicit and bounded. Disabled emits no scan access summaries. DebugOnly emits structured tracing::debug! summaries for decoded get_flight_info, get_schema, and do_get requests with only method, gRPC code, row/message counts, projection count, predicate presence, and which metadata guards were required. It intentionally excludes auth tokens, principal values, tenant/table identifiers, object paths, predicate values, and Arrow payloads so high-frequency read paths do not leak tenant data or flood INFO logs.

The reference HeaderToken auth policy is intentionally small. It validates a lowercase ASCII metadata header name, rejects binary metadata headers, requires a non-empty non-control-character token, and returns UNAUTHENTICATED for missing or mismatched scan request metadata. This is an auth-boundary contract for local harnesses and future adapters; it is not production tenant identity, ACL enforcement, TLS termination, or permission for gateway direct reads.

FlightScanScopePolicy adds an explicit tenant/namespace metadata guard for implemented scan calls. When enabled, the generated service decodes the scan command descriptor or ticket, then requires x-ehdb-tenant and x-ehdb-namespace metadata to match the decoded ScanLatestTableRequest before local scan execution. Missing scope metadata returns UNAUTHENTICATED; mismatched scope metadata returns PERMISSION_DENIED. This gives future catalog ACL checks a concrete request-scope contract without implementing an ACL engine yet.

FlightScanGrantPolicy adds the first local catalog-backed scan authorization bridge. When enabled, the generated service requires x-ehdb-principal metadata, validates it as a PrincipalId, resolves the requested table from replayed catalog state, and checks InMemoryCatalog::can_scan before scan execution. Missing or invalid principal metadata returns UNAUTHENTICATED; a principal without a matching CatalogScanGrant returns PERMISSION_DENIED. This enforces local reference ACL metadata without adding production IAM, policy composition, revocation, gateway direct reads, or a persistent tenant service.

LocalArrowFlightListener is the first bound reference harness, but it is deliberately loopback-only. It binds a configured or ephemeral loopback socket, reports the actual local address, serves the generated Flight service with configured message limits, and terminates through an explicit shutdown future. This is for local reference validation only: non-loopback exposure, production TLS/identity, request scheduling, gateway integration, SQL planning, predicate pushdown, and distributed execution remain future work.

The loopback client smoke path proves the harness over real gRPC transport. It starts the local listener, connects an Arrow Flight client, calls get_schema and get_flight_info with the existing scan command descriptor, validates returned FlightInfo against the expected scan ticket plus the schema returned by get_schema, extracts and revalidates the concrete endpoint ticket for do_get, and decodes Arrow record batches through a receiver helper that validates decoded data against the decoded get_schema schema, expected ticket, and returned FlightInfo row count, schema, and encoded byte count. Local service and server tests perform the same pairing through a single response-envelope helper for raw SchemaResult, returned FlightInfo, raw FlightData, and expected ticket; they also validate the concrete do_get ticket against the returned FlightInfo endpoint before rows are accepted. The authenticated smoke variant proves the same path with the header-token metadata policy enabled, and the scoped smoke variant proves tenant/namespace metadata over the same transport. The grant-enforced smoke variant proves replayed catalog scan grants through x-ehdb-principal metadata. It is a local-reference verification path only, not a gateway integration or platform exposure boundary.

Geo placement identifies the provider, region, and optional zone where the object currently belongs. Data-gravity shard identifies the tenant/workload/data-domain gravity center that should influence replication, write affinity, compaction locality, and read scheduling. These are storage-layer routing pointers, not permission for the gateway to touch data directly. Gateway remains the gatekeeper; read/write nodes and playbook steps perform bounded data-touch work through EHDB APIs.

PlacementPolicy makes this routing intent explicit before replication is implemented. A policy declares exactly one primary placement, a minimum copy count, and replica targets that all share the same data-gravity shard. The model rejects duplicate geo/shard targets and mixed-shard policies so future planners have a deterministic contract.

plan_replication is the local reference planner for that contract. It compares the current object and known replicas to the policy and returns copy-needed or already-satisfied actions. It does not execute copies; future bounded replicator workers can consume the plan behind EHDB APIs.

The object replica registry is the durable metadata source for those known replicas. It records object path, byte length, digest, placement, and data-gravity shard for each available copy, and rejects conflicting metadata for the same object. Registry mutations are replayable through the transaction log so local reference state can rebuild replica inventory before producing a replication plan. The registry remains metadata: actual copy execution belongs to bounded worker/playbook steps, never to gateway data-touch logic.

Storage metadata JSON decoding is strict for object refs, geo placements, placement policy targets, replica records, replication actions, and replication plans. Unknown persisted fields are rejected before replay or planning can treat them as valid routing metadata.

The local replication executor is the first bounded execution reference for that handoff. It consumes a deterministic ReplicationPlan, verifies the source object through the immutable object-store API, and appends RegisterReplica storage mutations for copy-needed targets. It does not schedule background loops, retain tenant process state, or move work into the gateway; it models the unit of work a future worker/playbook step can perform atomically.

Replay Reference Model

The transaction log must be reconstructive, not merely descriptive. A transaction record should carry enough durable facts to rebuild local reference state from replay without consulting side effects that happened before the append.

Replay-complete mutation domains:

  • Catalog create-table mutations include table schema.
  • Catalog commit-snapshot mutations include snapshot ID, parent snapshot, and object file references.
  • Stream publish mutations include subject, payload, and expected sequence.
  • Retrieval mutations include document source metadata, chunk text and checksum, and embedding vectors.
  • Local vector search is derived from replayed retrieval state; it does not add a new durable index log or sidecar service.
  • Local retrieval service search reads replayed runtime state in process; it does not add network serving or gateway data-touch behavior.
  • Local text search is derived from replayed chunk text and does not add a side index, search daemon, or external adapter.
  • Local hybrid search is a weighted score over replayed embeddings and chunk text; it does not add a planner, side index, or service process.
  • System library publish mutations include the full WASM manifest: entry export, target, object path, byte length, capabilities, revision, and digest.
  • Storage mutations include registered object replica metadata so placement inventory and replication planning can be reconstructed from log replay.

ehdb-reference applies replayed transaction records into the local catalog, stream, retrieval, and system-library catalogs. It validates derived state as it replays; for example, a stream publish whose durable sequence does not match reconstructed state fails deterministically instead of being silently repaired.

LocalReferenceRuntime is the first local end-to-end runtime boundary. It wraps LocalJsonlTransactionLog, previews the transaction record, applies it to cloned reference state, and only then performs the durable append. If projection fails, the JSONL log is not advanced. On reopen, the runtime reconstructs the same local catalog, stream, retrieval, system-library, and storage-replica state from transaction replay. The NoETL runtime surface fixture drives a worker/playbook-shaped flow only through LocalReferenceRuntime appends, then verifies catalog scan grants, stream consumer replay, retrieval lookup, system WASM library resolution, and storage replica inventory after reopen. It remains a local integration-readiness fixture, not a gateway route, direct gateway data path, persistent per-tenant service, production IAM surface, distributed executor, SQL planner, object mover, or external dependency replacement.

Local Transaction Log Reference

Before production consensus, EHDB uses a local reference transaction log to make replay and crash/restart behavior concrete. The LocalJsonlTransactionLog adapter writes one serialized TransactionRecord per line, calls sync_data after each append, and rebuilds ordered replay state on open.

The adapter enforces the same core invariants as the in-memory log:

  • Transactions must contain at least one mutation.
  • Transaction IDs are unique across process restarts.
  • Transaction sequences are contiguous and monotonic.
  • Transaction envelopes and mutation identifiers are revalidated before persisted records are accepted back into ordered replay state.
  • Unknown transaction record and mutation fields are rejected during replay instead of being silently ignored.
  • Corrupt JSONL records fail open instead of being skipped silently.

This is a local developer and test boundary, not the final distributed log. Raft/Paxos or another consensus engine should plug in behind the transaction-log boundary once the NoETL metadata and integration contracts stabilize.

Local Stream Journal Reference

Before production replicated stream storage, EHDB uses a local reference stream journal to make the NATS JetStream replacement path restartable. The LocalJsonlStreamLog adapter writes one serialized journal entry per line, calls sync_data after each state-changing operation, and rebuilds stream state on open. Persisted stream sequences are decoded through the same nonzero invariant as constructed sequences before retained records or consumer cursors are rebuilt. Unknown journal entry, stream config, stream record, and durable consumer fields are rejected instead of being silently ignored.

Journaled operations:

  • Create stream.
  • Create durable consumer.
  • Publish record.
  • Ack consumer cursor.

The adapter rebuilds retained records, durable consumer ack state, and the next stream sequence on open. Retention is replayed from the journal so restart behavior matches the in-memory stream model. Corrupt JSONL entries fail open instead of being skipped silently. Keep-all retention is supported, and bounded max-record retention must be positive before any stream is created or journaled. Published stream records use concrete Subject values that reject wildcard tokens and empty dot-delimited tokens; replay uses SubjectFilter values for exact matches, single-token * wildcards, and terminal > tail wildcards. Subject filters also reject empty tokens. Durable consumer replay can use the same filters over records pending after the consumer ack cursor without moving that cursor. JSONL replay revalidates persisted stream record subjects before rebuilding state. It also revalidates persisted stream coordinates, consumer names, and stream record transaction IDs before applying journal entries.

This is a local developer and test boundary, not the final distributed stream log. Production stream replication and any NATS compatibility bridge should plug in behind the stream boundary.

Compute Model

EHDB separates writes and reads:

  • Write nodes validate catalog mutations, commit transaction log entries, and publish immutable data files.
  • Read nodes serve snapshot-consistent scans, catalog lookups, and Arrow Flight streams.
  • Compaction and maintenance workers run as bounded jobs.
  • NoETL workers can call EHDB as atomic compute steps without becoming long-lived data services.
  • Stream materialization, retrieval indexing, and analytical maintenance run as bounded EHDB/NoETL jobs, not persistent per-tenant agent processes.

Protocols

Required early protocols:

  • Rust library API for internal crates.
  • gRPC control API for catalog and transaction operations.
  • Arrow Flight for columnar result movement.
  • Stream API for publish, subscribe, ack, replay, and cursor inspection.
  • Retrieval API for document/chunk registration, embedding metadata, and vector/text lookup.
  • System library API for publishing WASM module manifests, binding environment/channel aliases, and resolving NoETL-compatible module refs.

Candidate later protocols:

  • PostgreSQL wire compatibility for migration and tooling.
  • SQL query endpoint once the planner/executor boundary is clear.
  • Iceberg REST catalog compatibility if it does not force an external catalog dependency.
  • NATS JetStream compatibility bridge during migration only.

Tenant And Region Model

EHDB is multitenant by design.

Every durable object should carry enough identity to route and audit:

  • Tenant.
  • Namespace.
  • Region or placement scope.
  • Table/snapshot identity.
  • Data classification and residency policy.
  • Writer identity and lineage origin.
  • Stream subject and consumer identity where applicable.
  • Embedding model and retrieval policy where applicable.
  • System library path, release channel, environment, and digest where applicable.

Initial Crate Boundaries

The first Rust workspace should keep boundaries explicit:

  • ehdb-core: identifiers, error types, Arrow datatype helpers, transaction primitives.
  • ehdb-catalog: typed catalog model and in-memory reference catalog.
  • ehdb-reference: replay applier over the local reference catalogs.
  • ehdb-storage: object-store abstraction and local adapter.
  • ehdb-stream: typed stream log and durable consumer model.
  • ehdb-retrieval: RAG document/chunk/vector metadata model.
  • ehdb-system: system WASM library manifests and environment/channel bindings.
  • ehdb-transaction: transaction records, ordered replay, and local durable JSONL adapter.
  • ehdb-server: future service entry point.
  • ehdb-cli: development CLI for smoke tests and admin workflows.

Early code should prioritize type contracts and testable behavior over distributed consensus. Raft/Paxos integration belongs behind a transaction-log trait once the metadata model is stable.

Non-Goals For The First Milestone

  • Full SQL engine.
  • Production consensus implementation.
  • PostgreSQL wire protocol.
  • Full NATS JetStream compatibility.
  • Production vector ANN index.
  • Persistent retrieval daemon.
  • Multi-region replication.
  • NoETL production cutover.
  • Persistent per-tenant agent or MCP processes.

These are important later, but they should not block the first durable Rust model and local developer loop.

Clone this wiki locally