Skip to content

RFC External EHDB Driver

Kadyapam edited this page Jul 10, 2026 · 3 revisions

RFC: External EHDB Driver — outward-facing query access for third-party apps

Status: ACCEPTED — implementing (MVP) (2026-07-10). All five decision forks were resolved on 2026-07-10 (Alesha approved the recommended default on each — see Decisions resolved below). Phase 1 (MVP) is in progress. This document now records decisions being implemented, not a review gate.

Decisions resolved (2026-07-10)

# Fork Decision
1 Wire protocol Arrow Flight + Flight SQL first. Postgres-wire stays a deferred optional Phase-3 bridge.
2 Query surface Hybrid — Flight SQL over the relational projection tier; typed per-tier RPCs for event-log / KV / object / vector.
3 Endpoint placement A dedicated data-plane endpoint fronting the worker. Not the control-plane server; not the autoscaling worker pool directly.
4 Auth model Platform-session + scoped read-only API tokens (keychain/GSM-resolved). Tokens never enter the driver surface.
5 Read-only MVP Read-only, committed/materialized-only, per-request snapshot. Read-write deferred to a separate later RFC.

MVP slice: projection-tier reads over Flight SQL, served from the ehdb-service Arrow Flight scaffold, exposed via a dedicated data-plane endpoint fronting the worker, read-only + committed-only, behind a default-off flag, with a scoped read-only token seam resolved via keychain/GSM. Built on top of the #178 worker-side projection read contract (the ProjectionDriver read methods / tiers/{tier} handler) — not a forked read path. See noetl/ai-meta#184.

MVP status — built + kind-validated (2026-07-10)

Landed as three PRs (open for review, not merged): the Flight SQL read surface (noetl/ehdb#272), the dedicated worker endpoint (noetl/worker#180), and the kind rig + external client (noetl/ops#236).

Validated end-to-end in the local kind stack (LOCAL only — no GKE/prod; noetl.event untouched): an external pyarrow Flight SQL client connected to the dedicated noetl-ehdb-flightsql endpoint (Service :8092, fronting the worker — not the control-plane server) with a scoped keychain-resolved bearer token and ran SELECT * FROM executions → real projection rows with the secret-free Arrow schema. Proven along with it: SQL projection + WHERE filtering; fail-closed (endpoint refuses to start without the token alias); auth enforced (missing token → unauthenticated); read-only enforced (DELETE → only read-only SELECT statements are allowed). See the tracking issue for the full result. Tracks: noetl/ai-meta#184 (external-driver umbrella) · sibling of noetl/ai-meta#178 (internal query interface) · under noetl/ehdb#241 (completion program) / noetl/ehdb#234 (integration umbrella).

This RFC proposes an external driver / client surface that lets an application outside the NoETL platform open a connection and run read queries against EHDB's engine tiers (event-log, projection/read-model, KV, object/blob, vector). It is distinct from the internal query interface (#178), which is an operator-facing read API on the noetl server plus a noetl ehdb query CLI. Here the consumer is a third-party app, a BI tool, a data-science notebook, or another service — not a NoETL operator holding a session.

The platform boundary from the Completion-Program RFC is unchanged and load-bearing throughout:

gateway = gatekeeper
worker  = atomic compute
playbook = ephemeral blueprint
shared cache = state vehicle
event log = source of truth

The critical boundary applies here too — platform data only

EHDB serves NoETL platform data only, never business data. An external driver does not change that. What an external consumer can reach is the same set the internal interface exposes: executions, events, derived read-model state, platform KV/state, platform object/blob metadata, and platform vectors — the noetl.*-owned fabric. A tenant's Snowflake warehouse, their Postgres, their S3 bucket: those are external subsystems reached through playbook connectors and have nothing to do with EHDB or this driver.

So the external driver is not a general "query your business data in EHDB" surface — there is no business data in EHDB to query. It is an outward-facing read window onto NoETL's own operational fabric, for apps that today would have to poll the server API or scrape events.

Secrets never enter the driver surface. Like #178, responses carry only projected/secret-free columns and payloads; result / error / context / workload bodies that can carry credential material are never selectable. This is structural (the tier read-model views are already payload-free), not a post-hoc scrub, and the driver must not add a path that reintroduces raw payloads.


What already exists (do not rebuild it)

Two pieces of relevant scaffolding are already in the tree. The design below builds on them rather than starting cold.

1. #178 internal query interface — the read contract

Live in kind since 2026-07-06 (EHDB Query Interface):

  • Read-only /api/ehdb/* on the noetl server; projection/read-model served direct from the read-model store, raw data-plane tiers (eventlog | kv | object | vector) routed to the worker/system data plane behind the GET /api/ehdb/tiers/{tier} seam.
  • A control-plane guard: no EHDB data-plane engine is linked into the server binary. The server gatekeeps and read-serves the read-model; the worker owns tier storage.
  • Bounded (limit ≤ 1000 + forward after cursor), secret-free by construction, read-only.

This is the read contract the external driver should honor and, where possible, reuse — not a second, divergent read model.

2. ehdb-service crate — an Arrow Flight scaffold (loopback only)

crates/ehdb-service/src/lib.rs (~8.7k lines) already implements a first Arrow Flight (gRPC / tonic) service adapter. arrow-flight and tonic are workspace dependencies. What is present today:

  • LocalArrowFlightServer — a generated FlightService trait adapter implementing get_flight_info, get_schema, and do_get for latest-table Arrow scans; maps EhdbError → gRPC Status; streams FlightData; returns explicit UNIMPLEMENTED for non-scan methods.
  • ScanFlightTicket / ArrowScanResult — versioned scan-request ticket codec and result-stream codec (round-trips through Arrow Flight Ticket / FlightData / FlightInfo).
  • Retrieval-context (RAG) request/result payload codecs (AssembleRetrievalContextRequest, RetrievalContext, …) for a Flight do_action path.
  • Auth boundary contract: FlightAuthPolicy (DisabledForLocalReference / HeaderToken { header, token } / ExternalRequired), plus FlightScanScopePolicy (tenant / namespace scoping via x-ehdb-tenant / x-ehdb-namespace headers) and FlightScanGrantPolicy (principal / catalog-grant via x-ehdb-principal), and FlightAccessLogPolicy (bounded, DEBUG-only access summaries).
  • Bounded lifecycle: max_message_size, a fail-fast max_concurrent_requests semaphore returning gRPC RESOURCE_EXHAUSTED, and bind_loopback_listener.

What is deliberately not present (per the README's own framing — "future service surfaces"):

  • No external port bind, no TLS, no external identity federation, no ACL enforcement — HeaderToken is a test/harness contract, and unauthenticated mode is valid only for loopback.
  • No SQL planner, no predicate pushdown, no distributed executor.
  • Not linked into any deployed binary — neither the worker nor the server references ehdb-service. It compiles and is unit-tested as an in-process fixture only.

So the wire-format plumbing, the ticket/result codecs, and an auth/scope policy shape exist. The external-facing pieces — a network endpoint, real identity, the query surface beyond a single latest-table scan, and the wiring into a deployable role — do not.


Design goals

  1. Reuse the #178 read contract and the ehdb-service codecs. One read model, one secret-free discipline, one bounded-pagination story.
  2. Honor loose coupling. The external endpoint must not turn the control-plane server into an EHDB data plane, and must not pin a server replica to a storage node.
  3. Read-only by default, secret-free by construction. Read-write is out of scope for the MVP and gated behind a separate, later decision.
  4. Client ecosystem reuse where it is free. Prefer a wire choice that gives working clients (Python, JDBC/ODBC, BI tools) without us writing and maintaining a bespoke client per language.
  5. Fit EHDB's tiered model. The projection tier is relational; the event-log, KV, object, and vector tiers are not. The surface must not force the non-relational tiers through a SQL-shaped hole.

Decision 1 (fork) — wire protocol

The central fork. Options, with the trade-off that matters for EHDB:

Option Client ecosystem Fit to EHDB tiers Cost to us
A. Arrow Flight + Flight SQL (gRPC) Flight SQL has JDBC/ODBC drivers + pyarrow.flight; native columnar Best — Arrow is EHDB's native boundary; Flight SQL covers the relational projection tier, Flight do_action/do_get cover the non-relational tiers Lowest — ehdb-service already speaks Flight; extend it
B. PostgreSQL wire protocol (pgwire) Widest — every PG client / BI tool "just works" Poor for 4 of 5 tiers — forces event-log/KV/object/vector into a relational SQL shape they don't have; great only for the projection tier High — a new pgwire server + a SQL engine mapping catalog→PG system tables; duplicates Flight work
C. REST / HTTP (extend #178) Universal, but bespoke per-tier JSON; no BI-tool story OK for control-flow reads; weak for bulk columnar/vector Low-to-medium — #178 already exists; but it's operator-shaped, not a driver
D. Custom protocol None Whatever we want Highest — build + maintain clients ourselves

Recommendation: A — Arrow Flight, with Flight SQL for the projection tier.

Reasons: (1) ehdb-service already implements a Flight adapter and the scan/result codecs — this is the lowest-cost path by a wide margin; (2) Arrow is EHDB's declared native boundary (arrow-ipc / arrow-flight are already how columnar data crosses seams); (3) Flight SQL gives us the BI/JDBC/ODBC ecosystem for the relational projection tier without a separate pgwire stack, while Flight's do_get/do_action cleanly carry the non-relational tiers; (4) pyarrow.flight is a zero-custom-code Python client on day one.

Postgres-wire is not rejected outright — it is deferred as an optional bridge. If BI-tool reach over the projection tier proves to need more than Flight SQL's JDBC driver, a thin Flight-SQL→pgwire (or a standalone pgwire front for the projection tier only) can be added later in front of the same read model. It is a Phase-3 add-on, not the foundation. Building the foundation on pgwire would force the four non-relational tiers through SQL and duplicate the Flight work we already have.

→ Fork for Alesha: Flight-first (recommended) vs. Postgres-wire-first (maximize BI-tool reach at the cost of tier fit + duplicated effort).


Decision 2 (fork) — query surface / language

The tiers do not share one shape, so neither should the surface:

  • Projection / read-model tier (relational) → Flight SQL. A bounded, read-only SQL dialect over the execution/event read-model views (ExecutionStateView, EventReadModelView). This is the tier BI tools and analysts actually want, and it maps to SQL honestly.
  • Event-log tier → typed Flight do_get scan (by execution, or global-by-sequence with a cursor) — an ordered stream, not a table scan; SQL's unordered-set model fits it poorly.
  • KV tier → typed get(bucket, key) / scan(bucket, prefix) RPCs.
  • Object/blob tier → typed get / list / locate returning metadata only (id, digest, size, path) — never raw bytes that could carry secrets, and never business payloads.
  • Vector tier → typed query(collection, top_k, filter) retrieval, optionally the existing AssembleRetrievalContext Flight action.

So: SQL for the one relational tier, typed per-tier RPCs for the four non-relational tiers, all over the one Flight transport. This matches how #178 already split "projection served direct" vs "raw tiers routed".

→ Fork for Alesha: hybrid SQL+typed (recommended) vs. SQL-only (simpler mental model, but bends 4 tiers out of shape) vs. typed-only (loses the BI/JDBC ecosystem entirely).


Decision 3 (fork) — where the endpoint lives vs. #178 and the loose-coupling boundary

Three placements:

  1. On the noetl server. Rejected for the same reason #178 refuses it: the server is control-plane and the completion-program RFC bars an EHDB data plane in the server process. Bulk columnar scans and vector queries are data-plane work; putting them in the server breaks the boundary and couples edge scale to storage.
  2. A separate external data-plane endpoint fronting the worker/system pool (a new deployable role, e.g. ehdb-flight / an EHDB access gateway, co-located with the data plane that already owns tier storage). Honors the boundary: control-plane stays thin, the data-plane role serves bulk reads.
  3. Directly on the worker pool with an exposed port. Works, but couples the external surface's availability to autoscaling worker churn (KEDA scales the pool 1→N on backlog; an external endpoint wants stable addressing).

Recommendation: 2 — a dedicated external Flight endpoint that fronts the worker/system data plane, reusing #178's routing seam as its internal path.

Shape:

external app ──Flight/Flight SQL──►  ehdb external access endpoint  (NEW, data-plane role)
                                          │  (auth + scope + read-only + bounded)
                                          ├── projection/read-model ──► #178 read model (served DIRECT)
                                          └── raw tier (eventlog/kv/object/vector)
                                                   └── worker/system data plane
                                                          └── ehdb_reference tier drivers

The external endpoint reuses the same read contract as #178 (the /api/ehdb/tiers/{tier} routing seam becomes one internal consumer of the worker query handler; the external Flight endpoint becomes another). It is a separate surface (distinct auth model, distinct wire, distinct scaling), not a second read model. Reuse the contract and drivers; isolate the endpoint.

Note the worker-side query handler behind #178's 501 seam is a shared prerequisite: both the internal /api/ehdb/tiers/{tier} route and this external endpoint need it. Building it once serves both.

→ Fork for Alesha: new dedicated data-plane endpoint (recommended) vs. extend the server (violates loose coupling) vs. expose the worker directly (couples to autoscaling).


Decision 4 (fork) — auth & tenancy

External identity is the piece ehdb-service explicitly does not have yet (HeaderToken is a harness contract; ExternalRequired is a policy variant with no backing implementation). Options:

  • Reuse the platform's session auth (the gateway's Auth0/session posture). Good for first-party apps a user is already logged into; awkward for headless services and BI tools.
  • Scoped API tokens/keys, resolved through the keychain / GSM. Standard for machine clients and BI tools; issued per external consumer, carry a tenant/namespace scope + read-only grant. The token value lives in the keychain (GSM-backed), never in the driver surface or in events.

Recommendation: gateway-fronted auth for the endpoint, with scoped, read-only API tokens (keychain/GSM-resolved) as the machine-client credential.

Map external identity onto the existing ehdb-service policy shapes that are already coded: FlightAuthPolicy::ExternalRequired (require real metadata identity), FlightScanScopePolicy (tenant/namespace scoping via x-ehdb-tenant / x-ehdb-namespace), FlightScanGrantPolicy (principal/catalog-grant). What's missing and must be built: TLS, real token validation behind ExternalRequired, and per-tenant data scoping enforcement at the driver read (not just header presence).

Tenancy: every external session is scoped to a tenant/namespace; the driver never returns cross-tenant rows. Default read-only; the token grant has no write scope in the MVP.

Secrets rule (non-negotiable): the token is a keychain alias resolved at the endpoint, never surfaced back; response payloads stay secret-free by construction as above.

→ Fork for Alesha: platform-session + scoped API tokens (recommended) vs. API-tokens-only (simpler, drops first-party SSO) vs. reuse gateway-session-only (blocks headless/BI clients).


Decision 5 (fork) — read-only vs. read-write, and consistency

  • Read-only for the MVP. External write into the platform fabric (appending events, mutating KV) would cross authorship rules the completion-program RFC pins: application-event authorship stays gated through the gateway/server producer path; a data-plane external client must never fabricate events. Read-write is a separate, later RFC if it is ever wanted, and would almost certainly route through the server's authorized append path, not a direct driver write.
  • Consistency: committed / materialized reads only. An external consumer sees:
    • projection tier → the materialized read-model (committed projections only — no dirty/in-flight state);
    • event-log tier → the log up to the durable watermark (the segment-store's persisted, ordered, gapless sequence — see Durable Event-Log Backend);
    • KV/object → the latest committed version per key;
    • vector → the current committed index.
  • Snapshot semantics: per-request. Each request reads a consistent point; the MVP does not offer a pinned multi-request snapshot / cursor transaction. A forward cursor (after) gives monotonic pagination within a scan, matching #178.

Recommendation: read-only, committed/materialized-only, per-request snapshot for the MVP. Read-write and pinned snapshots are explicitly deferred.

→ Fork for Alesha: confirm read-only MVP (recommended) — or signal now that read-write is a near-term requirement so the auth/authorship design can accommodate it from the start rather than being retrofitted.


Decision 6 — client library / first client (recommendation, low-controversy)

  • Python first, for free: pyarrow.flight (and Flight SQL's ADBC / JDBC drivers) are working clients on day one for the Flight choice — no custom Python client to write or maintain. This is the single biggest argument for Flight and the fastest path to a usable MVP.
  • A thin Rust client crate (ehdb-client) wrapping the Flight endpoint with typed per-tier calls, reusing the existing ScanFlightTicket / ArrowScanResult codecs — for NoETL-internal Rust consumers and as the reference implementation.
  • JDBC/ODBC via Flight SQL for BI tools — no code, just the standard Flight SQL driver pointed at the endpoint.
  • If Decision 1 later adds the optional pgwire bridge, psql/BI clients get a second, even-more-universal path to the projection tier.

MVP client scope: pyarrow.flight verified against the endpoint + the Rust ehdb-client reference crate. No bespoke per-language clients.


Phased delivery plan

Design-first; nothing below is built until the forks are decided.

  • Phase 0 — this RFC + decisions. Lock Decisions 1–5. (No code.)
  • Phase 1 — MVP: read-only projection tier over Flight SQL. Stand up the external Flight endpoint (Decision 3) with TLS + real ExternalRequired token auth (Decision 4); expose the projection/ read-model tier via Flight SQL (Decisions 1–2); reuse #178's read model. Ship the pyarrow.flight smoke client. Kind-only, one tenant, one namespace. This is the smallest end-to-end slice that proves the wire + auth + boundary.
  • Phase 2 — non-relational tiers. Add typed Flight do_get/ do_action for event-log scan, KV get/scan, object list/locate (metadata only), and vector query — reusing the same worker-side query handler that #178's tiers/{tier} seam needs. Ship the Rust ehdb-client reference crate.
  • Phase 3 — ecosystem + scale. Optional pgwire bridge for the projection tier (if BI reach demands it beyond Flight SQL's JDBC), multi-tenant scoping enforcement hardening, sharded global-scan fan-out + k-way merge (the same one #178's /api/ehdb/events needs under multi-shard prod).
  • Phase 4 (deferred, separate RFC) — read-write. Only if a concrete need appears; routes through the authorized append path, never a direct driver write.

Explicit recommendation: approve Flight-first (Decision 1A), the hybrid surface (2), the dedicated data-plane endpoint (3), scoped read-only API tokens (4), and the read-only committed-only MVP (5); then build Phase 1 in kind.


Decisions (RESOLVED 2026-07-10 — all recommended defaults accepted)

  1. Wire protocol — ✅ Arrow Flight + Flight SQL first. (Decision 1)
  2. Query surface — ✅ Hybrid SQL-for-projection + typed-RPC-for-the-rest. (Decision 2)
  3. Endpoint placement — ✅ Dedicated data-plane Flight endpoint fronting the worker. (Decision 3)
  4. Auth model — ✅ Platform-session + scoped read-only API tokens (keychain/GSM). (Decision 4)
  5. Read-only MVP — ✅ Read-only, committed/materialized-only. Read-write deferred. (Decision 5)

Related

Clone this wiki locally