Skip to content

Sessions Log

Kadyapam edited this page Oct 10, 2026 · 185 revisions

2026-10-10 (later) — B3: the unsealed tail now leaves the node

v0.8.0 (#404) + v0.8.1 (#405).

A2/A3 made sealed history survive node loss. A part is local-only until it seals, so the window was the seal interval — B2 bounds it, cannot close it. Measured on prod while designing: SEAL_MAX_AGE=900 with oldest_unsealed_age=719 and unreplicated_oldest_age=0, i.e. nothing waiting to upload and ~12 minutes of records on one disk.

replicate_tail writes the pending records as write-once per-batch objects under tail/ (never parts/, or plan_retention could drop them), and cold_load_replicated replays them. Window: 900 s → the 15 s tick. Bounded, not zero — no remote ack gates the local append.

⭐ The RED control is the load-bearing test: with the flag off, 16 of 21 records come back, so the loss is real and the green results are falsifiable. With it on, destroying the local root and cold-loading from the replica alone returns 21 of 21 — and the same proof runs against a real GCS emulator in noetl/server.

⚠⚠ v0.8.1 fixes a defect I introduced in v0.8.0 and caught before enabling anything: replicate_tail performed the remote put while holding the engine lock, and in the server that lock is taken by shadow_append on the live write path — p50 77 ms / p99 234 ms added to any append colliding with a tick. The uploader thread had documented the rule all along ("substrate writes happen OUTSIDE the lock"). Split into prepare / upload / commit, with the no-I/O property pinned by a test whose RED control restores the defect.

2026-10-10 — the off-box replica was armed, and the acceptance proof found the defect

Two PRs: ehdb#401 (v0.6.0, replica backfill), #403 (v0.7.0, coordinator in-flight).

A2 armed on the prod server's embedded store. A genuine second failure domain — domains=[local-device:66320, remote-gcs] — and the aggregate gauges flipped: replica_set_size 1→2, survives_node_loss 0→1, single_point_of_failure 1→0. Over a 1853s window (stated deliberately: shorter than the ~900s seal period would have proved nothing) the age seal fired at t=934s and the part plus manifest reached the bucket, counted externally with gcloud.

⚠⚠ Then the acceptance proof refused to pass. Read back from the remote's own manifest: 40 parts listed, 39 at replica_count 1. Uploads are enqueued on seal and nowhere else, so arming a replica replicates nothing that already exists — and every durability gauge read healthy the whole time, because they describe the declared replica set, not the parts. survives_node_loss is a claim about configuration and was being read as a claim about data. Worse than an empty remote: the bucket held a manifest naming all 40 parts and the bytes of one, so it read as a complete store and was not a recoverable one.

backfill_under_replicated fixes it. The trap underneath was that durable_view nulls local_path, so after any restart a repair path that only knew fs::read would have compiled, run, reported success and copied nothing — the same defect one level down. Hence UploadSource::Replica. Backfills also get their own counters, because a part sealed in an earlier process has no seal→durable interval and folding it into upload_lag_micros_total would deflate a mean a capacity decision reads.

v0.7.0 came out of the worker's v0.3.2→v0.5.2 pin bump, whose only compile break was the new ShardLag.inflight. Refusing to default it to 0 was the point: render_snapshot publishes that field, and high lag with zero in-flight is specifically the stalled-consumer reading, so a hardcoded zero makes a healthy bus publish an outage signal every scrape.

2026-10-08 — the registry became discoverable, and five measurement gaps closed

Three PRs: ehdb#385 (north-star P2–P5), #386 (memory + concurrency), #387 (drain benches, state gauges, benches in CI). 1,188 tests passing, clippy -D warnings clean.

Registry. RuntimeKind + discover(kind, now, ttl) turn D8 from a worker table into a service registry; watch_since gives a resumable cursor; an Execution registers ephemerally and vanishes by not being renewed; SecretRef carries a pointer and never a value. A real bug surfaced: heartbeat reset the kind to Worker, so a renewed Execution stopped being discoverable one heartbeat after registering.

Memory is O(parts), not O(records) — exponent 0.52, ~16.6 KB + ~1.8 KB/part, so 1 M records holds ~1.7 MiB. Part count is the memory signal and merge is the control.

⚠⚠ The engine is a single writer, and a shared sequencer is not enough. Minting the sort key before taking the append lock reorders 34.2% of appends while losing and duplicating nothing — the engine-level reproduction of the mechanism that rolled back the v0.4.4 ramp (2026-09-30 entry below, noetl/ai-meta#362). It reproduces in 20 ms. My own first draft of the test asserted the sequencer was sufficient; the canary caught it.

⚠⚠ A standing replication deficit was invisible. The uploader records the successful replica writes only, so a part holding 1 of 2 copies is is_durable(), on_upload_done fires, and the durability window reports it done. New parts_under_replicated gauge; proven with a write-refusing substrate.

⚠⚠ Batching is worth 530x on the drain path, and the batch cap that fixed noetl/ai-meta#298 makes a drain above the cap quadratic (exponent 2.06). Both real; the tradeoff was never measured (ehdb#389). Separately, poll_assign without acking is 2.86x slower than with — the expiry scan walks the whole in-flight map every poll, so lazy acking is O(n²) in the consumer (ehdb#388).

Benches now run in CI with a shape guard (denominator + monotonicity), both negative controls fired. First run cross-machine: batch ordering 530.5x local vs 528.3x on the runner; per-poll cost 32.0x vs 16.4x. Gating the second would have failed CI on the runner rather than the code — which is why only the ordering is gated.

⚠ Five of my own instruments were wrong before they were right: a part count that matched the parts directory (1 part for 16,000 records), a println! whose labels did not match its arguments, a SERIES row count off by 3x from a line-oriented grep over wrapped Rust, a gauge called inert by a dataset with no dedupe_key, and a "not measured yet" list that understated the work already done.

ehdb#391 — the supersede hook, and ehdb became releasable

ehdb#392 + ehdb#393.

RED first: 2,400 records before, 2,400 after — the merge ran and collapsed nothing. Dataset::supersede_key plus a compacting merge_once took the 10x vector collection from 4,608 stored records / 65.02 ms to 500 records / 7.13 ms — 9.2x fewer records, 9.1x faster, converging to exactly the live set.

⚠⚠ Eligibility is narrower than the issue assumed. Only VectorDataset may opt in. RuntimeDataset (D8) is disqualified by watch_since — the op-log reader added in P2 — because compacting it would delete the history a watcher resumes from. One history reader disqualifies a dataset. D1EventLog is excluded by platform rule.

Mutation battery 5 of 6 caught, 1 verified benign; the keep-first mutant is caught by the tombstone test, which proves that guard fires.

ehdb is now releasable. It had no release automation at all — every v0.4.x tag was pushed by hand. semantic-release.yml + release-ehdb match the house pattern and avoid both known traps by construction: .releaserc.json omits @semantic-release/git (so GH006 cannot occur, and nothing asserts tag == Cargo.toml since every crate is 0.1.0), and the tag workflow is dispatched explicitly because a GITHUB_TOKEN-pushed tag does not trigger workflows.

⚠⚠ Releases are gated OFF behind EHDB_RELEASE_ENABLED, verified by merging it: Semantic Release ⚙️ ran on main and the job was skipped, and the tag read back is still v0.4.5. The next version computes as v0.5.0; cutting it is an owner decision because noetl/server pins both ehdb-feed and ehdb-l0 and they must move together.

⭐ The dry-run guard caught a real defect in its own first two runs: semantic-release takes its branch from env-ci (GITHUB_REF), not git HEAD, so on a pull_request event it reported refs/pull/N/merge, matched no --branches value, analysed nothing — and exited 0. The first fix (checking out the head ref) changed nothing, which is what located the mechanism.

⚠ And a tooling note: #393's first push produced zero workflow runs — no check-runs, no API entries — while every sibling PR ran. Several probes went into hunting a YAML defect that did not exist; an empty commit dispatched it immediately. Same symptom as catalog#28's "zero CI runs ever dispatched".

P6 — vectors, and a reframe

ehdb#390. top_k is an exact brute-force scan, so recall@k is 1.0 by construction — a reported 1.0 proves only that an exhaustive scan was exhaustive. It is kept as a correctness guard with its own control (complete 1.000 / damaged 0.333 / empty expected set undefined, so a known-answer set that failed to load cannot read as perfect).

⭐⭐ Query cost tracks op-log depth, not live points. 500 points re-embedded 10x (5,000 ops) costs 64.9 ms — within 2% of 5,000 points written once (66.1 ms). So re-embedding a collection is as expensive as growing it: 378x over the same 500 written once. Dimension is linear (exponent 0.90), as cosine being O(d) predicts.

⚠⚠ Compaction cannot fix it. A merge made the 10x case 9% worse (64.8 → 70.5 ms), and the code says why: live_points calls read_index_after(collection, 0) — every op ever written — then folds latest-wins; VectorDataset defines no dedupe_key; and the Dataset trait exposes no supersede or compaction hook at all. Merge combines parts, not keys, so it structurally cannot drop a superseded op. The prerequisite for vectors at scale is a key-level compaction primitive, not an ANN index — and it is broader than vectors, since RuntimeStore's D8 folds and the catalog's attribute/relation folds share the pattern (ehdb#391).

⚠ My first known-answer set was silently degenerate: 64 points over 16 dimensions, so the query vector for a high axis was the zero vector and recall@1 read 0.0 — looking like an engine defect when the fixture could not express the answer it claimed to know.

The catalog attachment needed no new dataset: (resource_type, path) maps onto (collection, point_id), pinned by set equality so cross-type leakage fails.


2026-09-30 — the v0.4.4 ramp is ROLLED BACK: a core assumption of the design is false

⚠⚠⚠ ROLLED BACK 2026-09-30 00:10Z — NOT armed in production. The 772/0 comparator reading was real but the window was too short; an hour later the same deployment read 13 divergences over 839 engagements (98.45%) and the three gates were removed. The cause is a THIRD mechanism, neither this release's staleness nor noetl/ai-meta#358's second writer: an event commits into the MIDDLE of an ORDER BY event_id read. A snowflake event_id is minted BEFORE the insert, so commit order is not id order, and the id-ordered log is not the append-only growing prefix populate_from_log assumes. Proven by position arithmetic — the store held at position 72 what the log holds at 73, one missing INTERIOR row. See noetl/ai-meta#362.

v0.4.4 itself stands: both defects it fixed are real and independently proven. It simply was not what was breaking prod.


2026-09-29 — v0.4.4: a behind snapshot is staleness, and the store went live (SUPERSEDED)

ehdb#376 → 720c1d2, tagged v0.4.4. FORMAT_VERSION stays 1, so no on-disk layout refusal.

Two defects in chain_populator, both of which made a healthy partition read as broken, and both proven by their own RED control:

  1. populate_from_log classed "the store holds more than the caller's snapshot" as Diverged. This store's only writer is populate_from_log, so a longer store means a fresher snapshot was already applied. New FromLog::StaleLog — benign, is_healthy(), partition stays servable.
  2. Coverage was not monotonic, so a stale snapshot could record a smaller total and make the partition unservable through the guarded read — fixing the classification alone would only have converted a false divergence into a false refusal.

⚠⚠ The crate deliberately does NOT decide case 1. The same shape is how a second writer looks (noetl/ai-meta#358, a measured 13.3% real divergence), and from one snapshot the two are indistinguishable. StaleLog is an unresolved report; the consumer re-reads the log and escalates when it does not catch up. My first version reasoned "sole writer, therefore always staleness" and the #358 regression test failed immediately — sound about this crate, false about the system.

The alarm is not off. A content conflict is still Diverged, including when the log is shorter, and it now also revokes serve-trust by dropping the coverage record — previously this function wrote nothing on a conflict, so a reader calling chain_if_authoritative alone was handed a chain we had just proved wrong. Events and the watermark are untouched: revoke trust, keep the evidence.

⭐ The monotonic-coverage mutant survived first. The test named for it populated 5 then read 2, which returns StaleLog early, before write_coverage is reached — it never touched the guard it was named for. Renamed and rewritten to construct the only path that reaches it.

Where it runs. Adopted by server v3.117.2 and ARMED IN PRODUCTION 2026-09-29 17:30Z: 772 comparisons, 0 failures, 100.00% agreement, up to 6 concurrent executions, 80/80 COMPLETED. An earlier ramp on v0.4.3 was rolled back at 13:28Z on the false divergence this release fixes.

31 tests in the populator suite; 45 suites, 0 FAILED; fmt + clippy clean.

2026-08-30 — F1–F5 remediation merged, all inert

Five PRs off #324, each merged without changing a byte of writer behavior:

finding PR posture
F4 lag-from-write metric #333 observability only
F3 age-based seal #334 flag, default off
F2 fencing enforcement #335 shadow — refuses nothing
F1 election + tokens #336 wired, not authoritative
F5 failure domains + 2nd substrate #337 no prod storage change

Plus noetl/worker#291 (v5.125.0) putting the window on the writer's scraped /metrics — the instrument the whole durability story reads from.

What the work turned up beyond the findings

  • ⭐ F5's second substrate exposed an unspecified trait contract the moment there was something to disagree with: get_range past an object's end errored in the filesystem impl and clamped in the new one. Resolved fail-closed — a short read would let a truncated part read as intact.
  • ⚠⚠ F3's mutation pass found that the "default off" test set the flag off explicitly, so flipping the default left everything green. It tested the wrong thing; two tests added.
  • ⚠ F2's scope was wider than the finding named — put_reclaim_watermark and delete_segment are mutations too, and a superseded writer deleting segments is strictly worse than one appending.
  • ⚠ The epoch could not ride the segment frame as the spec said: FRAME_HEADER_LEN is fixed and shared byte-identically with durable_eventlog.rs. It lives in a per-shard marker, which is what Invariant F actually needs.

19 mutations across the five, each caught by the intended guard.

Not done, deliberately: the four gates. Three of them touch the live writer path on an already-primary tier.

Sessions Log

2026-07-30 (L1 #205 + #208 PROD): 🏁 the bus overtakes NATS on latency and survives a writer restart

Headline. The #301 latency fix and the #302 writer-restart fix deployed to shastaratech prod in one rollout — engine 86a24f9, noetl-server v3.58.3, all worker pools + writers v5.81.3. Both T5 gates measured in the same window; both PASS.

Latency — the regression is not just gone, it is inverted. issued → claimed off noetl.event: p50 138.5 / p95 156.4 / p99 181.0 ms (n=60, unsaturated) against the 338 ms NATS baseline. The EHDB bus is now ~2.4× faster than the envelope it replaces, and 2–4× faster than it was at the T4 flip (285–557 ms). The loopback attribution — group commit + off-lock fsync via a duplicated fd + pipelined publish, 281 → 6.7 ms p50 at 64 publishers — holds on real infrastructure. ai-meta#205 CLOSED.

Two measurement rules came out of it. A saturating burst measures the pool, not the bus: 90 commands in 1.2 s gave p50 2123 / p99 3735 ms, all of it queueing behind the user pool's 2 replicas × 4 slots. And kind cannot resolve this at all — its worker pool peaks at 28–131 cmd/s against a ~1000 cmd/s bus ceiling, so the publish queue the fix removes never forms and run-to-run variance swamps any before/after signal. Prod, unsaturated, is the only rig that answers the question.

Writer restart — PASS. Writer pod deleted mid-stream under continuous load: 38 executions → 38 COMPLETED, 114 = 38 × 3 commands, 0 duplicates, 0 loss, redial in ~2.7 s (early eof at 18:18:22.056 → coordinator up 18:18:24.765), cursor-resume rather than shard replay. Steady-state latency unchanged across the restart. ai-meta#208 CLOSED.

⚠️ The finding that matters most: this was live and invisible. On arrival, before any deploy, user-playbook dispatch had been dead for ~2.4 days. The writer restarted 2026-07-28T07:49Z; both noetl-worker-rust pods' last log line was from that instant, then silence. noetl-worker-system-pool-shard1 happened to restart later and kept system commands completing, which masked the outage entirely. Two consequences for how this program is run: a cutover sign-off that never restarted the writer proves less than it looks like it does, and silent single-pool dispatch death is not observable today — which makes per-subject lag (#303) an observability fix, not just a scaler input.

Left open, both held for review. #304 — the resume logs from_cursor=0 origin="persisted" and ehdb_feed_shard_committed reset 408 → 165 (segment-relative, not a global offset), so the line cannot be told apart from a replay-from-0 even though the 114 = 38×3 accounting proves no replay happened. noetl/server#291 — the publish retry (3 × 250 ms) does not span a ~2.7 s pod restart; prod saw 2 × HTTP 500 during the gap (fail-closed, but not transparent).

T5 status. NATS untouched (ns nats, ns nats-supercluster). The go/no-go item is no longer latency but autoscaling: the user pool still triggers on nats-jetstream lag, and its ScaledObject is paused — which under KEDA 2.15 means the HPA is deleted and the scaler loop does not run, so the pool has had no autoscaling since 2026-07-26 (ai-meta#210). Also open: ai-meta#209 — a crash can still lose the unsealed tail, which is a loss class on a command bus and should land before T5 removes the fallback.

Monitoring correction. Prod monitoring is Google Managed Prometheus, not VictoriaMetrics — no vmservicescrape CRD, no vmstack namespace. The writer scrape is PodMonitoring/noetl-cmdbus-writer, live since the T4 cutover but uncommitted until noetl/ops#242. Because there is no in-cluster PromQL endpoint, the T5 autoscaler uses KEDA's metrics-api trigger straight against the writer's :9102.

2026-07-30 (L1 #208): 🔁 the command bus survives a writer restart — both defects fixed, a third found

Headline. Restarting only the per-shard writer left the bus not delivering, two ways — and both were recoverable only by flipping NOETL_COMMAND_BUS back to NATS, which T5 removes. Fixed off main at 03e94be (#205's group commit is the base), plus a third defect found in the server's publish path while validating.

Defect 1 — claimers never noticed. claim_next blocks in a read until a command is available, so a writer pod that goes away leaves the socket half-open: no data, no error, and every redial path hangs off an Err that never comes. Fixed with TCP keepalive on every ehdb-feed socket (~11 s detection with no FIN/RST) plus a negotiated coordinator heartbeat — one beat up front arms a 3-missed-beat read deadline for the connection, and it also catches a writer that is alive but stuck, which keepalive cannot. Heartbeats only go to a client that asked, so the wire stays backward compatible.

Defect 2 — the restarted writer replayed its whole shard (from_cursor = 0; kind lag 2738 draining at ~1/s, each stale record costing a control-plane round-trip). CursorStore persists committed_cursor() atomically beside the log; ClaimCoordinator::resume starts there; the cursor is clamped to the reopened tip (a log that lost an unsealed part reopens behind it, and an unclamped cursor no future key can exceed would take the bus permanently dark); the worker seals the log on SIGTERM. ehdb_feed_shard_committed{shard} is now on :9102 — lag reads 0 both when caught up and when a replay just finished, so the cursor is the series to watch.

Defect 3 — the server dropped one command per restart. The first publish after a restart discovers the dead socket by using it; the router was dropped for the next command but this one was lost (command.issued durable, nothing on the bus, HTTP 500). Now redialed and retried, 3 attempts 250 ms apart.

Kind, before/after in the same cluster:

Check Released worker:5.81.1 Fixed
Writer-only restart, fire, workers untouched 21 issued / 0 claimed, lag frozen 3094 for 100 s 61 issued / 60 claimed, claimers redialed in ~2 s
Lag right after the restart 3072 (whole log) 0, resumed from_cursor=3246 origin="persisted"
Writer restarted mid-burst (30 execs) — 30/30 completed, 0 dupes, lag → 0
Force-kill a worker mid-burst — 39 issued / 39 claimed / 0 dupes
Pool isolation + #166 — 0 cross-pool claims; 84 + 83 on the two system shards
Idle bus 10 min — 0 spurious redials

9 new tests in crates/ehdb-feed/tests/restart_recovery.rs + 5 in cursor.rs; the no-FIN wedge is driven deterministically through a relay that holds the socket open and stops forwarding (killing a listener sends a FIN, which a blocking read does surface, so it would not reproduce the defect). Gate green on rustc 1.97.1.

Posture. ehdb#302 + worker#196 + server#290 — all held for review, stacked so prod deploys once with #205. Kind restored to NOETL_COMMAND_BUS=nats; nothing in prod.

Follow-up. The SIGTERM seal races in-flight ingest — one command per restart was acked and still lost with the unsealed part, recovered by the orchestrator re-issuing ~30 s later (all 30 executions completed), so a latency artifact rather than loss. Clean fix: stop accepting ingest before sealing; the crash path needs L0-level recovery of local unsealed parts.

2026-07-28 (L1 #205): 🔬 command-bus dispatch latency ROOT-CAUSED + FIXED — group commit + off-lock fsync + pipelined publish

Headline. The post-T4 latency regression (issued → claimed p50 285–557 ms vs ~200 ms NATS) is not the poll-claim path and not the #203 writer-assigned key. It is posture-A fsync held inside the engine lock, multiplied by the control plane holding one publish mutex across the whole round-trip.

Attribution — new ehdb-feed/examples/dispatch_bench rebuilds the deployed topology on loopback and splits the wall time per hop:

Hop Measured
publisher-mutex wait 0 ms @1 pub → 64 ms @16 → 282 ms @64
publish RTT ~4.0 ms flat
append (fsync) ~4.0 ms flat → bus ceiling ~230 cmd/s
delivery (ack → member holds it) 16–33 µs, flat in member count

Delivery being microseconds and flat is what exonerates the claim path: the 250 ms poll interval is a cap on the tip_receiver park so ack_wait redeliveries surface, not a floor.

Fix (ehdb#301) — three parts, the first two inert without the third:

  1. FlushPolicy::CallerDriven + L0Engine::take_sync_handles — FeedWriter owns its commit points and fsyncs through a duplicated fd after releasing the engine lock, so the syscall stops stalling claimers.
  2. FeedWriter::append_batch group commit + a reader/committer split in serve_ingest; it takes what has already queued and never waits to fill a batch, so a lone record is unchanged.
  3. PipelinedPublishClient / PublishRouter::publish(&self) — without it only one record is ever in flight and nothing can be batched.

Durability, ordering (#203) and exactly-once unchanged by construction; out_of_order_appends asserted 0 against the real counter under 400 concurrent in-flight publishes.

Result — publish-submit → member holds the command: 281 ms → 6.7 ms p50 at 64 publishers / 24 claimers, flat in concurrency instead of growing, 225 → 6 983 cmd/s. Residual ~6 ms is the durable-ack fsync itself — a floor NATS JetStream does not pay per message.

Kind — correctness PASS on all dimensions (370/370 delivered, 0 unclaimed, 0 dup, pod-kill redelivery 0-loss, lag→0, 0 restarts, #166 intact). Latency is not resolvable there: per-fsync on the writer PVC is ~1 ms → bus ceiling ≈1 000 cmd/s, but the 3×4-slot worker pool on a 6-vCPU VM peaks at 28–131 cmd/s, so the bus is never the bottleneck (two identical passes differed 4× in p50). The prod re-measure is the real gate.

Two pre-existing defects found and filed (noetl/ai-meta#208), both blocking T5: workers never redial a restarted writer (blocked claim_next read, no keepalive → permanent wedge, 0/30 claimed), and a restarted writer replays its whole shard from cursor 0 (from_cursor = 0 hardcoded; observed lag 2 738 draining at ~1/s).

Adoption: server#289, worker#195 (draft until the ehdb pin can point at main). Tracking: noetl/ai-meta#205.

2026-07-27 (L1 T4): 🚀 EHDB IS THE PRODUCTION COMMAND BUS — full flip succeeded on shastaratech prod after fixing a silent feed-delivery loss

The L1 NATS takeover reached production. Commands now flow server → per-shard writer → worker over the EHDB feed on the shastaratech prod cluster; the server publishes zero commands to NATS. NATS is still installed purely as the rollback path, and T5 (delete NATS) is held on the human.

It took three rounds, and the middle one is the interesting part.

  • Round 1 (2026-07-26) — Step 0 + shadow PASS on a 2-shard cluster (0 divergence, 121 published == 121 in the feed). Held at canary: the EHDB bus partitions every command by execution_id including the shared pool, but the user pool is a single deployment claiming from a single writer, so flipping it would strand roughly half the shared commands. Rolled back clean. Filed as ai-meta#202.
  • Round 2 (2026-07-27) — chose Option 2 (collapse the command-bus axis to one shard; the two sharding axes are independent — sharding.rs:172, "correctness never depends on affinity"), kind-validated it, reconfigured prod with #166 state sharding intact, passed shadow and canary — and then the full flip failed: ~10% of commands were ingested into the writer feed and never delivered to any claimer, with feed lag = 0 throughout. Silent loss. Rolled back; prod clean.
  • Round 3 (2026-07-27) — with the fix, the same sequence succeeded: shadow 0-divergence → canary 30/30 (0 system claims by the user pool, 12/12 graceful-pod-kill redelivery) → full flip 30/30 + 50/50 soak = 80/80, 0 dup, 0 loss, feed lag → 0, writer 0 restarts, #166 state sharding intact throughout.

The bug — noetl/ehdb#300 → d4b6235

The command feed's producer (noetl-server) assigned each record's sort key — the command's snowflake event_id — and the writer trusted it. But the Dataset contract the feed cursor and range pruning depend on is "records are appended in ascending sort_key order within a partition". Snowflake ids are sparse, and under concurrent publish a lower id can reach the single writer after a higher one. When that happened:

  1. the follower cursor had already advanced past it (ChangeFeed::poll sets cursor = max sort_key read);
  2. refill reads read_partition_after(shard, cursor) → strictly > cursor → the late record never entered pending, never got delivered;
  3. lag() is cursor-relative too, so a record below the cursor was never counted — hence lag = 0 while commands vanished.

PartWriter::append had never enforced the ascending contract.

The fix moves key assignment to the writer (append_writer_assigned) — the only component that can actually guarantee the ordering the contract promises. A new out_of_order_appends counter records the condition (not yet exposed on the writer :9102 /metrics — ai-meta#206).

Kind before/after on an identical 40-concurrent burst, Option-2 topology:

delivered stuck (command.issued, no command.claimed) feed lag
pre-fix 17/40 23 permanent 0 (silent)
post-fix 41/41 0 0

Plus exactly-once 335=335 (0 dup) and pod-kill redelivery 25/25.

Releases carrying the fix: noetl-worker v5.81.1 (the writer host lives in the worker, so this is the image that carries the runtime fix) + noetl-server v3.58.1. Deployed to prod by digest after a crane copy ghcr → project Artifact Registry.

The one number that got worse

EHDB command dispatch (issued → claimed) is p50 285–557 ms (p99 ~1040, min ~145) against the ~200 ms NATS baseline — roughly 1.5–2.8×. Sub-second, and neither correctness nor completion was affected, but it is a real regression on the hot dispatch path and it is the open go/no-go item for T5: once NATS is deleted there is no cheap comparison left. Two suspects, not yet isolated — the 250 ms poll-claim interval versus NATS's push delivery (a push variant is already available via the existing FeedWriter::tip_receiver() seam), and the writer-side ordering-key append the fix introduced. Tracked in ai-meta#205.

Issues: #203 closed (prod-validated), #202 closed (Option 2 executed), #194 stays open for T5 + L2/L3. Full three-round execution log: Runbook — L1 Command-Bus Cutover. Prod IaC: noetl/ai-meta playbooks/194-l1-t4-prod-iac/.

2026-07-18 (L1): ✅ L1 NATS-takeover T0–T3 shadow arc COMPLETE — new ehdb-feed crate; T4 cutover human-gated

Built the L1 (NATS takeover) shadow delivery path on top of the finished L0 foundation, topology (c) per-shard-writer-as-broker, in a new ehdb-feed crate. Each slice its own PR, full CI gate green on 1.97.1, real proof. Additive, kind/local, NATS still authoritative, no prod.

  • T0 shadow feed (#290 Watch(shard,cursor) primitive + engine read_partition_after; #291 networked FeedWriter/serve/FeedSubscription + latency harness). Latency gate PASS: bus append→subscriber p50 57µs / p99 137µs (NATS-parity). The ~4ms first observed was diagnosed as posture-A fsync-per-append durability on the slow sandbox FS (sub-ms on NVMe; group-commit amortizes; NATS-JetStream shares it) — isolated and reported separately, not the delivery bus. Parity 0 missed / 0 spurious.
  • T1 consumer groups (#292) — ShardConsumerGroup: competing consumers, ack / ack_wait redelivery (at-least-once), crash-redelivery (0 loss), committed-cursor resume, shard routing (#166 equivalent). Clock-free deterministic coordinator.
  • T2 KEDA lag signal (#293) — per-shard backlog gauge (ShardConsumerGroup::lag) + Prometheus /metrics exposition. Ready before T4 (hard-ordering rule).
  • T3 gateway/SPA SSE feed (#294) — serve_sse text/event-stream; SSE id:/Last-Event-ID maps onto the ChangeFeed cursor, reconnect resumes 0-missed/0-dup.

T4 (command-bus cutover) + T5 (delete NATS) are human-gated and NOT executed. Prepared the reviewable cutover artifact: Runbook — L1 Command-Bus Cutover — the NOETL_COMMAND_BUS=nats|ehdb flag, the rolling-restart sequence with validation gates, the go/no-go checklist, and the one-command rollback.

ai-meta gitlink → ehdb 7e0b896 (ai-meta@5f059e75). Umbrella noetl/ai-meta#194 stays open (T4/T5 + L2/L3 ahead).

2026-07-18 (L0): ✅ L0 object-store foundation COMPLETE — predefined datasets D1–D10 all MERGED

The EHDB L0 replicated object store is finished: the engine slices (L0.1–L0.7) plus every predefined dataset D1–D10 are merged to main (0af99fc). D6–D10 landed this session, one PR each, each with the full CI gate green on rustc/clippy 1.97.1 and real proof — functional + cold-load + N-way replica-kill fallback (read_fallbacks > 0) — on the generic L0Engine<D>.

  • D6 vector/RAG (#285, eb8035e) — upsert / top-k cosine / delete.
  • D7 catalog (#286, ddf2b22) — register / get (latest+pinned) / snapshot / deregister (module registry.rs, avoids the internal meta-catalog name).
  • D8 runtime registration (#287, c7b3a82) — register / heartbeat / deregister / list-live with a wall-clock-free heartbeat-watermark eviction.
  • D9 system-WASM store (#288, 15f0077) — publish / bind (rollback) / resolve / unpublish over a (path,channel,env) triple; content-addressed bytes + op-log (D5 shape).
  • D10 provider-facts (#289, 0af99fc) — set-desired / set-observed (carry-forward fold) / drift-scan / forget; the credential-free provider_state fold behind noetl provider {plan,drift,orphans,adopt} (#189).

The shape: every mutable-state dataset is an append-only op-log over immutable parts; current state is a fold (last-op-wins per key, tombstones dropped). No in-place mutation; parts stay immutable; each dataset cold-loads

  • replica-fails-over for free. noetl-native throughout (no MinIO/S3 — noetl owns manifest/index/replication over the pluggable DurableSubstrate). kind/local only, no NATS, no prod; all NOETL_EHDB_* paths default off.

ai-meta gitlink → ehdb 0af99fc (ai-meta@4da4fae). Umbrella noetl/ai-meta#194 stays open — L1 (streaming / NATS takeover), L2, L3 are ahead, gated behind L0, and not starting without a go-ahead.

2026-07-12 (durable-backend): ✅ Alesha APPROVED "GO to prod durable-SHADOW" — sign-off recorded (no prod action)

Recorded the durability sign-off decision. Alesha approved GO to prod durable-SHADOW on 2026-07-12; the §C durability gate is signed off for the shadow stage. DOCS ONLY — no prod/GKE flag flipped by any agent; enabling durable-shadow is the prod team's explicit next operational step. Tracks ehdb#254.

  • Runbook — Durable Event-Log Prod Sign-off: header → SIGNED OFF; new §6 DECISION — APPROVED (Alesha, 2026-07-12) block (scope = durable-SHADOW only; durable-primary still separately gated); §5 refreshed for the prod team — Stage A′ target image updated stale v5.70.1 → current durable-capable release (worker merged main, carries GC + retention default-off), Stage B′ gained a recommended retention config to bound the shadow PVC (shadow has no consumer), Stage C′ gate reworded (R1 segment-GC decision now MET → set the retention policy). ehdb#254 slice-6 box CHECKED.
  • Runbook — Prod Cutover: Event-Log Tier: §C durability gate banner → RESOLVED (2026-07-12); §7 decision-log durability-gate row → checked (durable_segment, shadow scope).
  • What stays gated: Stage C durable-PRIMARY — ≥24 h clean shadow window + retention policy set on the writer pool + R2/R3 storage-class/topology, each its own explicit go.

2026-07-11 (durable-backend): prod-durability DECISION PACKET refresh — D11 GAP→MET + #261 perf re-run

Assembled the decision-ready packet for the slice-6 prod-durability sign-off (prep for @alesha's go/no-go — no prod flag flipped, decision stays with the owner). LOCAL kind only — no GKE/prod; noetl.event never purged; no secrets. Tracks ehdb#254 + ehdb#261.

  • #261 head-to-head RE-RUN on the merged-main image ehdb178184-merged (carries #264/#266/#267 + GC/retention). Deployed durable append ~740 ms → ~6 ms (~100–200×) — the O(segment) shared-publish bottleneck is gone. Layer-A authoritative (criterion): per-op-open+append flat at 5.11–5.35 ms driver / 13.9–14.5 ms full stack across S=100→10 000. Layer-B incumbent same-day: ~90 req/s p50 163 ms — not a regression, pure podman-VM contention (directional only; Layer A + the deployed re-measure are trusted). Design-Performance § #261 re-run.
  • Slice-6 sign-off refresh: D11 GAP → MET (segment GC + keep-last-N retention shipped + organic soak: shadow store 14→2 segments, held flat, reclaimed climbing, replay gapless-from-base, 0 restarts). Verdict + R1 (High Stage-C-blocker → Low operational) + R6 (perf de-risked) updated; no GAP remains. New §6 GO/NO-GO summary — recommend GO to prod durable-SHADOW; Stage-C durable-primary stays separately gated on the ≥24 h shadow window + setting the retention policy (operational) + R2/R3.
  • Open follow-ups (not blockers): prod shared/object-tier medium (R2), max-age retention deferred (mtime resets on hydrate — keep-last-N is robust), the read-write driver / prod-primary phase.

2026-07-10 (durable-backend): segment-GC limits-based retention — a shadow store self-bounds without a consumer

The optional follow-up after the GC arc closed: interest-based GC reclaims nothing without a durable consumer, so the deployed shadow event-log never self-bounds. Added keep-last-N limits-based retention that reclaims independent of consumer interest, composed with interest so it never reclaims ahead of a real consumer under primary. Tracks ehdb#254. LOCAL kind only — no GKE/prod; noetl.event never purged; no secrets. Detail: Design → Limits-based retention.

  • Merged ehdb#271 → ehdb main 49db503 (engine) + worker#178 (knob) + worker#179 → worker main 521fd10 (re-pin ehdb to the merged #271 SHA). plan_reclaim composes interest + retention into min(retention_target, interest_ceiling); the only engine change is plan_reclaim. Keep-last-N (count-based), not max-age (mtime resets on hydrate/cold-load). Config NOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS + NOETL_EHDB_EVENTLOG_SEGMENT_MAX_BYTES.
  • 251 tests (+3 retention: self-bounds a consumer-less store, never reclaims ahead of a lagging consumer, respects the min floor); clippy + fmt clean; bench variant.
  • Organic in-kind soak (v5.72.4-ehdbgcret254, user pool, shadow, no consumer, keep-3 + 1 KiB segments): the live store grew to 19 segments under appends, then the periodic task self-bounded it to 2 segments / 2,147 bytes (from ~22 MB) and held flat — eventlog_gc_ops_total{outcome="reclaimed"}=6, no error, last_degraded=0, 0 restarts; reclaim manifest committed (reclaimed_through_seq: 25331), read-only reopen replayed the retained log gapless-from-base (replay intact under organic reclamation).

2026-07-10 (durable-backend): segment-GC worker wiring + in-kind soak — D11's first deployed exercise

Wired the durable event-log segment GC to run periodically in the worker and proved it in the local kind cluster — the operational step that turns D11 from "mechanism proven" into "running in kind". Tracks ehdb#254. LOCAL kind only — no GKE/prod; noetl.event never purged; no secrets. Full detail: Design → Worker periodic invocation.

  • Worker wiring (noetl/worker, ehdb-reference pin → adca378): a periodic eventlog_gc task reclaims each owned shard (local + shared via reclaim_shard) on a spawn_blocking thread; a process-global per-shard advisory lock serializes GC against the per-op append path (resolving the two-writers-per-shard fork + closing a latent append↔append race). Spawns only when opted in on every axis (durable_segment + NOETL_EHDB_EVENTLOG_GC=consumer_ack
    • NOETL_EHDB_EVENTLOG_GC_INTERVAL_SECS>0 + data-plane contract); all default-off. noetl_ehdb_eventlog_gc_* metrics; ehdb-selfcheck durable-eventlog-gc [--drive 1] verb.
  • Fork surfaced — shadow has no consumer. No tail/ack site consumes the deployed shadow event-log, so interest-based GC correctly reclaims nothing organically (a noop, not a bug). Effective reclamation needs a durable consumer (natural under primary) or a limits-based retention mode for shadow (the follow-up).
  • In-kind soak (2026-07-10): worker v5.72.3-ehdbgc254 on the user pool (/ehdb-durable PVC, GC + interval=30s). Periodic task spawned + ran healthy (eventlog_gc_ops_total{outcome="noop"}, last_ok=1); --drive 1 on the real PVC reclaimed shared 40 → 10 at watermark 30, gc_holds=true (replay + cold-load + new-owner hydrate all coherent); 0 restarts. User + system pools rolled uniform to v5.72.3-ehdbgc254.

2026-07-09 (durable-backend): shared-tier segment GC — coherent reclamation across the shared medium

Made segment GC coherent for the shared-medium topology (the deployed shape: single-writer system pool + a shared PVC), the follow-up the same-day local-GC slice scoped. Tracks ehdb#254. Engine + CLI + bench in repos/ehdb only; repos/noetl + repos/server untouched; kind-only, no GKE/prod; noetl.event never purged; no secrets. Full detail: Design: Durable Event-Log Backend → Shared-tier reclamation.

  • The problem — local GC alone leaves reclaimed segments in the shared store, so cold_load / hydrate_owned_shard re-pull them and diverge from the owner (and a new owner un-reclaims on transfer).
  • The fix — a monotonic per-shard cross-replica reclaim watermark on the shared medium (SharedSegmentBackend::put_reclaim_watermark / reclaim_watermark / delete_segment; default-erroring so a non-capable backend refuses shared GC rather than losing coherence). SharedTierEventLog::reclaim_shard reclaims local + shared watermark-first (commit the watermark before any delete), so readers skip <= watermark and can never re-pull a segment mid-reclamation. reconcile_owned_shard realigns a crashed owner whose local segments are ahead of a committed watermark. Store refactored into plan_reclaim (decision) + apply_reclaim / reclaim_to_segment (mutation) to enable watermark-first.
  • Crash-safety — a crash after the watermark commit leaves orphan local/shared objects readers already skip: a bounded space leak the next reclaim_shard re-attempts, never a divergence. The only transient is a crashed owner briefly over-serving its own un-reclaimed segments until it reconciles — benign (shows more, never less; data also in noetl.event) + self-healing.
  • Validation — 248 tests pass (+5 shared-GC tests incl. a crash-window reconcile test proving owner+non-owner realign to a committed watermark); clippy -D warnings + fmt clean. CLI durable-eventlog-shared-gc (2 shards): shared store pruned 40 → 10 objects at watermark 30, non-owner cold-load + new-owner hydrate both coherent. eventlog_shared_tier_gc criterion group added.
  • D11 verdict: GAP → delivered + drive-tested for BOTH the single-writer local (PVC) AND the shared-medium topology — unblocks the Stage-C durable-primary consideration for the deployed shape. Remaining is operational: worker periodic-invocation wiring + an in-cluster GC soak.

2026-07-09 (durable-backend): interest-based segment GC — the R1 / D11 gap (Stage-C blocker)

Closed the segment-GC gap the prod-durability sign-off named as residual-risk R1 / D11 — the Stage-C blocker: slices 1–5 never reclaimed a sealed segment, so under KeepAll+primary the durable event-log store grew unbounded. Added interest-based segment GC (consumer-ack watermark) in ehdb-reference::durable_eventlog. Tracks ehdb#254. Engine + CLI + bench in repos/ehdb only; repos/noetl + repos/server untouched; kind-only, no GKE/prod; Postgres noetl.event never purged; no secret values. Full detail: Design: Durable Event-Log Backend → Segment GC.

  • The change (ehdb#269): a sealed segment is reclaimable once every durable consumer has acked past its last global sequence (JetStream-style interest), respecting a min_retained_segments floor. Behind NOETL_EHDB_EVENTLOG_GC=consumer_ack (off by default).
  • Invariants preserved — base offset (retained log gapless-from-reclaimed+1; reclaimed_through==0 default ⇒ byte-identical to pre-GC); durable fsync'd reclaim.json manifest = the crash-atomic commit point (crash mid-GC re-deletes below-watermark leftovers on the next open, both the replay and O(1) checkpoint-trust paths); write-forward of consumer Ack frames before unlink so replay-is-truth reconstructs consumer state. Fixed tail() to compute pending off the absolute tip, not the retained count.
  • Validation — 242 tests pass (+10 GC tests incl. crash-mid-GC-via-manifest, base-offset reopen, watermark, floor, RO refusal, fail-safe parse); clippy -D warnings + fmt clean. CLI drive (40 events, 220 B segments): 41 → 11 segments reclaimed at watermark 30, every invariant holds. eventlog_segment_gc criterion group added.
  • D11 verdict: GAP → mechanism delivered + proven for single-writer local (PVC). Unblocks the Stage-C durable-primary consideration for that topology. Shared-tier object reclamation + watermark coherence is the one remaining follow-up (else hydrate/cold_load re-pull reclaimed segments from shared); GC is local-only, off by default, and not auto-invoked, so it can't silently diverge.

2026-07-09 (perf): local replay-on-open fixed (#267) — deployed durable append ~0.5 s → ~4–16 ms/op

Fixed the dominant deployed cost the 2026-07-08 session isolated: the worker rebuilds the durable stack per op, so every mirrored append paid a fresh DurableSegmentStore::open that replayed every segment to rebuild the offset index (O(segment)). Tracked ehdb#267 (Refs #261). Engine fix in repos/ehdb only; repos/noetl + repos/server untouched; worker pin bump + rebuild; kind-only, no GKE/prod; no secret values. Full detail on Design: Performance & Load Testing → #267 fixed.

  • The fix (#268, merged f6fdaea, closes #267): a small checkpoint.json sidecar — {event_count, active_segment_id, active_len, consumers_seen, consumer_acks} — rewritten after each mutating op strictly after the frame fsync. Open-for-append loads it O(1) and skips the replay; the offset index is materialised lazily on the first read (ensure_index_loaded). Mirrors the #266 resumable-digest sidecar. scan_global / read_execution became &mut self; read-only cold-loads still replay eagerly.
  • Durability preserved — the checkpoint is a cache, never truth: replay-is- truth fallback on missing / stale / inconsistent checkpoint (strict active-segment length anchor → can never name more durable data than the segments hold); CRC integrity enforced on every read; torn-tail / fsync / single-writer / byte-exactly-once unchanged.
  • Micro-bench (authoritative): per-op-open+append flat — driver 5.11/5.21/5.35 ms, full stack 13.9/14.2/14.5 ms at S=100/2 000/10 000 (was O(segment)). 232 tests pass (+5 checkpoint tests); clippy -D warnings + fmt clean.
  • Deployed before→after (kind user pool, ~20 MB / 24 765-record store): noetl_ehdb_eventlog_last_duration_seconds ~0.509 s (ehdb266) → ~4–16 ms (ehdb267), a ~30–120× per-op reduction. First-ever op paid a one-time legacy replay + wrote the sidecar; every op after is O(1). hello_world end-to-end COMPLETED; both user + system pools rolled clean (0 restarts).
  • Merges + pins: ehdb#268 → main f6fdaea; worker #176 → 0597351 (pin a36484f→f6fdaea + Cargo.lock + 2 read-site let mut; version stays 5.70.3). Image localhost/noetl-worker:v5.72.0-ehdb267.

2026-07-08 (perf): shared-tier publish fix (#264 → #265+#266) shipped + validated in kind — bottleneck corrected to local replay-on-open (#267)

Fixed the shared-tier publish O(segment) bottleneck Layer B surfaced, validated it in kind, and corrected the Layer-B headline. Tracked ehdb#264 (Refs #261). Engine fix in repos/ehdb only; repos/noetl + repos/server untouched; worker pin bump + rebuild; kind-only, no GKE/prod; no secret values. Full detail on Design: Performance & Load Testing → Fix landed.

  • Harness #263 merged (→ ehdb main d5ed46d), making the Layer B load test reusable.
  • The fix (two PRs): #265 — publish_shard sends only the append-delta (not the whole segment) via a new SharedSegmentBackend::append_segment. #266 — the worker builds the stack per op, so #265's in-memory digest cache was empty every append (re-reading the whole prefix to re-seed); #266 persists the resumable XxHash64 state on the integrity sidecar (twox-hash serialize) so a fresh backend resumes the digest in O(delta). Guarantees preserved; whole-prefix digest unchanged. 227 ehdb-reference tests pass; micro-bench eventlog_shared_tier/append_at_size is flat ~12–13 ms at pre-fill 100 / 2 000 / 10 000 (no scaling with segment size).
  • Validated in kind — worker rebuilt on the #266 ehdb pin (localhost/noetl-worker:v5.72.0-ehdb266), user-pool + system-pool rolled to it (the two pre-wedged subscription pools couldn't roll — pre-existing broken init container — reverted to rc4). hello_world end-to-end green.
  • Corrected headline: with #266 the shared publish is provably O(delta), yet the deployed per-op mirror stayed ~0.6 s. Cause: the worker rebuilds the durable stack per op, and DurableSegmentStore::open replays every segment to rebuild the offset index — a second O(segment) cost, independent of the shared publish, that the earlier decompose conflated with it. Proven three ways: (a) fresh dir ~4 ms; (b) cost scales with local size (8 MB ~0.2 s / 16 MB ~0.65 s / 20 MB ~0.6 s) while shared is O(delta); (c) deployed ehdb-drive rate ~1.3 append/s, p99 ~1.65 s — unchanged. Filed as #267 (the dominant deployed cost).
  • Disposition: #264 (shared publish) closed — fixed + validated as O(delta); #267 (local replay-on-open) is the remaining deployed bottleneck; the deployed durable path stays shadow-only until #267 lands. All NOETL_EHDB_* flags default off.
  • Merge decision (Alesha): merge on the micro-bench flat-append evidence (the authoritative signal per methodology); don't gate on the slow in-kind rebuild. Shared-tier publish SLO now met at the engine level. ehdb #265+#266 merged (main a36484f); worker pin PR noetl/worker#175 merged (fc203b9 — ehdb pin → a36484f, pin+lockfile only). Deferred (not blocking): re-run Layer B to confirm the as-deployed rate/p99 improvement after #267 lands + when the VM isn't build-contended (add swap / build off-VM). The in-kind run this session already isolated #267 as the gate.

2026-07-08 (perf): Layer B — in-cluster load harness + EHDB-vs-incumbent head-to-head

Layer B of the perf/load-testing workstream (ehdb#261, PR ehdb#263). Ran in kind (kind-noetl, podman); no GKE/prod; no repos/noetl / repos/server / worker-runtime source changed — committed harness only. Full results on Design: Performance & Load Testing → Layer B.

  • Kind hygiene gate — NOETL_COMMANDS 0 pending on the shared/user pool; no durable orphan JetStream consumers (two random-named ones are live ephemerals with a 5 s inactive auto-reap); event-log path uniform at server v3.54.0-rc4 / worker v5.71.0-rc4; hello_world completed (7 s warm). (Pre-existing wedged subscription-pool/runtime rollout — broken init container curl: not found — is independent of the event-log path.)
  • Head-to-head (same VM, same ~400 B event envelope): incumbent Postgres+NATS append (POST /api/events) sustained ~2 150 ev/s, p50 12.6 ms / p95 33.9 ms / p99 52.5 ms, 0 fail; burst c=200 ~1 600 ev/s, p99 300 ms, 0 fail. EHDB durable local engine primitive (fresh segment, in-VM) ~3–10 ms (corroborates Layer A 3.9 ms). EHDB as-deployed shared-tier shadow mirror ~0.55–1.7 s/append (p50 ~740 ms, p99 ~1.17 s), throttling worker emission to ~1–2 append/s and building backlog.
  • Headline finding: durable_segment in the worker runtime is always the composed slice-3 shared tier; SharedTierEventLog::append calls publish_shard synchronously inside emit_event, re-reading+re-writing the whole active segment per append (O(active-segment-size)). In-VM decomposition: fresh empty segment ~3–10 ms vs live ~20 MB segment ~0.55–1.7 s — the bottleneck is the shared-publish strategy, not the durable engine or the VM fsync alone.
  • Corroborates Layer A on the load-bearing claim (local durable append is cheap + flat; reference O(n²) is real); the specific 2.7× ratio stays a Layer-A-only measurement by construction (worker never runs local-only or local_reference as a serving path).
  • SLO revisit: the single "durable append p99 ≤ 10 ms" target must be split by backend — the local primitive meets it, the shared tier (~1.2 s) cannot. Incumbent-parity would be ~50–100 ms p99.
  • Follow-up (revises Phase-1 conditional): priority = incremental shared publish (publish new/sealed bytes, not the whole active segment) + move it off the synchronous hot path; group-commit is secondary (the local fsync ceiling was not the in-cluster limiter).

2026-07-08 (perf): Phase 1 — engine micro-benchmarks + perf/load-test design doc

New perf/load-testing workstream (ehdb#261). Code + benchmarks only — no image build, no deploy, no GKE/prod; repos/noetl + repos/server untouched. Design doc: Design: Performance & Load Testing.

  • Deliverable 1 — design doc (new page). Goals + the headline question (quantify the event-log bottleneck fix: EHDB durable event-log vs the incumbent Postgres+NATS-JetStream log-and-store); the two test layers (A: engine micro-benches, reliable, this phase — B: in-cluster kind load + EHDB-vs-incumbent head-to-head, directional, next phase, podman-VM caveats); per-tier metrics; the baseline data as first data point; a proposed SLO strawman for @alesha to confirm; harness + reproducibility.
  • Deliverable 2 — benches (crates/ehdb-reference/benches/engine_micro.rs, cargo bench -p ehdb-reference --bench engine_micro). All five drivers + the durable segment backend, seeded xorshift64* payloads, isolated temp dirs, sample_size=10. No engine semantics changed; helpers bench-local.
  • Baseline (Apple M1 Max, 10 cores, 32 GiB, macOS 26.3.1, APFS SSD, rustc 1.92.0 — a dev box, not a bench rig):
    • Event-log headline — durable_segment beats local_reference at sustained append: 255 vs 96 ev/s at K=1000 (2.7×), widening with N (the reference driver replays the whole JSONL twice per append → O(n²)). Durable single-append flat at ~3.9 ms (fsync-bound → ~256/s/writer) vs local rising 9.3 → 15.9 ms with size. Segment rotation ~2% overhead (16 KiB vs 8 MiB segments). Cold replay ≈185 K ev/s (27 ms for 5000).
    • Projection fold ~1.48 K ev/s; checkpoint read 4.2 ms.
    • KV (size 1000): put 120–167/s (degrades), get 5.2 ms, scan 5.9 ms, CAS 5.2 ms.
    • Object: 1 MiB put ≈73 MiB/s (blob + SHA-256 bound), 4 KiB put 5.9 ms, get-verify/locate/list ~1.1 ms @ 200.
    • Vector (dim 1536): upsert @256 143 ms, cosine top-k=10 @128 28.9 ms / @512 115 ms — slowest tier; reference re-parses the whole float-array JSONL per op (O(catalog×dim)), not a serving index (HNSW/IVF = future).
  • Honest findings: only the event-log tier has a durable backend today; KV/object/vector reference drivers are O(n)-per-op shadow fixtures. Durable append is fsync-ceilinged at ~256/s/writer (group-commit is the lever if a hot shard needs more). Vector reference driver measured slowest, as expected for a replay-per-op fixture — reported honestly.
  • Trails: design page + _Sidebar + Home Current-Implementation-Notes + this log; ehdb#261 opened + on project 4; ai-meta MEMORY + topic file + AGENT-COORDINATION board. PR: perf/engine-micro-benchmarks → ehdb main.

2026-07-08 (kind-first): Phase 3.5 — #151 keychain PRs MERGED, event-log leak FIXED, Auth0 GREEN, OpenAI+IBKR scoped out

Merge + close-out of the #151 keychain work, with a real safety catch. No GKE/prod; no secret values printed or committed. Headline in Kind-Full-Functionality-Validation §0.

  • Leak-safety re-verification caught a real leak before merge. A fresh standalone keychain-http probe (hop-1) deferred cleanly, but a fresh two-step probe matching the ops/execution_ai_analyze openai_triage shape (http step referencing a prior step's output and a {{ keychain.* }} header) leaked Bearer sk-<key> into command.issued on rc3. So the leak was not a "heavily-rerun anomaly" — it reproduced on any hop≥2 drive-built keychain-http step.
  • Root cause: the off-server drive runs as an __orchestrate__ (tool_kind=wasm) command whose input embeds the whole playbook incl. follow-up steps' {{ keychain.* }}; the worker's generic dispatch ran inject_keychain_namespace over that input, resolving the secret into the drive's context, which the drive then persisted into the follow-up command.issued.
  • Fix (worker#174): skip keychain injection/render for the __orchestrate__ drive command; the control plane operates on deferred placeholders only, keychain resolves at the terminal user-pool dispatch (unchanged). Rebuilt noetl-worker:v5.71.0-rc4, rolled all 4 pools. Before/after: rc3 Bearer sk-<LEAKED> → rc4 Bearer {{ keychain.* }} (0 Bearer sk-); user pool still resolves at dispatch; http call succeeds. Regression test added.
  • Merged (squash): server#279 → v3.53.2 (fe500df9), worker#174 → v5.70.3 (2031a8b4, incl. leak fix), e2e#85 (252006f8), ops#235 (541e6b0d).
  • Auth0 password-grant fixture GREEN: test-user password stored in GSM (auth0-test-user-password), referenced via a provider: gcp keychain entry — never a plaintext workload input. Auth0 returned HTTP 200 + real access_token (expires_in 86400); playbook COMPLETED via success path; password absent from noetl.event. eid 333298747507216384. Verified by boolean/length only.
  • Scope: OpenAI (429 insufficient_quota, external billing) + IBKR (live gateway) formally scoped out per Alesha; ops-LLM trio + amadeus OpenAI steps scoped-out for their OpenAI-dependent portion only (keychain proven).
  • Pointer bumps: server fe500df9, worker 2031a8b4, e2e 252006f8, ops 541e6b0d, ehdb-wiki (this doc).

2026-07-08 (kind-first): Phase 3 — #151 keychain-template gap FIXED (platform) + durable GSM bridge

The last-mile pass on the kind full-functionality effort. #151 fixed as a platform change, GSM bridge made durable, and the remaining non-green reduced to external account limits (like IBKR). No GKE/prod; no secret values printed. Matrix headline in Kind-Full-Functionality-Validation §0.

  • #151 root cause (deeper than the title): {{ keychain.<alias>.<field> }} in a tool config is rendered by the drive (orchestrate-core::build_tool_command, in the system/orchestrate wasm plug-in) against a context with no keychain namespace → empty → 401.
  • Fix: orchestrate-core render_value_deferring_keychain defers keychain.* through the drive; the worker resolves it transiently at dispatch (secret never in noetl.event); the server now resolves kind: credential keychain entries. PRs server#279, worker#174, e2e#85 (fixture DSL drift: payload→json, auth0 endpoint→url, amadeus explicit token step), ops#235 (durable bridge). Kind images server:v3.54.0-rc4 / worker:v5.71.0-rc3, all 4 pools uniform.
  • Proven (real GSM): openai GET /v1/models 200, header deferred in command.issued; kind: credential probe deferred+resolved, no leak; Amadeus oauth2 token 200 + flight-offers 200.
  • Residual = external, not platform: OpenAI chat/completions 429 insufficient_quota (billing on the test key) gates ops-LLM + amadeus full-200; auth0 password grant needs real Auth0 user creds; IBKR needs a live gateway.
  • Durable GSM bridge (ops#235): fixed-clusterIP (10.96.0.53) relay + hostAliases + launchd host shim. Pod-restart proof PASSED.
  • Known follow-up (not overclaimed): heavily-rerun execution_ai_analyze openai_triage resolves the key drive-side (event-log exposure) while fresh probes defer — not root-caused; audit render_pipeline_config.
  • PRs OPEN, not merged — the drive-side leak-flag is surfaced for review; no ai-meta pointer bump this session.

2026-07-08 (kind-first): full-functionality validation Phase 2 — GSM external integration + Muno end-to-end

New capability (Alesha): kind CAN reach Google Secret Manager. Phase 2 stood up the external-integration tier for real in LOCAL kind and drove the Muno/travel planner end-to-end. No GKE/prod touched; no secret values printed; no repos/noetl/repos/server source change. Matrix on Kind-Full-Functionality-Validation (§6–§8).

  • GSM metadata bridge (deploy-time only): host-ADC token shim (gcloud auth application-default print-access-token on :48710) → in-cluster socat relay Deployment/Service gcp-metadata → hostAliases metadata.google.internal on all 4 worker pools; plus NOETL_GCP_METADATA_TOKEN_URL + GOOGLE_CLOUD_PROJECT on noetl-server. Unblocks both GSM paths (server keychain provider: gcp and the worker in-python metadata read the MCP playbooks use).
  • 9 external providers LIVE with real GSM creds: Duffel (10 offers), Google Places (10), HotelBeds hotels/activities/transfers (10/10/31), Firestore (query OK), Snowflake (rows), OpenAI + Anthropic (/v1/models 200). GCS already green (Phase-1 BUG-2).
  • Muno planner FULL end-to-end green in kind: flight-search (10 Duffel) → book (real Duffel TEST order) → hotels → activities → transfers → summary+calendar → confirm→map_view. All turns COMPLETED with live data.
  • Kafka + Pub/Sub subscription drains PASS (seed+drain+ack, count=5) after re-registering the Phase-1 SETUP-C creds. SETUP-D aliases (pg_tutorial, noetl_connection) registered.
  • Residual (not new platform bugs): #151 keychain-template gap (amadeus_ai_* / ops LLM / auth0 token complete but 401 — creds reachable, fixture uses broken {{keychain}}); IBKR needs a live gateway; 6 FIXTURE-E DSL-drift + 2 fixture-data bugs. GKE stays gated on Alesha's scope call.
  • Baseline unchanged: server v3.54.0-rc2, workers v5.71.0-rc1, EHDB shadow + durable_segment (user pool), NOETL_STATE_BUILDER=offserver.

2026-07-07 (kind-first): full-functionality validation Phase 1 — inventory, baseline, first run

Alesha directive: complete full functionality in LOCAL kind first — test all playbooks — before GKE. Phase 1 = inventory + baseline + first run. LOCAL kind only; no GKE/prod action. New page Kind-Full-Functionality-Validation is the living test matrix that scopes "done in kind".

  • Catalog: 268 playbook/automation files inventoried across e2e + noetl + ops + travel, categorised self-contained / in-cluster-infra / cred-blocked / infra-deploy / GKE-only.
  • Baseline standardised: all 4 worker pools → noetl-worker:v5.70.0 (was skewed v5.70.0 user / v5.69.0 others), server v3.53.0, EHDB tiers shadow, durable_segment event-log backend on the user pool PVC (byte- identical, non-primary). Reset cleared a system-pool orchestrate backlog (~720 stale __orchestrate__ re-drives draining ~0.6/s) + leaked NATS consumers (drifttest 23 992 msgs; 8 orphaned noetl_events consumers) that were starving new executions. Post-reset hello_world = 4 s. Postgres noetl.event never touched.
  • Runs: core self-contained gate 61 PASS / 4 FAIL (65); extended in-cluster-infra 9 PASS / 11 non-green (20); EHDB probes 2/2 (pft_sql_probe_v2, large_tabular_result_test). ~70 distinct green. EHDB projection/query/durable live; KV/object/vector shadow (segments 26.9→32.3 MB across runs).
  • Two platform bugs (each = 1 root cause, multi-fixture): BUG-1 pagination max_iterations/retry continuation wedge; BUG-2 large-result output_select/storage-tier resolve → /api/result/resolve 404 (3 fixtures). Rest non-green = cred setup (kafka/pubsub decryption, 2 missing aliases) + 6 fixture DSL-drift. Core execution model is healthy in kind.
  • ehdb-wiki: new validation page + _Sidebar + this log. repos/noetl + repos/server untouched; no prod/GKE.

2026-07-07 (sign-off): durable event-log backend — prod-durability sign-off PACKAGE (ehdb#254 slice 6)

Assembled the slice-6 prod-durability sign-off package — decision-ready evidence for clearing the tier-1 prod-cutover §C durability gate. DOCS/PLANNING ONLY — no prod/GKE action, no flag flip, no cutover; ehdb#254 slice-6 box left UNCHECKED (that's for after the actual prod sign-off). repos/noetl + repos/server untouched; no Rust/infra code change.

New page Runbook-Durable-EventLog-Prod-Signoff:

  • Go/no-go checklist (D1–D11): D1/D2/D6/D8/D10 MET; D3/D4/D5/D7/D9 mechanism-MET but NEEDS-PROD-VERIFICATION (real GKE PVC / topology / GMP, which the prod durable-shadow soak closes); D11 = GAP (no segment GC). Verdict: GO to prod durable-SHADOW; Stage C durable-PRIMARY gated on the soak passing + the segment-GC decision.
  • Evidence bundle — slices 1–4 PRs/commits + tests; the slice-5 kind soak (rotation seg-0001 8 387 133 B; metrics past 11 731 zero invalid/degraded; crash recovery sha256 c2ba3d33… byte-identical, replay 11 856, gapless).
  • Residual-risk register (R1 segment-GC = Stage-C blocker; R2 GKE storage class; R3 multi-replica affinity on the real event-writer pool; R4 disk pressure; R5 backup/DR; R6 perf; R7 selfcheck harness artifact) each with a prod-verification step / mitigation.
  • Extended durable rollout sequence (A′ deploy v5.70.1 flags-off → B′ prod durable-shadow on a PVC → C′ durable-primary, separately gated) with metrics/ alerts (PVC free-bytes alert since no GC) and per-step rollback — extends the tier-1 runbook rather than duplicating it.
  • Alignment deltas vs the tier-1 runbook: Stage A target image v5.66.0 → v5.70.1; new NOETL_EHDB_EVENTLOG_BACKEND=durable_segment + NOETL_EHDB_EVENTLOG_DURABLE_DIR (PVC) env axis; emptyDir → real PVC.

Cross-linked from the tier-1 runbook (§C update + Related), the design page, and _Sidebar. ehdb#254 slice-6 package comment posted (box unchecked). The gated action awaiting Alesha's go: Stage A′ + B′ (prod durable-shadow); Stage C′ NOT authorised by this sign-off.

2026-07-07 (soak): durable event-log backend — kind soak + real-pod-restart crash recovery (ehdb#254 slice 5)

Ran the durable durable_segment event-log backend under sustained real traffic in the local kind-noetl cluster and proved crash recovery on a real pod restart — LOCAL kind only, no GKE/prod, repos/noetl + repos/server untouched. Clears slice 5 of the durable-backend program (ehdb#254); slice 6 (prod §C sign-off) remains.

Deploy. Built localhost/noetl-worker:v5.70.0 (slice-4 durable wiring) — CARGO_BUILD_JOBS=2 temp Dockerfile guard (0 B swap on the 6 CPU/20 GB podman VM; reverted after), podman save + kind load image-archive. Rolled onto the noetl-worker-rust user pool (role worker, single-owner shard 0) with NOETL_EHDB_EVENTLOG_BACKEND=durable_segment + NOETL_EHDB_EVENTLOG_DURABLE_DIR=/ehdb-durable on a 2 Gi RWO PVC (ehdb-durable-soak), atop the existing EHDB shadow env.

Proven. (1) Real drive events land in CRC-framed /ehdb-durable/local/shard-0000/seg-*.eslog; slice-3 shared tier publishes each byte-identical to /ehdb-durable/shared/. (2) Segment rotation fired in cluster — seg-0001 sealed at 8 387 133 B (< 8 MiB) → seg-0002 opened. No JSONL written. (3) noetl_ehdb_eventlog_ops_total{outcome="mirrored"} advanced past 11 731, last_ok=1, last_degraded=0, zero invalid/degraded/ routed_away; 0 restarts/crashloops. (4) Crash recovery on kubectl delete pod --force → fresh pod on the same PVC: sealed segment byte-identical (sha256 c2ba3d33…), active tail replayed + continued (gapless sequence …11852,11853,11855), and ehdb-selfcheck durable-eventlog reopened the store read-only replaying durable_replay_records: 11856. Backend byte-identical between the soaked v5.70.0 and the concurrently-advanced v5.70.1 pointer (the ehdb#172 diff is KV/vector-only).

Trails: Design-Durable-EventLog-Backend (Kind-soak slice-5 section) + ehdb#254 slice-5 checkbox. Left running in SHADOW on v5.70.0 durable_segment for observation; rollback = kubectl set image … worker=localhost/noetl-worker:v5.69.0

  • kubectl set env … NOETL_EHDB_EVENTLOG_BACKEND-.

2026-07-08 (fix): KV + vector registry subjects hardened with SHA-256 digest tokens (ehdb#259 / worker#172)

Closed the last two tiers carrying the subject-length overflow that #256 fixed in the object tier. ehdb-reference built the KV subject noetl.kv.<bucket>.<hex(key)> and the vector subject noetl.vec.<hex(collection)>.<hex(point_id)> by hex-encoding the full ids (2 chars/byte); ehdb-stream::Subject::new caps a subject at 256 chars, so a long id overflows → InvalidIdentifier. Reproduced (math + in-code test): a ~150-byte KV key → 335-char subject; a real RAG collection (102 B) + point (117 B) → 449-char subject — both > 256. Not broken in live use only because the one live KV path uses short circuit.<id> keys and there's no live worker vector-upsert site.

Fix (ehdb#259, merged 52120a7): address each subject by a fixed-width SHA-256 digest token — noetl.kv.<bucket>.<sha256hex(key)> (bucket stays literal so the bucket scan is unchanged) and noetl.vec.<sha256hex(collection)>.<sha256hex(point_id)> (collection filter digests the collection identically). Full ids stay in the record payload; per-id reads filter the replay to the exact id and the vector query fold guards on the exact collection (digest-collision safety). Tracking issue ehdb#260 opened + closed (symmetric with the object #257 trail).

Validation (LOCAL only): cargo test -p ehdb-reference 220 passed (new long-key/long-id round-trips + in-code reproduction that the OLD hex subject is rejected by the cap); clippy + fmt clean; selfcheck kv-primary-serve + vector-primary-serve seeds switched to the real long coordinate form — both exit 0, served_by_ehdb=true, dual_run_holds=true; object regression clean. No worker image build. Worker ehdb-reference pin cca0d0d → 52120a7 (noetl/worker#172, pin+lockfile only, cargo check clean) — in-kind live re-proof deferred (no live long-key path exercises it yet). All NOETL_EHDB_* flags default off; prod worker v5.52.0 / server v3.52.0 untouched; repos/noetl + repos/server untouched. Refs noetl/ai-meta#241.

2026-07-07 (deploy + live-drive PROOF): worker v5.69.0 in kind — all four wired tiers PROVEN on live drives (OBJECT fix invalid→mirrored, PROJECTION, event-log, KV via spool-circuit runtime)

Built worker v5.69.0 (release 96a8b6b, pins ehdb-reference bbc5047) as a native-arm64 VM-safe image (CARGO_BUILD_JOBS=2), loaded it into the kind node (podman save + kind load image-archive), and rolled both data-plane pools — noetl-worker-system-pool (system role) and noetl-worker-rust (worker role) — from v5.68.0 to v5.69.0, keeping the all-5-tier shadow env (NOETL_EHDB_ENABLED=1, {EVENTLOG,PROJECTION,KV,OBJECT,VECTOR}=shadow, LOCAL_REFERENCE_LOG=/tmp/ehdb-shadow-ref.jsonl). LOCAL/kind only, prod untouched (repos/noetl + repos/server untouched; config/env only).

Baseline (v5.68.0, the bug): noetl_ehdb_object_ops_total{operation="mirror", outcome="invalid"} 2 on the system pool — the pre-fix state where ehdb-reference hex-encoded the full object key into the NATS subject and Subject::new's 256-char cap rejected ~150 B platform keys.

Real drive: registered + exec'd tests/large_tabular_result_test distributed (exec 333041450319089664, COMPLETED, 13 events) — a >inline-budget {columns,rows} result that the worker externalizes as an Arrow Feather object through the ControlPlaneClient::object_put chokepoint (real state-shard / result-tier platform keys), while normal event-log + off-server state-builder projection drain run.

PROVEN on v5.69.0 (both data-plane pools, fresh post-roll counters):

  • OBJECT — fix re-proven. object_ops_total{operation="mirror", outcome="mirrored"} = 2 (system) / 1 (user); zero outcome="invalid". The ref-store shows the real object at subject noetl.obj.055f20bc…9404 = noetl.obj.<sha256hex>, 74 chars (under the 256 cap). Clean flip from the invalid baseline — this is the ehdb#256 bbc5047 subject-length fix landing on real keys.
  • PROJECTION — live cadence hook proven. projection_ops_total{operation="materialize",outcome="materialized"} = 4 on both pools via the windowed state_builder::run_drain_loop → projection::mirror_live_window hook; parity held, no false key-divergence, no invalid.
  • EVENT-LOG — regression holds. eventlog_ops_total{operation="mirror", outcome="mirrored"} = 6 on both pools.
  • Normal behavior unaffected: drive COMPLETED, server keeps serving (/health ok), 0 pod restarts, shadow never serves (no serve/primary outcomes). The mirrored exec surfaces end-to-end via noetl ehdb query executions (server v3.53.0).

KV — PROVEN on the live circuit path (2026-07-07, later same day). The subscription pool is a plain command-consumer (no WORKER_MODE=subscription), so it never builds a SpoolRuntime — the live KV mirror lives in a dedicated subscription runtime. Stood one up on worker v5.69.0 + shadow env (noetl-subscription-runtime, NOETL_SUBSCRIPTION_PATH=subscriptions/spool_outage_stream) after registering the nats_e2e credential + sub_ingest_default playbook + the spool_outage_stream kind:Subscription (spool + one http-probed warehouse downstream), creating the SPOOL_OUTAGE NATS stream + spool-drain consumer, and deploying spool-downstream-echo. The runtime activated (subscription_id 333045155747598336, "spool runtime active", downstreams=1), and SpoolRuntime::persist_circuit → kv::mirror_live_put fires every ~2 s probe tick + on each circuit state change: noetl_ehdb_kv_ops_total{operation="mirror",outcome="mirrored"} climbed monotonically 10 → 50+ (kv_last_ok=1, degraded=0, zero invalid). An induced downstream outage (scale echo → 0) tripped subscription.circuit.opened (trips=1, "buffering to spool") and advanced KV further; recovery (echo → 1) closed the circuit (another persist_circuit). The ref-store holds the real circuit record at subject noetl.kv.noetl_subscription_circuit.<hex>, hex-decoding to circuit.333045155747598336 — a short key (~90-char subject, well under the 256 cap; the KV live path never hits the object subject-length trap). The runtime also mirrors event-log (mirrored 2 — its own subscription lifecycle events).

All four wired tiers now proven on live drives (event-log, object, projection, KV); vector stays deferred (no live upsert site). Whole stack left on v5.69.0 shadow (3 data-plane pools + the subscription runtime + echo + SPOOL_OUTAGE stream) so all four tiers keep mirroring live. Rollback: pools → set image …=v5.68.0 (subscription pool → p1-refs + drop EHDB env); delete the noetl-subscription-runtime + spool-downstream-echo deployments + the SPOOL_OUTAGE stream to remove the KV harness.

2026-07-07 (later still): projection runtime mirror LIVE-WIRED (windowed drain hook); vector mirror ready-but-unreachable (worker v5.69.0)

Wired the two remaining EHDB tier mirrors into the live worker runtime, each with its correct per-tier seam (worker#170, merged fa64e0a → release v5.69.0 96a8b6b; refs ehdb#234, ehdb#241).

projection — live-wired via a windowed cadence hook (NOT per-event). projection::shadow_project is a batch fold: it reads back the WHOLE projection log (list_executions) and compares against a full authoritative fold + offset. A naive per-event hook against the long-lived, unboundedly-accumulating KeepAll store would report persistent false key-divergence (the store keeps every execution; the worker index evicts terminals it keeps) — a hook that fires but lies. The faithful seam shipped: the off-server state-builder drain (state_builder::run_drain_loop) buffers each drained batch's real events and, at the natural post-batch checkpoint, calls projection::mirror_live_window once per batch (on a spawn_blocking thread so sync engine I/O never stalls the drain reactor). Each call windows the batch into a fresh throwaway per-window store (unique temp log) so the read-back sees exactly the window's executions — no cross-window accumulation ⇒ no false divergence — and parity-checks the EHDB engine's fold against an independent worker-side fold (fold_window_authoritative). Bounded + stateless. Advances noetl_ehdb_projection_* on real drives.

vector — ready hook, documented-unreachable. There is no platform vector-upsert in the worker's live loop today: RAG retrieval is read-only; RAG ingest writes a lexical fabric (RagChunk = text + checksum, not an embedding vector); vector::mirror_upsert is exercised only by ehdb-selfcheck. The ready-but-unwired vector::mirror_live_upsert + runtime_hook_env pair (tested to the same discipline) is provided so the future wire-up is one line, but is deliberately not invoked — fabricating a call site would create a hook that never mirrors a real upsert. Precise remaining seam: a future platform-RAG embed+upsert write site (executor/command.rs) calls it after the authoritative Qdrant upsert.

  • Discipline (both tiers): armed only when NOETL_EHDB_ENABLED + NOETL_EHDB_<TIER>=shadow + data-plane role on the bounded local_reference runtime; strict no-op off/disabled/primary/control-plane; errors panic-isolated (best-effort, metrics only, authoritative path unaffected); secret-free.
  • Tests: per tier — arms on shadow-enabled data-plane / no-op disabled+off / skipped control-plane / error-isolated; projection also proves the windowed fold produces no false divergence across repeated windows (18 new; all 178 ehdb + 32 state_builder unit tests green). clippy -D warnings + fmt clean.
  • Safety: LOCAL/KIND-only; no GKE, no prod, no image build inline. Prod stays worker v5.52.0, all NOETL_EHDB_* flags default off. object.rs subject-fix (v5.68.1) + durable segment-store untouched (sibling-owned). In-kind live-drive re-proof (redeploy + noetl_ehdb_projection_* advancing on a real drive) is PENDING redeploy.

2026-07-07 (later): object subject-length defect FIXED — SHA-256 digest token (ehdb#256 → worker v5.68.1)

Fixed the object-tier defect the v5.68.0 live-drive proof surfaced (below): the object mirror rejected every real platform object key. Root cause was in ehdb-reference::object — the registry subject hex-encoded the full logical key (noetl.obj.<hex(key)>, 2 chars/byte), and ehdb-stream::Subject::new caps a subject at 256 chars, so any key past ~123 bytes overflowed → InvalidIdentifier. Real ~140-150-byte coordinate keys (state-shard #166 / result-tier #104) all failed; state_materializer's 2 real object_puts were rejected outcome=invalid by the correctly-armed worker hook.

Fix (ehdb#256, merged bbc5047; closes ehdb#257): the subject is now a fixed-width SHA-256 digest token (noetl.obj.<sha256hex(key)>, a constant 74 chars). The full key already lives in the record payload (ObjectEnvelope.key), so get/list/locate resolve the real key without ever reversing the subject; latest_envelope filters the replay to the exact key for digest-collision safety. Content-addressed blob path unchanged.

  • Tests. New long_platform_key_subject_is_bounded_and_round_trips (140-byte real key, was rejected before; asserts subject == 74 chars < 256 and full put/get/locate/list/delete round-trip) + distinct_keys_get_distinct_bounded_subjects. 202 crate tests green; clippy -D warnings + fmt clean; selfcheck object-primary-serve seed keys switched to the real long coordinate form → served_by_ehdb: true, exit 0.
  • Other tiers. KV (noetl.kv.<bucket>.<hex(key)>) carries the same latent pattern but its only live path uses short circuit.<id> keys → unaffected; vector / projection / eventlog left untouched (short execution ids, or owned by the concurrent durable/mirror sibling). Follow-up: same digest fix for KV/vector.
  • Worker. ehdb-reference pin bumped 4c0df81 → bbc5047 (worker#169); semantic-release cut v5.68.1. Pin + lockfile only, no worker code change.
  • In-kind object-mirror re-proof PENDING — a follow-up worker redeploy (like the v5.68.0 one) will prove the object mirror now persists real keys.

Refs ai-meta#241. LOCAL/kind only — no image build, no prod/GKE this session.

2026-07-07 (evening): v5.68.0 live-drive proof — eventlog PASS, object BLOCKED by 256-char subject cap, KV unexercised

Redeployed the kind data-plane worker pools (noetl-worker-rust = user, noetl-worker-system-pool = system) from v5.67.0 to v5.68.0 (image localhost/noetl-worker:v5.68.0, sha f39a724abf8b, native arm64; built with podman then podman save → kind load image-archive — the podman-provider-safe load path), keeping the full 5-tier shadow env unchanged (NOETL_EHDB_ENABLED=1, {EVENTLOG,PROJECTION,KV,OBJECT,VECTOR}=shadow, LOCAL_REFERENCE_LOG=/tmp/ehdb-shadow-ref.jsonl). 0 restarts. Reversible to v5.67.0 (prior image + saved deploy specs).

Drove automation/pft_sql_probe_v2 via POST /api/execute → execution 332960196592668672, COMPLETED, 13 events, 0 failed. Fresh-pod baseline: all mirror counters empty.

  • eventlog — PASS (regression). noetl_ehdb_eventlog_ops_total{operation= "mirror",outcome="mirrored"} advanced 0→6 on both pools; last_ok=1, last_degraded=0. Ref-store holds 6 entries for the exec (subject noetl.event.exec.332960196592668672, stream noetl_event_log). Also surfaced through the read-only query API (GET /api/ehdb/executions → the exec, COMPLETED, event_count 13).
  • object — HOOK FIRES LIVE but every real put is REJECTED (blocking defect). The state_materializer (system pool, NOETL_STATE_SHARD_WRITE=true) did 2 real object_puts for this exec (.../state/open.feather 146 B, .../state/sealed.feather 148 B; authoritative writes succeeded, errors=0). The armed mirror hook intercepted both → noetl_ehdb_object_ops_total{operation="mirror",outcome="invalid"} 2, last_ok=0 — 0 mirrored, 2 invalid, 0 in the ref-store. Root cause: ehdb-reference object::key_subject hex-encodes the full key into the registry subject noetl.obj.<hex(key)>, and ehdb-stream::Subject::new caps subject length at 256 chars. A real platform object key is ~140-150 B → hex ~290-300 + noetl.obj. → subject ~300-310 > 256 → InvalidIdentifier. Any object key > 123 bytes fails. Both state_materializer and result_producer_stage build keys via coords.physical_key(...) in the long noetl/env=…/execution=…/state|result/… shape, so the object mirror cannot persist any real platform object. Needs a fix in ehdb-reference (hash the key into the subject instead of hex-encoding it, or raise the cap) before the object tier can mirror real drives.
  • KV — NOT exercised by this drive shape. The KV mirror lives in SpoolRuntime::persist_circuit (bucket noetl_subscription_circuit), which runs only in the subscription runtime. noetl-worker-rust-subscription-pool is still on p1-refs with no EHDB env — outside the system+user roll — and a pft /api/execute drive doesn't touch the subscription circuit, so nothing mirrored. Predicted to mirror cleanly once exercised: the circuit key is circuit.<subscription_id> (~40 B → subject ~115 chars < 256). Proving KV live requires rolling the subscription pool to v5.68.0 + shadow env (a bigger version jump with login-path risk; deferred).

Net: eventlog live-mirror proven again (v5.68.0 regression); object hook is in the live path but blocked by the reference engine's subject-length cap on real keys; KV still unproven on a real drive; projection/vector remain deferred. Cluster left on v5.68.0 shadow. LOCAL kind only; no GKE/prod; repos/noetl + repos/server untouched (config/env only). Refs ehdb#234, ehdb#241.

2026-07-07: KV + object runtime mirrors wired into the live worker path (v5.68.0)

Extended the event-log live-append hook (v5.67.0, worker#167) to two more tiers, replicating its env-armed / guarded / error-isolated shape. All shadow-only, best-effort, disabled-by-default, reversible. LOCAL/KIND-only — no GKE, no prod, no image build inline.

  • KV — wired live. kv::runtime_hook_env + kv::mirror_live_put, armed once in SpoolRuntime::build, invoked in persist_circuit after the authoritative NATS-KV circuit-state put (bucket noetl_subscription_circuit — the only live platform NATS-KV write in the worker).
  • object — wired live. object::runtime_hook_env + object::mirror_live_put, armed once in ControlPlaneClient::new, invoked in object_put after the server durably accepts the object. object_put is the single chokepoint every platform object tier (result-tier, state-shard, plugin intent) funnels through, so this mirrors ALL live object puts with digest parity (bytes cloned only when armed).
  • projection — deferred (documented in src/ehdb/mod.rs). shadow_project is a batch materialize that reads back the whole accumulating projection log and compares against a full authoritative fold; a per-event hook would report persistent false key-divergence. Faithful seam = bounded windowed batch drive feeding the incumbent fold — larger than a call-site hook.
  • vector — deferred. No live platform vector-upsert exists in the worker loop today (platform RAG is in-process, read-only; vector::mirror_upsert is exercised only by ehdb-selfcheck). A hook now would never fire.

Each hook arms only when NOETL_EHDB_ENABLED + NOETL_EHDB_<TIER>=shadow + a data-plane role (worker/playbook/system) on the bounded local_reference runtime; strict no-op when off/disabled/primary or a control-plane role; engine errors panic-caught → Unavailable, metered, never propagated; secret-free noetl_ehdb_{kv,object}_* metrics advance on real ops.

Delivery: worker#168 merged (squash 2ec2e2b) → v5.68.0 (release 3927bdf). 16 new unit tests (arms / no-op-off / no-op-tier-off-primary / skip-control-plane / error-isolated per tier); all 158 ehdb unit tests green; cargo check + clippy clean (no new lints); ehdb-selfcheck still builds.

Live-drive proof — PENDING (redeploy). The live drive pool is on v5.67.0; v5.68.0 carries these hooks. Follow-up: redeploy the worker image and drive a real playbook, confirming noetl_ehdb_kv_* (on a circuit persist) and noetl_ehdb_object_* (on a state-shard / result-tier put) shadow metrics advance. Refs ehdb#234, ehdb#241.

2026-07-06: EHDB program docs/hygiene consolidation (Roadmap + Home + MEMORY + issues to ground truth)

Documentation-only reconciliation across the whole EHDB program — no code, no deploy, no prod/GKE, live kind stack (server v3.53.0 + worker v5.67.0 shadow) left running. Brought the tracking surfaces to one coherent current state:

  • Roadmap — Phase 6 status now states the shadow mirror is wired into the LIVE ControlPlaneClient::emit_event path and proven on real in-kind drives (v5.67.0, worker#167 d310c7b); added the runtime-mirror "Landed" bullet with the two proven drives (332760742153424896 + 332760854506246144) and the eventlog-only-so-far caveat.
  • Home — Phase 6 note reflects the live/proven mirror; added a Live in kind bullet (runtime mirror + query interface end-to-end via server v3.53.0 + durable-backend slice 1 ehdb#253) and an explicit What's NOT done yet bullet (prod cutovers gated, durable later slices, other-tier runtime mirrors, raw-tier query seam).
  • MEMORY.md (ai-meta agent memory) — compacted the EHDB index block into a program-status header + coherent bullets, flipped the resolved "v5.66 mirror not wired → FIXED" gotcha to past tense, all 26 topic-file pointers + PR/SHA references preserved.
  • Issues — posted the state-of-program summary on the completion umbrella ehdb#241; cross-linked ehdb#234 (dev umbrella), ai-meta#178 (query interface), and ehdb#254 (durable backend). No issue closed — the program continues (prod cutovers + durable later slices remain).

Ground-truth cross-check confirmed against git/GitHub: ehdb HEAD 99c4570 (Phases 6-10 + durable slice 1 merged), ehdb-wiki eef435c, prod still worker v5.52.0 / server v3.52.0 with all NOETL_EHDB_* flags off.

2026-07-06: v5.67.0 event-log mirror PROVEN firing on a real in-kind drive (VM resized 6/20)

Built worker v5.67.0 (release ae4164d, PR #167 d310c7b) and ran it live in the kind stack in SHADOW on both data-plane pools, then drove a real playbook and confirmed the event-log mirror fires on real events — the runtime evidence the code-only entry below was missing.

Environment: podman machine noetl-dev resized to 6 CPU / 20 GB (was 16 GB) and restarted; recovered a transient post-boot ssh/socket-gateway stall (podman socket + noetl-control-plane came up on their own within ~2 min). kind-noetl context, podman provider — no GKE, no prod touched.

Deploy (LOCAL kind only): localhost/noetl-worker:v5.67.0 (built native arm64, CARGO_BUILD_JOBS=2, podman save + kind load image-archive) rolled onto noetl-worker-rust (role=worker) + noetl-worker-system-pool (role=system), keeping the existing EHDB shadow env (NOETL_EHDB_ENABLED=1, all 5 tiers =shadow, LOCAL_REFERENCE_LOG=/tmp/ehdb-shadow-ref.jsonl). Server/gateway stayed control-plane (auth-sync-167, v3.49.0). No tier set to primary. Reversible: prior image v5.66.0-p10, NOETL_EHDB_ENABLED- ⇒ pure incumbent.

Real drive: automation/pft_sql_probe_v2 (postgres SELECT 1 AS ok → python end) via POST /api/execute → execution 332760742153424896, status COMPLETED, 0 failed steps.

Mirror fired on the real drive (the whole point):

  • noetl_ehdb_eventlog_ops_total{operation="mirror",outcome="mirrored"} advanced across the drive — worker-rust 6 → 12 (the drive's exactly 6 events), system-pool 254 → 288. last_ok=1, last_degraded=0, only outcome="mirrored" (no error/degraded outcomes ever). Contrast: v5.66.0 left these flat because the runtime hook wasn't wired.
  • /tmp/ehdb-shadow-ref.jsonl on both pools received exactly 6 entries keyed to execution 332760742153424896 (subject noetl.event.exec.332760742153424896). Decoded worker-rust seq 8 = the call.done/step=start event carrying the real SQL result (rows:[{ok:1}], status:success) — real execution data mirrored into EHDB.
  • Incumbent unaffected: Postgres noetl.event persisted all 13 events (playbook_started→playbook.completed); shadow never served; 0 restarts on either pool; readiness 1.

Bonus (query interface) DEFERRED: /api/ehdb/executions 404s on the running server (v3.49.0 predates the v3.53.0 /api/ehdb/* surface). Needs the v3.53.0 server image loaded — out of scope for this worker-mirror validation.

Left in SHADOW on v5.67.0 so the user can watch real mirroring. Prod unchanged (worker v5.52.0 / server v3.52.0, all EHDB flags off).

Re-confirmed 2026-07-07 (independent second run): a later session found the live pools had drifted back to v5.66.0-p10 (pods recreated on the last set image value) — the mirror had gone dark again (/tmp/ehdb-shadow-ref.jsonl absent, only the noetl_ehdb_readiness_* family on /metrics, no noetl_ehdb_eventlog_*). Rebuilt v5.67.0 (native arm64; the podman VM swap-died mid-build on wasmtime-cranelift and was recovered via podman machine stop/start + podman start noetl-control-plane, then rebuilt with the noetl server+worker pools scaled to 0 for headroom), re-rolled both data-plane pools, and re-proved the mirror with a fresh drive: execution 332760854506246144 (pft_sql_probe_v2, COMPLETED) → 6 mirrored events (seq 301–306, subject noetl.event.exec.332760854506246144) in the ref store, counter mirror/mirrored climbing (273 → 344 → 378), last_ok=1, last_degraded=0, 0 restarts, incumbent served all 13 events. Takeaway: the pool is not sticky on v5.67.0 — a redeploy/recreate reverts it and the mirror goes dark until re-rolled.

2026-07-06: Event-log shadow mirror wired into the LIVE event-emit path

Closed the "runtime not wired" gap the live-in-kind session below surfaced: the per-tier mirror engines only ran via ehdb-selfcheck + tests, so a real drive did not dual-write into EHDB. Now the event-log tier's shadow mirror fires on real events.

Change (worker v5.67.0, noetl/worker#167 → merged d310c7b):

  • Hook site = ControlPlaneClient::emit_event, the single authoritative event-append chokepoint every worker path funnels through (EventEmitter, emit_event_with_retry, spool_runtime, subscription, plugin). After the control plane accepts an event, the same eventlog::mirror_event shadow dual-write + parity path the selfcheck drives mirrors it into the derived EHDB fabric.
  • Armed once at construction by eventlog::runtime_hook_env: only NOETL_EHDB_ENABLED + NOETL_EHDB_EVENTLOG=shadow + data-plane role + local-reference log. Disabled / tier off|primary / control-plane role ⇒ strict per-event no-op, byte-identical /metrics.
  • mirror_live_event is panic-isolated + best-effort: a mirror failure surfaces as a metered non-ok outcome and never propagates into the authoritative event path. authoritative_sequence=None (count + order parity). Shadow never serves; event authorship untouched.

Validation: cargo test --lib 457 pass (7 new hook tests: arms only for enabled+shadow+data-plane; no-op when disabled/off/primary; skips control-plane; fires on shadow; guard-refused for control-plane role; engine error isolated). cargo clippy --lib clean. LOCAL/code only — no GKE, no image build inline.

PENDING — redeploy: the in-kind live-drive proof (build v5.67.0 image, redeploy to the shadow pool, run a pft drive, confirm noetl_ehdb_eventlog_* advances on a real execution) is a follow-up deploy step — the live pool is on v5.66.0-p10 shadow; the next image carries this hook.

Follow-up slices: projection / kv / object / vector tier mirrors plug into the same emit_event seam. See Design: Event-Log Core Engine → Runtime integration.

Refs noetl/ehdb#241, noetl/ai-meta#234.

2026-07-06: EHDB deployed LIVE in the kind stack — SHADOW across all 5 tiers

First time EHDB runs inside the running kind noetl stack (not a one-shot Job). Deployed worker v5.66.0 (Phase 10 config verb) as localhost/noetl-worker:v5.66.0-p10 onto both data-plane pools with EHDB in shadow across all five tiers, dual-writing + parity-checking against the incumbents while real executions run, fully reversible. LOCAL kind only — no GKE/gcloud, no prod (prod stays v5.52.0, all flags off).

Deploy (control plane untouched):

  • noetl-worker-system-pool — NOETL_EHDB_CLIENT_ROLE=system
  • noetl-worker-rust (user/shared pool) — NOETL_EHDB_CLIENT_ROLE=worker
  • env: NOETL_EHDB_ENABLED=1, NOETL_EHDB_MODE=local_reference, NOETL_EHDB_LOCAL_REFERENCE_LOG=/tmp/ehdb-shadow-ref.jsonl, NOETL_EHDB_{EVENTLOG,PROJECTION,KV,OBJECT,VECTOR}=shadow.

Evidence (all in-cluster):

  • Boot readiness preflight ok, outcome=empty, role guard correct; /metrics renders noetl_ehdb_readiness_* (disabled ⇒ zero ehdb lines).
  • ehdb-selfcheck config in the live pod: coherent, 5×shadow, backend=external (shadow keeps the incumbent serving), secret-free, exit 0.
  • All 5 tier suites in the live pod (eventlog/projection/kv/object/vector): ok=true — dual-write mirrored/materialized, parity holds.
  • Real drive automation/pft_sql_probe_v2: postgres start step succeeds end-to-end; drive leaves shadow inert (only readiness on /metrics), events persist to Postgres normally, 0 restarts.
  • Reversibility: NOETL_EHDB_ENABLED=0 ⇒ byte-identical /metrics; back to =1 ⇒ live shadow again. Left in shadow so it can be observed.

Realizes Stage A/B of the prod-cutover runbook in kind. Note the runtime EHDB surface is the readiness preflight + /metrics renderer; the per-tier mirror/parity engines are exercised by the in-image ehdb-selfcheck binary (kind validation / operator preflight), not the worker's request path — so a real drive proves non-interference, and the shadow dual-write/parity is demonstrated via the in-pod selfcheck suites. GKE/prod cutover remains gated on the user, per tier.

Observe: kubectl --context kind-noetl exec <pool-pod> -c noetl-worker -- /app/ehdb-selfcheck config; port-forward the pod's :9090 and curl /metrics | grep noetl_ehdb.

2026-07-06: Durable event-log backend — first slice (segment store + crash recovery)

Landed the durable substrate the Phase-6 note deferred as "production disk format" and the prod-cutover runbook §C durability gate names as the hard blocker for Stage C (primary cutover) — the only backend today, local_reference, is a pod-local JSONL file lost on restart + divergent across replicas, so it is not production-durable as the authoritative store under primary. LOCAL/code only — no GKE/gcloud, no worker image build; validated via cargo test + clippy -D warnings + the built ehdb-local-reference selfcheck. In-container kind validation is PENDING — follow-up deploy. No repos/noetl or repos/server changes.

  • ehdb — #253: ehdb-reference::durable_eventlog — DurableSegmentStore (append-only CRC32-framed seg-*.eslog segments + size rollover, in-memory offset index global_seq → (segment, offset) + per-execution index + durable-consumer ack cursors, fsync-per-append, crash-recovery replay with torn-tail discard+truncate and complete-bad-CRC hard error; payloads cold-loaded via the index → bounded index memory, the noetl/ai-meta#166 property). DurableEventLogDriver implements the same EventLogDriver contract as LocalReferenceEventLogDriver; EventLogStorageBackend (local_reference default | durable_segment) fail-safe selector. exercise_durable_recovery + ehdb-local-reference durable-eventlog-recovery verb prove zero-loss + gapless ordering + per-execution scope + payload fidelity + durable-cursor survival across a simulated restart.
  • Evidence — durable-eventlog-recovery → {"recovered":true, "zero_loss":true,"ordering_ok":true,"scope_ok":true,"payloads_match":true, "cursor_survived":true,"pending_after_restart":2}, exit 0; a durable seg-*.eslog persists on disk. 17 tests incl. rollover across files, torn-tail discard, bit-rot hard error, and parity vs LocalReference over identical ops. cargo fmt --all --check + clippy --workspace --all-targets -D warnings + cargo test --workspace all green.
  • Design note — Durable Event-Log Backend (PVC-vs-object-store recommendation + flagged deploy assumptions).
  • Tracking — #254 (program + 6-slice checklist); remaining slices: execution-affinity single-writer routing → shared object-store tier → worker wiring → kind soak → prod-durability sign-off.

2026-07-06: Phase 10 (Tunable-Backend Config Surface) — IMPLEMENTED → EHDB program (Phases 6–10) CODE-COMPLETE

The final phase of the EHDB completion program landed: the consolidated per-tier backend-selection config surface. It folds the scattered NOETL_EHDB_* per-tier flags into one coherent, documented schema — no breaking rename, the existing env stays the source of truth. With Phase 10 the whole EHDB program (Phases 6–10) is code-complete and kind-validated; the only remaining step is the prod/GKE cutover, gated on the user, per tier. LOCAL/code only — no GKE/gcloud, no image build; the ehdb-selfcheck config matrix was validated as a batched in-kind check. No repos/noetl or repos/server changes (Codex-active).

  • ehdb — #252 (merged 4c0df81): ehdb-reference::backends — PlatformTier (the five tiers, each with its runtime env var + external incumbent), TierMode {off|shadow|primary}, Backend {ehdb|external} + backend_for_mode (primary ⇒ EHDB, else the incumbent), and BackendMatrix with coherence validate() + secret-free to_json(). Pure data (reads no env, opens no engine). 10 unit tests; cargo fmt/test/clippy -D warnings clean.
  • worker — #166 (released v5.66.0): src/ehdb/backends.rs resolve(&EnvMap) maps the process env into the matrix through the same <Tier>Mode::from_env parsers the runtime dispatch uses (backward-compatible by construction). New ehdb-selfcheck config verb (alias backends) prints the resolved 5-tier backend+mode matrix (secret-free; exit 0 coherent, 4 incoherent). Pins ehdb-reference to the Phase-10 SHA. 9 unit tests; cargo test --lib = 449 passed; clippy clean.
  • Selfcheck evidence (ehdb-selfcheck config): all-external default (no env) ⇒ 5×backend:external, exit 0; all-EHDB (enabled worker, 5×primary) ⇒ 5×backend:ehdb, exit 0; mixed (log+vector primary, projection shadow) ⇒ per-tier ehdb/external, exit 0; primary without enable ⇒ coherent:false
    • clear error, exit 4; data-plane tier on gateway role ⇒ coherent:false (2 errors), exit 4; sensitive-keyed env present ⇒ secret_free:true, 0 value leaks.
  • Boundaries — platform-only (business data never in EHDB); no behavior change (disabled-by-default strict no-op intact); EHDB default, not lock-in.
  • Reference: Backend Configuration · Roadmap Phase 10.

2026-07-06: Phase 9 tiers 3–5 (KV / object / vector) — in-kind dual-run VALIDATED → Phase 9 CODE-COMPLETE

The combined tier 3–5 in-cluster dual-run — the last validation step before Phase 9 is code-complete — passed. Tiers 1–2 were already kind-validated (2026-07-05); this run closes tiers 3, 4, 5, so all five per-tier primary cutovers are now IMPLEMENTED + MERGED and in-kind dual-run VALIDATED. Phase 9 is CODE-COMPLETE. LOCAL kind only — no GKE/gcloud, no prod cutover; the prod/GKE cutover stays gated on the user, per tier. Nothing in prod changed (prod worker v5.52.0; all NOETL_EHDB_* flags default off). No repos/noetl or repos/server changes (Codex-active).

  • Build — worker v5.65.0 (ceedbba, the latest release, contains all five tiers from sequential merges) built native-arm64 from repos/worker/Dockerfile → localhost/noetl-worker:v5.65.0-p9 (7d30b250d261, ships ehdb-selfcheck); cargo-chef cook re-ran because the ehdb-reference pin bumped d08013c5 → 0f47fe3f (DuckDB C++ amalgamation was the long pole). Loaded into the kind-noetl node's containerd (ctr -n k8s.io images import).
  • In-cluster run — Job ehdb-p9-t345 (ns ehdb-p9-validate) on node noetl-control-plane; container kernel 6.19.7-200.fc43.aarch64, Alpine 3.22 — the host kernel is the container's, NOT the macOS host (context kind-noetl, API https://127.0.0.1:61866). Pod Succeeded, 0 restarts. The umbrella switch NOETL_EHDB_ENABLED=1 + per-tier primary + worker-role NOETL_EHDB_LOCAL_REFERENCE_LOG drive the served path; server role stays guard_refused.
  • Tier 3 (KV) — served_primary + served_by_ehdb:true + reversible:true
    • dual_run_holds:true (5 parities present/value/ttl vs NATS-KV) + put/get/scan/CAS/delete/TTL ok + keys_after_revert:3; server ⇒ guard_refused.
  • Tier 4 (object) — served_primary + dual_run_holds:true (4 parities digest/length/present/retrievable — content-addressed digest integrity vs the object store) + put/get/list/locate/delete ok + keys_after_revert:3; server ⇒ guard_refused.
  • Tier 5 (vector) — served_primary + dual_run_holds:true (3 top-k parities ids/order/monotonic vs Qdrant) + upsert/query-topk/delete ok + query_returned:3 (bounded top-k cap) + candidates_after_revert:3; server ⇒ guard_refused.
  • All three tiers: metrics_secret_free:true; off ⇒ disabled byte-identical no-op (exit 0). No divergence, no failures.

2026-07-06: Phase 9 tier 5 (vector) — primary-serve IMPLEMENTED + MERGED (ALL FIVE TIERS DONE)

The fifth and final per-tier primary cutover — serving NoETL's internal platform vector tier from EHDB in place of the internal Qdrant retrieval path (the platform RAG / catalog embeddings reached in-process via the Phase-E retrieval path) — is built and merged, activated behind NOETL_EHDB_VECTOR=primary, reversible, and dual-run-verified. Mirrors the tier-1 event-log, tier-2 projection, tier-3 KV, and tier-4 object patterns exactly. With this tier, all five Phase-9 primary-serve activations are implemented. LOCAL only — no prod/GKE cutover; that stays gated on the user, per tier. No repos/noetl or repos/server changes (Codex-active).

  • ehdb — #251 (merged 0f47fe3): ehdb-reference::vector gains exercise_primary_serve + VectorPrimaryInput + VectorPrimaryServeReport::served_by_ehdb() — one authoritative cycle through upsert → served cosine top-k query → tombstone delete → fresh-driver replay, preserving the Qdrant retrieval semantics and dual-run parity-checking each served query (id-set + rank-order + score-monotonicity) against a Qdrant mirror ranked in lockstep via compare_vector_parity. CLI verb vector-primary-serve. 155 crate tests green; clippy -D warnings + fmt clean.
  • worker — #165 (merged 681782c, released v5.65.0 ceedbba): src/ehdb/vector.rs flips PRIMARY_SERVE_ACTIVATED false → true; mirror_upsert serves authoritatively under primary (served_primary / primary_divergence); new serve_primary_cycle runs the cycle + demonstrates reversibility (flip back to shadow, mirror one more point, confirm the whole live set serves). ehdb-selfcheck vector-primary-serve verb proves served-by-EHDB + reversibility
    • secret-free metrics. Pins ehdb-reference to 0f47fe3.

Local proof (ehdb-selfcheck, debug host): off ⇒ no-op + empty metrics (exit 0); primary ⇒ served_by_ehdb:true + reversible:true + candidates_after_revert:3 + dual_run_holds:true (3 parities all hold) + secret-free noetl_ehdb_vector_* (exit 0); control-plane server ⇒ guard_refused (exit 4); shadow ⇒ primary_unavailable (exit 4); over-limit dims ⇒ rejected (exit 3).

Kind dual-run: BATCHED — combined tier 3-5 run to follow. No worker image built this session. Reversible with two levers (runtime NOETL_EHDB_VECTOR=primary→shadow/off restores Qdrant instantly, zero data loss; compile-time PRIMARY_SERVE_ACTIVATED=false is the structural kill switch). Remaining Phase-9 work = the combined tier 3-5 in-kind dual-run (tiers 1–2 already kind-validated), then Phase 9 is code-complete — prod/GKE cutover gated on user.

2026-07-05: Phase 9 tier 4 (object / blob) — primary-serve IMPLEMENTED + MERGED

The fourth per-tier primary cutover — serving NoETL's internal platform object/blob tier from EHDB in place of the internal external object store (the state shards #166 + result tier #104 reached through the server's /api/internal/objects/{key} API) — is built and merged, activated behind NOETL_EHDB_OBJECT=primary, reversible, and dual-run-verified. Mirrors the tier-1 event-log, tier-2 projection, and tier-3 KV patterns exactly. LOCAL only — no prod/GKE cutover; that stays gated on the user. No repos/noetl or repos/server changes (Codex-active).

  • ehdb — #250 (merged bb42b8d): ehdb-reference::object gains exercise_primary_serve + ObjectPrimaryInput + ObjectPrimaryServeReport::served_by_ehdb() — one authoritative cycle through put → per-key digest-verified served get → prefix list → in-cluster locate → tombstone delete → fresh-driver replay, preserving the external-store semantics and dual-run digest-parity-checking each served read against an external-store mirror via compare_object_parity. CLI verb object-primary-serve.
  • worker — #164 (merged a100adf, released v5.64.0 369e4c1): src/ehdb/object.rs flips PRIMARY_SERVE_ACTIVATED false → true; primary serves the object op authoritatively (served_primary / primary_divergence); serve_primary_cycle drives the cycle + demonstrates reversibility (flip back to shadow, mirror one more object over the same store, confirm the whole live set serves). ehdb-selfcheck object-primary-serve verb.

Local ehdb-selfcheck proof: off ⇒ byte-identical no-op (exit 0); primary ⇒ served_primary + served_by_ehdb:true + reversible:true + keys_after_revert:3 + secret-free noetl_ehdb_object_* metrics (exit 0); control-plane server ⇒ guard_refused (exit 4).

  • Kind dual-run: BATCHED — combined tier 3-5 in-cluster run to follow (one worker image build validates the three sequential releases). No worker image built for tier 4. prod/GKE cutover stays GATED on the user.

Reversibility: runtime NOETL_EHDB_OBJECT=primary → shadow/off restores the external object store instantly (zero data loss); compile-time PRIMARY_SERVE_ACTIVATED = false is the structural kill switch. Platform object tier only (state shards + result tier); never authors the event log.

Remaining after this: vector (1 tier) + the combined tier 3-5 in-cluster dual-run.

2026-07-05: Phase 9 tier 3 (KV / state) — primary-serve IMPLEMENTED + MERGED

The third per-tier primary cutover — serving NoETL's internal platform KV/state tier from EHDB in place of the internal NATS-KV bucket — is built and merged, activated behind NOETL_EHDB_KV=primary, reversible, and dual-run-verified. Mirrors the tier-1 event-log and tier-2 projection patterns exactly. LOCAL only — no prod/GKE cutover; that stays gated on the user. No repos/noetl or repos/server changes (Codex-active).

  • ehdb — #249 (merged 73b1446): ehdb-reference::kv gains exercise_primary_serve + KvPrimaryInput + KvPrimaryServeReport::served_by_ehdb() — one authoritative cycle through put → per-key served get → bucket scan → optimistic CAS (versioned swap + create-only conflict) → tombstone delete → absolute-TTL lease → fresh-driver replay, preserving the NATS-KV semantics and dual-run parity-checking each served read against a NATS-KV mirror via compare_kv_parity. CLI verb kv-primary-serve.
  • worker — #163 (merged ba9f829, released v5.63.0 a7925e0): src/ehdb/kv.rs flips PRIMARY_SERVE_ACTIVATED false → true; primary serves the KV op authoritatively (served_primary / primary_divergence); serve_primary_cycle drives the cycle + demonstrates reversibility (flip back to shadow, mirror one more key over the same store, confirm the whole live set serves). ehdb-selfcheck kv-primary-serve verb.
  • Local proof (ehdb-selfcheck kv-primary-serve, host debug build): off ⇒ ehdb:disabled + metrics_empty:true (exit 0); primary ⇒ outcome:served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (5 served-read parities, present_ok/value_ok/ttl_ok, divergence:null) + put_ok/get_ok/scan_ok/cas_ok/delete_ok/ttl_ok/ replay_matches all true + keys_after_revert:3 (NATS-KV path restored whole on flip-back, zero data loss) + secret-free noetl_ehdb_kv_* metrics (exit 0); server role ⇒ guard_refused (exit 4).
  • Kind dual-run: BATCHED — combined tier 3-5 in-cluster run to follow (one worker image build validates the three sequential releases). No worker image built this session.
  • Rollback (two levers, zero data loss): runtime NOETL_EHDB_KV=primary → shadow/off restores NATS-KV instantly; compile-time PRIMARY_SERVE_ACTIVATED = false in src/ehdb/kv.rs is the structural kill switch. Documented in Roadmap → Phase 9 → Tier-3 rollback procedure.

Remaining after this: object + vector (2 tiers) + the combined tier 3-5 in-cluster kind dual-run.

2026-07-05: Phase 9 tiers 1 & 2 — in-cluster kind dual-run VALIDATED

The deferred in-cluster proof for both Phase 9 primary cutovers — event-log (tier 1) and projection / read-model (tier 2) — ran and passed on the local kind cluster. This closes the "kind dual-run PENDING/BATCHED" that the prior two sessions left open when the podman VM's ssh socket was wedged. LOCAL/kind only — no prod/GKE cutover; that stays gated on the user. No new ehdb/worker code — pure validation of the already-merged v5.62.0.

  • Environment recovery — the podman VM (noetl-dev) ssh socket was wedged (handshake reset though the machine reported running); a podman machine stop/start cycle reset it. Kind cluster noetl restarted (its control-plane container had stopped). Proof it is kind, not GKE: kubectl context kind-noetl, API server https://127.0.0.1:61866, node noetl-control-plane (containerd, arm64 Fedora podman VM).
  • Image — worker v5.62.0 (36875e3; ships both tiers + the ehdb-selfcheck binary) built native-arm64 from the repos/worker Dockerfile, saved to an archive, and loaded into the kind node's containerd (localhost/noetl-worker:v5.62.0-p9, sha256:9d222db6, linux/arm64).
  • In-cluster run — a batch/v1 Job (ehdb-p9-dualrun, namespace ehdb-p9-validate, imagePullPolicy: Never) ran the shipped ehdb-selfcheck verbs in a pod on node noetl-control-plane. Pod host kernel 6.19.7-...fc43.aarch64 confirms the verbs executed inside the kind container, not on the host. Pod Succeeded, 0 restarts, no crashloop.
  • Tier 1 (event-log) — off ⇒ ehdb:disabled + metrics_empty:true (exit 0); primary ⇒ served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (3× count_ok/order_ok/ sequence_ok, divergence:null) + replay_matches:true + scope_ok:true + records_after_revert:4 (incumbent JetStream+Postgres path restored whole on flip-back to shadow) + secret-free noetl_ehdb_eventlog_* metrics (exit 0); server role ⇒ guard_refused (exit 4).
  • Tier 2 (projection) — off ⇒ ehdb:disabled + metrics_empty:true (exit 0); primary ⇒ served_primary + served_by_ehdb:true + reversible:true + dual_run_holds:true (key_ok/value_ok/checkpoint_ok, checkpoint_lag:0) + list_ok:true + read_event_ok:true + replay_idempotent:true + replay_matches:true + scope_ok:true + rows_after_revert:3 (Postgres-materializer read path restored whole on flip-back) + secret-free noetl_ehdb_projection_* metrics (exit 0); server role ⇒ guard_refused (exit 4).
  • Boundaries held — platform-only; event authorship unchanged (the cutover changes only the serving engine, never who appends); loose coupling; control-plane guard-refused; event log stays the source of truth (replay-is-truth proven). repos/noetl + repos/server untouched (Codex lane).
  • Roadmap tier-1 + tier-2 rows flipped PENDING/BATCHED → VALIDATED; rollback procedures unchanged (both levers still documented).

2026-07-05: Phase 9 tier 2 (projection / read-model) — primary-serve implemented + merged (kind dual-run batched)

The second per-tier primary cutover of the completion program — the projection / read-model tier — is built, activated, reversible, and merged, mirroring the tier-1 event-log pattern exactly. Under NOETL_EHDB_PROJECTION=primary the EHDB projection engine now serves the materialized read-models the control plane queries (list_executions, per-execution read_execution_state, read_event) authoritatively in place of the PostgreSQL materializer, dual-run parity-checked against the incumbent. No prod/GKE cutover was performed — that is a separate later step gated on the user. LOCAL/kind scope only.

  • ehdb — #248 (merged d08013c): projection::exercise_primary_serve + ProjectionPrimaryInput + ProjectionPrimaryServeReport::served_by_ehdb() (served-by-EHDB proof over the full authoritative cycle: apply → the three read-model query contracts → checkpoint → idempotent re-apply → fresh-engine replay, dual-run parity-checked) + projection-primary-serve CLI verb. Reversible + additive toward the incumbent (KeepAll projection store, replay-is-truth proven).
  • worker — #162 (merged a56583c, released v5.62.0 36875e3): PRIMARY_SERVE_ACTIVATED false → true (reversible via the runtime flag); serve_primary_cycle + ProjectionServeResult + ehdb-selfcheck projection-primary-serve verb that proves served-by-EHDB and reversibility (flip back to shadow, one more execution materialized, read-models replay whole). Pins ehdb d08013c.
  • Local proof (same binary the image ships): off ⇒ byte-identical no-op; primary ⇒ served_by_ehdb:true + reversible:true + secret-free metrics; control-plane ⇒ guard_refused (exit 4).
  • Kind dual-run: VALIDATED 2026-07-05 — see the 2026-07-05 in-cluster validation session entry at the top. ehdb-selfcheck projection-primary-serve ran in a Job pod on kind node noetl-control-plane (image v5.62.0 sha256:9d222db6): served_by_ehdb:true + reversible:true + dual_run_holds:true + secret-free metrics (exit 0); control-plane ⇒ guard_refused (exit 4).
  • Boundaries held: platform-only; a projection is a derived read-model built by consuming the append-only event log — it never authors an event; loose coupling; control-plane guard-refused; event log stays the source of truth. repos/noetl + repos/server untouched (Codex lane).
  • Rollback procedure documented on Roadmap.

2026-07-05: Phase 9 tier 1 (event log) — primary-serve implemented + merged (kind dual-run pending)

The first per-tier primary cutover of the completion program — the event-log tier, the bottleneck fix that motivated the whole program — is built, activated, reversible, and merged. Under NOETL_EHDB_EVENTLOG=primary the EHDB event-log engine now serves the platform event log authoritatively (append

  • read + tail + ack + replay), dual-run parity-checked against the incumbent. No prod/GKE cutover was performed — that is a separate later step gated on the user. LOCAL/kind scope only.
  • ehdb — #247 (merged 7f014c9): exercise_primary_serve + EventLogPrimaryEvent + EventLogPrimaryServeReport (served-by-EHDB proof over the full authoritative cycle) + eventlog-primary-serve CLI verb. Reversible + additive toward the incumbent (KeepAll log, replay-is-truth proven).
  • worker — #161 (merged ddf41de, released v5.61.0 7e98538): PRIMARY_SERVE_ACTIVATED false → true (reversible via the runtime flag); serve_primary_cycle + EventLogServeResult + ehdb-selfcheck eventlog-primary-serve verb that proves served-by-EHDB and reversibility (flip back to shadow, log replays whole). Pins ehdb 7f014c9.
  • Local proof (same binary the image ships): off ⇒ byte-identical no-op; primary ⇒ served_by_ehdb:true + reversible:true + secret-free metrics; control-plane ⇒ guard_refused.
  • Kind dual-run: VALIDATED 2026-07-05 — see the 2026-07-05 in-cluster validation session entry at the top. ehdb-selfcheck eventlog-primary-serve ran in a Job pod on kind node noetl-control-plane (image v5.62.0 sha256:9d222db6): served_by_ehdb:true + reversible:true + dual_run_holds:true (3× append parity) + secret-free metrics (exit 0); control-plane ⇒ guard_refused (exit 4).
  • Boundaries held: platform-only; event authorship unchanged (gateway/server stay the gatekeeper — primary changes only the serving engine underneath); loose coupling; control-plane guard-refused; never authors a NoETL event.
  • Rollback procedure documented on Roadmap.

2026-07-05: Phase 8 vector engine slice landed — Phase 8 shadow-complete (all three engines)

Third and final Phase-8 engine slice (Roadmap Phase 8, #241): EHDB's vector engine — the durable vector engine underneath NoETL's internal platform vector tier (RAG / catalog embeddings), formalizing the already-in-process Phase-E retrieval path behind a first-class VectorDriver. Design note: KV/State + Object/Blob + Vector Engines (Phase 8). With this slice, all three Phase-8 engines (KV + object + vector) are shadow-complete. No tier is cut over to primary — per-tier primary cutover is the separately-gated Phase 9 step, so Phase 8 stays in progress overall.

  • ehdb — PR #246 (merged f6c5a6f). New ehdb-reference::vector: the VectorDriver trait (upsert / query (top-k) / delete) and LocalReferenceVectorDriver over one canonical noetl_vector_index stream (per-point subject noetl.vec.<hex(collection)>.<hex(point_id)>, KeepAll retention → replay-is-truth). Query is a bounded cosine top-k filtered to the query's model + dimensionality, ranked descending with a deterministic point-id tie-break — the same scoring shape as ehdb_retrieval::InMemoryRetrievalCatalog::search_similar (the Phase-E path). Bounded + secret-free (MAX_VECTOR_DIMENSIONS 4096, MAX_VECTOR_QUERY_TOP_K 64, MAX_VECTOR_PAYLOAD_BYTES 16 KiB; over-cap → rejected, empty/non-finite/zero → invalid). compare_vector_parity = pure id-set + rank-order + score-monotonicity parity (tolerance for cross-engine float slack). 17 vector engine tests; clippy -D warnings + fmt.
  • worker — PR #160 (merged c7b8872, v5.60.0 0f57625). src/ehdb/vector.rs shadow behind NOETL_EHDB_VECTOR=off|shadow|primary (default off): shadow dual-writes one platform vector into the EHDB engine alongside the authoritative Qdrant path, then reads back a self-retrieval top-k query and compares id-set/rank-order/monotonicity parity without serving reads or touching the authoritative Qdrant retrieval path; primary recognised but inert (compile-time PRIMARY_SERVE_ACTIVATED = false). Secret-free noetl_ehdb_vector_* metrics; control-plane guard; structural no-NoETL-event-writer test. ehdb-selfcheck gains mirror-vector / vector-suite. Pins ehdb-reference to the merged #246 rev f6c5a6f. 14 unit tests; clippy -D warnings + fmt.
  • Boundary preserved: the vector engine indexes derived platform embeddings — it never authors an event; platform-only (business vector collections stay external on their own Qdrant); bounded; secret-free; loose coupling (driver-selectable back to Qdrant, Phase 10).
  • Validation: engine side — 17 vector engine tests + clippy/fmt clean. Worker side — 14 ehdb::vector + 2 ehdb::metrics tests pass, clippy -D warnings + fmt clean, and the built ehdb-selfcheck binary proves: off → byte-identical no-op + empty metrics (exit 0); shadow vector-suite → 4/4 engine steps (upsert/query-parity/top-k-truncate/delete) + secret-free noetl_ehdb_vector_* metrics (exit 0); shadow mirror-vector → mirrored, top-k parity holds (exit 0); control-plane (server) role → guard_refused (exit 4); primary → primary_unavailable, inert (exit 4); over-limit dimensionality → rejected (exit 3); zero vector → invalid (exit 4). In-container kind run deferred — env build eviction (the worker-rust image is a ~110-min cold build; same standing-evidence pattern as Phase 6/7/E). No GKE. repos/noetl + repos/server untouched (Codex-active).
  • Phase 8 → Phase 9 handoff: Phase 8 is now engine-complete across all five platform tiers (event-log, projection, KV, object, vector — the latter three from Phase 8, the first two from Phases 6–7), each in disabled-by-default shadow. Phase 9 is the five independent per-tier primary cutovers (shadow→primary, EHDB serves reads, incumbent retired for that tier), each gated + dual-run verified + rollback-documented + kind-before-GKE. See the Phase-9 per-tier table in Roadmap.

2026-07-05: Phase 7 started — projection / read-model engine + driver interface + disabled-by-default shadow

Second implementation slice of the completion program (Roadmap Phase 7, #241): EHDB's projection / read-model engine — the engine that builds + serves the materialized read-models the control plane queries, off the Phase-6 event-log tail, retiring the PostgreSQL materializer. Design note: Projection / Read-Model Engine (Phase 7). Design + engine slice + disabled-by-default shadow only — reads are NOT cut over off Postgres (a later gated step, Phase 9).

  • ehdb — PR #243. New ehdb-reference::projection: the ProjectionDriver trait (apply / read_execution_state / read_event / list_executions / checkpoint) and LocalReferenceProjectionEngine, composing the append-only stream primitives over one noetl_projection_log store (per-execution scope via noetl.projection.exec.<id> subject). Materializes the event read-model (keyed on event_id, the ON CONFLICT DO NOTHING twin), the folded execution-state read-model (the projection_snapshot twin), and a durable consumer checkpoint (the event_stream.position twin). Apply is idempotent / exactly-once keyed on the Phase-6 global sequence (skip <= checkpoint + event_id dedup); rebuild-from-log is deterministic (batch-boundary independent). ProjectionEventInput::from_event_log_record bridges the Phase-6 tail; compare_projection_parity is the pure key/value/checkpoint parity vs the authoritative materializer. ehdb-local-reference gains projection-apply / projection-read-exec / projection-read-event / projection-list / projection-checkpoint / projection-from-eventlog (eventlog→projection bridge) / projection-suite (distinct exit codes 0/3/4/5). 19 unit tests (73 in the crate); clippy -D warnings + fmt.
  • worker — PR #157 (merged eadc3a5). src/ehdb/projection.rs shadow behind NOETL_EHDB_PROJECTION=off|shadow|primary (default off): shadow dual-materializes the read-models from the event-log tail alongside the Postgres materializer + compares parity (compare_projection_parity) without serving reads or touching the authoritative materializer; primary recognised but not activated (compile-time PRIMARY_SERVE_ACTIVATED = false). Secret-free noetl_ehdb_projection_* metrics; control-plane guard; bounded apply batch. Pins ehdb-reference to the merged #243 rev e0f1c0f. 13 unit tests; clippy -D warnings + fmt.
  • Boundary preserved: the projection engine materializes read-models by consuming the log — it never authors an event (the #103 sole-writer stays the only noetl.event writer), platform-only, bounded, secret-free.
  • Validation: engine side — the built ehdb-local-reference drives projection-suite (exit 0, parity holds), idempotent re-apply (all skipped-below-checkpoint), invalid id → exit 4, and the eventlog→projection bridge folds correctly. Worker side — 63 ehdb:: unit tests pass (13 in projection.rs), clippy -D warnings + fmt clean, and the built ehdb-selfcheck binary proves: projection-suite disabled → byte-identical no-op + empty metrics (exit 0); shadow (worker role) → apply_1 materialized 3 / apply_2_replay applied 0, parity holds, secret-free noetl_ehdb_projection_* metrics render (exit 0); control-plane (server) role → guard_refused (exit 4); primary → primary_unavailable, inert (exit 4). In-container kind run deferred — env build eviction (the worker-rust image is a ~110-min cold build; same standing-evidence pattern as Phase 6/E). No GKE.

2026-07-05: Phase 6 started — event-log core engine + driver interface + disabled-by-default shadow

First implementation slice of the completion program (Roadmap Phase 6, #241): EHDB's event-log core engine — the durable persistence + ordering + serving layer that replaces the JetStream + Postgres log-and-store path underneath the append-only producer path. Design note: Event-Log Core Engine (Phase 6). Design + disabled-by-default shadow only — the log is NOT cut over.

  • ehdb — PR #242 MERGED squash 9a9b28d. New ehdb-reference::eventlog: the EventLogDriver trait (append / scan_global / read_execution / tail / ack) and LocalReferenceEventLogDriver, composing the append-only stream primitives over one canonical noetl_event_log stream so its sequence is the global, monotonic, gapless event-log sequence. Per-execution scope via noetl.event.exec.<execution_id> subject; durable-consumer tail/ack; compare_shadow_parity. ehdb-local-reference gains eventlog-append / eventlog-scan / eventlog-read-exec / eventlog-tail / eventlog-ack / eventlog-suite (distinct exit codes 0/3/4/5). 12 unit tests; clippy -D warnings + fmt; CI green.
  • worker — PR noetl/worker#156 MERGED squash 43c8f0f. New src/ehdb/eventlog.rs shadow behind NOETL_EHDB_EVENTLOG=off|shadow|primary (default off): shadow dual-writes each already-authored event into the engine + compares sequence/count/order parity without serving reads or touching the authoritative path; primary recognised but not activated (compile-time PRIMARY_SERVE_ACTIVATED = false). Parity uses the engine's gapless invariant (global_sequence == log_record_count) — concurrency-safe, no shared bookkeeping. Secret-free noetl_ehdb_eventlog_* metrics; control-plane guard; ehdb-selfcheck mirror-eventlog / eventlog-suite. Pins ehdb-reference c2aaad5 → 9a9b28d. 10 unit tests; clippy + fmt.
  • Validation — local ehdb-selfcheck drive against the built binary: off ⇒ metrics_empty:true exit 0 (byte-identical no-op); shadow ⇒ 3 mirrors parity-holds, secret-free metrics, exit 0; control-plane (server) ⇒ guard_refused exit 4, no write; primary ⇒ primary_unavailable exit 4 (never serves); parity mismatch ⇒ detected exit 5. In-container kind-noetl run deferred — env build eviction (~110-min cold worker-rust image; same standing evidence pattern as the Phase E first/second slices). No GKE. repos/noetl + repos/server untouched (Codex-active).
  • Event authorship unchanged — gateway/server still gatekeep what is appended; the shadow only mirrors already-authored events and never writes noetl.event. What changed is the engine underneath the producer path.
  • Remaining Phase 6 — production segmented disk format + offset index
    • compaction; sharded/multi-stream ordering; a JetStream+Postgres EventLogDriver (Phase 10 tunable surface); the projection engine (Phase 7) attaching to the tail/ack/read-execution surface; primary-serve cutover (dual-run verify + rollback, kind before GKE).

2026-07-05: RFC DECIDED — EHDB completion program (loose coupling + noetl self-sufficiency; EHDB is the event-log engine)

Docs/RFC session. Recorded the decided EHDB completion program on the wiki page RFC: EHDB Completion Program — Server↔EHDB Coupling + noetl Self-Sufficiency (status decided, applying — not a review gate), grounded against the current Architecture + Roadmap pages (Phase 5 integration A–E complete) and #234. Supersedes the earlier "pluggable substrate, up for review" framing.

  • Headline motivation — fix the event-sourcing log bottleneck. EHDB IS a distributed event-log database: it writes, orders, persists, and serves the noetl.event log and builds/serves the projections, replacing the JetStream + Postgres-materializer path that is today's scaling pressure point (off-server state builder; unbounded WAL index ai-meta#166; materializer soak ai-meta#104).
  • What EHDB is — a family of core engines under one catalog/URN namespace: event-log, projection/read-model, KV/state, object/blob, vector — potentially separate engines per workload type.
  • Critical boundary — EHDB serves NoETL platform functionality only (event log, projections, platform KV/state, catalog, system-WASM store, platform artifacts/vector). Business data is never in EHDB — tenant data stays in real business systems (Postgres, Snowflake, Cassandra, Kafka, Elasticsearch, ClickHouse, object stores) via playbook connectors. "EHDB replaces Postgres/Qdrant/NATS/object store" = NoETL's internal platform uses only.
  • Decision 1 — coupling LOOSE (unchanged). Server/API/gateway control-plane-only + stateless; EHDB data-plane in-process in worker/system; seam = event log + catalog/URN, logical mapping not 1:1 pinning. Server MAY expose a thin control surface + route; MUST NOT embed a data plane. Narrow control-plane-read-cache exception.
  • Decision 2 — EHDB as noetl's self-sufficient internal substrate. End-state: NoETL runs on k8s with no external infra dependency for platform functionality; EHDB default across platform tiers, every tier tunable back to the incumbent (JetStream+Postgres / NATS KV / object store / Qdrant). Default, not lock-in.
  • Reconciliation (not deleted, scoped). Event authorship rules unchanged: gateway/server gatekeep what is appended through the append-only producer path; non-log data-plane roles (worker bounded step, RAG, system store) never fabricate application events. The storage/ordering/projection engine underneath that path is now EHDB (replacing JetStream + the Postgres materializer). Authorship unchanged; engine is EHDB.
  • Roadmap — restructured post-integration phases into the completion trajectory Phases 6–10: 6 event-log core engine (bottleneck fix) → 7 projection/read-model engine → 8 KV/object/vector engines → 9 external-dependency retirement (k8s-only) → 10 tunable-backend config surface. Old Phase 6 (Postgres replacement) + Phase 7 (dependency collapse, #6) absorbed. Architecture + Home + Sidebar link the RFC.
  • Applying, not reviewing. Tracking issue #241 retitled to the completion program (decided, implementing) with the phase checklist; pointer comment on #234. No code repos touched (repos/noetl / repos/server untouched — Codex-active). No GKE.

2026-07-05: EHDB Phase E second slice — bounded RAG retrieval wired into worker-rust (Phase E COMPLETE)

The second (and final) Phase E slice — bounded RAG retrieval — is DONE end to end, unblocking the RAG direction the first slice had deferred:

  • ehdb side: #240 MERGED (squash c2aaad5) → ehdb main. Adds the bounded, read-only retrieve_local_reference_context helper to ehdb-reference (runs ehdb-retrieval's search_text under the hood), an ingest_local_reference_retrieval_document companion, and ehdb-local-reference ingest-doc / retrieve CLI verbs (retrieve exit 0=hit/empty, 3=rejected, 4=invalid). Three caps enforced in the helper: top-k (ceiling 64), per-hit size (ceiling 64 KiB, char-boundary truncation), wall-clock budget (ceiling 60 s, time_capped). 9 unit/integration tests; clippy -D warnings + fmt clean.
  • worker side: noetl/worker#155 MERGED (squash d1ebaf2). New in-process src/ehdb/rag.rs bridges the helpers (retrieve + ingest, matching the Phase C/D/E pattern), a noetl_ehdb_rag_* metric family, and ehdb-selfcheck ingest-rag / retrieve-rag / rag-suite (A→E+RAG driver). Pins ehdb-reference 9bb5928→c2aaad5. 8 unit tests.
  • Boundaries (unit + integration tested): disabled-by-default no-op (byte-identical /metrics), control-plane guard (gateway/api/server refused), bounded (NOETL_EHDB_RAG_TOP_K 8/64, NOETL_EHDB_RAG_MAX_CHUNK_BYTES 4 KiB/64 KiB, NOETL_EHDB_RAG_TIME_BUDGET_MS 5 s/60 s; over-limit Rejected), stateless, event-log-authoritative (retrieve read-only; ingest writes only the private JSONL fabric, never noetl.event).
  • Validated via the built ehdb-selfcheck A→E+RAG drive: disabled no-op (empty metrics), enabled rag-suite (ingest 3 chunks → retrieve hit with top-k truncation → empty → over-limit rejected, ok:true, metrics_secret_free), over-limit Rejected (exit 3), control-plane server refused (exit 4, retrieve:null), secret-free noetl_ehdb_rag_* metrics. In-container kind-noetl validation deferred (image localhost/local/noetl:ehdb-rag, distinct ehdb-rag-selfcheck pod): the worker-rust image is a ~110-min cold build (the ehdb-reference rev bump busts the cargo-chef dependency cache) that timed out in this environment — the built ehdb-selfcheck drive above is the standing evidence. No GKE.
  • ai-meta pointers (gitlink-only, each commit only its own gitlink): repos/ehdb → c2aaad5, repos/worker → d1ebaf2, repos/ehdb-wiki → this commit.
  • Phase E is COMPLETE (both slices — system WASM store + RAG retrieval). The broad EHDB-integration umbrella #234 stays OPEN with Phase E marked done; remaining roadmap work is Phase 6 (PostgreSQL replacement path) and Phase 7 (dependency collapse).

2026-07-05: EHDB Phase E first slice — system WASM store wired into worker-rust

Phase E (system-WASM store → RAG) starts, Rust-first. The first slice — the system WASM library store — is DONE end to end:

  • ehdb side already merged: #239 → ehdb main 9bb5928 — bounded publish / bind / resolve *_system_module helpers in ehdb-reference (immutable module manifests + environment/channel bindings).
  • worker side: noetl/worker#154 MERGED (squash 3162b1d). New in-process src/ehdb/systemstore.rs bridges the helpers (matches the Phase C/D pattern), a noetl_ehdb_systemstore_* metric family, and ehdb-selfcheck publish-system/bind-system/resolve-system + a system-suite A→E driver. Bumps ehdb-reference 3cefba9→9bb5928.
  • Boundaries (unit + integration tested): disabled-by-default no-op (byte-identical /metrics), control-plane guard (gateway/api/server refused), bounded (NOETL_EHDB_SYSTEM_MAX_MODULE_BYTES 16 MiB/256 MiB, NOETL_EHDB_SYSTEM_MAX_CAPABILITIES 16/64; over-bound Rejected), stateless, event-log-authoritative (private JSONL, never noetl.event). WASM execution stays host-side/sandboxed — EHDB catalogs the module ref only.
  • Kind-validated against the worker-rust image (localhost/local/noetl:ehdb-phase-e, kind-noetl, distinct ehdb-phase-e-selfcheck pod on node noetl-control-plane, provider kind://podman/...): disabled no-op (empty metrics) A–D + Phase E, enabled A–D drive, enabled Phase-E system-suite (absent→publish→bind→resolve rev1→publish→rebind→resolve rev2, ok:true), control-plane server/api refused (exit 4, resolve:null), oversized publish Rejected (exit 3), unsupported target Invalid (exit 4), 0 noetl.event writes (only PublishLibrary/BindLibrary in the private log), secret-free noetl_ehdb_systemstore_* metrics, Rust stack 0 restarts. No GKE.
  • ai-meta pointers (gitlink-only, each commit only its own gitlink): repos/worker → 3162b1d (subsumes #153 d6226a2, which never reached ai-meta main), repos/ehdb → 9bb5928, repos/ehdb-wiki → this commit.
  • Next slice — bounded RAG retrieval — DEFERRED: no bounded retrieval helper in the merged ehdb-reference slice; needs a follow-up ehdb slice. Umbrella #234 stays OPEN.

2026-07-05: EHDB re-home COMPLETE — both PRs merged, #238 closed

The Python-retire review gate cleared (approved), so the re-home is done:

  • noetl/noetl#691 MERGED (merge commit ff3a920f) — the Python EHDB path is retired; noetl main carries zero noetl/core/ehdb_* modules (guard test in place). Pairs with noetl/worker#153 (d6226a2, merged 2026-07-04), which owns the integration in worker-rust in process.
  • Re-validated in kind against the worker-rust image (localhost/noetl-worker:ehdb-test, kind-noetl, distinct ehdb-* Job name): disabled no-op (metrics empty), enabled worker full drive, cross-process durable cursor (fresh process sees acked_sequence advanced), control-plane guard refused (no write), secret-free metrics, event-log-authoritative. Rust stack 0 restarts. No GKE.
  • ai-meta pointers: repos/noetl → ff3a920f, repos/ehdb-wiki → this commit (repos/worker → d6226a2 bumped 2026-07-04). Each pointer-bump commit carries only its own gitlink.
  • noetl/ehdb#238 CLOSED; umbrella #234 stays open for Phase E (Rust-first).

2026-07-04: EHDB worker integration re-homed into worker-rust (Rust-only); Python path retired

Implements the Rust-first decision below (owner directive) — closes the integration debt: the EHDB worker/playbook integration now lives in worker-rust, in process, and the Python EHDB path is retired.

  • worker-rust in-process integration — noetl/worker#153 MERGED (squash d6226a2). New src/ehdb module: contract + guard (control-plane boundary), readiness (non-fatal bootstrap preflight over summarize_local_reference), dataplane (bounded append/read), eventstream (bounded project/consume/ack durable-consumer drain), metrics (secret-free noetl_ehdb_* appended to the worker /metrics). Depends on the ehdb-reference crate as a git dep pinned to 3cefba9 — no subprocess. Ships an ehdb-selfcheck bin for in-image validation. Disabled-by-default strict no-op; control-plane guard; bounded + stateless; secret-free metrics; event-log-authoritative (no NoETL event-writer import, structurally asserted).
  • Python path retired — noetl/noetl#691 deletes the Python EHDB modules (noetl.core.ehdb_*), the worker bootstrap wiring, the step CLIs / smoke scripts, and stops bundling the ehdb-local-reference binary in the Python images; adds a guard test. Open, awaiting required review (protected main) — not merged.
  • Kind validation against the worker-rust image (kind-noetl, localhost/noetl-worker:ehdb-test): disabled no-op (byte-identical, metrics empty), enabled worker full drive (readiness ready → append → read → project → consume → ack), cross-process durable cursor (fresh process sees the advanced ack cursor), control-plane guard refused (exit 4, no log written), secret-free metrics, event-log-authoritative (only the EHDB JSONL written). No GKE. Running Rust stack undisturbed (server-rust + system-pool 0 restarts).
  • Pointers: ai-meta repos/worker → d6226a2, repos/ehdb-wiki → this commit. repos/noetl deferred until noetl#691 merges.
  • Tracks noetl/ehdb#238 (closes on the Python-retire merge) under umbrella #234.

2026-07-04: Rust-first EHDB worker-integration decision + Phase D pointer finalized

Documentation / decision session (no product code).

  • Phase D pointer finalized. noetl/noetl#690 merged (merge commit 04d3e273); the ai-meta repos/noetl pointer was bumped to it, completing the Phase D pointer set (ehdb 3cefba9 + ehdb-wiki 064b29a were already bumped). #234 stays OPEN for Phase E.
  • Architecture Decision (owner directive): Rust-first EHDB worker integration; Python as thin wrapper. The EHDB storage engine is already Rust (the ehdb crate: append/read/project/consume/ack). The Phase B–D worker/playbook integration glue (readiness hook, data-plane step, event-stream drain) was added in the legacy Python worker runtime (noetl/core), shelling out to the Rust binary as a subprocess — but prod runs worker-rust, so those disabled-by-default Python hooks do not execute in prod (they only exercise against the Python worker, as a stopgap). Target: the integration lives in worker-rust invoking the ehdb crate in-process (no subprocess shell-out), with Python reduced to a thin binding over the Rust core (the "Polars model"). Captured on the Architecture page (Architecture Decision: Rust-first EHDB worker integration) and the Roadmap (Integration debt item + Phase E marked Rust-first).
  • Tracking issue: noetl/ehdb#238 — re-home Phase B–D worker integration into worker-rust; tracks #234.

2026-07-04: NoETL Integration Phase D — Worker/Playbook Event-Stream Drain

The event-stream integration path: a bounded, disabled-by-default worker/playbook/system drain that mirrors already-emitted NoETL events into a derived EHDB stream and consumes them through a durable consumer with explicit ack-after-materialize semantics.

  • EHDB adds consume / ack subcommands + consume_local_reference_event_records / ack_local_reference_event_consumer to the ehdb-local-reference helper (composing with the Phase C append project leg), and a read-only InMemoryStreamLog::consumer cursor getter. consume creates the durable consumer on first pull and returns pending records after its cursor without moving it; ack advances the cursor after materialize and rejects backwards / unknown / zero sequences atomically.
  • NoETL adds noetl.core.ehdb_eventstream (project / consume / ack), adapter LocalReferenceConsumeResult / LocalReferenceAckResult, worker /metrics (noetl_ehdb_eventstream_*), a worker/playbook-local step CLI (scripts/ehdb_eventstream_step.py), and a kind smoke (scripts/smoke_ehdb_eventstream.py). The Dockerfiles bump EHDB_REF to the Phase D helper.
  • Event-log-authoritative invariant: the NoETL event log stays the append-only source of truth; EHDB is a derived, auxiliary consumer that never writes back to it (no event-writer import; a unit test asserts it structurally). Still disabled-by-default (strict no-op, byte-identical /metrics), worker/playbook/system-only with the code-level control-plane guard (assert_event_stream_access_allowed), bounded (payload + consume-limit + ack sequence ≥ 1 + time caps), stateless, secret-free metrics.

Landed:

Validation:

  • cargo test -p ehdb-reference → 30 passed; workspace green; clippy -D warnings clean; cargo fmt --all --check clean.
  • pytest tests/core/test_ehdb_eventstream.py tests/scripts/test_ehdb_eventstream_step.py → 27 passed (incl. a real-binary project→consume→ack cursor-restart drain); 131 EHDB tests green, no regression.
  • Local kind (kind-noetl, Phase D image with the real linux binary): a Job ran all checks green — helper carries consume/ack verbs, disabled no-op + byte-identical /metrics, enabled worker drain + durable cursor restart, control-plane gateway guard refused (exit 4, no write), secret-free metrics, Phase C append/read no-regression. Nothing on GKE/prod.

Next: Phase E — system WASM store integration, then RAG retrieval.

2026-07-04: NoETL Integration Phase C — Worker/Playbook Data-Plane Step

First bounded data-plane integration (product code in noetl/noetl, built over an EHDB append/read helper):

  • noetl.core.ehdb_dataplane (append_ehdb_domain_record / read_ehdb_domain_records) appends and reads a single domain record through the local-reference adapter — the first data-plane op, past the readiness summary of Phase B. Adapter gains typed LocalReferenceAppendResult / LocalReferenceReadResult; payload reaches the helper verbatim (byte-exact record).
  • Still disabled-by-default (strict no-op, byte-identical /metrics), worker/playbook/system-only. Code-level control-plane guard (assert_data_plane_access_allowed) refuses gateway/api/server before any helper runs → guard_refused, no write.
  • Bounded (payload byte cap + read-limit cap + short time cap; over-bound → rejected) and stateless (helper opened+dropped per call). Secret-free noetl_ehdb_dataplane_ops_total{operation,outcome} metrics on the worker /metrics surface.
  • Worker/playbook-local step CLI (scripts/ehdb_dataplane_step.py) / kind smoke (scripts/smoke_ehdb_dataplane.py) — not a server endpoint.

Landed:

Validation:

  • pytest tests/core/test_ehdb_dataplane.py tests/scripts/test_ehdb_dataplane_step.py → 23 passed (incl. a real-binary append/read roundtrip); 84 existing EHDB tests still green.
  • Local kind (kind-noetl, Phase C overlay image, real linux binary): a Job ran six checks green — disabled no-op + byte-identical /metrics, worker/system/playbook append→read roundtrip, control-plane gateway guard refused (exit 4, no write), secret-free metrics. Nothing on GKE/prod.

See the Architecture page's Worker/Playbook Data-Plane Step (Integration Phase C) section and the Roadmap integration phases.

2026-07-04: NoETL Integration Phase B — Worker/Playbook Readiness Hook

NoETL-side integration milestone (product code in noetl/noetl, not the EHDB repo):

  • Bounded, stateless readiness preflight (noetl.core.ehdb_readiness.evaluate_ehdb_readiness) over the Phase-A read_ehdb_local_reference_summary_from_env summary. Wired for worker/playbook/system data-plane roles only; code-level guard (assert_data_plane_read_allowed) refuses a data-plane read for any control-plane role. Time-bounded, holds no state.
  • Observability noetl_ehdb_readiness_* on the worker /metrics surface (no secret values). Disabled-by-default → strict no-op (byte-identical, no metric recorded).
  • Worker/playbook-local preflight command (scripts/ehdb_readiness_preflight.py) / kind smoke step — not a server endpoint.

Landed:

Validation:

  • pytest tests/core/test_ehdb_readiness.py (+ full EHDB core suite, 81 passed).
  • Local kind (kind-noetl) Job on an overlay image (Phase A helper image + this code, real ehdb-local-reference helper): disabled no-op (rc=0), worker local_reference bounded read (rc=0), gateway control-plane guard (rc=4) — all passed. No GKE/prod changes.

See the Architecture page's Worker/Playbook Readiness Hook (Integration Phase B) section and the Roadmap integration phases.

2026-07-04: Local Reference Summary Helper

Implementation branch:

  • kadyapam/ehdb-local-reference-summary-helper

Design target:

  • Add a concrete local helper binary for NoETL-side integration tests and worker/playbook diagnostics.
  • Open the local JSONL transaction log through LocalReferenceRuntime and report deterministic JSON counts from replayed state.
  • Include transaction, catalog, stream, retrieval, system-library, and storage counts so the helper reflects the current NoETL-domain storage surfaces.
  • Keep this bounded to local reference inspection: no daemon, network API, gateway route, SQL planner, distributed executor, production IAM, storage mutation behavior, background worker, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt
  • cargo test -p ehdb-stream -p ehdb-retrieval -p ehdb-reference

2026-07-03: Embedded NoETL Role Capability Policy

Implementation branch:

  • kadyapam/ehdb-embedded-role-policy

Design target:

  • Make EHDB's embedded distributed database direction explicit without weakening the NoETL execution model.
  • Add typed NoETL embedded roles for gateway, API, worker, playbook, and system contexts.
  • Add typed EHDB capabilities that distinguish control-plane embedding from catalog, transaction, stream, object, retrieval, replication, and system-library data-plane access.
  • Keep gateway/API roles control-plane only by default; workers, playbooks, and system jobs carry explicit data-plane capabilities for bounded work.
  • Keep this as policy/modeling only: no daemon, network API, gateway route, SQL planner, distributed executor, production IAM, storage mutation behavior, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt
  • cargo test -p ehdb-core
  • cargo test --workspace
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 258 Rust tests across unit, integration, and doc-test targets.

2026-06-20: Bootstrap EHDB Into NoETL

Context:

  • noetl/ehdb was created as the future Event Horizon Database.
  • ai-meta added the code repository and wiki as submodules.
  • EHDB is positioned as the core database storage layer for the NoETL multitenant distributed operating-system cloud platform.

Design decisions:

  • EHDB stores both operational metadata/catalog data and historical analytical data.
  • The catalog lives inside EHDB as first-class transactional metadata.
  • Object storage holds immutable data files: Arrow IPC, Parquet, Iceberg-compatible tables, and blobs.
  • Metadata is native EHDB state, not an external PostgreSQL dependency once EHDB reaches self-hosting.
  • Early implementation starts with Rust workspace boundaries and local reference adapters before distributed consensus.

Initial mission statement:

EHDB is an Arrow-native distributed database and catalog platform that stores metadata transactionally and data in multi-cloud object storage, providing a unified foundation for operational, analytical, and historical workloads.

Next actions:

  • Create and track bootstrap/design issues in noetl/ehdb.
  • Scaffold Rust crates and local tests.
  • Keep the wiki current as the code shape changes.

Issues opened:

2026-06-20: Domain-Specific NoETL Storage Direction

Direction update:

  • EHDB should be a NoETL-domain-specific storage system, not a generic database that happens to serve NoETL.
  • EHDB should grow to absorb the NoETL platform roles currently served by PostgreSQL, NATS JetStream, external object stores, Qdrant, and ClickHouse.
  • RAG support is first-class: documents, chunks, embedding metadata, vector index metadata, retrieval policies, tenant context, and execution lineage belong in the EHDB model.
  • NATS JetStream functionality should be represented as EHDB-native stream logs, durable consumers, replay cursors, ack state, and retention policy. A NATS bridge may exist for migration, but the target state is EHDB-owned stream durability.

Architectural implication:

NoETL should see EHDB as the storage system. Cloud object APIs, vector-index files, analytical columnar layouts, and stream persistence may exist internally, but they should not remain separate permanent NoETL platform dependencies.

Tracking issue:

2026-06-20: Stream And Retrieval Foundation Branch

Implementation branch:

  • kadyapam/ehdb-stream-retrieval-foundation

Implemented first local reference models:

  • ehdb-stream: typed NoETL stream records, subjects, retention policy, durable consumers, replay cursors, and ack state.
  • ehdb-retrieval: documents, chunks, embedding metadata, model identity, dimension validation, and tenant/namespace-scoped text lookup fixtures.
  • ehdb-transaction: replayable transaction records for catalog, stream, and retrieval mutations, with duplicate transaction ID rejection and ordered replay.
  • ehdb-core: shared stream, consumer, document, chunk, and embedding model identifiers.
  • Integration coverage ties catalog create-table, stream publish, and retrieval registration through one transaction log.

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 25 Rust tests across unit, integration, and doc-test targets.
  • Edge coverage includes duplicate catalog/stream/consumer/document/ chunk/embedding records, missing lookup behavior, unsafe object paths, invalid subjects, stream retention, cursor replay, ack monotonicity, embedding dimension validation, tenant/namespace retrieval isolation, duplicate transaction IDs, empty transactions, and replay after cursor.

Benchmark baseline:

Benchmark Workload Baseline
stream_publish_replay_1000 1000 stream publishes + full replay ~466 us
transaction_append_replay_1000 1000 transaction appends + full replay ~843 us

This is intentionally pre-distributed and pre-service. The point is to make the NoETL-native replacement surfaces concrete before adding network protocols, consensus, or persistence adapters.

2026-06-21: Local Durable Transaction Log

Implementation branch:

  • kadyapam/ehdb-local-persistence-foundation

Implemented Phase 3 local durability:

  • Added serde support for EHDB durable identifiers and transaction records.
  • Added LocalJsonlTransactionLog, a fsynced append-only JSONL adapter for local developer and crash/restart tests.
  • Reused in-memory duplicate transaction ID and contiguous sequence checks while rebuilding state from disk.
  • Added restart, duplicate-after-reopen, and corrupt-record tests.
  • Added a Criterion benchmark for the durable local append/reopen path.

Tracking issue:

Validation:

  • cargo fmt --all
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 28 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
stream_publish_replay_1000 1000 stream publishes + full replay ~489 us
transaction_append_replay_1000 1000 transaction appends + full replay ~1.04 ms
local_transaction_jsonl/append_reopen_100 100 fsynced JSONL appends + reopen + full replay ~456 ms

This remains a local reference implementation. Production replicated metadata durability still belongs behind a consensus-backed transaction log boundary.

2026-06-21: Local Durable Stream Journal

Implementation branch:

  • kadyapam/ehdb-local-stream-journal

Implemented Phase 3.5 local stream durability:

  • Added serde support for EHDB stream configs, subjects, sequences, records, retention policies, and durable consumers.
  • Added LocalJsonlStreamLog, a fsynced append-only JSONL stream journal for local developer and restart tests.
  • Journaled create-stream, create-consumer, publish, and ack operations.
  • Rebuilt retained records, durable consumer ack cursors, and next sequence on open.
  • Added restart/cursor, retention-after-reopen, and corrupt-entry tests.
  • Added a Criterion benchmark for durable local stream publish/reopen.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 31 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
stream_publish_replay_1000 1000 stream publishes + full replay ~629 us
transaction_append_replay_1000 1000 transaction appends + full replay ~1.04 ms
local_transaction_jsonl/append_reopen_100 100 fsynced JSONL appends + reopen + full replay ~454 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~457 ms

This remains a local reference implementation. Production stream replication and any NATS compatibility bridge should plug in behind the stream boundary.

2026-06-21: System WASM Library Registry

Implementation branch:

  • kadyapam/ehdb-system-wasm-library-registry

Implemented Phase 3.75 system-library cataloging:

  • Added ehdb-system for NoETL system WASM library manifests and environment/channel bindings.
  • Modeled immutable module manifests with logical path, revision, digest, entry export, target, object path, byte length, capability grants, and transaction provenance.
  • Modeled mutable bindings by tenant, namespace, environment, release channel, and logical path.
  • Added deterministic resolution of the active module for a requested environment/channel/path.
  • Added hot replacement by rebinding a stable channel to a new digest/revision while retaining prior immutable manifests.
  • Added transaction-log SystemMutation records for publish/bind events.
  • Extended the cross-domain NoETL surface test to include system library publish/bind replay alongside catalog, stream, and retrieval mutations.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 36 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
stream_publish_replay_1000 1000 stream publishes + full replay ~626 us
transaction_append_replay_1000 1000 transaction appends + full replay ~1.04 ms
local_transaction_jsonl/append_reopen_100 100 fsynced JSONL appends + reopen + full replay ~448 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~456 ms

This is the catalog/storage side of the NoETL WASM plug-in model. EHDB does not embed Wasmtime in this slice; worker/system-pool execution and host capability enforcement remain separate runtime concerns.

2026-06-21: Local Durable System Library Journal

Implementation branch:

  • kadyapam/ehdb-system-library-journal

Implemented restartable system-library registry state:

  • Added LocalJsonlSystemLibraryCatalog, a fsynced append-only JSONL journal for system WASM library publish and bind operations.
  • Rebuilt immutable module manifests and mutable environment/channel bindings on open.
  • Preserved hot-replacement behavior across restart by replaying stable channel rebindings to new digest/revision values.
  • Added restart, hot-replacement-after-reopen, and corrupt-record tests.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 39 Rust tests across unit, integration, and doc-test targets.

This remains a local reference implementation. Production replicated system-library registry durability belongs behind the same EHDB transaction/replication boundary as catalog metadata.

2026-06-21: Replay-Complete Transaction Mutations

Implementation branch:

  • kadyapam/ehdb-replay-complete-mutations

Implemented reconstructive transaction replay:

  • Enabled serde support for Arrow schema datatypes and EHDB table schemas.
  • Expanded transaction mutations so replay carries the facts needed to rebuild reference state:
    • catalog create-table includes schema;
    • stream publish includes subject, payload, and expected sequence;
    • retrieval mutations include document source metadata, chunk text and checksum, and embedding vectors;
    • system library publish includes the full WASM module manifest.
  • Added ehdb-reference, a replay applier over the local catalog, stream, retrieval, and system-library reference models.
  • Added coverage proving transaction replay rebuilds all current reference domains from the log alone.
  • Added mismatch detection for durable stream sequence values that do not match reconstructed state.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 41 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
stream_publish_replay_1000 1000 stream publishes + full replay ~626 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.21 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~488 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~486 ms

This moves the reference model closer to the NoETL execution boundary: the transaction log is the durable source of truth, while local catalogs are reconstructable projections.

2026-06-21: Local Reference Runtime

Implementation branch:

  • kadyapam/ehdb-local-reference-runtime

Implemented the first local runtime boundary over replay-complete transactions:

  • Added LocalReferenceRuntime, which opens a local JSONL transaction log and rebuilds reference state by replaying transaction records.
  • Added projection validation before durable append: the runtime previews the transaction record, applies it to cloned reference state, and only appends when the projection succeeds.
  • Kept invalid projected commits from advancing the durable JSONL log.
  • Added clone support to the in-memory catalog, stream, and retrieval reference stores so projected state can be validated without mutating the live state.
  • Hardened local storage temp paths used by tests to avoid collisions in fast repeated runs.
  • Added a Criterion benchmark for local runtime append/reopen.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 43 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
stream_publish_replay_1000 1000 stream publishes + full replay ~626 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.18 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~486 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~482 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~477 ms

This makes the local developer runtime stricter: the transaction log is only advanced after the current reference projections agree the mutation set is valid, while restart still treats replay as the source of truth.

2026-06-21: Content-Checked Immutable Objects

Implementation branch:

  • kadyapam/ehdb-content-addressed-objects

Implemented the first content-checked object reference boundary:

  • Added ObjectDigest as a SHA-256 digest wrapper for stored object bytes.
  • Extended ObjectRef to carry path, byte length, and digest.
  • Added ImmutableObjectStore::get_verified, which reads through an object reference and rejects length or digest mismatches.
  • Added a deterministic table/snapshot object path helper: {tenant}/{namespace}/tables/{table}/snapshots/{snapshot}/{file}.
  • Added corruption detection and path-layout coverage for the local object-store adapter.
  • Added a Criterion benchmark for local object put plus verified read.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 46 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~14.8 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~640 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.17 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~476 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~454 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~452 ms

This closes an important Phase 2 gap: local object references are now content-checked and deterministic enough to be attached to future catalog snapshots and manifests.

2026-06-21: Catalog Table Snapshots

Implementation branch:

  • kadyapam/ehdb-catalog-snapshots

Implemented immutable table snapshot metadata in the catalog reference:

  • Added CatalogSnapshot and CommitSnapshot with snapshot ID, parent snapshot, content-checked object file refs, and committing transaction ID.
  • Added latest snapshot tracking per table.
  • Rejected missing tables, empty file sets, duplicate snapshots, and parent-chain mismatches.
  • Added CatalogMutation::CommitSnapshot so snapshot metadata is replayable through the transaction log.
  • Extended ehdb-reference and LocalReferenceRuntime coverage so replay rebuilds catalog snapshot state.
  • Added a Criterion benchmark for catalog snapshot commits.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 49 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~2.06 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~16.9 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~656 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.37 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~491 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~508 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~534 ms

This connects the Phase 1 catalog model and Phase 2 object layer: content-checked object references can now become durable table snapshot metadata, and the latest snapshot projection is reconstructable from the transaction log.

2026-06-21: Geo Placement And Data-Gravity Shard Pointers

Implementation branch:

  • kadyapam/ehdb-placement-pointers

Added explicit distributed-storage placement pointers to object references before implementing production replication:

  • Added CloudProvider, GeoLocation, DataGravityShard, and ObjectPlacement to ehdb-storage.
  • Extended ObjectRef so catalog snapshots carry path, byte length, digest, geo placement, and data-gravity shard metadata together.
  • Kept the local filesystem adapter on deterministic local-dev placement.
  • Added validation coverage for geo region/zone and data-gravity shard identifiers.
  • Documented that these are storage-layer routing pointers for future read/write nodes, replicators, and placement planners, not a gateway data-touch path.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 50 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~1.94 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~14.6 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~613 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.14 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~486 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~571 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~517 ms

2026-06-21: Storage Placement Policy

Implementation branch:

  • kadyapam/ehdb-placement-policy

Added a declarative placement policy model over geo-location and data-gravity shard metadata:

  • Added PlacementRole, PlacementTarget, and PlacementPolicy to ehdb-storage.
  • Enforced exactly one primary placement per policy.
  • Enforced a minimum copy count before a policy is accepted.
  • Enforced one shared data-gravity shard across all targets.
  • Rejected duplicate geo/shard targets.
  • Added deterministic local-dev placement policy support.
  • Added validation tests and a Criterion benchmark for policy validation.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 53 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
placement_policy_validate_1000 1000 three-target placement policy validations ~1.19 ms
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~2.12 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~14.7 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~614 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.16 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~443 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~542 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~534 ms

This turns placement pointers into an explicit contract for future replication planners while keeping execution aligned with NoETL's gateway/worker/playbook boundary.

2026-06-21: Deterministic Replication Planning

Implementation branch:

  • kadyapam/ehdb-replication-plan

Added a deterministic replication planning model over placement policy:

  • Added ObjectReplica, ReplicationAction, and ReplicationPlan.
  • Added plan_replication, which compares a source object and known replicas with a PlacementPolicy.
  • Emits AlreadySatisfied actions for existing policy targets.
  • Emits CopyNeeded actions for missing policy targets.
  • Rejects source/policy data-gravity shard mismatches.
  • Rejects known replicas with mismatched digest, length, or shard.
  • Added unit coverage and a Criterion benchmark for the planner.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 56 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
replication_plan_1000 1000 three-target replication plans ~2.92 ms
placement_policy_validate_1000 1000 three-target placement policy validations ~1.18 ms
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~1.98 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~15.7 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~626 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.16 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~516 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~554 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~540 ms

This gives EHDB a planner-level contract for future bounded replication workers while keeping object copy execution out of the gateway and out of this reference model.

2026-06-22: Durable Object Replica Registry

Implementation branch:

  • kadyapam/ehdb-replica-registry

Design target:

  • Add an EHDB-owned object replica registry so replication plans are derived from durable metadata instead of caller-supplied arrays.
  • Record object path, byte length, digest, placement, and data-gravity shard for each available copy.
  • Reject conflicting digest, length, or shard metadata for the same object.
  • Make replica registration replayable through the transaction log and local reference runtime.
  • Keep copy execution out of the registry. Replication remains a future bounded worker/playbook operation, not gateway behavior.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 61 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
replication_plan_from_registry_1000 1000 three-target replication plans from registry state ~3.82 ms
replication_plan_1000 1000 three-target replication plans ~2.95 ms
replica_registry_register_1000 1000 object replica registrations ~1.10 ms
placement_policy_validate_1000 1000 three-target placement policy validations ~1.17 ms
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~1.98 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~16.0 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~635 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.16 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~509 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~507 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~516 ms

This makes replica inventory replayable EHDB metadata while keeping object-copy execution out of gateway request handling.

2026-06-22: Bounded Local Replication Executor

Implementation branch:

  • kadyapam/ehdb-local-replication-executor

Design target:

  • Consume deterministic replication plans from a bounded local executor.
  • Verify source object bytes through the immutable object-store API before recording replica success.
  • Append StorageMutation::RegisterReplica mutations for copy-needed targets through LocalReferenceRuntime.
  • Treat already-satisfied plans as no-op work.
  • Keep scheduling, long-lived processes, cloud transfer adapters, and gateway data-touch behavior out of scope.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 64 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
local_replication_executor/register_25 25 verified source reads + fsynced replica-registration transactions + reopen ~139 ms
replication_plan_from_registry_1000 1000 three-target replication plans from registry state ~3.75 ms
replication_plan_1000 1000 three-target replication plans ~2.81 ms
replica_registry_register_1000 1000 object replica registrations ~1.08 ms
placement_policy_validate_1000 1000 three-target placement policy validations ~1.15 ms
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~2.00 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~20.2 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~630 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.17 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~441 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~541 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~523 ms

This is the first local execution reference for replication work. It models an atomic worker/playbook step and keeps scheduling, cloud copy adapters, and gateway data-touch behavior outside the slice.

2026-06-22: Arrow IPC Table Write/Read Fixture

Implementation branch:

  • kadyapam/ehdb-arrow-ipc-fixture

Design target:

  • Write Arrow RecordBatch values into immutable object storage as IPC files.
  • Commit catalog snapshots over content-checked ObjectRef file sets.
  • Read the latest snapshot back through verified object reads and Arrow IPC decoding.
  • Preserve tenant, namespace, table, snapshot, and transaction provenance in the reference runtime.
  • Keep Arrow Flight, distributed query execution, Parquet adapters, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 66 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
local_arrow_ipc_table/write_read_10 10 Arrow IPC write + catalog snapshot + verified read cycles ~123 ms
local_replication_executor/register_25 25 verified source reads + fsynced replica-registration transactions + reopen ~155 ms
replication_plan_from_registry_1000 1000 three-target replication plans from registry state ~3.76 ms
replication_plan_1000 1000 three-target replication plans ~2.82 ms
replica_registry_register_1000 1000 object replica registrations ~1.08 ms
placement_policy_validate_1000 1000 three-target placement policy validations ~1.13 ms
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~1.96 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~15.0 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~609 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.17 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~530 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~491 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~502 ms

This proves the local catalog/object boundary for Arrow-native analytical data while leaving Arrow Flight and distributed query execution as explicit future surfaces.

2026-06-22: Arrow Snapshot Scan Fixture

Implementation branch:

  • kadyapam/ehdb-arrow-scan-fixture

Design target:

  • Scan the latest catalog snapshot through verified Arrow IPC object reads.
  • Return decoded Arrow RecordBatch output.
  • Support optional named column projection in caller-specified order.
  • Fail deterministically for missing projection columns.
  • Keep predicate pushdown, SQL planning, Arrow Flight, distributed query execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 68 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
local_arrow_scan/project_latest_100 100 verified latest-snapshot scans with two-column projection ~13.1 ms
local_arrow_ipc_table/write_read_10 10 Arrow IPC write + catalog snapshot + verified read cycles ~107 ms
local_replication_executor/register_25 25 verified source reads + fsynced replica-registration transactions + reopen ~133 ms
replication_plan_from_registry_1000 1000 three-target replication plans from registry state ~3.86 ms
replication_plan_1000 1000 three-target replication plans ~2.97 ms
replica_registry_register_1000 1000 object replica registrations ~1.12 ms
placement_policy_validate_1000 1000 three-target placement policy validations ~1.21 ms
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~2.06 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~16.4 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~793 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.22 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~562 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~443 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~517 ms

This moves the local analytical path from write/read proof to a minimal scan operation while leaving predicates, SQL planning, Flight, and distributed execution out of scope.

2026-06-22: Arrow Equality Filter Fixture

Implementation branch:

  • kadyapam/ehdb-arrow-filter-fixture

Design target:

  • Extend local Arrow snapshot scanning with optional single-column equality predicates.
  • Support UTF-8 and Int64 equality first.
  • Apply filtering after verified object reads and Arrow IPC decode, and before optional projection.
  • Fail deterministically for missing predicate columns and unsupported predicate/column type combinations.
  • Keep SQL planning, predicate pushdown, object statistics, Arrow Flight, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-reference --bench local_runtime
  • cargo bench -p ehdb-storage --bench local_store
  • cargo bench -p ehdb-catalog --bench snapshots
  • cargo bench -p ehdb-transaction --bench reference_models

Coverage snapshot:

  • 72 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
local_arrow_scan/filter_project_latest_100 100 verified latest-snapshot scans with equality filter and two-column projection ~13.0 ms
local_arrow_ipc_table/write_read_10 10 Arrow IPC write + catalog snapshot + verified read cycles ~111 ms
local_replication_executor/register_25 25 verified source reads + fsynced replica-registration transactions + reopen ~159 ms
replication_plan_from_registry_1000 1000 three-target replication plans from registry state ~3.77 ms
replication_plan_1000 1000 three-target replication plans ~2.89 ms
replica_registry_register_1000 1000 object replica registrations ~1.09 ms
placement_policy_validate_1000 1000 three-target placement policy validations ~1.36 ms
catalog_commit_snapshots_1000 1000 catalog snapshot commits + latest lookup ~2.03 ms
local_object_store/put_get_verified_100 100 immutable 4 KiB local object puts + verified reads ~15.6 ms
stream_publish_replay_1000 1000 stream publishes + full replay ~640 us
transaction_append_replay_1000 1000 replay-complete transaction appends + full replay ~1.15 ms
local_reference_runtime/append_reopen_100 create stream + 100 projection-validated fsynced transaction appends + reopen + replay ~543 ms
local_transaction_jsonl/append_reopen_100 100 fsynced replay-complete JSONL appends + reopen + full replay ~461 ms
local_stream_jsonl/publish_reopen_100 100 fsynced stream publishes + reopen + full replay ~464 ms

This adds a first local predicate path while keeping SQL planning, predicate pushdown, Flight, distributed execution, and gateway reads out of scope.

2026-06-22: Arrow Scan Service Boundary

Implementation branch:

  • kadyapam/ehdb-arrow-scan-service-boundary

Design target:

  • Add the first Phase 4 service-facing scan API boundary.
  • Create ehdb-service with typed latest-table scan request/result structures.
  • Wrap LocalArrowSnapshotScanner in LocalArrowScanService.
  • Return Arrow schema, batches, and row count from scan results.
  • Preserve projection and equality-filter behavior from the local scanner.
  • Keep Arrow Flight networking, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run
  • cargo bench -p ehdb-service --bench local_scan_service

Coverage snapshot:

  • 76 Rust tests across unit, integration, and doc-test targets.

Benchmark baseline:

Benchmark Workload Baseline
local_arrow_scan_service/filter_project_latest_100 100 service-boundary latest-snapshot scans with equality filter and two-column projection ~12.0 ms

This establishes the service API shape that a future Arrow Flight server can expose while keeping data-touch work behind EHDB-owned APIs and out of the NoETL gateway.

2026-06-22: Arrow Flight Scan Ticket Codec

Implementation branch:

  • kadyapam/ehdb-arrow-flight-ticket-codec

Design target:

  • Add a versioned Arrow Flight ticket payload for latest-table scans.
  • Preserve tenant, namespace, table, projection, and equality predicate fields in the encoded request.
  • Round-trip through Arrow Flight Ticket bytes.
  • Build command FlightDescriptor values for the future Flight read API.
  • Reject unsupported ticket versions and malformed payloads before scan execution.
  • Keep the Arrow Flight network server, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 82 Rust tests across unit, integration, and doc-test targets.

This gives Phase 4 a concrete Flight request contract without starting a long-lived service or changing the NoETL gateway boundary.

2026-06-22: Arrow Flight Scan Result Stream Codec

Implementation branch:

  • kadyapam/ehdb-arrow-flight-result-codec

Design target:

  • Add a pre-network Arrow Flight result stream codec for local scan outputs.
  • Encode ArrowScanResult batches into Arrow Flight FlightData messages.
  • Decode FlightData messages back into validated ArrowScanResult values with schema and row count.
  • Preserve projected schemas and filtered rows.
  • Reject empty or malformed streams deterministically.
  • Keep the Arrow Flight network server/client, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 86 Rust tests across unit, integration, and doc-test targets.

This proves the local response side of the future Flight do_get path while preserving the NoETL gateway and bounded-worker boundaries.

2026-06-22: Arrow Flight Scan Info Fixture

Implementation branch:

  • kadyapam/ehdb-arrow-flight-info-fixture

Design target:

  • Add a pre-network Arrow Flight FlightInfo fixture for latest-table scans.
  • Build schema IPC bytes from the scan result schema.
  • Attach the command descriptor generated from ScanFlightTicket.
  • Return one ordered endpoint with the encoded scan ticket.
  • Report total records and encoded FlightData byte count.
  • Keep the Arrow Flight network server/client, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 88 Rust tests across unit, integration, and doc-test targets.

This proves the future get_flight_info metadata shape without starting a network service or changing NoETL gateway data-touch boundaries.

2026-06-22: Local Arrow Flight Scan Service Facade

Implementation branch:

  • kadyapam/ehdb-local-flight-service-facade

Design target:

  • Add an in-process local facade for future Arrow Flight scan behavior.
  • Provide get_flight_info from typed latest-table scan requests.
  • Provide do_get from Arrow Flight tickets to FlightData result streams.
  • Reuse ScanFlightTicket, ArrowScanResult::to_flight_info, and ArrowScanResult::to_flight_data.
  • Reject malformed tickets before scan execution and propagate missing table errors deterministically.
  • Keep the Arrow Flight network server/client, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 91 Rust tests across unit, integration, and doc-test targets.

This moves Phase 4 from static protocol fixtures to local service behavior while still avoiding a long-lived network process.

2026-06-22: Arrow Flight Scan Service Trait Adapter

Implementation branch:

  • kadyapam/ehdb-flight-service-trait-adapter

Design target:

  • Implement the generated Arrow Flight FlightService trait for the local scan path.
  • Accept command FlightDescriptor requests for get_flight_info by decoding the existing ScanFlightTicket payload.
  • Accept Arrow Flight Ticket requests for do_get and stream FlightData through the tonic response type.
  • Map EHDB errors to deterministic gRPC statuses.
  • Return explicit UNIMPLEMENTED statuses for non-scan Flight methods.
  • Keep port binding, long-lived server runtime lifecycle, TLS/auth, request concurrency policy, access-log policy, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 95 Rust tests across unit, integration, and doc-test targets.

This establishes the first network-facing trait boundary while still avoiding a bound listener or persistent service process.

2026-06-22: Arrow Flight Server Lifecycle Config

Implementation branch:

  • kadyapam/ehdb-flight-server-lifecycle-config

Design target:

  • Add a validated lifecycle/config surface for the local Arrow Flight service adapter.
  • Track intended bind address, max decode/encode message sizes, max concurrent requests, auth policy, and access-log policy.
  • Default to loopback local-reference use with bounded message sizes, bounded concurrency, disabled local auth, and DEBUG-only access logs.
  • Reject zero bounds and unauthenticated non-loopback binds.
  • Build the generated FlightServiceServer with message limits applied.
  • Keep actual socket binding, long-lived runtime management, TLS/auth implementation, request scheduling, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 99 Rust tests across unit, integration, and doc-test targets.

This adds server lifecycle guardrails before EHDB exposes any bound Flight listener.

2026-06-22: Loopback Arrow Flight Listener Harness

Implementation branch:

  • kadyapam/ehdb-loopback-flight-listener

Design target:

  • Add a loopback-only listener harness behind LocalArrowFlightServerConfig.
  • Bind a configured or ephemeral loopback socket and expose the actual bound local address.
  • Serve the generated Arrow Flight service with configured message limits.
  • Terminate cleanly through an explicit shutdown future.
  • Reject non-loopback listener binds even when external auth policy is selected.
  • Keep non-loopback service exposure, TLS/auth implementation, gateway integration, request scheduling, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 101 Rust tests across unit, integration, and doc-test targets.

This proves a bounded local listener lifecycle without exposing EHDB as a networked platform service yet.

2026-06-22: Loopback Arrow Flight Client Smoke Test

Implementation branch:

  • kadyapam/ehdb-loopback-flight-client-smoke

Design target:

  • Start LocalArrowFlightListener on loopback with explicit shutdown.
  • Connect with the Arrow Flight client over real tonic/gRPC transport.
  • Call get_flight_info using the existing scan command descriptor.
  • Follow the returned endpoint ticket with do_get.
  • Decode the returned Arrow record batches and assert filtered rows.
  • Keep this local-reference only: no gateway integration, non-loopback exposure, TLS/auth implementation, request scheduler, SQL planning, predicate pushdown, distributed execution, or gateway direct data access.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 102 Rust tests across unit, integration, and doc-test targets.

This verifies the first end-to-end Flight read path over local gRPC transport while preserving the NoETL execution-model boundary.

2026-06-22: Arrow Flight Auth Header Policy Contract

Implementation branch:

  • kadyapam/ehdb-flight-auth-header-policy

Design target:

  • Extend FlightAuthPolicy with a validated header-token mode for the local Arrow Flight reference service.
  • Reject invalid metadata header names, binary metadata headers, empty tokens, oversized tokens, and control-character tokens at config validation time.
  • Enforce request metadata auth on implemented scan methods: get_flight_info and do_get.
  • Keep default local-reference behavior unauthenticated and loopback only.
  • Prove the policy directly through the generated service trait adapter and over the loopback listener with the Arrow Flight client.
  • Keep this as an auth-boundary contract only: no non-loopback exposure, production TLS/identity, ACL enforcement, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 106 Rust tests across unit, integration, and doc-test targets.

This gives the local Flight harness an explicit request metadata boundary before EHDB grows a production service exposure model.

2026-06-22: Arrow Flight Scan Scope Metadata Guard

Implementation branch:

  • kadyapam/ehdb-flight-scope-metadata-guard

Design target:

  • Add FlightScanScopePolicy for tenant/namespace scope metadata on implemented Arrow Flight scan calls.
  • Keep default local-reference behavior scope-disabled for existing in-process tests and local harnesses.
  • When enabled, require x-ehdb-tenant and x-ehdb-namespace metadata to match the decoded ScanLatestTableRequest before local scan execution.
  • Return UNAUTHENTICATED for missing scan scope metadata and PERMISSION_DENIED for mismatched scope metadata.
  • Prove the guard directly through the generated service trait adapter and over the loopback listener with the Arrow Flight client.
  • Keep this as a scope-boundary contract only: no ACL engine, non-loopback exposure, production TLS/identity, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 110 Rust tests across unit, integration, and doc-test targets.

This gives future catalog ACL work a concrete tenant/namespace request scope contract without changing the NoETL execution-model boundary.

2026-06-22: Catalog Scan Grant Reference Model

Implementation branch:

  • kadyapam/ehdb-catalog-scan-grants

Design target:

  • Add PrincipalId as a typed EHDB identifier.
  • Add CatalogScanGrant and GrantScan to the local catalog reference.
  • Record tenant, namespace, table ID, principal, and granting transaction ID for scan grants.
  • Reject scan grants for missing tables and duplicate table/principal grants.
  • Add InMemoryCatalog::can_scan and scan_grant_count for future service authorization checks.
  • Add CatalogMutation::GrantScan so grants are replayable through the transaction log, ehdb-reference, and LocalReferenceRuntime.
  • Keep this as catalog ACL metadata only: no production IAM, policy composition, revocation, service enforcement, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 113 Rust tests across unit, integration, and doc-test targets.

This adds durable catalog-side scan authorization metadata for future ACL enforcement while preserving the NoETL execution-model boundary.

2026-06-22: Flight Catalog Scan Grant Enforcement

Implementation branch:

  • kadyapam/ehdb-flight-catalog-scan-grants

Design target:

  • Add FlightScanGrantPolicy for catalog-backed local Arrow Flight scan authorization.
  • Keep default local-reference behavior grant-disabled for existing harnesses.
  • When enabled, require x-ehdb-principal metadata on implemented Arrow Flight scan calls.
  • Validate the metadata value as a PrincipalId, resolve the requested table from replayed catalog state, and call InMemoryCatalog::can_scan before local scan execution.
  • Return UNAUTHENTICATED for missing or invalid principal metadata and PERMISSION_DENIED when the principal lacks a replayed CatalogScanGrant for the requested table.
  • Prove the guard directly through the generated service trait adapter and over the loopback listener with the Arrow Flight client.
  • Preserve the previous new_with_policies constructor shape and add a fuller constructor for auth, scan-scope, and scan-grant policy combinations.
  • Keep this as local reference enforcement only: no production IAM, policy composition, revocation, non-loopback exposure, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 117 Rust tests across unit, integration, and doc-test targets.

This connects replayed catalog ACL metadata to the local Flight read path while keeping gateway as gatekeeper and EHDB scans inside explicit service/playbook boundaries.

2026-06-22: Bounded Flight Scan Access Logs

Implementation branch:

  • kadyapam/ehdb-flight-access-log-policy

Design target:

  • Add bounded local Arrow Flight scan access summaries for decoded scan requests.
  • Keep default local-reference behavior DEBUG-only, never INFO-level.
  • Add FlightScanAccessLogEntry with method, gRPC code, row/message counts, projection count, predicate presence, and which metadata guards were required.
  • Add disabled mode that emits no scan access summaries.
  • Exclude auth tokens, principal values, tenant/table identifiers, object paths, predicate values, and Arrow payloads from the summary contract.
  • Wire the configured policy through LocalArrowFlightServerConfig into implemented get_flight_info and do_get paths.
  • Preserve existing constructor shapes and add a fuller runtime-policy constructor for auth, scope, grant, and access-log combinations.
  • Keep this as local reference observability only: no non-loopback exposure, production IAM, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 120 Rust tests across unit, integration, and doc-test targets.

This completes the first log-hygiene contract for the local Flight read path without introducing high-volume INFO logs or tenant data leakage.

2026-06-22: Arrow Flight Schema Adapter

Implementation branch:

  • kadyapam/ehdb-flight-get-schema-adapter

Design target:

  • Add get_schema support for local Arrow Flight scan command descriptors.
  • Reuse the existing versioned ScanFlightTicket command descriptor contract.
  • Add LocalArrowFlightService::get_schema to produce Arrow Flight SchemaResult values from the projected latest-table scan schema.
  • Implement generated FlightService::get_schema with the same request metadata auth, tenant/namespace scan scope, catalog scan grant, and bounded access-log policies used by get_flight_info and do_get.
  • Extend the loopback client smoke path to call FlightClient::get_schema over tonic/gRPC before reading data.
  • Keep this as local reference schema discovery only: no non-loopback exposure, production IAM, gateway direct reads, SQL planning, predicate pushdown, distributed execution, request scheduler, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 120 Rust tests across unit, integration, and doc-test targets.

This completes the first schema-discovery endpoint for the local Flight read path without changing the NoETL gateway/worker/playbook boundary.

2026-06-22: Flight Request Concurrency Guard

Implementation branch:

  • kadyapam/ehdb-flight-concurrency-guard

Design target:

  • Enforce LocalArrowFlightServerConfig::max_concurrent_requests for implemented local Arrow Flight scan calls.
  • Add a fail-fast semaphore to LocalArrowFlightServer and pass the configured request budget when building services from config.
  • Preserve existing constructors by defaulting them to the local reference request budget.
  • Return deterministic gRPC RESOURCE_EXHAUSTED when all local request slots are occupied.
  • Cover get_flight_info, get_schema, and do_get.
  • Keep this as a local reference lifecycle guard only: no request queue, scheduler, non-loopback exposure, production IAM, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 121 Rust tests across unit, integration, and doc-test targets.

This makes the validated local concurrency budget executable without introducing a long-lived scheduler or changing the NoETL execution-model boundary.

2026-06-22: Local Retrieval Vector Similarity Fixture

Implementation branch:

  • kadyapam/ehdb-retrieval-vector-search

Design target:

  • Add a typed VectorSearch request and VectorSearchHit result boundary to ehdb-retrieval.
  • Scope candidates by tenant, namespace, and embedding model.
  • Compute exact cosine similarity over registered chunk embeddings.
  • Validate finite non-zero embedding and query vectors.
  • Apply dimension compatibility and deterministic score/tie ordering.
  • Keep this as a local RAG correctness fixture only: no ANN index, retrieval daemon, gateway data path, production IAM, external Qdrant adapter, distributed query engine, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 124 Rust tests across unit, integration, and doc-test targets.

This is the first local vector lookup primitive on replayed EHDB retrieval state, moving the RAG path one step closer to replacing a permanent Qdrant dependency without changing the NoETL execution-model boundary.

2026-06-22: Local Retrieval Search Service Boundary

Implementation branch:

  • kadyapam/ehdb-retrieval-service-boundary

Design target:

  • Add LocalRetrievalSearchService in ehdb-service.
  • Add typed SearchSimilarChunksRequest and SearchSimilarChunksHit service-facing boundaries.
  • Read replayed LocalReferenceRuntime retrieval state and call the exact local VectorSearch fixture.
  • Return ranked chunk identity, document identity, ordinal, text, checksum, embedding model, dimensions, and score while excluding raw embedding vectors from service results.
  • Cover replayed-state search, tenant/namespace/model scoping, score ordering, empty results, and validation propagation.
  • Keep this as an in-process reference boundary only: no network service, gateway route, production IAM, ANN index, external Qdrant adapter, distributed query engine, retrieval daemon, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 126 Rust tests across unit, integration, and doc-test targets.

This gives NoETL a stable local RAG lookup boundary for future worker/playbook integration while preserving the gateway/worker/playbook execution model.

2026-06-22: Local Retrieval Text Search Boundary

Implementation branch:

  • kadyapam/ehdb-retrieval-text-service

Design target:

  • Add typed TextSearch and TextSearchHit boundaries to ehdb-retrieval.
  • Validate non-empty queries and positive limits.
  • Scope exact case-insensitive substring matching by tenant and namespace.
  • Rank by match count with deterministic document/ordinal/chunk tie ordering.
  • Add SearchTextChunksRequest, SearchTextChunksHit, and LocalRetrievalSearchService::search_text in ehdb-service.
  • Cover local catalog search, replayed-state service search, limits, empty results, and validation propagation.
  • Keep this as an in-process reference boundary only: no full-text index, BM25 engine, network service, gateway route, production IAM, external search adapter, distributed query engine, retrieval daemon, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 130 Rust tests across unit, integration, and doc-test targets.

This complements exact vector lookup with text retrieval semantics while keeping the first RAG service surface local and worker/playbook-shaped.

2026-06-22: Local Hybrid Retrieval Search Boundary

Implementation branch:

  • kadyapam/ehdb-retrieval-hybrid-search

Design target:

  • Add typed HybridSearch and HybridSearchHit boundaries to ehdb-retrieval.
  • Validate vector query, text query, positive limit, finite non-negative weights, and at least one positive weight.
  • Scope candidates by tenant, namespace, embedding model, and vector dimension.
  • Combine exact cosine similarity and exact text match counts with caller-provided weights.
  • Add SearchHybridChunksRequest, SearchHybridChunksHit, and LocalRetrievalSearchService::search_hybrid in ehdb-service.
  • Cover local catalog search, replayed-state service search, limits, empty results, deterministic ordering, and validation propagation.
  • Keep this as an in-process reference boundary only: no ANN index, full-text index, BM25 engine, query planner, network service, gateway route, production IAM, external Qdrant/search adapter, distributed query engine, retrieval daemon, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 135 Rust tests across unit, integration, and doc-test targets.

This gives the local RAG surface a deterministic hybrid scoring fixture without changing the NoETL execution-model boundary.

2026-06-22: Local Retrieval Context Assembly Boundary

Implementation branch:

  • kadyapam/ehdb-retrieval-context-assembly

Design target:

  • Add AssembleRetrievalContextRequest, RetrievalContextBlock, and RetrievalContext to ehdb-service.
  • Build context blocks from the existing replayed local hybrid search path, preserving ordering and score metadata.
  • Enforce positive hit limits through hybrid search plus positive per-block and total text budgets at context assembly.
  • Return chunk id, document id, ordinal, checksum, model id, dimensions, vector score, text match count, combined score, clipped text, original text length, and truncation metadata.
  • Cover ordered assembly, budget clipping, empty results, validation, and tenant/namespace scoping.
  • Keep this as an in-process worker/playbook-shaped boundary only: no ANN index, BM25 engine, learned ranker, prompt template engine, LLM invocation, network service, gateway route, production IAM, external search adapter, distributed query engine, retrieval daemon, gateway direct data path, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 139 Rust tests across unit, integration, and doc-test targets.

This gives NoETL RAG tests a deterministic local context shape without moving retrieval state or data access into the gateway.

2026-06-22: Local Retrieval Context Payload Codec

Implementation branch:

  • kadyapam/ehdb-retrieval-context-codec

Design target:

  • Add versioned RetrievalContextRequestPayload and RetrievalContextResultPayload wrappers to ehdb-service.
  • Encode and decode local retrieval context assembly requests/results as deterministic JSON bytes.
  • Reject malformed JSON and unsupported request/result payload versions before execution or handoff.
  • Derive serialization for AssembleRetrievalContextRequest, RetrievalContextBlock, and RetrievalContext.
  • Cover request payload round-trip, assembled result payload round-trip, malformed payloads, and unsupported versions.
  • Keep this as a local worker/playbook payload boundary only: no network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 143 Rust tests across unit, integration, and doc-test targets.

This gives NoETL RAG context assembly a stable local serialization boundary without turning retrieval into a standalone service.

2026-06-22: Local Retrieval Context Payload Executor

Implementation branch:

  • kadyapam/ehdb-retrieval-context-executor

Design target:

  • Add LocalRetrievalSearchService::execute_context_payload.
  • Decode versioned retrieval context request payload bytes.
  • Assemble context from replayed LocalReferenceRuntime retrieval state using the existing local context assembly boundary.
  • Encode the assembled context as a versioned result payload.
  • Propagate malformed payload, unsupported version, and invalid search/budget errors deterministically.
  • Cover happy-path payload execution, malformed request payloads, unsupported request versions, invalid assembly inputs, and empty result payloads.
  • Keep this as an in-process worker/playbook payload executor only: no network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 146 Rust tests across unit, integration, and doc-test targets.

This gives NoETL RAG context assembly a complete local payload-in / payload-out handoff without moving retrieval into a standalone service.

2026-06-22: Bounded Local Retrieval Context Payload Execution

Implementation branch:

  • kadyapam/ehdb-retrieval-context-bounds

Design target:

  • Add RetrievalContextPayloadExecutorConfig to ehdb-service.
  • Validate positive max request and max result payload byte limits.
  • Keep LocalRetrievalSearchService::execute_context_payload on a conservative default config.
  • Add config-aware local payload execution for worker/playbook tests.
  • Reject oversized request payloads before JSON decode.
  • Reject oversized encoded result payloads before returning bytes.
  • Cover default config, invalid config, oversized request payloads, oversized result payloads, and happy-path configured execution.
  • Keep this as an in-process worker/playbook payload guard only: no network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 150 Rust tests across unit, integration, and doc-test targets.

This bounds the local RAG payload handoff without making retrieval a service or moving data access into the gateway.

2026-06-22: Local Retrieval Context Payload Scope Guard

Implementation branch:

  • kadyapam/ehdb-retrieval-context-scope

Design target:

  • Add RetrievalContextPayloadScope to ehdb-service.
  • Validate decoded context assembly requests against an expected tenant and namespace before context assembly.
  • Add LocalRetrievalSearchService::execute_context_payload_with_scope that composes existing byte bounds with the local scope check.
  • Keep default and config-aware payload execution unchanged.
  • Cover matching scope, tenant mismatch, namespace mismatch, malformed payload propagation, and oversized request propagation.
  • Keep this as an in-process worker/playbook correctness guard only: no production IAM, policy engine, ACL integration, network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 153 Rust tests across unit, integration, and doc-test targets.

This adds local tenant/namespace guardrails for RAG context payload execution without turning them into production authorization.

2026-06-22: Redacted Local Retrieval Context Execution Summaries

Implementation branch:

  • kadyapam/ehdb-retrieval-context-summary

Design target:

  • Add RetrievalContextPayloadExecutionSummary to ehdb-service.
  • Add RetrievalContextPayloadExecution for summary-returning local payload execution APIs.
  • Keep existing byte-returning default, config-aware, and scope-aware execution APIs behavior-compatible by delegating through the summary path.
  • Report request/result byte counts, context block count, total text chars, truncation status, and whether a local scope guard was required.
  • Explicitly exclude tenant IDs, namespace values, query text, chunk text, tokens, vectors, payload bytes, object paths, and principals from the summary.
  • Cover happy path, truncation metadata, scoped execution metadata, and request/result bound propagation.
  • Keep this as redacted metrics/audit metadata for local worker/playbook tests only: no logging sink, network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 157 Rust tests across unit, integration, and doc-test targets.

This gives the local RAG payload handoff a safe summary shape for future audit/metrics plumbing without exposing retrieval-sensitive content or creating a service boundary.

2026-06-22: Redacted Retrieval Context Execution Receipt Codec

Implementation branch:

  • kadyapam/ehdb-retrieval-context-receipt

Design target:

  • Add RETRIEVAL_CONTEXT_EXECUTION_RECEIPT_VERSION.
  • Add RetrievalContextPayloadExecutionReceiptPayload to wrap RetrievalContextPayloadExecutionSummary in a versioned JSON byte codec.
  • Derive serialization for the redacted summary only.
  • Reject malformed receipt JSON and unsupported receipt versions deterministically.
  • Prove encoded receipts exclude tenant IDs, namespace values, query text, chunk text, tokens, vectors, payload bytes, object paths, and principals.
  • Keep existing retrieval context payload executor behavior unchanged.
  • Keep this as a durable receipt shape for future event-log/audit plumbing only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 160 Rust tests across unit, integration, and doc-test targets.

This makes the local RAG payload summary durable and replay-friendly without turning it into logging, event publication, or a service boundary.

2026-06-22: Local Retrieval Context Execution Receipt Helper

Implementation branch:

  • kadyapam/ehdb-retrieval-context-receipt-helper

Design target:

  • Add RetrievalContextPayloadExecution::encode_receipt_payload.
  • Produce versioned RetrievalContextPayloadExecutionReceiptPayload bytes directly from the redacted execution summary.
  • Keep existing executor and receipt codec behavior unchanged.
  • Prove helper-produced receipts decode to the same summary returned by execution.
  • Prove helper-produced receipt bytes exclude tenant IDs, namespace values, query text, chunk text, tokens, vectors, payload bytes, object paths, and principals.
  • Keep this as local worker/playbook helper wiring only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 162 Rust tests across unit, integration, and doc-test targets.

This lets a local retrieval context execution return result bytes and receipt bytes from one execution object while preserving the redacted receipt boundary.

2026-06-22: Retrieval Context Receipt Summary Validation

Implementation branch:

  • kadyapam/ehdb-retrieval-context-receipt-validation

Design target:

  • Add RetrievalContextPayloadExecutionSummary::validate.
  • Require positive request and result payload byte counts in receipt summaries.
  • Reject non-zero total text chars when context block count is zero.
  • Apply summary validation during RetrievalContextPayloadExecutionReceiptPayload::encode and decode.
  • Keep execution-produced receipts and helper-produced receipts behavior-compatible.
  • Cover valid empty-context receipts, invalid counter combinations, decode-time validation, and helper-produced receipt compatibility.
  • Keep this as local receipt contract hardening only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 165 Rust tests across unit, integration, and doc-test targets.

This prevents malformed hand-built receipt summaries from becoming durable fixtures while preserving the redacted local worker/playbook receipt boundary.

2026-06-23: Bounded Retrieval Context Execution Artifacts

Implementation branch:

  • kadyapam/ehdb-retrieval-context-artifacts

Design target:

  • Add DEFAULT_RETRIEVAL_CONTEXT_MAX_RECEIPT_PAYLOAD_BYTES.
  • Add max_receipt_payload_bytes to RetrievalContextPayloadExecutorConfig with positive validation.
  • Add RetrievalContextPayloadExecutionArtifacts carrying result payload bytes and redacted receipt payload bytes.
  • Add artifact helpers for default, configured, and scope-aware local retrieval context payload execution.
  • Keep existing byte-returning, summary-returning, scope, and receipt APIs behavior-compatible.
  • Cover default config, invalid receipt limit, happy-path artifact emission, scoped artifact emission, and oversized receipt rejection.
  • Keep this as a local worker/playbook handoff shape only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 168 Rust tests across unit, integration, and doc-test targets.

This gives local worker/playbook tests one bounded artifact object for result bytes and redacted receipt bytes without publishing events or opening a service boundary.

2026-06-23: Retrieval Context Execution Artifact Consistency

Implementation branch:

  • kadyapam/ehdb-retrieval-context-artifact-validation

Design target:

  • Add RetrievalContextPayloadExecutionArtifacts::receipt_summary.
  • Add RetrievalContextPayloadExecutionArtifacts::validate.
  • Decode and validate the redacted receipt payload through the existing receipt codec.
  • Reject empty result payload bytes and empty receipt payload bytes.
  • Reject artifacts whose receipt summary result byte count does not match the actual result payload length.
  • Ensure artifact helpers return artifacts that pass consistency validation.
  • Cover valid helper artifacts, malformed receipts, empty payloads, and result-length mismatches.
  • Keep this as local artifact contract hardening only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 171 Rust tests across unit, integration, and doc-test targets.

This makes the local result+receipt artifact pair self-checking before future audit/event plumbing treats it as a coherent handoff.

2026-06-23: Retrieval Context Receipt Event Payload

Implementation branch:

  • kadyapam/ehdb-retrieval-context-receipt-event

Design target:

  • Add RetrievalContextPayloadExecutionReceiptEventPayload.
  • Add the stable subject ehdb.retrieval.context.execution.receipt.
  • Encode and decode a versioned JSON event envelope containing only validated redacted receipt bytes.
  • Build event payloads from validated RetrievalContextPayloadExecutionArtifacts.
  • Reject empty receipt payloads, malformed receipts, and unsupported event envelope versions.
  • Preserve the redaction boundary by excluding result payload/context bytes from the event envelope.
  • Keep this as local stream-ready payload modeling only: no automatic stream publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 174 Rust tests across unit, integration, and doc-test targets.

This gives future EHDB stream/audit plumbing a stable local payload shape while keeping publication under explicit worker/playbook control.

2026-06-23: Retrieval Receipt Event Stream Publisher

Implementation branch:

  • kadyapam/ehdb-retrieval-receipt-event-publisher

Design target:

  • Add RetrievalContextReceiptEventStreamTarget.
  • Add RetrievalContextReceiptEventStreamLog.
  • Implement explicit local publisher support for InMemoryStreamLog and LocalJsonlStreamLog.
  • Publish only caller-supplied validated receipt event payloads to the stable subject ehdb.retrieval.context.execution.receipt.
  • Require caller-owned tenant, namespace, stream name, mutable stream log, and transaction id.
  • Cover in-memory publish/replay, JSONL persist/reopen/replay, missing stream errors, and malformed artifact rejection.
  • Keep this as explicit local worker/playbook publication only: no automatic stream publication, background task, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 177 Rust tests across unit, integration, and doc-test targets.

This gives worker/playbook tests an explicit path to append redacted retrieval receipt events into EHDB streams without making publication implicit or gateway-driven.

2026-06-23: Retrieval Receipt Event Stream Replay

Implementation branch:

  • kadyapam/ehdb-retrieval-receipt-event-replay

Design target:

  • Add RetrievalContextReceiptEventStreamRecord.
  • Add RetrievalContextReceiptEventStreamReadLog.
  • Decode receipt event stream records by validating the stable subject ehdb.retrieval.context.execution.receipt.
  • Validate payloads through RetrievalContextPayloadExecutionReceiptEventPayload.
  • Preserve stream sequence and transaction id for local audit assertions.
  • Replay from caller-supplied InMemoryStreamLog and LocalJsonlStreamLog instances, including cursor replay.
  • Cover ordered replay, cursor replay, JSONL reopen/replay, wrong subject rejection, and malformed payload rejection.
  • Keep this as explicit local worker/playbook replay only: no background consumer, subscription loop, automatic processing, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 180 Rust tests across unit, integration, and doc-test targets.

This gives local audit tests the matching read side for explicit receipt event publication while keeping consumption caller-controlled.

2026-06-23: Retrieval Receipt Event Durable Consumers

Implementation branch:

  • kadyapam/ehdb-retrieval-receipt-event-consumer

Design target:

  • Add RetrievalContextReceiptEventDurableConsumerLog.
  • Add explicit target helpers to create durable consumers, replay pending validated receipt events for a consumer, and ack receipt event sequences.
  • Support caller-supplied InMemoryStreamLog and LocalJsonlStreamLog instances.
  • Keep replay validation on the stable subject ehdb.retrieval.context.execution.receipt and receipt event payload.
  • Cover consumer resume/ack behavior, ack rollback rejection, missing consumer rejection, and JSONL reopen cursor behavior.
  • Keep this as explicit local worker/playbook consumer control only: no background consumer, subscription loop, scheduler, automatic processing, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 183 Rust tests across unit, integration, and doc-test targets.

This gives local audit tests resumable receipt event consumption while keeping cursor advancement under explicit caller control.

2026-06-23: Retrieval Receipt Event Stream Setup

Implementation branch:

  • kadyapam/ehdb-retrieval-receipt-stream-setup

Design target:

  • Add explicit stream setup helpers to RetrievalContextReceiptEventStreamTarget.
  • Build StreamConfig from target tenant, namespace, stream, and caller-selected retention policy.
  • Create receipt event streams through caller-supplied InMemoryStreamLog and LocalJsonlStreamLog instances.
  • Keep setup explicit: publish helpers do not auto-create streams.
  • Cover setup plus publish/replay, duplicate stream rejection, and JSONL stream persistence after reopen.
  • Keep this as explicit local worker/playbook stream setup only: no auto-create-on-publish, scheduler, automatic processing, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 186 Rust tests across unit, integration, and doc-test targets.

This gives worker/playbook tests an explicit setup step for receipt event streams before publication.

2026-06-23: Retrieval Receipt Event Stream Retention Setup

Implementation branch:

  • kadyapam/ehdb-retrieval-receipt-stream-retention

Design target:

  • Add create_keep_all_stream to RetrievalContextReceiptEventStreamTarget.
  • Add create_bounded_stream to RetrievalContextReceiptEventStreamTarget.
  • Reject zero bounded retention before touching the stream log.
  • Cover bounded retention replay behavior, zero-bound rejection, and JSONL bounded stream reopen behavior.
  • Keep this as explicit local worker/playbook stream setup only: no auto-create-on-publish, scheduler, automatic processing, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 189 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Stream Positive Retention Validation

Implementation branch:

  • kadyapam/ehdb-stream-positive-retention

Design target:

  • Reject RetentionPolicy::MaxRecords(0) at the stream log boundary.
  • Keep in-memory and JSONL stream creation aligned on the same config validation.
  • Reject zero max-record retention before JSONL stream setup writes a journal entry.
  • Preserve keep-all retention and positive bounded retention behavior.
  • Keep this as local stream log validation only: no scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 191 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Stream Subject-Filtered Replay

Implementation branch:

  • kadyapam/ehdb-stream-subject-filtered-replay

Design target:

  • Add subject matching for exact subjects, single-token * wildcards, and terminal > tail wildcards.
  • Keep misplaced > filters from matching unrelated subjects.
  • Add explicit subject-filtered replay for InMemoryStreamLog.
  • Add matching delegated subject-filtered replay for LocalJsonlStreamLog.
  • Cover exact matching, wildcard matching, cursor behavior, and JSONL reopen behavior.
  • Keep this as local explicit stream-log replay only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 194 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Stream Filtered Durable Consumer Replay

Implementation branch:

  • kadyapam/ehdb-stream-filtered-consumer-replay

Design target:

  • Add explicit subject-filtered durable consumer replay for InMemoryStreamLog.
  • Add matching delegated subject-filtered durable consumer replay for LocalJsonlStreamLog.
  • Filter records pending after the durable consumer ack cursor without moving that cursor.
  • Preserve missing consumer EhdbError::NotFound behavior.
  • Cover wildcard filtering, cursor behavior, missing consumers, and JSONL reopen behavior.
  • Keep this as local explicit stream-log replay only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 196 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Stream Journal Subject Replay Validation

Implementation branch:

  • kadyapam/ehdb-stream-journal-subject-validation

Design target:

  • Validate StreamRecord.subject before inserting replayed journal records.
  • Reject persisted publish entries with wildcard concrete subjects on JSONL reopen.
  • Reject persisted publish entries with empty-token concrete subjects on JSONL reopen.
  • Preserve valid JSONL replay behavior.
  • Keep this as local stream journal replay validation only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 198 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Stream Journal Identifier Replay Validation

Implementation branch:

  • kadyapam/ehdb-stream-journal-identifier-validation

Design target:

  • Revalidate persisted stream config tenant, namespace, and stream identifiers during JSONL replay.
  • Revalidate create-consumer, publish, and ack tenant/namespace/stream coordinates before applying journal entries.
  • Revalidate persisted consumer names and stream record transaction IDs before rebuilding consumer cursor or retained-record state.
  • Preserve valid JSONL replay behavior.
  • Keep this as local stream journal replay validation only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 202 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Stream Sequence JSONL Replay Validation

Implementation branch:

  • kadyapam/ehdb-stream-sequence-serde-validation

Design target:

  • Ensure StreamSequence serde deserialization preserves the same nonzero invariant as StreamSequence::new.
  • Reject zero persisted publish record sequences during JSONL stream journal replay.
  • Reject zero persisted ack cursor sequences during JSONL stream journal replay.
  • Preserve the existing compact numeric JSON representation for valid stream sequences.
  • Keep this as local stream journal replay validation only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 241 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Stream Journal Unknown Field Replay Validation

Implementation branch:

  • kadyapam/ehdb-strict-stream-journal-fields

Design target:

  • Make local JSONL stream journal entry replay strict.
  • Reject unknown create-stream, create-consumer, publish, and ack operation fields during replay.
  • Reject unknown persisted stream config fields during replay.
  • Reject unknown persisted stream record fields during replay.
  • Keep valid stream journal replay behavior unchanged.
  • Keep this as local stream journal replay validation only: no stream publication behavior changes, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 242 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Transaction Journal Identifier Replay Validation

Implementation branch:

  • kadyapam/ehdb-transaction-journal-identifier-validation

Design target:

  • Revalidate persisted transaction record transaction ID, tenant, and namespace identifiers during JSONL replay.
  • Revalidate catalog, stream, retrieval, system-library, and storage mutation identifiers before inserting replayed records.
  • Revalidate stream publish subjects in transaction mutations as concrete subjects.
  • Preserve valid JSONL transaction-log replay behavior.
  • Keep this as local transaction-log replay validation only: no consensus engine, background processing, network API, gateway data-touch behavior, production replication, scheduler behavior, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 208 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Transaction Journal Unknown Field Replay Validation

Implementation branch:

  • kadyapam/ehdb-strict-transaction-journal-fields

Design target:

  • Make local JSONL transaction journal replay strict.
  • Reject unknown persisted transaction record fields during replay.
  • Reject unknown catalog, stream, retrieval, system-library, and storage mutation fields during replay.
  • Keep valid transaction journal replay behavior unchanged.
  • Keep this as local transaction journal replay validation only: no network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed transaction coordinator, production replication, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 243 Rust tests across unit, integration, and doc-test targets.

2026-06-25: System Library Journal Unknown Field Replay Validation

Implementation branch:

  • kadyapam/ehdb-strict-system-library-journal-fields

Design target:

  • Reject unknown top-level system-library journal fields during JSONL replay.
  • Reject unknown persisted publish request fields during JSONL replay.
  • Reject unknown persisted bind request fields during JSONL replay.
  • Preserve valid publish/bind replay behavior and hot-replacement bindings.
  • Keep this as local system-library journal replay validation only: no WASM execution, background processing, network API, gateway data-touch behavior, production replication, scheduler behavior, distributed transaction coordinator, object transfer execution, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 244 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Storage Metadata Unknown Field Decode Validation

Implementation branch:

  • kadyapam/ehdb-strict-storage-metadata-fields

Design target:

  • Reject unknown JSON fields when decoding object refs.
  • Reject unknown JSON fields in geo placements and placement policy targets.
  • Reject unknown JSON fields in replica records, replication actions, and replication plans.
  • Preserve existing placement validation, replica registry behavior, and deterministic replication planning.
  • Keep this as storage metadata JSON decoding validation only: no object movement, cloud adapters, network API, gateway data-touch behavior, production replication, scheduler, background worker, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 245 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Catalog Metadata Unknown Field Decode Validation

Implementation branch:

  • kadyapam/ehdb-strict-catalog-metadata-fields

Design target:

  • Reject unknown JSON fields when decoding catalog tables.
  • Reject unknown JSON fields in catalog snapshots and nested object references.
  • Reject unknown JSON fields in catalog scan grants.
  • Reject unknown JSON fields in create-table, snapshot commit, and scan grant request shapes.
  • Reject unknown JSON fields in table schema and column schema metadata.
  • Preserve existing in-memory catalog behavior for table creation, snapshot parent-chain validation, and scan grant checks.
  • Keep this as catalog metadata JSON decoding validation only: no network API, gateway route, production ACL/IAM engine, query planner, distributed transaction coordinator, background worker, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 246 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Retrieval Metadata Unknown Field Decode Validation

Implementation branch:

  • kadyapam/ehdb-strict-retrieval-metadata-fields

Design target:

  • Reject unknown JSON fields when decoding retrieval documents.
  • Reject unknown JSON fields in chunks and embeddings.
  • Reject unknown JSON fields in document, chunk, and embedding registration requests.
  • Reject unknown JSON fields in vector, text, and hybrid search request shapes.
  • Reject unknown JSON fields in local vector, text, and hybrid search hit metadata, including nested chunk and embedding payloads.
  • Preserve existing local retrieval catalog behavior for registration, exact vector search, exact text search, and hybrid scoring.
  • Keep this as retrieval metadata JSON decoding validation only: no ANN index, full-text index, retrieval daemon, network API, gateway route, prompt engine, LLM invocation, production IAM, query planner, background worker, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 247 Rust tests across unit, integration, and doc-test targets.

2026-06-25: System Library Metadata Unknown Field Decode Validation

Implementation branch:

  • kadyapam/ehdb-strict-system-library-metadata-fields

Design target:

  • Reject unknown JSON fields when decoding resolved WASM system library manifests.
  • Reject unknown JSON fields when decoding NoETL WASM plugin references used for worker handoff metadata.
  • Reject unknown JSON fields when decoding environment/channel system library bindings.
  • Preserve existing publish, bind, resolve, local JSONL replay, and hot replacement behavior.
  • Keep this as system-library metadata JSON decoding validation only: no WASM execution, background processing, network API, gateway data-touch behavior, production replication, scheduler behavior, distributed transaction coordinator, object transfer execution, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 248 Rust tests across unit, integration, and doc-test targets.

2026-07-02: Local Arrow Scan Projection Shape Validation

Implementation branch:

  • kadyapam/ehdb-local-arrow-scan-projection-validation

Design target:

  • Reject empty direct local Arrow scan projection lists before object reads.
  • Reject duplicate projection columns in direct local Arrow scan requests before object reads.
  • Preserve ordered valid projections and existing missing-column errors.
  • Keep this as local Arrow IPC scan request validation only: no SQL planning, predicate pushdown, distributed execution, gateway direct reads, Arrow Flight protocol changes, production IAM/ACL behavior, request scheduling, object movement, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 249 Rust tests across unit, integration, and doc-test targets.

2026-07-02: Local Arrow Scan Selector Identifier Validation

Implementation branch:

  • kadyapam/ehdb-local-arrow-scan-selector-validation

Design target:

  • Validate projection-column selector identifiers in direct local Arrow scan requests before object reads.
  • Validate equality-predicate column selector identifiers in direct local Arrow scan requests before object reads.
  • Preserve existing NotFound behavior for valid-but-missing projection and predicate columns.
  • Keep this as direct local Arrow IPC scan selector validation only: no SQL planning, predicate pushdown, distributed execution, gateway direct reads, Arrow Flight protocol changes, production IAM/ACL behavior, request scheduling, object movement, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 250 Rust tests across unit, integration, and doc-test targets.

2026-07-02: Table Schema Duplicate Column Validation

Implementation branch:

  • kadyapam/ehdb-reject-duplicate-schema-columns

Design target:

  • Reject duplicate TableSchema column names in ehdb-core.
  • Preserve existing non-empty schema and per-column identifier validation.
  • Keep Arrow projection and predicate selectors unambiguous before catalog state is created.
  • Keep this as table schema validation only: no schema evolution, type coercion, SQL planning, predicate pushdown, distributed execution, gateway direct reads, production IAM/ACL behavior, object movement, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 251 Rust tests across unit, integration, and doc-test targets.

2026-07-02: Table Schema Column Identifier Revalidation

Implementation branch:

  • kadyapam/ehdb-revalidate-schema-column-identifiers

Design target:

  • Revalidate every TableSchema column identifier in ehdb-core, including preconstructed or decoded ColumnSchema values.
  • Preserve existing non-empty schema and duplicate-column validation.
  • Keep Arrow projection and predicate selectors unambiguous before catalog state is created.
  • Keep this as table schema validation only: no schema evolution, type coercion, SQL planning, predicate pushdown, distributed execution, gateway direct reads, production IAM/ACL behavior, object movement, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 252 Rust tests across unit, integration, and doc-test targets.

2026-07-02: Table Schema JSON Decode Validation

Implementation branch:

  • kadyapam/ehdb-schema-json-decode-validation

Design target:

  • Route ColumnSchema JSON decode through ColumnSchema::new.
  • Route TableSchema JSON decode through TableSchema::new.
  • Preserve strict unknown-field behavior and the existing schema JSON shape.
  • Reject invalid column identifiers and duplicate table schema columns during JSON decode before metadata is accepted.
  • Keep this as table/column schema JSON decode validation only: no schema evolution, type coercion, SQL planning, predicate pushdown, distributed execution, gateway direct reads, production IAM/ACL behavior, object movement, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 253 Rust tests across unit, integration, and doc-test targets.

2026-07-02: Core Identifier JSON Decode Validation

Implementation branch:

  • kadyapam/ehdb-core-identifier-json-decode-validation

Design target:

  • Route core identifier newtype JSON decode through each identifier constructor.
  • Preserve existing string JSON shape for identifiers.
  • Reject malformed tenant, namespace, table, transaction, stream, retrieval, and related identifiers during JSON decode.
  • Keep this as core identifier JSON decode validation only: no schema evolution, type coercion, SQL planning, predicate pushdown, distributed execution, gateway direct reads, production IAM/ACL behavior, object movement, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 254 Rust tests across unit, integration, and doc-test targets.

2026-07-02: NoETL Runtime Surface Replay Fixture

Implementation branch:

  • kadyapam/ehdb-noetl-runtime-surface-replay-fixture

Design target:

  • Add a local NoETL-shaped runtime surface fixture over LocalReferenceRuntime.
  • Drive catalog, stream, retrieval, system WASM, and storage replica metadata only through appended CommitTransaction records.
  • Reopen the runtime and verify catalog scan grants, stream consumer replay, retrieval text lookup, system library resolution, and storage replica inventory from transaction replay.
  • Keep this as a local worker/playbook-shaped integration-readiness fixture only: no gateway route, direct gateway data access, persistent per-tenant service, production IAM, distributed execution, SQL planner, object movement, or external dependency replacement behavior.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 255 Rust tests across unit, integration, and doc-test targets.

2026-06-23: System Library Journal Identifier Replay Validation

Implementation branch:

  • kadyapam/ehdb-system-journal-identifier-validation

Design target:

  • Revalidate persisted system-library publish entries during JSONL replay: library path, revision, digest, object path, and transaction ID.
  • Revalidate persisted system-library bind entries during JSONL replay: tenant, namespace, environment, channel, path, revision, digest, and transaction ID.
  • Preserve valid publish/bind replay behavior and hot-replacement bindings.
  • Keep this as local system-library journal replay validation only: no WASM execution, background processing, network API, gateway data-touch behavior, production replication, scheduler behavior, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 210 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Retrieval Context Payload Canonical Encoding

Implementation branch:

  • kadyapam/ehdb-canonical-retrieval-context-payload-encoding

Design target:

  • Require local retrieval context request/result payload bytes to match the canonical EHDB encoding produced by each payload's encode method.
  • Reject pretty-printed or otherwise non-canonical JSON request payload bytes before local worker/playbook execution or handoff.
  • Reject pretty-printed or otherwise non-canonical JSON result payload bytes before local worker/playbook execution or handoff.
  • Preserve valid payloads produced by encode.
  • Keep existing local payload execution paths on the stricter decode path.
  • Keep this as local retrieval context worker/playbook payload byte-contract validation only: no network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 234 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Retrieval Receipt Payload Unknown Field Rejection

Implementation branch:

  • kadyapam/ehdb-strict-retrieval-receipt-payload-fields

Design target:

  • Make local retrieval context execution receipt JSON payload decoding strict.
  • Reject unknown top-level receipt payload fields.
  • Reject unknown redacted execution summary fields.
  • Preserve valid receipt payload round trips.
  • Keep existing artifact and receipt-event helper paths on the stricter receipt decoder.
  • Keep this as local retrieval context receipt payload validation only: no stream publication behavior, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 235 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Retrieval Receipt Payload Canonical Encoding

Implementation branch:

  • kadyapam/ehdb-canonical-retrieval-receipt-payload-encoding

Design target:

  • Require local retrieval context execution receipt payload bytes to match the canonical EHDB encoding produced by RetrievalContextPayloadExecutionReceiptPayload::encode.
  • Reject pretty-printed or otherwise non-canonical receipt JSON bytes before artifact validation, receipt decode handoff, or receipt-event helper use.
  • Preserve valid receipt payloads produced by encode.
  • Keep existing artifact and receipt-event helper paths on the stricter receipt decoder.
  • Keep this as local retrieval context receipt payload byte-contract validation only: no stream publication behavior, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 236 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Retrieval Receipt Event Payload Unknown Field Rejection

Implementation branch:

  • kadyapam/ehdb-strict-retrieval-receipt-event-payload-fields

Design target:

  • Make local retrieval context execution receipt event JSON payload decoding strict.
  • Reject unknown top-level event envelope fields before replay decode, publisher helper use, or consumer handoff.
  • Preserve valid event payload round trips from execution artifacts.
  • Keep nested receipt payload bytes on the existing strict canonical receipt decoder.
  • Keep this as local retrieval context receipt event payload validation only: no stream publication behavior changes, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 237 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Retrieval Receipt Event Payload Canonical Encoding

Implementation branch:

  • kadyapam/ehdb-canonical-retrieval-receipt-event-payload-encoding

Design target:

  • Require local retrieval context execution receipt event payload bytes to match the canonical EHDB encoding produced by RetrievalContextPayloadExecutionReceiptEventPayload::encode.
  • Reject pretty-printed or otherwise non-canonical event JSON bytes before replay decode, publisher helper use, or consumer handoff.
  • Preserve valid event payload round trips from execution artifacts.
  • Keep nested receipt payload bytes on the existing strict canonical receipt decoder.
  • Keep this as local retrieval context receipt event payload byte-contract validation only: no stream publication behavior changes, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 238 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Retrieval Context Payload Unknown Field Rejection

Implementation branch:

  • kadyapam/ehdb-strict-retrieval-context-payload-fields

Design target:

  • Make local retrieval context request/result JSON payload decoding strict.
  • Reject unknown top-level request payload fields.
  • Reject unknown embedded context assembly request fields.
  • Reject unknown top-level result payload fields.
  • Reject unknown context object and context block fields.
  • Preserve valid request/result payload round trips.
  • Keep this as local retrieval context worker/playbook payload validation only: no network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 233 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Retrieval Context Payload Identifier Validation

Implementation branch:

  • kadyapam/ehdb-retrieval-payload-identifier-validation

Design target:

  • Revalidate decoded RetrievalContextRequestPayload identifiers: tenant, namespace, and embedding model.
  • Revalidate decoded RetrievalContextResultPayload context-block identifiers: chunk, document, and embedding model.
  • Validate identifiers on encode as well as decode.
  • Preserve valid request/result payload round trips.
  • Keep this as local RAG payload codec validation only: no ANN index, retrieval daemon, RPC protocol, Arrow Flight retrieval endpoint, gateway data-touch behavior, prompt/LLM invocation, background processing, scheduler behavior, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 212 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Arrow Flight Scan Ticket Identifier Validation

Implementation branch:

  • kadyapam/ehdb-scan-ticket-identifier-validation

Design target:

  • Revalidate decoded ScanFlightTicket tenant, namespace, and table-name identifiers.
  • Validate the same identifiers on encode, Arrow Ticket, and command-descriptor paths.
  • Preserve valid scan-ticket round trips.
  • Keep this as local Arrow Flight scan ticket codec validation only: no SQL planner, predicate pushdown, distributed execution, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 213 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Arrow Flight Scan Selector Identifier Validation

Implementation branch:

  • kadyapam/ehdb-scan-selector-identifier-validation

Design target:

  • Revalidate decoded ScanFlightTicket projection-column identifiers.
  • Revalidate decoded ScanFlightTicket equality-predicate column identifiers.
  • Validate the same selectors on encode, Arrow Ticket, and command-descriptor paths.
  • Preserve valid projection and predicate scan-ticket round trips.
  • Keep this as local Arrow Flight scan ticket codec validation only: no SQL planner, predicate pushdown implementation, distributed execution, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 214 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Result Stream Metadata Validation

Implementation branch:

  • kadyapam/ehdb-scan-result-stream-metadata

Design target:

  • Mark locally produced ArrowScanResult FlightData streams with an EHDB scan-result stream version.
  • Reject empty streams before Arrow decode.
  • Reject missing result-stream metadata before Arrow decode.
  • Reject unsupported result-stream metadata before Arrow decode.
  • Preserve valid result stream round trips, including projected schemas.
  • Keep this as local Arrow Flight scan result stream codec validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 215 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Result Metadata Envelope

Implementation branch:

  • kadyapam/ehdb-scan-result-metadata-envelope

Design target:

  • Accept the EHDB scan-result stream version only on the first FlightData message.
  • Keep locally produced later FlightData message app metadata empty.
  • Reject non-empty later-message app metadata before Arrow decode.
  • Preserve valid result stream round trips.
  • Keep this as local Arrow Flight scan result stream codec validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 216 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Fixture Validation

Implementation branch:

  • kadyapam/ehdb-scan-info-fixture-validation

Design target:

  • Add a local validator for EHDB-produced scan FlightInfo fixtures.
  • Accept valid locally produced scan FlightInfo.
  • Reject unsupported FlightInfo app metadata.
  • Reject unordered scan FlightInfo.
  • Reject negative record or byte counts.
  • Reject missing endpoint tickets and multiple endpoints.
  • Keep this as local Arrow Flight scan FlightInfo fixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 217 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Descriptor-Ticket Consistency

Implementation branch:

  • kadyapam/ehdb-scan-info-ticket-consistency

Design target:

  • Decode scan FlightInfo command descriptors and endpoint tickets in the local validator.
  • Reject a valid command descriptor paired with a valid but different endpoint ticket.
  • Preserve valid locally produced FlightInfo fixtures.
  • Keep this as local Arrow Flight scan FlightInfo fixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 218 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Endpoint Envelope

Implementation branch:

  • kadyapam/ehdb-scan-info-endpoint-envelope

Design target:

  • Keep locally produced scan FlightInfo endpoint locations empty.
  • Reject endpoint locations in the local FlightInfo validator.
  • Reject endpoint expiration timestamps in the local FlightInfo validator.
  • Reject endpoint app metadata in the local FlightInfo validator.
  • Preserve valid locally produced FlightInfo fixtures.
  • Keep this as local Arrow Flight scan FlightInfo fixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 219 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Schema Metadata

Implementation branch:

  • kadyapam/ehdb-scan-info-schema-metadata

Design target:

  • Require non-empty schema IPC bytes in the local scan FlightInfo validator.
  • Reject missing or empty schema metadata.
  • Preserve valid locally produced FlightInfo fixtures.
  • Keep this as local Arrow Flight scan FlightInfo fixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 220 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Schema IPC

Implementation branch:

  • kadyapam/ehdb-scan-info-schema-ipc

Design target:

  • Decode scan FlightInfo schema IPC bytes during local fixture validation.
  • Reject non-empty malformed schema IPC metadata.
  • Preserve valid locally produced FlightInfo fixtures.
  • Keep this as local Arrow Flight scan FlightInfo fixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 221 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Byte Count

Implementation branch:

  • kadyapam/ehdb-scan-info-byte-count

Design target:

  • Require positive total_bytes metadata in local scan FlightInfo validation.
  • Reject zero byte-count metadata while preserving negative byte-count rejection.
  • Preserve valid locally produced FlightInfo fixtures.
  • Keep this as local Arrow Flight scan FlightInfo fixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 222 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Result Consistency

Implementation branch:

  • kadyapam/ehdb-scan-info-result-consistency

Design target:

  • Add local result-bound scan FlightInfo validation.
  • Reject FlightInfo fixtures whose schema metadata does not match the producing ArrowScanResult schema.
  • Reject row-count and byte-count metadata that does not match the producing result.
  • Reject descriptor/ticket pairs that are internally consistent but do not match the expected scan ticket.
  • Keep this as local Arrow Flight scan FlightInfo fixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 223 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Expected Ticket

Implementation branch:

  • kadyapam/ehdb-scan-info-expected-ticket

Design target:

  • Add receiver-side scan FlightInfo validation on ScanFlightTicket.
  • Accept returned scan FlightInfo that matches the expected ticket.
  • Reject internally consistent FlightInfo descriptor/ticket pairs that belong to a different scan request.
  • Validate returned FlightInfo in the loopback client smoke path before following the endpoint ticket.
  • Keep this as local Arrow Flight scan FlightInfo receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 224 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Info Schema Response

Implementation branch:

  • kadyapam/ehdb-scan-info-schema-response

Design target:

  • Add receiver-side scan FlightInfo schema validation on ScanFlightTicket.
  • Accept returned scan FlightInfo whose schema metadata matches the expected Arrow schema.
  • Reject well-formed scan FlightInfo whose schema metadata differs from the expected schema.
  • Validate get_schema and get_flight_info schema consistency in the loopback client smoke path before following the endpoint ticket.
  • Keep this as local Arrow Flight scan FlightInfo receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 225 Rust tests across unit, integration, and doc-test targets.

2026-06-24: Arrow Flight Scan Data Info Consistency

Implementation branch:

  • kadyapam/ehdb-scan-data-info-consistency

Design target:

  • Add receiver-side scan data validation against returned FlightInfo.
  • Accept matching FlightData streams and FlightInfo metadata.
  • Reject row-count, byte-count, and schema mismatches.
  • Validate decoded loopback client data against returned FlightInfo before treating batches as coherent scan output.
  • Keep this as local Arrow Flight scan result receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Scan Ticket Canonical Encoding

Implementation branch:

  • kadyapam/ehdb-canonical-scan-ticket-encoding

Design target:

  • Require decoded local Arrow Flight scan ticket bytes to match the canonical EHDB encoding produced by ScanFlightTicket::encode.
  • Reject pretty-printed or otherwise non-canonical JSON ticket bytes before scan execution or Flight handoff.
  • Preserve valid tickets produced by ScanFlightTicket::encode, to_arrow_ticket, and command_descriptor.
  • Keep server get_flight_info, get_schema, and do_get on the stricter decode path.
  • Keep this as local Arrow Flight scan ticket byte-contract validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 231 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Scan Ticket Unknown Field Rejection

Implementation branch:

  • kadyapam/ehdb-strict-scan-ticket-fields

Design target:

  • Make the local Arrow Flight scan ticket JSON contract strict.
  • Reject unknown top-level ScanFlightTicket fields before execution or handoff.
  • Reject unknown embedded latest-table scan request fields before execution or handoff.
  • Reject unknown equality predicate fields before execution or handoff.
  • Preserve valid ticket encode/decode, command descriptor, and local scan paths.
  • Keep this as local Arrow Flight scan ticket/request payload validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 230 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Scan Command Descriptor Path Validation

Implementation branch:

  • kadyapam/ehdb-scan-descriptor-path-validation

Design target:

  • Tighten local Arrow Flight scan descriptor decoding so descriptor shape fails before scan execution.
  • Reject direct scan get_flight_info and get_schema command descriptors that carry non-empty path entries.
  • Preserve valid command descriptors produced by ScanFlightTicket::command_descriptor.
  • Keep this as local Arrow Flight scan descriptor request validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 229 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Scan Projection Selector Shape Validation

Implementation branch:

  • kadyapam/ehdb-scan-projection-selector-validation

Design target:

  • Tighten local Arrow scan request validation so invalid projection shapes fail before scan execution or Arrow Flight ticket use.
  • Reject projection: Some([]) at the service and ticket request validation boundary.
  • Reject duplicate projection column selectors before local scan execution or Flight ticket encode/decode succeeds.
  • Preserve valid projection/filter scan requests.
  • Keep this as local scan request selector validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 228 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight DoGet Endpoint Ticket Binding

Implementation branch:

  • kadyapam/ehdb-do-get-endpoint-ticket-binding

Design target:

  • Add receiver-side validation for the concrete Arrow Flight endpoint Ticket used for do_get.
  • Bind that supplied endpoint ticket back to the returned scan FlightInfo, decoded schema, and expected ScanFlightTicket.
  • Accept the locally produced endpoint ticket from matching scan FlightInfo.
  • Reject a well-formed endpoint ticket for a different scan request before do_get is treated as coherent.
  • Use the helper in local service/server and loopback client receiver paths before issuing or accepting do_get results.
  • Keep this as local Arrow Flight endpoint-ticket receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Scan Response Envelope Validation

Implementation branch:

  • kadyapam/ehdb-scan-response-envelope-validation

Design target:

  • Add an ArrowScanResult helper for the complete receiver-side scan response envelope: raw SchemaResult, returned scan FlightInfo, raw FlightData, and expected ScanFlightTicket.
  • Decode and validate SchemaResult, validate returned FlightInfo, validate raw FlightData, and return the decoded schema plus coherent result.
  • Accept locally produced schema/info/FlightData responses for the expected scan ticket.
  • Reject well-formed schema responses whose decoded schema differs from returned scan FlightInfo before accepting raw scan data.
  • Use the helper in local service and server receiver tests that consume raw FlightData.
  • Keep this as local Arrow Flight scan response receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Raw Data Schema Validation

Implementation branch:

  • kadyapam/ehdb-raw-flight-data-schema-validation

Design target:

  • Add an ArrowScanResult helper for raw Arrow Flight FlightData plus the decoded get_schema schema, returned scan FlightInfo, and expected ScanFlightTicket.
  • Accept locally produced schema/info/FlightData responses for the expected scan ticket.
  • Reject well-formed decoded schemas that differ from returned scan FlightInfo metadata before decoding or accepting raw scan data.
  • Use the helper in local service and server receiver paths before accepting returned FlightData from do_get.
  • Keep this as local Arrow Flight raw scan data receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Decoded Client Response Validation

Implementation branch:

  • kadyapam/ehdb-decoded-client-response-validation

Design target:

  • Add an ArrowScanResult helper for already-decoded Arrow batches plus the decoded get_schema schema, returned scan FlightInfo, and expected ScanFlightTicket.
  • Accept locally produced schema/info/batch responses for the expected scan ticket.
  • Reject well-formed decoded schemas that differ from returned scan FlightInfo metadata before treating returned batches as coherent.
  • Use the helper in loopback client smoke paths before accepting returned batches from do_get.
  • Keep this as local Arrow Flight decoded client response receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Schema Result Endpoint Ticket Helper

Implementation branch:

  • kadyapam/ehdb-schema-result-endpoint-ticket

Design target:

  • Add receiver-side schema-result plus endpoint-ticket extraction on ScanFlightTicket.
  • Decode returned Arrow Flight SchemaResult, validate returned scan FlightInfo against the expected scan ticket and decoded schema, and return the decoded schema plus validated endpoint ticket for do_get.
  • Accept locally produced schema results and scan info for the expected ticket.
  • Reject well-formed schema results whose decoded schema differs from returned scan FlightInfo metadata before returning an endpoint ticket.
  • Use the helper in local service and server tests before do_get.
  • Keep this as local Arrow Flight schema/scan-info endpoint-ticket receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Schema Result Scan Info Validation

Implementation branch:

  • kadyapam/ehdb-schema-result-info-validation

Design target:

  • Add receiver-side schema-result validation on ScanFlightTicket.
  • Decode returned Arrow Flight SchemaResult and validate returned scan FlightInfo against the expected scan ticket and decoded schema.
  • Accept locally produced schema results and scan info for the expected ticket.
  • Reject well-formed schema results whose decoded schema differs from returned scan FlightInfo metadata.
  • Use the helper in local service and server tests before treating get_schema and get_flight_info as coherent; keep client paths validating already-decoded schemas against returned FlightInfo.
  • Keep this as local Arrow Flight schema/scan-info receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Scan Data Expected Ticket Validation

Implementation branch:

  • kadyapam/ehdb-scan-data-ticket-validation

Design target:

  • Add receiver-side helpers on ArrowScanResult for validating returned scan data against both returned FlightInfo and the expected ScanFlightTicket.
  • Support both raw FlightData streams and decoded Arrow batches for client paths that already receive batches.
  • Accept locally produced scan data and returned scan info for the expected ticket.
  • Reject internally valid scan info/data when the expected ticket differs.
  • Use the helpers in local service, server, and loopback client smoke paths before treating returned scan data as coherent.
  • Keep this as local Arrow Flight scan data receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Scan Info Schema Endpoint Ticket Helper

Implementation branch:

  • kadyapam/ehdb-scan-info-schema-endpoint-ticket

Design target:

  • Add schema-aware receiver-side endpoint-ticket extraction on ScanFlightTicket.
  • Return the endpoint ticket only after validating returned scan FlightInfo against the expected scan ticket and the expected Arrow schema.
  • Reject well-formed scan FlightInfo whose schema metadata differs from the expected get_schema response before returning an endpoint ticket.
  • Use the schema-aware endpoint-ticket helper in local service, server, and loopback client smoke paths before do_get.
  • Keep this as local Arrow Flight scan FlightInfo receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-25: Arrow Flight Scan Info Endpoint Ticket Helper

Implementation branch:

  • kadyapam/ehdb-scan-info-endpoint-ticket-helper

Design target:

  • Add receiver-side endpoint-ticket extraction on ScanFlightTicket.
  • Return the endpoint ticket only after validating returned scan FlightInfo against the expected scan ticket.
  • Reject internally consistent scan FlightInfo that belongs to a different scan request before returning an endpoint ticket.
  • Use the validated endpoint-ticket helper in the loopback client smoke path before do_get.
  • Keep this as local Arrow Flight scan FlightInfo receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 226 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Stream Subject Filter Type

Implementation branch:

  • kadyapam/ehdb-stream-subject-filters

Design target:

  • Keep Subject concrete for published stream record subjects.
  • Reject wildcard tokens * and > in Subject::new.
  • Add SubjectFilter for replay selectors.
  • Accept exact filters, single-token * wildcards, and terminal > tail wildcards in SubjectFilter.
  • Reject misplaced > and partial wildcard tokens in SubjectFilter::new.
  • Move filtered replay APIs to SubjectFilter.
  • Cover concrete wildcard rejection, valid wildcard filters, invalid filter placement, filtered replay, filtered durable consumer replay, and JSONL reopen behavior.
  • Keep this as local stream log validation and replay only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 196 Rust tests across unit, integration, and doc-test targets.

2026-06-23: Stream Non-Empty Subject Tokens

Implementation branch:

  • kadyapam/ehdb-stream-nonempty-subject-tokens

Design target:

  • Reject leading, trailing, and double-dot empty tokens in Subject::new.
  • Reject leading, trailing, and double-dot empty tokens in SubjectFilter::new.
  • Preserve valid concrete subjects and valid exact/wildcard filters.
  • Cover concrete subject rejection, subject filter rejection, valid exact/wildcard filters, and filtered replay behavior.
  • Keep this as local stream log validation only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.

Tracking issue:

Validation:

  • cargo fmt --all --check
  • cargo test --workspace
  • cargo clippy --workspace --all-targets -- -D warnings
  • cargo bench --workspace --no-run

Coverage snapshot:

  • 196 Rust tests across unit, integration, and doc-test targets.

Clone this wiki locally