Repository navigation
Sessions Log
v0.8.0 (#404) + v0.8.1 (#405).
A2/A3 made sealed history survive node loss. A part is local-only until it
seals, so the window was the seal interval — B2 bounds it, cannot close it.
Measured on prod while designing: SEAL_MAX_AGE=900 with
oldest_unsealed_age=719 and unreplicated_oldest_age=0, i.e. nothing waiting
to upload and ~12 minutes of records on one disk.
replicate_tail writes the pending records as write-once per-batch objects
under tail/ (never parts/, or plan_retention could drop them), and
cold_load_replicated replays them. Window: 900 s → the 15 s tick. Bounded,
not zero — no remote ack gates the local append.
⭐ The RED control is the load-bearing test: with the flag off, 16 of 21 records
come back, so the loss is real and the green results are falsifiable. With it
on, destroying the local root and cold-loading from the replica alone returns
21 of 21 — and the same proof runs against a real GCS emulator in
noetl/server.
⚠⚠ v0.8.1 fixes a defect I introduced in v0.8.0 and caught before enabling
anything: replicate_tail performed the remote put while holding the engine
lock, and in the server that lock is taken by shadow_append on the live write
path — p50 77 ms / p99 234 ms added to any append colliding with a tick. The
uploader thread had documented the rule all along ("substrate writes happen
OUTSIDE the lock"). Split into prepare / upload / commit, with the no-I/O
property pinned by a test whose RED control restores the defect.
Two PRs: ehdb#401 (v0.6.0, replica backfill), #403 (v0.7.0, coordinator in-flight).
A2 armed on the prod server's embedded store. A genuine second failure domain —
domains=[local-device:66320, remote-gcs] — and the aggregate gauges flipped:
replica_set_size 1→2, survives_node_loss 0→1, single_point_of_failure 1→0. Over a
1853s window (stated deliberately: shorter than the ~900s seal period would have proved
nothing) the age seal fired at t=934s and the part plus manifest reached the bucket,
counted externally with gcloud.
⚠⚠ Then the acceptance proof refused to pass. Read back from the remote's own
manifest: 40 parts listed, 39 at replica_count 1. Uploads are enqueued on seal and
nowhere else, so arming a replica replicates nothing that already exists — and every
durability gauge read healthy the whole time, because they describe the declared replica
set, not the parts. survives_node_loss is a claim about configuration and was being
read as a claim about data. Worse than an empty remote: the bucket held a manifest naming
all 40 parts and the bytes of one, so it read as a complete store and was not a
recoverable one.
backfill_under_replicated fixes it. The trap underneath was that durable_view nulls
local_path, so after any restart a repair path that only knew fs::read would have
compiled, run, reported success and copied nothing — the same defect one level down. Hence
UploadSource::Replica. Backfills also get their own counters, because a part sealed in an
earlier process has no seal→durable interval and folding it into upload_lag_micros_total
would deflate a mean a capacity decision reads.
v0.7.0 came out of the worker's v0.3.2→v0.5.2 pin bump, whose only compile break was
the new ShardLag.inflight. Refusing to default it to 0 was the point: render_snapshot
publishes that field, and high lag with zero in-flight is specifically the stalled-consumer
reading, so a hardcoded zero makes a healthy bus publish an outage signal every scrape.
Three PRs: ehdb#385 (north-star P2–P5),
#386 (memory + concurrency),
#387 (drain benches, state gauges, benches in CI).
1,188 tests passing, clippy -D warnings clean.
Registry. RuntimeKind + discover(kind, now, ttl) turn D8 from a worker table into a
service registry; watch_since gives a resumable cursor; an Execution registers
ephemerally and vanishes by not being renewed; SecretRef carries a pointer and never a
value. A real bug surfaced: heartbeat reset the kind to Worker, so a renewed
Execution stopped being discoverable one heartbeat after registering.
Memory is O(parts), not O(records) — exponent 0.52, ~16.6 KB + ~1.8 KB/part, so 1 M records holds ~1.7 MiB. Part count is the memory signal and merge is the control.
⚠⚠ The engine is a single writer, and a shared sequencer is not enough. Minting the sort key before taking the append lock reorders 34.2% of appends while losing and duplicating nothing — the engine-level reproduction of the mechanism that rolled back the v0.4.4 ramp (2026-09-30 entry below, noetl/ai-meta#362). It reproduces in 20 ms. My own first draft of the test asserted the sequencer was sufficient; the canary caught it.
⚠⚠ A standing replication deficit was invisible. The uploader records the successful
replica writes only, so a part holding 1 of 2 copies is is_durable(), on_upload_done
fires, and the durability window reports it done. New parts_under_replicated gauge;
proven with a write-refusing substrate.
⚠⚠ Batching is worth 530x on the drain path, and the batch cap that fixed
noetl/ai-meta#298 makes a drain above the cap quadratic (exponent 2.06). Both real; the
tradeoff was never measured (ehdb#389). Separately, poll_assign without acking is
2.86x slower than with — the expiry scan walks the whole in-flight map every poll, so
lazy acking is O(n²) in the consumer (ehdb#388).
Benches now run in CI with a shape guard (denominator + monotonicity), both negative controls fired. First run cross-machine: batch ordering 530.5x local vs 528.3x on the runner; per-poll cost 32.0x vs 16.4x. Gating the second would have failed CI on the runner rather than the code — which is why only the ordering is gated.
⚠ Five of my own instruments were wrong before they were right: a part count that matched
the parts directory (1 part for 16,000 records), a println! whose labels did not match
its arguments, a SERIES row count off by 3x from a line-oriented grep over wrapped Rust,
a gauge called inert by a dataset with no dedupe_key, and a "not measured yet" list that
understated the work already done.
RED first: 2,400 records before, 2,400 after — the merge ran and collapsed nothing.
Dataset::supersede_key plus a compacting merge_once took the 10x vector collection from
4,608 stored records / 65.02 ms to 500 records / 7.13 ms — 9.2x fewer records, 9.1x
faster, converging to exactly the live set.
⚠⚠ Eligibility is narrower than the issue assumed. Only VectorDataset may opt in.
RuntimeDataset (D8) is disqualified by watch_since — the op-log reader added in P2 —
because compacting it would delete the history a watcher resumes from. One history reader
disqualifies a dataset. D1EventLog is excluded by platform rule.
Mutation battery 5 of 6 caught, 1 verified benign; the keep-first mutant is caught by the tombstone test, which proves that guard fires.
ehdb is now releasable. It had no release automation at all — every v0.4.x tag was
pushed by hand. semantic-release.yml + release-ehdb match the house pattern and avoid
both known traps by construction: .releaserc.json omits @semantic-release/git (so GH006
cannot occur, and nothing asserts tag == Cargo.toml since every crate is 0.1.0), and the
tag workflow is dispatched explicitly because a GITHUB_TOKEN-pushed tag does not
trigger workflows.
⚠⚠ Releases are gated OFF behind EHDB_RELEASE_ENABLED, verified by merging it:
Semantic Release ⚙️ ran on main and the job was skipped, and the tag read back is
still v0.4.5. The next version computes as v0.5.0; cutting it is an owner decision
because noetl/server pins both ehdb-feed and ehdb-l0 and they must move together.
⭐ The dry-run guard caught a real defect in its own first two runs: semantic-release
takes its branch from env-ci (GITHUB_REF), not git HEAD, so on a pull_request event
it reported refs/pull/N/merge, matched no --branches value, analysed nothing — and
exited 0. The first fix (checking out the head ref) changed nothing, which is what
located the mechanism.
⚠ And a tooling note: #393's first push produced zero workflow runs — no check-runs, no API entries — while every sibling PR ran. Several probes went into hunting a YAML defect that did not exist; an empty commit dispatched it immediately. Same symptom as catalog#28's "zero CI runs ever dispatched".
ehdb#390. top_k is an exact brute-force
scan, so recall@k is 1.0 by construction — a reported 1.0 proves only that an
exhaustive scan was exhaustive. It is kept as a correctness guard with its own control
(complete 1.000 / damaged 0.333 / empty expected set undefined, so a
known-answer set that failed to load cannot read as perfect).
⭐⭐ Query cost tracks op-log depth, not live points. 500 points re-embedded 10x (5,000
ops) costs 64.9 ms — within 2% of 5,000 points written once (66.1 ms). So
re-embedding a collection is as expensive as growing it: 378x over the same 500 written
once. Dimension is linear (exponent 0.90), as cosine being O(d) predicts.
⚠⚠ Compaction cannot fix it. A merge made the 10x case 9% worse (64.8 → 70.5 ms),
and the code says why: live_points calls read_index_after(collection, 0) — every op ever
written — then folds latest-wins; VectorDataset defines no dedupe_key; and the
Dataset trait exposes no supersede or compaction hook at all. Merge combines parts,
not keys, so it structurally cannot drop a superseded op. The prerequisite for vectors at
scale is a key-level compaction primitive, not an ANN index — and it is broader than
vectors, since RuntimeStore's D8 folds and the catalog's attribute/relation folds share
the pattern (ehdb#391).
⚠ My first known-answer set was silently degenerate: 64 points over 16 dimensions, so
the query vector for a high axis was the zero vector and recall@1 read 0.0 — looking
like an engine defect when the fixture could not express the answer it claimed to know.
The catalog attachment needed no new dataset: (resource_type, path) maps onto
(collection, point_id), pinned by set equality so cross-type leakage fails.
⚠⚠⚠ ROLLED BACK 2026-09-30 00:10Z — NOT armed in production. The 772/0 comparator reading was real but the window was too short; an hour later the same deployment read 13 divergences over 839 engagements (98.45%) and the three gates were removed. The cause is a THIRD mechanism, neither this release's staleness nor noetl/ai-meta#358's second writer: an event commits into the MIDDLE of an ORDER BY event_id read. A snowflake event_id is minted BEFORE the insert, so commit order is not id order, and the id-ordered log is not the append-only growing prefix populate_from_log assumes. Proven by position arithmetic — the store held at position 72 what the log holds at 73, one missing INTERIOR row. See noetl/ai-meta#362.
v0.4.4 itself stands: both defects it fixed are real and independently proven. It simply was not what was breaking prod.
ehdb#376 → 720c1d2, tagged v0.4.4.
FORMAT_VERSION stays 1, so no on-disk layout refusal.
Two defects in chain_populator, both of which made a healthy partition read
as broken, and both proven by their own RED control:
-
populate_from_logclassed "the store holds more than the caller's snapshot" asDiverged. This store's only writer ispopulate_from_log, so a longer store means a fresher snapshot was already applied. NewFromLog::StaleLog— benign,is_healthy(), partition stays servable. - Coverage was not monotonic, so a stale snapshot could record a smaller total and make the partition unservable through the guarded read — fixing the classification alone would only have converted a false divergence into a false refusal.
⚠⚠ The crate deliberately does NOT decide case 1. The same shape is how a
second writer looks (noetl/ai-meta#358,
a measured 13.3% real divergence), and from one snapshot the two are
indistinguishable. StaleLog is an unresolved report; the consumer re-reads the
log and escalates when it does not catch up. My first version reasoned "sole writer,
therefore always staleness" and the #358 regression test failed immediately — sound
about this crate, false about the system.
The alarm is not off. A content conflict is still Diverged, including when the
log is shorter, and it now also revokes serve-trust by dropping the coverage
record — previously this function wrote nothing on a conflict, so a reader calling
chain_if_authoritative alone was handed a chain we had just proved wrong. Events
and the watermark are untouched: revoke trust, keep the evidence.
⭐ The monotonic-coverage mutant survived first. The test named for it populated
5 then read 2, which returns StaleLog early, before write_coverage is reached
— it never touched the guard it was named for. Renamed and rewritten to construct
the only path that reaches it.
Where it runs. Adopted by server v3.117.2 and ARMED IN PRODUCTION 2026-09-29 17:30Z: 772 comparisons, 0 failures, 100.00% agreement, up to 6 concurrent executions, 80/80 COMPLETED. An earlier ramp on v0.4.3 was rolled back at 13:28Z on the false divergence this release fixes.
31 tests in the populator suite; 45 suites, 0 FAILED; fmt + clippy clean.
Five PRs off #324, each merged without changing a byte of writer behavior:
| finding | PR | posture |
|---|---|---|
| F4 lag-from-write metric | #333 | observability only |
| F3 age-based seal | #334 | flag, default off |
| F2 fencing enforcement | #335 | shadow — refuses nothing |
| F1 election + tokens | #336 | wired, not authoritative |
| F5 failure domains + 2nd substrate | #337 | no prod storage change |
Plus noetl/worker#291 (v5.125.0) putting the window on the writer's scraped
/metrics — the instrument the whole durability story reads from.
What the work turned up beyond the findings
- ⭐ F5's second substrate exposed an unspecified trait contract the moment there
was something to disagree with:
get_rangepast an object's end errored in the filesystem impl and clamped in the new one. Resolved fail-closed — a short read would let a truncated part read as intact. - ⚠⚠ F3's mutation pass found that the "default off" test set the flag off explicitly, so flipping the default left everything green. It tested the wrong thing; two tests added.
- ⚠ F2's scope was wider than the finding named —
put_reclaim_watermarkanddelete_segmentare mutations too, and a superseded writer deleting segments is strictly worse than one appending. - ⚠ The epoch could not ride the segment frame as the spec said:
FRAME_HEADER_LENis fixed and shared byte-identically withdurable_eventlog.rs. It lives in a per-shard marker, which is what Invariant F actually needs.
19 mutations across the five, each caught by the intended guard.
Not done, deliberately: the four gates. Three of them touch the live writer
path on an already-primary tier.
Headline. The #301 latency fix
and the #302 writer-restart fix
deployed to shastaratech prod in one rollout — engine 86a24f9,
noetl-server v3.58.3, all worker pools + writers v5.81.3. Both T5
gates measured in the same window; both PASS.
Latency — the regression is not just gone, it is inverted. issued → claimed off noetl.event: p50 138.5 / p95 156.4 / p99 181.0 ms (n=60,
unsaturated) against the 338 ms NATS baseline. The EHDB bus is now ~2.4×
faster than the envelope it replaces, and 2–4× faster than it was at the T4
flip (285–557 ms). The loopback attribution — group commit + off-lock fsync
via a duplicated fd + pipelined publish, 281 → 6.7 ms p50 at 64 publishers —
holds on real infrastructure. ai-meta#205
CLOSED.
Two measurement rules came out of it. A saturating burst measures the pool, not the bus: 90 commands in 1.2 s gave p50 2123 / p99 3735 ms, all of it queueing behind the user pool's 2 replicas × 4 slots. And kind cannot resolve this at all — its worker pool peaks at 28–131 cmd/s against a ~1000 cmd/s bus ceiling, so the publish queue the fix removes never forms and run-to-run variance swamps any before/after signal. Prod, unsaturated, is the only rig that answers the question.
Writer restart — PASS. Writer pod deleted mid-stream under continuous
load: 38 executions → 38 COMPLETED, 114 = 38 × 3 commands, 0 duplicates,
0 loss, redial in ~2.7 s (early eof at 18:18:22.056 → coordinator up
18:18:24.765), cursor-resume rather than shard replay. Steady-state
latency unchanged across the restart.
ai-meta#208 CLOSED.
2026-07-28T07:49Z; both noetl-worker-rust pods' last
log line was from that instant, then silence.
noetl-worker-system-pool-shard1 happened to restart later and kept system
commands completing, which masked the outage entirely. Two consequences
for how this program is run: a cutover sign-off that never restarted the
writer proves less than it looks like it does, and silent single-pool
dispatch death is not observable today — which makes per-subject lag
(#303) an observability fix, not
just a scaler input.
Left open, both held for review.
#304 — the resume logs
from_cursor=0 origin="persisted" and ehdb_feed_shard_committed reset
408 → 165 (segment-relative, not a global offset), so the line cannot be told
apart from a replay-from-0 even though the 114 = 38×3 accounting proves no
replay happened. noetl/server#291
— the publish retry (3 × 250 ms) does not span a ~2.7 s pod restart; prod saw
2 × HTTP 500 during the gap (fail-closed, but not transparent).
T5 status. NATS untouched (ns nats, ns nats-supercluster). The go/no-go
item is no longer latency but autoscaling: the user pool still triggers on
nats-jetstream lag, and its ScaledObject is paused — which under KEDA 2.15
means the HPA is deleted and the scaler loop does not run, so the pool has had
no autoscaling since 2026-07-26
(ai-meta#210). Also open:
ai-meta#209 — a crash can
still lose the unsealed tail, which is a loss class on a command bus and
should land before T5 removes the fallback.
Monitoring correction. Prod monitoring is Google Managed Prometheus, not
VictoriaMetrics — no vmservicescrape CRD, no vmstack namespace. The
writer scrape is PodMonitoring/noetl-cmdbus-writer, live since the T4
cutover but uncommitted until noetl/ops#242.
Because there is no in-cluster PromQL endpoint, the T5 autoscaler uses KEDA's
metrics-api trigger straight against the writer's :9102.
2026-07-30 (L1 #208): 🔁 the command bus survives a writer restart — both defects fixed, a third found
Headline. Restarting only the per-shard writer left the bus not
delivering, two ways — and both were recoverable only by flipping
NOETL_COMMAND_BUS back to NATS, which T5 removes. Fixed off main at
03e94be (#205's group commit is the base), plus a third defect found in
the server's publish path while validating.
Defect 1 — claimers never noticed. claim_next blocks in a read until a
command is available, so a writer pod that goes away leaves the socket
half-open: no data, no error, and every redial path hangs off an Err that
never comes. Fixed with TCP keepalive on every ehdb-feed socket (~11 s
detection with no FIN/RST) plus a negotiated coordinator heartbeat — one
beat up front arms a 3-missed-beat read deadline for the connection, and it also
catches a writer that is alive but stuck, which keepalive cannot. Heartbeats
only go to a client that asked, so the wire stays backward compatible.
Defect 2 — the restarted writer replayed its whole shard (from_cursor = 0;
kind lag 2738 draining at ~1/s, each stale record costing a control-plane
round-trip). CursorStore persists committed_cursor() atomically beside the
log; ClaimCoordinator::resume starts there; the cursor is clamped to the
reopened tip (a log that lost an unsealed part reopens behind it, and an
unclamped cursor no future key can exceed would take the bus permanently dark);
the worker seals the log on SIGTERM. ehdb_feed_shard_committed{shard} is
now on :9102 — lag reads 0 both when caught up and when a replay just
finished, so the cursor is the series to watch.
Defect 3 — the server dropped one command per restart. The first publish
after a restart discovers the dead socket by using it; the router was dropped
for the next command but this one was lost (command.issued durable, nothing
on the bus, HTTP 500). Now redialed and retried, 3 attempts 250 ms apart.
Kind, before/after in the same cluster:
| Check | Released worker:5.81.1
|
Fixed |
|---|---|---|
| Writer-only restart, fire, workers untouched | 21 issued / 0 claimed, lag frozen 3094 for 100 s | 61 issued / 60 claimed, claimers redialed in ~2 s |
| Lag right after the restart | 3072 (whole log) |
0, resumed from_cursor=3246 origin="persisted"
|
| Writer restarted mid-burst (30 execs) | — | 30/30 completed, 0 dupes, lag → 0 |
| Force-kill a worker mid-burst | — | 39 issued / 39 claimed / 0 dupes |
| Pool isolation + #166 | — | 0 cross-pool claims; 84 + 83 on the two system shards |
| Idle bus 10 min | — | 0 spurious redials |
9 new tests in crates/ehdb-feed/tests/restart_recovery.rs + 5 in cursor.rs;
the no-FIN wedge is driven deterministically through a relay that holds the
socket open and stops forwarding (killing a listener sends a FIN, which a
blocking read does surface, so it would not reproduce the defect). Gate green
on rustc 1.97.1.
Posture. ehdb#302 +
worker#196 +
server#290 — all held for
review, stacked so prod deploys once with #205. Kind restored to
NOETL_COMMAND_BUS=nats; nothing in prod.
Follow-up. The SIGTERM seal races in-flight ingest — one command per restart was acked and still lost with the unsealed part, recovered by the orchestrator re-issuing ~30 s later (all 30 executions completed), so a latency artifact rather than loss. Clean fix: stop accepting ingest before sealing; the crash path needs L0-level recovery of local unsealed parts.
2026-07-28 (L1 #205): 🔬 command-bus dispatch latency ROOT-CAUSED + FIXED — group commit + off-lock fsync + pipelined publish
Headline. The post-T4 latency regression (issued → claimed p50
285–557 ms vs ~200 ms NATS) is not the poll-claim path and not the
#203 writer-assigned key. It is posture-A fsync held inside the engine
lock, multiplied by the control plane holding one publish mutex across the
whole round-trip.
Attribution — new ehdb-feed/examples/dispatch_bench rebuilds the
deployed topology on loopback and splits the wall time per hop:
| Hop | Measured |
|---|---|
| publisher-mutex wait | 0 ms @1 pub → 64 ms @16 → 282 ms @64 |
| publish RTT | ~4.0 ms flat |
append (fsync) |
~4.0 ms flat → bus ceiling ~230 cmd/s |
| delivery (ack → member holds it) | 16–33 µs, flat in member count |
Delivery being microseconds and flat is what exonerates the claim path:
the 250 ms poll interval is a cap on the tip_receiver park so
ack_wait redeliveries surface, not a floor.
Fix (ehdb#301) — three parts, the first two inert without the third:
-
FlushPolicy::CallerDriven+L0Engine::take_sync_handles—FeedWriterowns its commit points andfsyncs through a duplicated fd after releasing the engine lock, so the syscall stops stalling claimers. -
FeedWriter::append_batchgroup commit + a reader/committer split inserve_ingest; it takes what has already queued and never waits to fill a batch, so a lone record is unchanged. -
PipelinedPublishClient/PublishRouter::publish(&self)— without it only one record is ever in flight and nothing can be batched.
Durability, ordering (#203) and exactly-once unchanged by construction;
out_of_order_appends asserted 0 against the real counter under 400
concurrent in-flight publishes.
Result — publish-submit → member holds the command: 281 ms → 6.7 ms
p50 at 64 publishers / 24 claimers, flat in concurrency instead of
growing, 225 → 6 983 cmd/s. Residual ~6 ms is the durable-ack fsync
itself — a floor NATS JetStream does not pay per message.
Kind — correctness PASS on all dimensions (370/370 delivered, 0
unclaimed, 0 dup, pod-kill redelivery 0-loss, lag→0, 0 restarts, #166
intact). Latency is not resolvable there: per-fsync on the writer PVC
is ~1 ms → bus ceiling ≈1 000 cmd/s, but the 3×4-slot worker pool on a
6-vCPU VM peaks at 28–131 cmd/s, so the bus is never the bottleneck (two
identical passes differed 4× in p50). The prod re-measure is the real gate.
Two pre-existing defects found and filed
(noetl/ai-meta#208), both
blocking T5: workers never redial a restarted writer (blocked
claim_next read, no keepalive → permanent wedge, 0/30 claimed), and a
restarted writer replays its whole shard from cursor 0
(from_cursor = 0 hardcoded; observed lag 2 738 draining at ~1/s).
Adoption: server#289,
worker#195 (draft until the
ehdb pin can point at main). Tracking:
noetl/ai-meta#205.
2026-07-27 (L1 T4): 🚀 EHDB IS THE PRODUCTION COMMAND BUS — full flip succeeded on shastaratech prod after fixing a silent feed-delivery loss
The L1 NATS takeover reached production. Commands now flow server → per-shard writer → worker over the EHDB feed on the shastaratech prod cluster; the server publishes zero commands to NATS. NATS is still installed purely as the rollback path, and T5 (delete NATS) is held on the human.
It took three rounds, and the middle one is the interesting part.
-
Round 1 (2026-07-26) — Step 0 + shadow PASS on a 2-shard cluster
(0 divergence, 121 published == 121 in the feed). Held at canary: the
EHDB bus partitions every command by
execution_idincluding thesharedpool, but the user pool is a single deployment claiming from a single writer, so flipping it would strand roughly half the shared commands. Rolled back clean. Filed as ai-meta#202. -
Round 2 (2026-07-27) — chose Option 2 (collapse the command-bus axis
to one shard; the two sharding axes are independent —
sharding.rs:172, "correctness never depends on affinity"), kind-validated it, reconfigured prod with #166 state sharding intact, passed shadow and canary — and then the full flip failed: ~10% of commands were ingested into the writer feed and never delivered to any claimer, with feedlag = 0throughout. Silent loss. Rolled back; prod clean. - Round 3 (2026-07-27) — with the fix, the same sequence succeeded: shadow 0-divergence → canary 30/30 (0 system claims by the user pool, 12/12 graceful-pod-kill redelivery) → full flip 30/30 + 50/50 soak = 80/80, 0 dup, 0 loss, feed lag → 0, writer 0 restarts, #166 state sharding intact throughout.
The bug — noetl/ehdb#300 → d4b6235
The command feed's producer (noetl-server) assigned each record's sort key
— the command's snowflake event_id — and the writer trusted it. But the
Dataset contract the feed cursor and range pruning depend on is "records are
appended in ascending sort_key order within a partition". Snowflake ids are
sparse, and under concurrent publish a lower id can reach the single
writer after a higher one. When that happened:
- the follower cursor had already advanced past it
(
ChangeFeed::pollsetscursor = max sort_key read); -
refillreadsread_partition_after(shard, cursor)→ strictly> cursor→ the late record never enteredpending, never got delivered; -
lag()is cursor-relative too, so a record below the cursor was never counted — hencelag = 0while commands vanished.
PartWriter::append had never enforced the ascending contract.
The fix moves key assignment to the writer (append_writer_assigned) —
the only component that can actually guarantee the ordering the contract
promises. A new out_of_order_appends counter records the condition
(not yet exposed on the writer :9102 /metrics —
ai-meta#206).
Kind before/after on an identical 40-concurrent burst, Option-2 topology:
| delivered | stuck (command.issued, no command.claimed) |
feed lag | |
|---|---|---|---|
| pre-fix | 17/40 | 23 permanent | 0 (silent) |
| post-fix | 41/41 | 0 | 0 |
Plus exactly-once 335=335 (0 dup) and pod-kill redelivery 25/25.
Releases carrying the fix: noetl-worker v5.81.1 (the writer host lives in
the worker, so this is the image that carries the runtime fix) +
noetl-server v3.58.1. Deployed to prod by digest after a crane copy
ghcr → project Artifact Registry.
EHDB command dispatch (issued → claimed) is p50 285–557 ms
(p99 ~1040, min ~145) against the ~200 ms NATS baseline — roughly
1.5–2.8×. Sub-second, and neither correctness nor completion was affected,
but it is a real regression on the hot dispatch path and it is the open
go/no-go item for T5: once NATS is deleted there is no cheap comparison
left. Two suspects, not yet isolated — the 250 ms poll-claim interval versus
NATS's push delivery (a push variant is already available via the existing
FeedWriter::tip_receiver() seam), and the writer-side ordering-key append
the fix introduced. Tracked in
ai-meta#205.
Issues: #203 closed
(prod-validated), #202
closed (Option 2 executed), #194
stays open for T5 + L2/L3. Full three-round execution log:
Runbook — L1 Command-Bus Cutover. Prod IaC:
noetl/ai-meta playbooks/194-l1-t4-prod-iac/.
2026-07-18 (L1): ✅ L1 NATS-takeover T0–T3 shadow arc COMPLETE — new ehdb-feed crate; T4 cutover human-gated
Built the L1 (NATS takeover) shadow delivery path on top of the finished L0
foundation, topology (c) per-shard-writer-as-broker, in a new
ehdb-feed crate. Each slice its own PR, full CI gate green on 1.97.1,
real proof. Additive, kind/local, NATS still authoritative, no prod.
-
T0 shadow feed (#290
Watch(shard,cursor)primitive + engineread_partition_after; #291 networkedFeedWriter/serve/FeedSubscription+ latency harness). Latency gate PASS: bus append→subscriber p50 57µs / p99 137µs (NATS-parity). The ~4ms first observed was diagnosed as posture-A fsync-per-append durability on the slow sandbox FS (sub-ms on NVMe; group-commit amortizes; NATS-JetStream shares it) — isolated and reported separately, not the delivery bus. Parity 0 missed / 0 spurious. -
T1 consumer groups (#292) —
ShardConsumerGroup: competing consumers, ack / ack_wait redelivery (at-least-once), crash-redelivery (0 loss), committed-cursor resume, shard routing (#166 equivalent). Clock-free deterministic coordinator. -
T2 KEDA lag signal (#293) — per-shard backlog gauge (
ShardConsumerGroup::lag) + Prometheus/metricsexposition. Ready before T4 (hard-ordering rule). -
T3 gateway/SPA SSE feed (#294) —
serve_ssetext/event-stream; SSEid:/Last-Event-IDmaps onto the ChangeFeed cursor, reconnect resumes 0-missed/0-dup.
T4 (command-bus cutover) + T5 (delete NATS) are human-gated and NOT
executed. Prepared the reviewable cutover artifact:
Runbook — L1 Command-Bus Cutover — the
NOETL_COMMAND_BUS=nats|ehdb flag, the rolling-restart sequence with
validation gates, the go/no-go checklist, and the one-command rollback.
ai-meta gitlink → ehdb 7e0b896 (ai-meta@5f059e75). Umbrella
noetl/ai-meta#194 stays open
(T4/T5 + L2/L3 ahead).
The EHDB L0 replicated object store is finished: the engine slices
(L0.1–L0.7) plus every predefined dataset D1–D10 are merged to main
(0af99fc). D6–D10 landed this session, one PR each, each with the full CI
gate green on rustc/clippy 1.97.1 and real proof — functional +
cold-load + N-way replica-kill fallback (read_fallbacks > 0) — on
the generic L0Engine<D>.
-
D6 vector/RAG (#285,
eb8035e) — upsert / top-k cosine / delete. -
D7 catalog (#286,
ddf2b22) — register / get (latest+pinned) / snapshot / deregister (moduleregistry.rs, avoids the internal meta-catalog name). -
D8 runtime registration (#287,
c7b3a82) — register / heartbeat / deregister / list-live with a wall-clock-free heartbeat-watermark eviction. -
D9 system-WASM store (#288,
15f0077) — publish / bind (rollback) / resolve / unpublish over a(path,channel,env)triple; content-addressed bytes + op-log (D5 shape). -
D10 provider-facts (#289,
0af99fc) — set-desired / set-observed (carry-forward fold) / drift-scan / forget; the credential-freeprovider_statefold behindnoetl provider {plan,drift,orphans,adopt}(#189).
The shape: every mutable-state dataset is an append-only op-log over immutable parts; current state is a fold (last-op-wins per key, tombstones dropped). No in-place mutation; parts stay immutable; each dataset cold-loads
- replica-fails-over for free. noetl-native throughout (no MinIO/S3 — noetl
owns manifest/index/replication over the pluggable
DurableSubstrate). kind/local only, no NATS, no prod; allNOETL_EHDB_*paths default off.
ai-meta gitlink → ehdb 0af99fc (ai-meta@4da4fae). Umbrella
noetl/ai-meta#194 stays open
— L1 (streaming / NATS takeover), L2, L3 are ahead, gated behind L0, and
not starting without a go-ahead.
2026-07-12 (durable-backend): ✅ Alesha APPROVED "GO to prod durable-SHADOW" — sign-off recorded (no prod action)
Recorded the durability sign-off decision. Alesha approved GO to prod durable-SHADOW on 2026-07-12; the §C durability gate is signed off for the shadow stage. DOCS ONLY — no prod/GKE flag flipped by any agent; enabling durable-shadow is the prod team's explicit next operational step. Tracks ehdb#254.
- Runbook — Durable Event-Log Prod Sign-off: header → SIGNED OFF; new §6 DECISION — APPROVED (Alesha, 2026-07-12) block (scope = durable-SHADOW only; durable-primary still separately gated); §5 refreshed for the prod team — Stage A′ target image updated stale v5.70.1 → current durable-capable release (worker merged main, carries GC + retention default-off), Stage B′ gained a recommended retention config to bound the shadow PVC (shadow has no consumer), Stage C′ gate reworded (R1 segment-GC decision now MET → set the retention policy). ehdb#254 slice-6 box CHECKED.
- Runbook — Prod Cutover: Event-Log Tier: §C durability gate banner → RESOLVED (2026-07-12); §7 decision-log durability-gate row → checked (durable_segment, shadow scope).
- What stays gated: Stage C durable-PRIMARY — ≥24 h clean shadow window + retention policy set on the writer pool + R2/R3 storage-class/topology, each its own explicit go.
2026-07-11 (durable-backend): prod-durability DECISION PACKET refresh — D11 GAP→MET + #261 perf re-run
Assembled the decision-ready packet for the slice-6 prod-durability sign-off
(prep for @alesha's go/no-go — no prod flag flipped, decision stays with the
owner). LOCAL kind only — no GKE/prod; noetl.event never purged; no
secrets. Tracks ehdb#254 +
ehdb#261.
-
#261 head-to-head RE-RUN on the merged-main image
ehdb178184-merged(carries #264/#266/#267 + GC/retention). Deployed durable append ~740 ms → ~6 ms (~100–200×) — the O(segment) shared-publish bottleneck is gone. Layer-A authoritative (criterion): per-op-open+append flat at 5.11–5.35 ms driver / 13.9–14.5 ms full stack across S=100→10 000. Layer-B incumbent same-day: ~90 req/s p50 163 ms — not a regression, pure podman-VM contention (directional only; Layer A + the deployed re-measure are trusted). Design-Performance § #261 re-run. -
Slice-6 sign-off refresh: D11 GAP → MET (segment GC + keep-last-N
retention shipped + organic soak: shadow store 14→2 segments, held flat,
reclaimedclimbing, replay gapless-from-base, 0 restarts). Verdict + R1 (High Stage-C-blocker → Low operational) + R6 (perf de-risked) updated; no GAP remains. New §6 GO/NO-GO summary — recommend GO to prod durable-SHADOW; Stage-C durable-primary stays separately gated on the ≥24 h shadow window + setting the retention policy (operational) + R2/R3. - Open follow-ups (not blockers): prod shared/object-tier medium (R2), max-age retention deferred (mtime resets on hydrate — keep-last-N is robust), the read-write driver / prod-primary phase.
2026-07-10 (durable-backend): segment-GC limits-based retention — a shadow store self-bounds without a consumer
The optional follow-up after the GC arc closed: interest-based GC reclaims
nothing without a durable consumer, so the deployed shadow event-log never
self-bounds. Added keep-last-N limits-based retention that reclaims
independent of consumer interest, composed with interest so it never reclaims
ahead of a real consumer under primary. Tracks
ehdb#254. LOCAL kind only — no
GKE/prod; noetl.event never purged; no secrets. Detail:
Design → Limits-based retention.
-
Merged ehdb#271 → ehdb main
49db503(engine) + worker#178 (knob) + worker#179 → worker main521fd10(re-pin ehdb to the merged #271 SHA).plan_reclaimcomposes interest + retention intomin(retention_target, interest_ceiling); the only engine change isplan_reclaim. Keep-last-N (count-based), not max-age (mtime resets on hydrate/cold-load). ConfigNOETL_EHDB_EVENTLOG_GC_MAX_RETAINED_SEGMENTS+NOETL_EHDB_EVENTLOG_SEGMENT_MAX_BYTES. - 251 tests (+3 retention: self-bounds a consumer-less store, never reclaims ahead of a lagging consumer, respects the min floor); clippy + fmt clean; bench variant.
-
Organic in-kind soak (
v5.72.4-ehdbgcret254, user pool,shadow, no consumer, keep-3 + 1 KiB segments): the live store grew to 19 segments under appends, then the periodic task self-bounded it to 2 segments / 2,147 bytes (from ~22 MB) and held flat —eventlog_gc_ops_total{outcome="reclaimed"}=6, noerror,last_degraded=0, 0 restarts; reclaim manifest committed (reclaimed_through_seq: 25331), read-only reopen replayed the retained log gapless-from-base (replay intact under organic reclamation).
2026-07-10 (durable-backend): segment-GC worker wiring + in-kind soak — D11's first deployed exercise
Wired the durable event-log segment GC to run periodically in the worker and
proved it in the local kind cluster — the operational step that turns D11 from
"mechanism proven" into "running in kind". Tracks
ehdb#254. LOCAL kind only — no
GKE/prod; noetl.event never purged; no secrets. Full detail:
Design → Worker periodic invocation.
-
Worker wiring (noetl/worker, ehdb-reference pin →
adca378): a periodiceventlog_gctask reclaims each owned shard (local + shared viareclaim_shard) on aspawn_blockingthread; a process-global per-shard advisory lock serializes GC against the per-op append path (resolving the two-writers-per-shard fork + closing a latent append↔append race). Spawns only when opted in on every axis (durable_segment+NOETL_EHDB_EVENTLOG_GC=consumer_ack-
NOETL_EHDB_EVENTLOG_GC_INTERVAL_SECS>0+ data-plane contract); all default-off.noetl_ehdb_eventlog_gc_*metrics;ehdb-selfcheck durable-eventlog-gc [--drive 1]verb.
-
-
Fork surfaced — shadow has no consumer. No
tail/acksite consumes the deployedshadowevent-log, so interest-based GC correctly reclaims nothing organically (anoop, not a bug). Effective reclamation needs a durable consumer (natural underprimary) or a limits-based retention mode for shadow (the follow-up). -
In-kind soak (2026-07-10): worker
v5.72.3-ehdbgc254on the user pool (/ehdb-durablePVC, GC + interval=30s). Periodic task spawned + ran healthy (eventlog_gc_ops_total{outcome="noop"},last_ok=1);--drive 1on the real PVC reclaimed shared 40 → 10 at watermark 30,gc_holds=true(replay + cold-load + new-owner hydrate all coherent); 0 restarts. User + system pools rolled uniform tov5.72.3-ehdbgc254.
2026-07-09 (durable-backend): shared-tier segment GC — coherent reclamation across the shared medium
Made segment GC coherent for the shared-medium topology (the deployed shape:
single-writer system pool + a shared PVC), the follow-up the same-day local-GC
slice scoped. Tracks ehdb#254.
Engine + CLI + bench in repos/ehdb only; repos/noetl + repos/server
untouched; kind-only, no GKE/prod; noetl.event never purged; no secrets.
Full detail:
Design: Durable Event-Log Backend → Shared-tier reclamation.
-
The problem — local GC alone leaves reclaimed segments in the shared store,
so
cold_load/hydrate_owned_shardre-pull them and diverge from the owner (and a new owner un-reclaims on transfer). -
The fix — a monotonic per-shard cross-replica reclaim watermark on the
shared medium (
SharedSegmentBackend::put_reclaim_watermark/reclaim_watermark/delete_segment; default-erroring so a non-capable backend refuses shared GC rather than losing coherence).SharedTierEventLog::reclaim_shardreclaims local + shared watermark-first (commit the watermark before any delete), so readers skip<= watermarkand can never re-pull a segment mid-reclamation.reconcile_owned_shardrealigns a crashed owner whose local segments are ahead of a committed watermark. Store refactored intoplan_reclaim(decision) +apply_reclaim/reclaim_to_segment(mutation) to enable watermark-first. -
Crash-safety — a crash after the watermark commit leaves orphan local/shared
objects readers already skip: a bounded space leak the next
reclaim_shardre-attempts, never a divergence. The only transient is a crashed owner briefly over-serving its own un-reclaimed segments until it reconciles — benign (shows more, never less; data also innoetl.event) + self-healing. -
Validation — 248 tests pass (+5 shared-GC tests incl. a crash-window
reconcile test proving owner+non-owner realign to a committed watermark);
clippy
-D warnings+ fmt clean. CLIdurable-eventlog-shared-gc(2 shards): shared store pruned 40 → 10 objects at watermark 30, non-owner cold-load + new-owner hydrate both coherent.eventlog_shared_tier_gccriterion group added. - D11 verdict: GAP → delivered + drive-tested for BOTH the single-writer local (PVC) AND the shared-medium topology — unblocks the Stage-C durable-primary consideration for the deployed shape. Remaining is operational: worker periodic-invocation wiring + an in-cluster GC soak.
Closed the segment-GC gap the prod-durability sign-off
named as residual-risk R1 / D11 — the Stage-C blocker: slices 1–5 never
reclaimed a sealed segment, so under KeepAll+primary the durable event-log
store grew unbounded. Added interest-based segment GC (consumer-ack
watermark) in ehdb-reference::durable_eventlog. Tracks
ehdb#254. Engine + CLI + bench in
repos/ehdb only; repos/noetl + repos/server untouched; kind-only, no
GKE/prod; Postgres noetl.event never purged; no secret values. Full detail:
Design: Durable Event-Log Backend → Segment GC.
-
The change (ehdb#269): a sealed
segment is reclaimable once every durable consumer has acked past its last
global sequence (JetStream-style interest), respecting a
min_retained_segmentsfloor. BehindNOETL_EHDB_EVENTLOG_GC=consumer_ack(off by default). -
Invariants preserved — base offset (retained log gapless-from-
reclaimed+1;reclaimed_through==0default ⇒ byte-identical to pre-GC); durablefsync'dreclaim.jsonmanifest = the crash-atomic commit point (crash mid-GC re-deletes below-watermark leftovers on the next open, both the replay and O(1) checkpoint-trust paths); write-forward of consumerAckframes before unlink so replay-is-truth reconstructs consumer state. Fixedtail()to compute pending off the absolute tip, not the retained count. -
Validation — 242 tests pass (+10 GC tests incl. crash-mid-GC-via-manifest,
base-offset reopen, watermark, floor, RO refusal, fail-safe parse); clippy
-D warnings+ fmt clean. CLI drive (40 events, 220 B segments): 41 → 11 segments reclaimed at watermark 30, every invariant holds.eventlog_segment_gccriterion group added. -
D11 verdict: GAP → mechanism delivered + proven for single-writer local (PVC).
Unblocks the Stage-C durable-primary consideration for that topology.
Shared-tier object reclamation + watermark coherence is the one remaining
follow-up (else
hydrate/cold_loadre-pull reclaimed segments from shared); GC is local-only, off by default, and not auto-invoked, so it can't silently diverge.
Fixed the dominant deployed cost the 2026-07-08 session isolated: the worker
rebuilds the durable stack per op, so every mirrored append paid a fresh
DurableSegmentStore::open that replayed every segment to rebuild the offset
index (O(segment)). Tracked ehdb#267
(Refs #261). Engine fix in
repos/ehdb only; repos/noetl + repos/server untouched; worker pin bump +
rebuild; kind-only, no GKE/prod; no secret values. Full detail on
Design: Performance & Load Testing → #267 fixed.
-
The fix (#268, merged
f6fdaea, closes #267): a smallcheckpoint.jsonsidecar —{event_count, active_segment_id, active_len, consumers_seen, consumer_acks}— rewritten after each mutating op strictly after the framefsync. Open-for-append loads it O(1) and skips the replay; the offset index is materialised lazily on the first read (ensure_index_loaded). Mirrors the #266 resumable-digest sidecar.scan_global/read_executionbecame&mut self; read-only cold-loads still replay eagerly. - Durability preserved — the checkpoint is a cache, never truth: replay-is- truth fallback on missing / stale / inconsistent checkpoint (strict active-segment length anchor → can never name more durable data than the segments hold); CRC integrity enforced on every read; torn-tail / fsync / single-writer / byte-exactly-once unchanged.
-
Micro-bench (authoritative): per-op-open+append flat — driver
5.11/5.21/5.35 ms, full stack 13.9/14.2/14.5 ms at S=100/2 000/10 000 (was
O(segment)). 232 tests pass (+5 checkpoint tests); clippy
-D warnings+ fmt clean. -
Deployed before→after (kind user pool, ~20 MB / 24 765-record store):
noetl_ehdb_eventlog_last_duration_seconds~0.509 s (ehdb266) → ~4–16 ms (ehdb267), a ~30–120× per-op reduction. First-ever op paid a one-time legacy replay + wrote the sidecar; every op after is O(1).hello_worldend-to-end COMPLETED; both user + system pools rolled clean (0 restarts). -
Merges + pins: ehdb#268 → main
f6fdaea; worker #176 →0597351(pina36484f→f6fdaea+ Cargo.lock + 2 read-sitelet mut; version stays 5.70.3). Imagelocalhost/noetl-worker:v5.72.0-ehdb267.
2026-07-08 (perf): shared-tier publish fix (#264 → #265+#266) shipped + validated in kind — bottleneck corrected to local replay-on-open (#267)
Fixed the shared-tier publish O(segment) bottleneck Layer B surfaced, validated
it in kind, and corrected the Layer-B headline. Tracked
ehdb#264 (Refs
#261). Engine fix in repos/ehdb
only; repos/noetl + repos/server untouched; worker pin bump + rebuild;
kind-only, no GKE/prod; no secret values. Full detail on
Design: Performance & Load Testing → Fix landed.
-
Harness #263 merged (→ ehdb main
d5ed46d), making the Layer B load test reusable. -
The fix (two PRs): #265 —
publish_shardsends only the append-delta (not the whole segment) via a newSharedSegmentBackend::append_segment. #266 — the worker builds the stack per op, so #265's in-memory digest cache was empty every append (re-reading the whole prefix to re-seed); #266 persists the resumableXxHash64state on the integrity sidecar (twox-hashserialize) so a fresh backend resumes the digest in O(delta). Guarantees preserved; whole-prefixdigestunchanged. 227ehdb-referencetests pass; micro-bencheventlog_shared_tier/append_at_sizeis flat ~12–13 ms at pre-fill 100 / 2 000 / 10 000 (no scaling with segment size). -
Validated in kind — worker rebuilt on the #266 ehdb pin
(
localhost/noetl-worker:v5.72.0-ehdb266), user-pool + system-pool rolled to it (the two pre-wedged subscription pools couldn't roll — pre-existing broken init container — reverted to rc4).hello_worldend-to-end green. -
Corrected headline: with #266 the shared publish is provably O(delta), yet
the deployed per-op mirror stayed ~0.6 s. Cause: the worker rebuilds the durable
stack per op, and
DurableSegmentStore::openreplays every segment to rebuild the offset index — a second O(segment) cost, independent of the shared publish, that the earlier decompose conflated with it. Proven three ways: (a) fresh dir ~4 ms; (b) cost scales with local size (8 MB ~0.2 s / 16 MB ~0.65 s / 20 MB ~0.6 s) while shared is O(delta); (c) deployedehdb-driverate ~1.3 append/s, p99 ~1.65 s — unchanged. Filed as #267 (the dominant deployed cost). -
Disposition: #264 (shared publish) closed — fixed + validated as O(delta);
#267 (local replay-on-open) is the remaining deployed bottleneck; the deployed
durable path stays shadow-only until #267 lands. All
NOETL_EHDB_*flags default off. -
Merge decision (Alesha): merge on the micro-bench flat-append evidence
(the authoritative signal per methodology); don't gate on the slow in-kind
rebuild. Shared-tier publish SLO now met at the engine level. ehdb #265+#266
merged (main
a36484f); worker pin PR noetl/worker#175 merged (fc203b9— ehdb pin →a36484f, pin+lockfile only). Deferred (not blocking): re-run Layer B to confirm the as-deployed rate/p99 improvement after #267 lands + when the VM isn't build-contended (add swap / build off-VM). The in-kind run this session already isolated #267 as the gate.
Layer B of the perf/load-testing workstream (ehdb#261,
PR ehdb#263). Ran in kind
(kind-noetl, podman); no GKE/prod; no repos/noetl / repos/server /
worker-runtime source changed — committed harness only. Full results on
Design: Performance & Load Testing → Layer B.
-
Kind hygiene gate —
NOETL_COMMANDS0 pending on the shared/user pool; no durable orphan JetStream consumers (two random-named ones are live ephemerals with a 5 s inactive auto-reap); event-log path uniform at serverv3.54.0-rc4/ workerv5.71.0-rc4;hello_worldcompleted (7 s warm). (Pre-existing wedged subscription-pool/runtime rollout — broken init containercurl: not found— is independent of the event-log path.) -
Head-to-head (same VM, same ~400 B event envelope): incumbent
Postgres+NATS append (
POST /api/events) sustained ~2 150 ev/s, p50 12.6 ms / p95 33.9 ms / p99 52.5 ms, 0 fail; burst c=200 ~1 600 ev/s, p99 300 ms, 0 fail. EHDB durable local engine primitive (fresh segment, in-VM) ~3–10 ms (corroborates Layer A 3.9 ms). EHDB as-deployed shared-tier shadow mirror ~0.55–1.7 s/append (p50 ~740 ms, p99 ~1.17 s), throttling worker emission to ~1–2 append/s and building backlog. -
Headline finding:
durable_segmentin the worker runtime is always the composed slice-3 shared tier;SharedTierEventLog::appendcallspublish_shardsynchronously insideemit_event, re-reading+re-writing the whole active segment per append (O(active-segment-size)). In-VM decomposition: fresh empty segment ~3–10 ms vs live ~20 MB segment ~0.55–1.7 s — the bottleneck is the shared-publish strategy, not the durable engine or the VMfsyncalone. -
Corroborates Layer A on the load-bearing claim (local durable append is
cheap + flat; reference
O(n²)is real); the specific 2.7× ratio stays a Layer-A-only measurement by construction (worker never runs local-only orlocal_referenceas a serving path). - SLO revisit: the single "durable append p99 ≤ 10 ms" target must be split by backend — the local primitive meets it, the shared tier (~1.2 s) cannot. Incumbent-parity would be ~50–100 ms p99.
-
Follow-up (revises Phase-1 conditional): priority = incremental
shared publish (publish new/sealed bytes, not the whole active segment) +
move it off the synchronous hot path; group-commit is secondary (the
local
fsyncceiling was not the in-cluster limiter).
New perf/load-testing workstream (ehdb#261).
Code + benchmarks only — no image build, no deploy, no GKE/prod;
repos/noetl + repos/server untouched. Design doc:
Design: Performance & Load Testing.
- Deliverable 1 — design doc (new page). Goals + the headline question (quantify the event-log bottleneck fix: EHDB durable event-log vs the incumbent Postgres+NATS-JetStream log-and-store); the two test layers (A: engine micro-benches, reliable, this phase — B: in-cluster kind load + EHDB-vs-incumbent head-to-head, directional, next phase, podman-VM caveats); per-tier metrics; the baseline data as first data point; a proposed SLO strawman for @alesha to confirm; harness + reproducibility.
-
Deliverable 2 — benches (
crates/ehdb-reference/benches/engine_micro.rs,cargo bench -p ehdb-reference --bench engine_micro). All five drivers + the durable segment backend, seeded xorshift64* payloads, isolated temp dirs,sample_size=10. No engine semantics changed; helpers bench-local. -
Baseline (Apple M1 Max, 10 cores, 32 GiB, macOS 26.3.1, APFS SSD,
rustc1.92.0 — a dev box, not a bench rig):-
Event-log headline —
durable_segmentbeatslocal_referenceat sustained append: 255 vs 96 ev/s at K=1000 (2.7×), widening with N (the reference driver replays the whole JSONL twice per append →O(n²)). Durable single-append flat at ~3.9 ms (fsync-bound → ~256/s/writer) vs local rising 9.3 → 15.9 ms with size. Segment rotation ~2% overhead (16 KiB vs 8 MiB segments). Cold replay ≈185 K ev/s (27 ms for 5000). - Projection fold ~1.48 K ev/s; checkpoint read 4.2 ms.
- KV (size 1000): put 120–167/s (degrades), get 5.2 ms, scan 5.9 ms, CAS 5.2 ms.
- Object: 1 MiB put ≈73 MiB/s (blob + SHA-256 bound), 4 KiB put 5.9 ms, get-verify/locate/list ~1.1 ms @ 200.
- Vector (dim 1536): upsert @256 143 ms, cosine top-k=10 @128 28.9 ms /
@512 115 ms — slowest tier; reference re-parses the whole float-array
JSONL per op (
O(catalog×dim)), not a serving index (HNSW/IVF = future).
-
Event-log headline —
-
Honest findings: only the event-log tier has a durable backend today;
KV/object/vector reference drivers are
O(n)-per-op shadow fixtures. Durable append isfsync-ceilinged at ~256/s/writer (group-commit is the lever if a hot shard needs more). Vector reference driver measured slowest, as expected for a replay-per-op fixture — reported honestly. -
Trails: design page +
_Sidebar+ Home Current-Implementation-Notes + this log; ehdb#261 opened + on project 4; ai-meta MEMORY + topic file + AGENT-COORDINATION board. PR:perf/engine-micro-benchmarks→ ehdb main.
2026-07-08 (kind-first): Phase 3.5 — #151 keychain PRs MERGED, event-log leak FIXED, Auth0 GREEN, OpenAI+IBKR scoped out
Merge + close-out of the #151 keychain work, with a real safety catch. No
GKE/prod; no secret values printed or committed. Headline in
Kind-Full-Functionality-Validation §0.
-
Leak-safety re-verification caught a real leak before merge. A fresh
standalone keychain-http probe (hop-1) deferred cleanly, but a fresh
two-step probe matching the
ops/execution_ai_analyzeopenai_triageshape (http step referencing a prior step's output and a{{ keychain.* }}header) leakedBearer sk-<key>intocommand.issuedon rc3. So the leak was not a "heavily-rerun anomaly" — it reproduced on any hop≥2 drive-built keychain-http step. -
Root cause: the off-server drive runs as an
__orchestrate__(tool_kind=wasm) command whose input embeds the whole playbook incl. follow-up steps'{{ keychain.* }}; the worker's generic dispatch raninject_keychain_namespaceover that input, resolving the secret into the drive's context, which the drive then persisted into the follow-upcommand.issued. -
Fix (worker#174): skip keychain injection/render for the
__orchestrate__drive command; the control plane operates on deferred placeholders only, keychain resolves at the terminal user-pool dispatch (unchanged). Rebuiltnoetl-worker:v5.71.0-rc4, rolled all 4 pools. Before/after: rc3Bearer sk-<LEAKED>→ rc4Bearer {{ keychain.* }}(0Bearer sk-); user pool still resolves at dispatch; http call succeeds. Regression test added. -
Merged (squash): server#279 → v3.53.2 (
fe500df9), worker#174 → v5.70.3 (2031a8b4, incl. leak fix), e2e#85 (252006f8), ops#235 (541e6b0d). -
Auth0 password-grant fixture GREEN: test-user password stored in GSM
(
auth0-test-user-password), referenced via aprovider: gcpkeychain entry — never a plaintext workload input. Auth0 returned HTTP 200 + real access_token (expires_in 86400); playbook COMPLETED via success path; password absent fromnoetl.event. eid333298747507216384. Verified by boolean/length only. -
Scope: OpenAI (
429 insufficient_quota, external billing) + IBKR (live gateway) formally scoped out per Alesha; ops-LLM trio + amadeus OpenAI steps scoped-out for their OpenAI-dependent portion only (keychain proven). - Pointer bumps: server
fe500df9, worker2031a8b4, e2e252006f8, ops541e6b0d, ehdb-wiki (this doc).
The last-mile pass on the kind full-functionality effort. #151 fixed as a platform change, GSM bridge made durable, and the remaining non-green reduced to external account limits (like IBKR). No GKE/prod; no secret values printed. Matrix headline in Kind-Full-Functionality-Validation §0.
-
#151 root cause (deeper than the title):
{{ keychain.<alias>.<field> }}in a tool config is rendered by the drive (orchestrate-core::build_tool_command, in thesystem/orchestratewasm plug-in) against a context with nokeychainnamespace → empty → 401. -
Fix: orchestrate-core
render_value_deferring_keychaindeferskeychain.*through the drive; the worker resolves it transiently at dispatch (secret never innoetl.event); the server now resolveskind: credentialkeychain entries. PRs server#279, worker#174, e2e#85 (fixture DSL drift:payload→json, auth0endpoint→url, amadeus explicit token step), ops#235 (durable bridge). Kind imagesserver:v3.54.0-rc4/worker:v5.71.0-rc3, all 4 pools uniform. -
Proven (real GSM): openai
GET /v1/models200, header deferred incommand.issued;kind: credentialprobe deferred+resolved, no leak; Amadeus oauth2 token 200 + flight-offers 200. -
Residual = external, not platform: OpenAI chat/completions 429
insufficient_quota(billing on the test key) gates ops-LLM + amadeus full-200; auth0 password grant needs real Auth0 user creds; IBKR needs a live gateway. -
Durable GSM bridge (
ops#235): fixed-clusterIP (10.96.0.53) relay +hostAliases+ launchd host shim. Pod-restart proof PASSED. -
Known follow-up (not overclaimed): heavily-rerun
execution_ai_analyzeopenai_triageresolves the key drive-side (event-log exposure) while fresh probes defer — not root-caused; auditrender_pipeline_config. - PRs OPEN, not merged — the drive-side leak-flag is surfaced for review; no ai-meta pointer bump this session.
2026-07-08 (kind-first): full-functionality validation Phase 2 — GSM external integration + Muno end-to-end
New capability (Alesha): kind CAN reach Google Secret Manager. Phase 2
stood up the external-integration tier for real in LOCAL kind and drove the
Muno/travel planner end-to-end. No GKE/prod touched; no secret values
printed; no repos/noetl/repos/server source change. Matrix on
Kind-Full-Functionality-Validation
(§6–§8).
-
GSM metadata bridge (deploy-time only): host-ADC token shim
(
gcloud auth application-default print-access-tokenon:48710) → in-clustersocatrelayDeployment/Service gcp-metadata→hostAliases metadata.google.internalon all 4 worker pools; plusNOETL_GCP_METADATA_TOKEN_URL+GOOGLE_CLOUD_PROJECTon noetl-server. Unblocks both GSM paths (server keychainprovider: gcpand the worker in-python metadata read the MCP playbooks use). -
9 external providers LIVE with real GSM creds: Duffel (10 offers),
Google Places (10), HotelBeds hotels/activities/transfers (10/10/31),
Firestore (query OK), Snowflake (rows), OpenAI + Anthropic (
/v1/models200). GCS already green (Phase-1 BUG-2). -
Muno planner FULL end-to-end green in kind: flight-search (10 Duffel)
→ book (real Duffel TEST order) → hotels → activities → transfers →
summary+calendar → confirm→
map_view. All turns COMPLETED with live data. -
Kafka + Pub/Sub subscription drains PASS (seed+drain+ack, count=5)
after re-registering the Phase-1 SETUP-C creds. SETUP-D aliases
(
pg_tutorial,noetl_connection) registered. -
Residual (not new platform bugs):
#151keychain-template gap (amadeus_ai_* / ops LLM / auth0 token complete but 401 — creds reachable, fixture uses broken{{keychain}}); IBKR needs a live gateway; 6 FIXTURE-E DSL-drift + 2 fixture-data bugs. GKE stays gated on Alesha's scope call. - Baseline unchanged: server
v3.54.0-rc2, workersv5.71.0-rc1, EHDB shadow + durable_segment (user pool),NOETL_STATE_BUILDER=offserver.
Alesha directive: complete full functionality in LOCAL kind first — test all playbooks — before GKE. Phase 1 = inventory + baseline + first run. LOCAL kind only; no GKE/prod action. New page Kind-Full-Functionality-Validation is the living test matrix that scopes "done in kind".
- Catalog: 268 playbook/automation files inventoried across e2e + noetl + ops + travel, categorised self-contained / in-cluster-infra / cred-blocked / infra-deploy / GKE-only.
-
Baseline standardised: all 4 worker pools →
noetl-worker:v5.70.0(was skewed v5.70.0 user / v5.69.0 others), server v3.53.0, EHDB tiers shadow, durable_segment event-log backend on the user pool PVC (byte- identical, non-primary). Reset cleared a system-pool orchestrate backlog (~720 stale__orchestrate__re-drives draining ~0.6/s) + leaked NATS consumers (drifttest23 992 msgs; 8 orphanednoetl_eventsconsumers) that were starving new executions. Post-resethello_world= 4 s. Postgresnoetl.eventnever touched. -
Runs: core self-contained gate 61 PASS / 4 FAIL (65); extended
in-cluster-infra 9 PASS / 11 non-green (20); EHDB probes 2/2
(
pft_sql_probe_v2,large_tabular_result_test). ~70 distinct green. EHDB projection/query/durable live; KV/object/vector shadow (segments 26.9→32.3 MB across runs). -
Two platform bugs (each = 1 root cause, multi-fixture): BUG-1 pagination
max_iterations/retrycontinuation wedge; BUG-2 large-resultoutput_select/storage-tier resolve →/api/result/resolve404 (3 fixtures). Rest non-green = cred setup (kafka/pubsub decryption, 2 missing aliases) + 6 fixture DSL-drift. Core execution model is healthy in kind. - ehdb-wiki: new validation page +
_Sidebar+ this log.repos/noetl+repos/serveruntouched; no prod/GKE.
2026-07-07 (sign-off): durable event-log backend — prod-durability sign-off PACKAGE (ehdb#254 slice 6)
Assembled the slice-6 prod-durability sign-off package — decision-ready
evidence for clearing the tier-1 prod-cutover §C durability gate. DOCS/PLANNING
ONLY — no prod/GKE action, no flag flip, no cutover; ehdb#254 slice-6 box left
UNCHECKED (that's for after the actual prod sign-off). repos/noetl +
repos/server untouched; no Rust/infra code change.
New page Runbook-Durable-EventLog-Prod-Signoff:
- Go/no-go checklist (D1–D11): D1/D2/D6/D8/D10 MET; D3/D4/D5/D7/D9 mechanism-MET but NEEDS-PROD-VERIFICATION (real GKE PVC / topology / GMP, which the prod durable-shadow soak closes); D11 = GAP (no segment GC). Verdict: GO to prod durable-SHADOW; Stage C durable-PRIMARY gated on the soak passing + the segment-GC decision.
-
Evidence bundle — slices 1–4 PRs/commits + tests; the slice-5 kind soak
(rotation
seg-00018 387 133 B; metrics past 11 731 zero invalid/degraded; crash recoverysha256 c2ba3d33…byte-identical, replay 11 856, gapless). - Residual-risk register (R1 segment-GC = Stage-C blocker; R2 GKE storage class; R3 multi-replica affinity on the real event-writer pool; R4 disk pressure; R5 backup/DR; R6 perf; R7 selfcheck harness artifact) each with a prod-verification step / mitigation.
- Extended durable rollout sequence (A′ deploy v5.70.1 flags-off → B′ prod durable-shadow on a PVC → C′ durable-primary, separately gated) with metrics/ alerts (PVC free-bytes alert since no GC) and per-step rollback — extends the tier-1 runbook rather than duplicating it.
-
Alignment deltas vs the tier-1 runbook: Stage A target image v5.66.0 →
v5.70.1; new
NOETL_EHDB_EVENTLOG_BACKEND=durable_segment+NOETL_EHDB_EVENTLOG_DURABLE_DIR(PVC) env axis; emptyDir → real PVC.
Cross-linked from the tier-1 runbook (§C update + Related), the design page,
and _Sidebar. ehdb#254 slice-6 package comment posted (box unchecked).
The gated action awaiting Alesha's go: Stage A′ + B′ (prod durable-shadow);
Stage C′ NOT authorised by this sign-off.
2026-07-07 (soak): durable event-log backend — kind soak + real-pod-restart crash recovery (ehdb#254 slice 5)
Ran the durable durable_segment event-log backend under sustained real traffic
in the local kind-noetl cluster and proved crash recovery on a real pod
restart — LOCAL kind only, no GKE/prod, repos/noetl + repos/server
untouched. Clears slice 5 of the durable-backend program
(ehdb#254); slice 6 (prod §C sign-off)
remains.
Deploy. Built localhost/noetl-worker:v5.70.0 (slice-4 durable wiring) —
CARGO_BUILD_JOBS=2 temp Dockerfile guard (0 B swap on the 6 CPU/20 GB podman
VM; reverted after), podman save + kind load image-archive. Rolled onto the
noetl-worker-rust user pool (role worker, single-owner shard 0) with
NOETL_EHDB_EVENTLOG_BACKEND=durable_segment +
NOETL_EHDB_EVENTLOG_DURABLE_DIR=/ehdb-durable on a 2 Gi RWO PVC
(ehdb-durable-soak), atop the existing EHDB shadow env.
Proven. (1) Real drive events land in CRC-framed
/ehdb-durable/local/shard-0000/seg-*.eslog; slice-3 shared tier publishes each
byte-identical to /ehdb-durable/shared/. (2) Segment rotation fired in
cluster — seg-0001 sealed at 8 387 133 B (< 8 MiB) → seg-0002 opened. No
JSONL written. (3) noetl_ehdb_eventlog_ops_total{outcome="mirrored"} advanced
past 11 731, last_ok=1, last_degraded=0, zero invalid/degraded/
routed_away; 0 restarts/crashloops. (4) Crash recovery on kubectl delete pod --force → fresh pod on the same PVC: sealed segment byte-identical
(sha256 c2ba3d33…), active tail replayed + continued (gapless sequence
…11852,11853,11855), and ehdb-selfcheck durable-eventlog reopened the store
read-only replaying durable_replay_records: 11856. Backend byte-identical
between the soaked v5.70.0 and the concurrently-advanced v5.70.1 pointer
(the ehdb#172 diff is KV/vector-only).
Trails: Design-Durable-EventLog-Backend (Kind-soak slice-5 section) + ehdb#254
slice-5 checkbox. Left running in SHADOW on v5.70.0 durable_segment for
observation; rollback = kubectl set image … worker=localhost/noetl-worker:v5.69.0
-
kubectl set env … NOETL_EHDB_EVENTLOG_BACKEND-.
2026-07-08 (fix): KV + vector registry subjects hardened with SHA-256 digest tokens (ehdb#259 / worker#172)
Closed the last two tiers carrying the subject-length overflow that #256 fixed in
the object tier. ehdb-reference built the KV subject noetl.kv.<bucket>.<hex(key)>
and the vector subject noetl.vec.<hex(collection)>.<hex(point_id)> by hex-encoding
the full ids (2 chars/byte); ehdb-stream::Subject::new caps a subject at 256 chars,
so a long id overflows → InvalidIdentifier. Reproduced (math + in-code test):
a ~150-byte KV key → 335-char subject; a real RAG collection (102 B) + point
(117 B) → 449-char subject — both > 256. Not broken in live use only because the
one live KV path uses short circuit.<id> keys and there's no live worker
vector-upsert site.
Fix (ehdb#259, merged 52120a7):
address each subject by a fixed-width SHA-256 digest token —
noetl.kv.<bucket>.<sha256hex(key)> (bucket stays literal so the bucket scan is
unchanged) and noetl.vec.<sha256hex(collection)>.<sha256hex(point_id)> (collection
filter digests the collection identically). Full ids stay in the record payload;
per-id reads filter the replay to the exact id and the vector query fold guards on
the exact collection (digest-collision safety). Tracking issue
ehdb#260 opened + closed (symmetric with
the object #257 trail).
Validation (LOCAL only): cargo test -p ehdb-reference 220 passed (new
long-key/long-id round-trips + in-code reproduction that the OLD hex subject is
rejected by the cap); clippy + fmt clean; selfcheck kv-primary-serve +
vector-primary-serve seeds switched to the real long coordinate form — both exit 0,
served_by_ehdb=true, dual_run_holds=true; object regression clean. No worker image
build. Worker ehdb-reference pin cca0d0d → 52120a7
(noetl/worker#172, pin+lockfile only,
cargo check clean) — in-kind live re-proof deferred (no live long-key path
exercises it yet). All NOETL_EHDB_* flags default off; prod worker v5.52.0 /
server v3.52.0 untouched; repos/noetl + repos/server untouched. Refs
noetl/ai-meta#241.
2026-07-07 (deploy + live-drive PROOF): worker v5.69.0 in kind — all four wired tiers PROVEN on live drives (OBJECT fix invalid→mirrored, PROJECTION, event-log, KV via spool-circuit runtime)
Built worker v5.69.0 (release 96a8b6b, pins ehdb-reference bbc5047) as a
native-arm64 VM-safe image (CARGO_BUILD_JOBS=2), loaded it into the kind node
(podman save + kind load image-archive), and rolled both data-plane pools
— noetl-worker-system-pool (system role) and noetl-worker-rust (worker role)
— from v5.68.0 to v5.69.0, keeping the all-5-tier shadow env
(NOETL_EHDB_ENABLED=1, {EVENTLOG,PROJECTION,KV,OBJECT,VECTOR}=shadow,
LOCAL_REFERENCE_LOG=/tmp/ehdb-shadow-ref.jsonl). LOCAL/kind only, prod
untouched (repos/noetl + repos/server untouched; config/env only).
Baseline (v5.68.0, the bug): noetl_ehdb_object_ops_total{operation="mirror", outcome="invalid"} 2 on the system pool — the pre-fix state where
ehdb-reference hex-encoded the full object key into the NATS subject and
Subject::new's 256-char cap rejected ~150 B platform keys.
Real drive: registered + exec'd tests/large_tabular_result_test distributed
(exec 333041450319089664, COMPLETED, 13 events) — a >inline-budget
{columns,rows} result that the worker externalizes as an Arrow Feather object
through the ControlPlaneClient::object_put chokepoint (real state-shard /
result-tier platform keys), while normal event-log + off-server state-builder
projection drain run.
PROVEN on v5.69.0 (both data-plane pools, fresh post-roll counters):
-
OBJECT — fix re-proven.
object_ops_total{operation="mirror", outcome="mirrored"}= 2 (system) / 1 (user); zerooutcome="invalid". The ref-store shows the real object at subjectnoetl.obj.055f20bc…9404=noetl.obj.<sha256hex>, 74 chars (under the 256 cap). Clean flip from theinvalidbaseline — this is the ehdb#256bbc5047subject-length fix landing on real keys. -
PROJECTION — live cadence hook proven.
projection_ops_total{operation="materialize",outcome="materialized"}= 4 on both pools via the windowedstate_builder::run_drain_loop→projection::mirror_live_windowhook; parity held, no false key-divergence, noinvalid. -
EVENT-LOG — regression holds.
eventlog_ops_total{operation="mirror", outcome="mirrored"}= 6 on both pools. -
Normal behavior unaffected: drive COMPLETED, server keeps serving
(
/healthok), 0 pod restarts, shadow never serves (noserve/primaryoutcomes). The mirrored exec surfaces end-to-end vianoetl ehdb query executions(serverv3.53.0).
KV — PROVEN on the live circuit path (2026-07-07, later same day). The
subscription pool is a plain command-consumer (no WORKER_MODE=subscription),
so it never builds a SpoolRuntime — the live KV mirror lives in a dedicated
subscription runtime. Stood one up on worker v5.69.0 + shadow env
(noetl-subscription-runtime, NOETL_SUBSCRIPTION_PATH=subscriptions/spool_outage_stream)
after registering the nats_e2e credential + sub_ingest_default playbook + the
spool_outage_stream kind:Subscription (spool + one http-probed warehouse
downstream), creating the SPOOL_OUTAGE NATS stream + spool-drain consumer, and
deploying spool-downstream-echo. The runtime activated (subscription_id 333045155747598336, "spool runtime active", downstreams=1), and
SpoolRuntime::persist_circuit → kv::mirror_live_put fires every ~2 s probe
tick + on each circuit state change:
noetl_ehdb_kv_ops_total{operation="mirror",outcome="mirrored"} climbed
monotonically 10 → 50+ (kv_last_ok=1, degraded=0, zero invalid). An
induced downstream outage (scale echo → 0) tripped
subscription.circuit.opened (trips=1, "buffering to spool") and advanced KV
further; recovery (echo → 1) closed the circuit (another persist_circuit). The
ref-store holds the real circuit record at subject
noetl.kv.noetl_subscription_circuit.<hex>, hex-decoding to
circuit.333045155747598336 — a short key (~90-char subject, well under the
256 cap; the KV live path never hits the object subject-length trap). The runtime
also mirrors event-log (mirrored 2 — its own subscription lifecycle events).
All four wired tiers now proven on live drives (event-log, object,
projection, KV); vector stays deferred (no live upsert site). Whole stack
left on v5.69.0 shadow (3 data-plane pools + the subscription runtime + echo +
SPOOL_OUTAGE stream) so all four tiers keep mirroring live. Rollback: pools →
set image …=v5.68.0 (subscription pool → p1-refs + drop EHDB env); delete the
noetl-subscription-runtime + spool-downstream-echo deployments + the
SPOOL_OUTAGE stream to remove the KV harness.
2026-07-07 (later still): projection runtime mirror LIVE-WIRED (windowed drain hook); vector mirror ready-but-unreachable (worker v5.69.0)
Wired the two remaining EHDB tier mirrors into the live worker runtime, each with
its correct per-tier seam (worker#170,
merged fa64e0a → release v5.69.0 96a8b6b; refs
ehdb#234,
ehdb#241).
projection — live-wired via a windowed cadence hook (NOT per-event).
projection::shadow_project is a batch fold: it reads back the WHOLE projection
log (list_executions) and compares against a full authoritative fold + offset. A
naive per-event hook against the long-lived, unboundedly-accumulating KeepAll
store would report persistent false key-divergence (the store keeps every
execution; the worker index evicts terminals it keeps) — a hook that fires but
lies. The faithful seam shipped: the off-server state-builder drain
(state_builder::run_drain_loop) buffers each drained batch's real events and, at
the natural post-batch checkpoint, calls projection::mirror_live_window once per
batch (on a spawn_blocking thread so sync engine I/O never stalls the drain
reactor). Each call windows the batch into a fresh throwaway per-window store
(unique temp log) so the read-back sees exactly the window's executions — no
cross-window accumulation ⇒ no false divergence — and parity-checks the EHDB
engine's fold against an independent worker-side fold (fold_window_authoritative).
Bounded + stateless. Advances noetl_ehdb_projection_* on real drives.
vector — ready hook, documented-unreachable. There is no platform
vector-upsert in the worker's live loop today: RAG retrieval is read-only; RAG
ingest writes a lexical fabric (RagChunk = text + checksum, not an embedding
vector); vector::mirror_upsert is exercised only by ehdb-selfcheck. The
ready-but-unwired vector::mirror_live_upsert + runtime_hook_env pair (tested to
the same discipline) is provided so the future wire-up is one line, but is
deliberately not invoked — fabricating a call site would create a hook that
never mirrors a real upsert. Precise remaining seam: a future platform-RAG
embed+upsert write site (executor/command.rs) calls it after the authoritative
Qdrant upsert.
-
Discipline (both tiers): armed only when
NOETL_EHDB_ENABLED+NOETL_EHDB_<TIER>=shadow+ data-plane role on the bounded local_reference runtime; strict no-op off/disabled/primary/control-plane; errors panic-isolated (best-effort, metrics only, authoritative path unaffected); secret-free. -
Tests: per tier — arms on shadow-enabled data-plane / no-op disabled+off /
skipped control-plane / error-isolated; projection also proves the windowed
fold produces no false divergence across repeated windows (18 new; all 178
ehdb + 32 state_builder unit tests green). clippy
-D warnings+ fmt clean. -
Safety: LOCAL/KIND-only; no GKE, no prod, no image build inline. Prod stays
worker
v5.52.0, allNOETL_EHDB_*flags default off. object.rs subject-fix (v5.68.1) + durable segment-store untouched (sibling-owned). In-kind live-drive re-proof (redeploy +noetl_ehdb_projection_*advancing on a real drive) is PENDING redeploy.
2026-07-07 (later): object subject-length defect FIXED — SHA-256 digest token (ehdb#256 → worker v5.68.1)
Fixed the object-tier defect the v5.68.0 live-drive proof surfaced (below): the
object mirror rejected every real platform object key. Root cause was in
ehdb-reference::object — the registry subject hex-encoded the full logical
key (noetl.obj.<hex(key)>, 2 chars/byte), and ehdb-stream::Subject::new caps a
subject at 256 chars, so any key past ~123 bytes overflowed → InvalidIdentifier.
Real ~140-150-byte coordinate keys (state-shard #166 / result-tier #104) all
failed; state_materializer's 2 real object_puts were rejected outcome=invalid
by the correctly-armed worker hook.
Fix (ehdb#256, merged
bbc5047; closes ehdb#257): the subject
is now a fixed-width SHA-256 digest token (noetl.obj.<sha256hex(key)>, a
constant 74 chars). The full key already lives in the record payload
(ObjectEnvelope.key), so get/list/locate resolve the real key without ever
reversing the subject; latest_envelope filters the replay to the exact key for
digest-collision safety. Content-addressed blob path unchanged.
-
Tests. New
long_platform_key_subject_is_bounded_and_round_trips(140-byte real key, was rejected before; asserts subject== 74chars< 256and full put/get/locate/list/delete round-trip) +distinct_keys_get_distinct_bounded_subjects. 202 crate tests green; clippy-D warnings+ fmt clean; selfcheckobject-primary-serveseed keys switched to the real long coordinate form →served_by_ehdb: true, exit 0. -
Other tiers. KV (
noetl.kv.<bucket>.<hex(key)>) carries the same latent pattern but its only live path uses shortcircuit.<id>keys → unaffected; vector / projection / eventlog left untouched (short execution ids, or owned by the concurrent durable/mirror sibling). Follow-up: same digest fix for KV/vector. -
Worker. ehdb-reference pin bumped
4c0df81→bbc5047(worker#169); semantic-release cut v5.68.1. Pin + lockfile only, no worker code change. - In-kind object-mirror re-proof PENDING — a follow-up worker redeploy (like the v5.68.0 one) will prove the object mirror now persists real keys.
Refs ai-meta#241. LOCAL/kind only — no image build, no prod/GKE this session.
2026-07-07 (evening): v5.68.0 live-drive proof — eventlog PASS, object BLOCKED by 256-char subject cap, KV unexercised
Redeployed the kind data-plane worker pools (noetl-worker-rust = user,
noetl-worker-system-pool = system) from v5.67.0 to v5.68.0 (image
localhost/noetl-worker:v5.68.0, sha f39a724abf8b, native arm64; built with
podman then podman save → kind load image-archive — the podman-provider-safe
load path), keeping the full 5-tier shadow env unchanged
(NOETL_EHDB_ENABLED=1, {EVENTLOG,PROJECTION,KV,OBJECT,VECTOR}=shadow,
LOCAL_REFERENCE_LOG=/tmp/ehdb-shadow-ref.jsonl). 0 restarts. Reversible to
v5.67.0 (prior image + saved deploy specs).
Drove automation/pft_sql_probe_v2 via POST /api/execute → execution
332960196592668672, COMPLETED, 13 events, 0 failed. Fresh-pod baseline: all
mirror counters empty.
-
eventlog — PASS (regression).
noetl_ehdb_eventlog_ops_total{operation= "mirror",outcome="mirrored"}advanced 0→6 on both pools;last_ok=1,last_degraded=0. Ref-store holds 6 entries for the exec (subjectnoetl.event.exec.332960196592668672, streamnoetl_event_log). Also surfaced through the read-only query API (GET /api/ehdb/executions→ the exec, COMPLETED, event_count 13). -
object — HOOK FIRES LIVE but every real put is REJECTED (blocking defect).
The
state_materializer(system pool,NOETL_STATE_SHARD_WRITE=true) did 2 realobject_puts for this exec (.../state/open.feather146 B,.../state/sealed.feather148 B; authoritative writes succeeded, errors=0). The armed mirror hook intercepted both →noetl_ehdb_object_ops_total{operation="mirror",outcome="invalid"} 2,last_ok=0— 0 mirrored, 2 invalid, 0 in the ref-store. Root cause:ehdb-referenceobject::key_subjecthex-encodes the full key into the registry subjectnoetl.obj.<hex(key)>, andehdb-stream::Subject::newcaps subject length at 256 chars. A real platform object key is ~140-150 B → hex ~290-300 +noetl.obj.→ subject ~300-310 > 256 →InvalidIdentifier. Any object key > 123 bytes fails. Bothstate_materializerandresult_producer_stagebuild keys viacoords.physical_key(...)in the longnoetl/env=…/execution=…/state|result/…shape, so the object mirror cannot persist any real platform object. Needs a fix inehdb-reference(hash the key into the subject instead of hex-encoding it, or raise the cap) before the object tier can mirror real drives. -
KV — NOT exercised by this drive shape. The KV mirror lives in
SpoolRuntime::persist_circuit(bucketnoetl_subscription_circuit), which runs only in the subscription runtime.noetl-worker-rust-subscription-poolis still onp1-refswith no EHDB env — outside the system+user roll — and a pft/api/executedrive doesn't touch the subscription circuit, so nothing mirrored. Predicted to mirror cleanly once exercised: the circuit key iscircuit.<subscription_id>(~40 B → subject ~115 chars < 256). Proving KV live requires rolling the subscription pool tov5.68.0+ shadow env (a bigger version jump with login-path risk; deferred).
Net: eventlog live-mirror proven again (v5.68.0 regression); object hook is
in the live path but blocked by the reference engine's subject-length cap on real
keys; KV still unproven on a real drive; projection/vector remain deferred.
Cluster left on v5.68.0 shadow. LOCAL kind only; no GKE/prod; repos/noetl +
repos/server untouched (config/env only). Refs
ehdb#234,
ehdb#241.
Extended the event-log live-append hook (v5.67.0, worker#167) to two more
tiers, replicating its env-armed / guarded / error-isolated shape. All
shadow-only, best-effort, disabled-by-default, reversible. LOCAL/KIND-only — no
GKE, no prod, no image build inline.
-
KV — wired live.
kv::runtime_hook_env+kv::mirror_live_put, armed once inSpoolRuntime::build, invoked inpersist_circuitafter the authoritative NATS-KV circuit-stateput(bucketnoetl_subscription_circuit— the only live platform NATS-KV write in the worker). -
object — wired live.
object::runtime_hook_env+object::mirror_live_put, armed once inControlPlaneClient::new, invoked inobject_putafter the server durably accepts the object.object_putis the single chokepoint every platform object tier (result-tier, state-shard, plugin intent) funnels through, so this mirrors ALL live object puts with digest parity (bytes cloned only when armed). -
projection — deferred (documented in
src/ehdb/mod.rs).shadow_projectis a batch materialize that reads back the whole accumulating projection log and compares against a full authoritative fold; a per-event hook would report persistent false key-divergence. Faithful seam = bounded windowed batch drive feeding the incumbent fold — larger than a call-site hook. -
vector — deferred. No live platform vector-upsert exists in the worker loop
today (platform RAG is in-process, read-only;
vector::mirror_upsertis exercised only byehdb-selfcheck). A hook now would never fire.
Each hook arms only when NOETL_EHDB_ENABLED + NOETL_EHDB_<TIER>=shadow + a
data-plane role (worker/playbook/system) on the bounded local_reference
runtime; strict no-op when off/disabled/primary or a control-plane role; engine
errors panic-caught → Unavailable, metered, never propagated; secret-free
noetl_ehdb_{kv,object}_* metrics advance on real ops.
Delivery: worker#168 merged
(squash 2ec2e2b) → v5.68.0 (release 3927bdf). 16 new unit tests (arms /
no-op-off / no-op-tier-off-primary / skip-control-plane / error-isolated per
tier); all 158 ehdb unit tests green; cargo check + clippy clean (no new
lints); ehdb-selfcheck still builds.
Live-drive proof — PENDING (redeploy). The live drive pool is on v5.67.0;
v5.68.0 carries these hooks. Follow-up: redeploy the worker image and drive a
real playbook, confirming noetl_ehdb_kv_* (on a circuit persist) and
noetl_ehdb_object_* (on a state-shard / result-tier put) shadow metrics
advance. Refs ehdb#234,
ehdb#241.
2026-07-06: EHDB program docs/hygiene consolidation (Roadmap + Home + MEMORY + issues to ground truth)
Documentation-only reconciliation across the whole EHDB program — no code, no
deploy, no prod/GKE, live kind stack (server v3.53.0 + worker v5.67.0
shadow) left running. Brought the tracking surfaces to one coherent current
state:
-
Roadmap — Phase 6 status now states the shadow mirror is wired into the
LIVE
ControlPlaneClient::emit_eventpath and proven on real in-kind drives (v5.67.0, worker#167d310c7b); added the runtime-mirror "Landed" bullet with the two proven drives (332760742153424896+332760854506246144) and the eventlog-only-so-far caveat. -
Home — Phase 6 note reflects the live/proven mirror; added a Live in
kind bullet (runtime mirror + query interface end-to-end via server
v3.53.0+ durable-backend slice 1ehdb#253) and an explicit What's NOT done yet bullet (prod cutovers gated, durable later slices, other-tier runtime mirrors, raw-tier query seam). - MEMORY.md (ai-meta agent memory) — compacted the EHDB index block into a program-status header + coherent bullets, flipped the resolved "v5.66 mirror not wired → FIXED" gotcha to past tense, all 26 topic-file pointers + PR/SHA references preserved.
-
Issues — posted the state-of-program summary on the completion umbrella
ehdb#241; cross-linkedehdb#234(dev umbrella),ai-meta#178(query interface), andehdb#254(durable backend). No issue closed — the program continues (prod cutovers + durable later slices remain).
Ground-truth cross-check confirmed against git/GitHub: ehdb HEAD 99c4570
(Phases 6-10 + durable slice 1 merged), ehdb-wiki eef435c, prod still worker
v5.52.0 / server v3.52.0 with all NOETL_EHDB_* flags off.
Built worker v5.67.0 (release ae4164d, PR #167 d310c7b) and ran it live
in the kind stack in SHADOW on both data-plane pools, then drove a real
playbook and confirmed the event-log mirror fires on real events — the runtime
evidence the code-only entry below was missing.
Environment: podman machine noetl-dev resized to 6 CPU / 20 GB (was
16 GB) and restarted; recovered a transient post-boot ssh/socket-gateway stall
(podman socket + noetl-control-plane came up on their own within ~2 min).
kind-noetl context, podman provider — no GKE, no prod touched.
Deploy (LOCAL kind only): localhost/noetl-worker:v5.67.0 (built native
arm64, CARGO_BUILD_JOBS=2, podman save + kind load image-archive) rolled
onto noetl-worker-rust (role=worker) + noetl-worker-system-pool
(role=system), keeping the existing EHDB shadow env
(NOETL_EHDB_ENABLED=1, all 5 tiers =shadow, LOCAL_REFERENCE_LOG=/tmp/ehdb-shadow-ref.jsonl).
Server/gateway stayed control-plane (auth-sync-167, v3.49.0). No tier set to
primary. Reversible: prior image v5.66.0-p10, NOETL_EHDB_ENABLED- ⇒ pure
incumbent.
Real drive: automation/pft_sql_probe_v2 (postgres SELECT 1 AS ok →
python end) via POST /api/execute → execution 332760742153424896,
status COMPLETED, 0 failed steps.
Mirror fired on the real drive (the whole point):
-
noetl_ehdb_eventlog_ops_total{operation="mirror",outcome="mirrored"}advanced across the drive — worker-rust 6 → 12 (the drive's exactly 6 events), system-pool 254 → 288.last_ok=1,last_degraded=0, onlyoutcome="mirrored"(no error/degraded outcomes ever). Contrast: v5.66.0 left these flat because the runtime hook wasn't wired. -
/tmp/ehdb-shadow-ref.jsonlon both pools received exactly 6 entries keyed to execution332760742153424896(subjectnoetl.event.exec.332760742153424896). Decoded worker-rust seq 8 = thecall.done/step=startevent carrying the real SQL result (rows:[{ok:1}],status:success) — real execution data mirrored into EHDB. - Incumbent unaffected: Postgres
noetl.eventpersisted all 13 events (playbook_started→playbook.completed); shadow never served; 0 restarts on either pool; readiness1.
Bonus (query interface) DEFERRED: /api/ehdb/executions 404s on the
running server (v3.49.0 predates the v3.53.0 /api/ehdb/* surface). Needs the
v3.53.0 server image loaded — out of scope for this worker-mirror validation.
Left in SHADOW on v5.67.0 so the user can watch real mirroring. Prod unchanged (worker v5.52.0 / server v3.52.0, all EHDB flags off).
Re-confirmed 2026-07-07 (independent second run): a later session found the
live pools had drifted back to v5.66.0-p10 (pods recreated on the last
set image value) — the mirror had gone dark again (/tmp/ehdb-shadow-ref.jsonl
absent, only the noetl_ehdb_readiness_* family on /metrics, no
noetl_ehdb_eventlog_*). Rebuilt v5.67.0 (native arm64; the podman VM
swap-died mid-build on wasmtime-cranelift and was recovered via
podman machine stop/start + podman start noetl-control-plane, then rebuilt
with the noetl server+worker pools scaled to 0 for headroom), re-rolled both
data-plane pools, and re-proved the mirror with a fresh drive: execution
332760854506246144 (pft_sql_probe_v2, COMPLETED) → 6 mirrored events
(seq 301–306, subject noetl.event.exec.332760854506246144) in the ref store,
counter mirror/mirrored climbing (273 → 344 → 378), last_ok=1,
last_degraded=0, 0 restarts, incumbent served all 13 events. Takeaway:
the pool is not sticky on v5.67.0 — a redeploy/recreate reverts it and the
mirror goes dark until re-rolled.
Closed the "runtime not wired" gap the live-in-kind session
below surfaced: the per-tier mirror engines only ran via ehdb-selfcheck +
tests, so a real drive did not dual-write into EHDB. Now the event-log
tier's shadow mirror fires on real events.
Change (worker v5.67.0, noetl/worker#167 → merged d310c7b):
- Hook site =
ControlPlaneClient::emit_event, the single authoritative event-append chokepoint every worker path funnels through (EventEmitter,emit_event_with_retry, spool_runtime, subscription, plugin). After the control plane accepts an event, the sameeventlog::mirror_eventshadow dual-write + parity path the selfcheck drives mirrors it into the derived EHDB fabric. - Armed once at construction by
eventlog::runtime_hook_env: onlyNOETL_EHDB_ENABLED+NOETL_EHDB_EVENTLOG=shadow+ data-plane role + local-reference log. Disabled / tier off|primary / control-plane role ⇒ strict per-event no-op, byte-identical/metrics. -
mirror_live_eventis panic-isolated + best-effort: a mirror failure surfaces as a metered non-ok outcome and never propagates into the authoritative event path.authoritative_sequence=None(count + order parity). Shadow never serves; event authorship untouched.
Validation: cargo test --lib 457 pass (7 new hook tests: arms only for
enabled+shadow+data-plane; no-op when disabled/off/primary; skips
control-plane; fires on shadow; guard-refused for control-plane role; engine
error isolated). cargo clippy --lib clean. LOCAL/code only — no GKE, no
image build inline.
PENDING — redeploy: the in-kind live-drive proof (build v5.67.0 image,
redeploy to the shadow pool, run a pft drive, confirm noetl_ehdb_eventlog_*
advances on a real execution) is a follow-up deploy step — the live pool is
on v5.66.0-p10 shadow; the next image carries this hook.
Follow-up slices: projection / kv / object / vector tier mirrors plug into
the same emit_event seam. See
Design: Event-Log Core Engine →
Runtime integration.
Refs noetl/ehdb#241, noetl/ai-meta#234.
First time EHDB runs inside the running kind noetl stack (not a one-shot
Job). Deployed worker v5.66.0 (Phase 10 config verb) as
localhost/noetl-worker:v5.66.0-p10 onto both data-plane pools with EHDB
in shadow across all five tiers, dual-writing + parity-checking against the
incumbents while real executions run, fully reversible. LOCAL kind only — no
GKE/gcloud, no prod (prod stays v5.52.0, all flags off).
Deploy (control plane untouched):
-
noetl-worker-system-pool—NOETL_EHDB_CLIENT_ROLE=system -
noetl-worker-rust(user/shared pool) —NOETL_EHDB_CLIENT_ROLE=worker - env:
NOETL_EHDB_ENABLED=1,NOETL_EHDB_MODE=local_reference,NOETL_EHDB_LOCAL_REFERENCE_LOG=/tmp/ehdb-shadow-ref.jsonl,NOETL_EHDB_{EVENTLOG,PROJECTION,KV,OBJECT,VECTOR}=shadow.
Evidence (all in-cluster):
- Boot readiness preflight
ok,outcome=empty, role guard correct;/metricsrendersnoetl_ehdb_readiness_*(disabled ⇒ zero ehdb lines). -
ehdb-selfcheck configin the live pod: coherent, 5×shadow,backend=external(shadow keeps the incumbent serving), secret-free, exit 0. - All 5 tier suites in the live pod (
eventlog/projection/kv/object/vector):ok=true— dual-write mirrored/materialized, parity holds. - Real drive
automation/pft_sql_probe_v2: postgresstartstep succeeds end-to-end; drive leaves shadow inert (only readiness on/metrics), events persist to Postgres normally, 0 restarts. - Reversibility:
NOETL_EHDB_ENABLED=0⇒ byte-identical/metrics; back to=1⇒ live shadow again. Left in shadow so it can be observed.
Realizes Stage A/B of the
prod-cutover runbook in kind. Note the
runtime EHDB surface is the readiness preflight + /metrics renderer; the
per-tier mirror/parity engines are exercised by the in-image ehdb-selfcheck
binary (kind validation / operator preflight), not the worker's request path —
so a real drive proves non-interference, and the shadow dual-write/parity is
demonstrated via the in-pod selfcheck suites. GKE/prod cutover remains gated on
the user, per tier.
Observe: kubectl --context kind-noetl exec <pool-pod> -c noetl-worker -- /app/ehdb-selfcheck config; port-forward the pod's :9090 and
curl /metrics | grep noetl_ehdb.
Landed the durable substrate the Phase-6 note deferred as "production disk
format" and the prod-cutover runbook §C durability gate
names as the hard blocker for Stage C (primary cutover) — the only backend
today, local_reference, is a pod-local JSONL file lost on restart +
divergent across replicas, so it is not production-durable as the
authoritative store under primary. LOCAL/code only — no GKE/gcloud, no
worker image build; validated via cargo test + clippy -D warnings + the
built ehdb-local-reference selfcheck. In-container kind validation is
PENDING — follow-up deploy. No repos/noetl or repos/server changes.
-
ehdb — #253:
ehdb-reference::durable_eventlog—DurableSegmentStore(append-only CRC32-framedseg-*.eslogsegments + size rollover, in-memory offset indexglobal_seq → (segment, offset)+ per-execution index + durable-consumer ack cursors,fsync-per-append, crash-recovery replay with torn-tail discard+truncate and complete-bad-CRC hard error; payloads cold-loaded via the index → bounded index memory, the noetl/ai-meta#166 property).DurableEventLogDriverimplements the sameEventLogDrivercontract asLocalReferenceEventLogDriver;EventLogStorageBackend(local_referencedefault |durable_segment) fail-safe selector.exercise_durable_recovery+ehdb-local-reference durable-eventlog-recoveryverb prove zero-loss + gapless ordering + per-execution scope + payload fidelity + durable-cursor survival across a simulated restart. -
Evidence —
durable-eventlog-recovery→{"recovered":true, "zero_loss":true,"ordering_ok":true,"scope_ok":true,"payloads_match":true, "cursor_survived":true,"pending_after_restart":2}, exit 0; a durableseg-*.eslogpersists on disk. 17 tests incl. rollover across files, torn-tail discard, bit-rot hard error, and parity vsLocalReferenceover identical ops.cargo fmt --all --check+clippy --workspace --all-targets -D warnings+cargo test --workspaceall green. - Design note — Durable Event-Log Backend (PVC-vs-object-store recommendation + flagged deploy assumptions).
- Tracking — #254 (program + 6-slice checklist); remaining slices: execution-affinity single-writer routing → shared object-store tier → worker wiring → kind soak → prod-durability sign-off.
2026-07-06: Phase 10 (Tunable-Backend Config Surface) — IMPLEMENTED → EHDB program (Phases 6–10) CODE-COMPLETE
The final phase of the EHDB completion program landed: the consolidated
per-tier backend-selection config surface. It folds the scattered
NOETL_EHDB_* per-tier flags into one coherent, documented schema — no
breaking rename, the existing env stays the source of truth. With Phase 10
the whole EHDB program (Phases 6–10) is code-complete and kind-validated;
the only remaining step is the prod/GKE cutover, gated on the user, per tier.
LOCAL/code only — no GKE/gcloud, no image build; the ehdb-selfcheck config
matrix was validated as a batched in-kind check. No repos/noetl or
repos/server changes (Codex-active).
-
ehdb — #252 (merged
4c0df81):ehdb-reference::backends—PlatformTier(the five tiers, each with its runtime env var + external incumbent),TierMode {off|shadow|primary},Backend {ehdb|external}+backend_for_mode(primary⇒ EHDB, else the incumbent), andBackendMatrixwith coherencevalidate()+ secret-freeto_json(). Pure data (reads no env, opens no engine). 10 unit tests;cargo fmt/test/clippy -D warningsclean. -
worker — #166 (released
v5.66.0):src/ehdb/backends.rsresolve(&EnvMap)maps the process env into the matrix through the same<Tier>Mode::from_envparsers the runtime dispatch uses (backward-compatible by construction). Newehdb-selfcheck configverb (aliasbackends) prints the resolved 5-tier backend+mode matrix (secret-free; exit 0 coherent, 4 incoherent). Pinsehdb-referenceto the Phase-10 SHA. 9 unit tests;cargo test --lib= 449 passed; clippy clean. -
Selfcheck evidence (
ehdb-selfcheck config): all-external default (no env) ⇒ 5×backend:external, exit 0; all-EHDB (enabled worker, 5×primary) ⇒ 5×backend:ehdb, exit 0; mixed (log+vectorprimary, projectionshadow) ⇒ per-tier ehdb/external, exit 0;primarywithout enable ⇒coherent:false- clear error, exit 4; data-plane tier on
gatewayrole ⇒coherent:false(2 errors), exit 4; sensitive-keyed env present ⇒secret_free:true, 0 value leaks.
- clear error, exit 4; data-plane tier on
- Boundaries — platform-only (business data never in EHDB); no behavior change (disabled-by-default strict no-op intact); EHDB default, not lock-in.
- Reference: Backend Configuration · Roadmap Phase 10.
2026-07-06: Phase 9 tiers 3–5 (KV / object / vector) — in-kind dual-run VALIDATED → Phase 9 CODE-COMPLETE
The combined tier 3–5 in-cluster dual-run — the last validation step before
Phase 9 is code-complete — passed. Tiers 1–2 were already kind-validated
(2026-07-05); this run closes tiers 3, 4, 5, so all five per-tier primary
cutovers are now IMPLEMENTED + MERGED and in-kind dual-run VALIDATED. Phase 9
is CODE-COMPLETE. LOCAL kind only — no GKE/gcloud, no prod cutover; the
prod/GKE cutover stays gated on the user, per tier. Nothing in prod changed
(prod worker v5.52.0; all NOETL_EHDB_* flags default off). No repos/noetl
or repos/server changes (Codex-active).
-
Build — worker
v5.65.0(ceedbba, the latest release, contains all five tiers from sequential merges) built native-arm64 fromrepos/worker/Dockerfile→localhost/noetl-worker:v5.65.0-p9(7d30b250d261, shipsehdb-selfcheck); cargo-chef cook re-ran because theehdb-referencepin bumpedd08013c5 → 0f47fe3f(DuckDB C++ amalgamation was the long pole). Loaded into thekind-noetlnode's containerd (ctr -n k8s.io images import). -
In-cluster run — Job
ehdb-p9-t345(nsehdb-p9-validate) on nodenoetl-control-plane; container kernel6.19.7-200.fc43.aarch64, Alpine 3.22 — the host kernel is the container's, NOT the macOS host (contextkind-noetl, APIhttps://127.0.0.1:61866). Pod Succeeded, 0 restarts. The umbrella switchNOETL_EHDB_ENABLED=1+ per-tierprimary+ worker-roleNOETL_EHDB_LOCAL_REFERENCE_LOGdrive the served path; server role staysguard_refused. -
Tier 3 (KV) —
served_primary+served_by_ehdb:true+reversible:true-
dual_run_holds:true(5 parities present/value/ttl vs NATS-KV) + put/get/scan/CAS/delete/TTL ok +keys_after_revert:3; server ⇒guard_refused.
-
-
Tier 4 (object) —
served_primary+dual_run_holds:true(4 parities digest/length/present/retrievable — content-addressed digest integrity vs the object store) + put/get/list/locate/delete ok +keys_after_revert:3; server ⇒guard_refused. -
Tier 5 (vector) —
served_primary+dual_run_holds:true(3 top-k parities ids/order/monotonic vs Qdrant) + upsert/query-topk/delete ok +query_returned:3(bounded top-k cap) +candidates_after_revert:3; server ⇒guard_refused. - All three tiers:
metrics_secret_free:true;off ⇒ disabledbyte-identical no-op (exit 0). No divergence, no failures.
The fifth and final per-tier primary cutover — serving NoETL's internal
platform vector tier from EHDB in place of the internal Qdrant retrieval path
(the platform RAG / catalog embeddings reached in-process via the Phase-E
retrieval path) — is built and merged, activated behind
NOETL_EHDB_VECTOR=primary, reversible, and dual-run-verified. Mirrors the
tier-1 event-log, tier-2 projection, tier-3 KV, and tier-4 object patterns
exactly. With this tier, all five Phase-9 primary-serve activations are
implemented. LOCAL only — no prod/GKE cutover; that stays gated on the user,
per tier. No repos/noetl or repos/server changes (Codex-active).
-
ehdb — #251 (merged
0f47fe3):ehdb-reference::vectorgainsexercise_primary_serve+VectorPrimaryInput+VectorPrimaryServeReport::served_by_ehdb()— one authoritative cycle through upsert → served cosine top-k query → tombstone delete → fresh-driver replay, preserving the Qdrant retrieval semantics and dual-run parity-checking each served query (id-set + rank-order + score-monotonicity) against a Qdrant mirror ranked in lockstep viacompare_vector_parity. CLI verbvector-primary-serve. 155 crate tests green; clippy-D warnings+ fmt clean. -
worker — #165 (merged
681782c, releasedv5.65.0ceedbba):src/ehdb/vector.rsflipsPRIMARY_SERVE_ACTIVATEDfalse → true;mirror_upsertserves authoritatively underprimary(served_primary/primary_divergence); newserve_primary_cycleruns the cycle + demonstrates reversibility (flip back toshadow, mirror one more point, confirm the whole live set serves).ehdb-selfcheck vector-primary-serveverb proves served-by-EHDB + reversibility- secret-free metrics. Pins
ehdb-referenceto0f47fe3.
- secret-free metrics. Pins
Local proof (ehdb-selfcheck, debug host): off ⇒ no-op + empty metrics
(exit 0); primary ⇒ served_by_ehdb:true + reversible:true +
candidates_after_revert:3 + dual_run_holds:true (3 parities all hold) +
secret-free noetl_ehdb_vector_* (exit 0); control-plane server ⇒
guard_refused (exit 4); shadow ⇒ primary_unavailable (exit 4); over-limit
dims ⇒ rejected (exit 3).
Kind dual-run: BATCHED — combined tier 3-5 run to follow. No worker image
built this session. Reversible with two levers (runtime
NOETL_EHDB_VECTOR=primary→shadow/off restores Qdrant instantly, zero data loss;
compile-time PRIMARY_SERVE_ACTIVATED=false is the structural kill switch).
Remaining Phase-9 work = the combined tier 3-5 in-kind dual-run (tiers 1–2 already
kind-validated), then Phase 9 is code-complete — prod/GKE cutover gated on user.
The fourth per-tier primary cutover — serving NoETL's internal platform
object/blob tier from EHDB in place of the internal external object store
(the state shards #166 + result tier #104 reached through the server's
/api/internal/objects/{key} API) — is built and merged, activated behind
NOETL_EHDB_OBJECT=primary, reversible, and dual-run-verified. Mirrors the
tier-1 event-log, tier-2 projection, and tier-3 KV patterns exactly. LOCAL
only — no prod/GKE cutover; that stays gated on the user. No repos/noetl or
repos/server changes (Codex-active).
-
ehdb — #250 (merged
bb42b8d):ehdb-reference::objectgainsexercise_primary_serve+ObjectPrimaryInput+ObjectPrimaryServeReport::served_by_ehdb()— one authoritative cycle through put → per-key digest-verified served get → prefix list → in-cluster locate → tombstone delete → fresh-driver replay, preserving the external-store semantics and dual-run digest-parity-checking each served read against an external-store mirror viacompare_object_parity. CLI verbobject-primary-serve. -
worker — #164 (merged
a100adf, releasedv5.64.0369e4c1):src/ehdb/object.rsflipsPRIMARY_SERVE_ACTIVATEDfalse → true;primaryserves the object op authoritatively (served_primary/primary_divergence);serve_primary_cycledrives the cycle + demonstrates reversibility (flip back toshadow, mirror one more object over the same store, confirm the whole live set serves).ehdb-selfcheck object-primary-serveverb.
Local ehdb-selfcheck proof: off ⇒ byte-identical no-op (exit 0); primary ⇒
served_primary + served_by_ehdb:true + reversible:true +
keys_after_revert:3 + secret-free noetl_ehdb_object_* metrics (exit 0);
control-plane server ⇒ guard_refused (exit 4).
- Kind dual-run: BATCHED — combined tier 3-5 in-cluster run to follow (one worker image build validates the three sequential releases). No worker image built for tier 4. prod/GKE cutover stays GATED on the user.
Reversibility: runtime NOETL_EHDB_OBJECT=primary → shadow/off restores the
external object store instantly (zero data loss); compile-time
PRIMARY_SERVE_ACTIVATED = false is the structural kill switch. Platform object
tier only (state shards + result tier); never authors the event log.
Remaining after this: vector (1 tier) + the combined tier 3-5 in-cluster dual-run.
The third per-tier primary cutover — serving NoETL's internal platform
KV/state tier from EHDB in place of the internal NATS-KV bucket — is built
and merged, activated behind NOETL_EHDB_KV=primary, reversible, and
dual-run-verified. Mirrors the tier-1 event-log and tier-2 projection patterns
exactly. LOCAL only — no prod/GKE cutover; that stays gated on the user. No
repos/noetl or repos/server changes (Codex-active).
-
ehdb — #249 (merged
73b1446):ehdb-reference::kvgainsexercise_primary_serve+KvPrimaryInput+KvPrimaryServeReport::served_by_ehdb()— one authoritative cycle through put → per-key served get → bucket scan → optimistic CAS (versioned swap + create-only conflict) → tombstone delete → absolute-TTL lease → fresh-driver replay, preserving the NATS-KV semantics and dual-run parity-checking each served read against a NATS-KV mirror viacompare_kv_parity. CLI verbkv-primary-serve. -
worker — #163 (merged
ba9f829, releasedv5.63.0a7925e0):src/ehdb/kv.rsflipsPRIMARY_SERVE_ACTIVATEDfalse → true;primaryserves the KV op authoritatively (served_primary/primary_divergence);serve_primary_cycledrives the cycle + demonstrates reversibility (flip back toshadow, mirror one more key over the same store, confirm the whole live set serves).ehdb-selfcheck kv-primary-serveverb. -
Local proof (
ehdb-selfcheck kv-primary-serve, host debug build):off ⇒ehdb:disabled+metrics_empty:true(exit 0);primary ⇒outcome:served_primary+served_by_ehdb:true+reversible:true+dual_run_holds:true(5 served-read parities,present_ok/value_ok/ttl_ok,divergence:null) +put_ok/get_ok/scan_ok/cas_ok/delete_ok/ttl_ok/replay_matchesall true +keys_after_revert:3(NATS-KV path restored whole on flip-back, zero data loss) + secret-freenoetl_ehdb_kv_*metrics (exit 0);serverrole⇒guard_refused(exit 4). - Kind dual-run: BATCHED — combined tier 3-5 in-cluster run to follow (one worker image build validates the three sequential releases). No worker image built this session.
-
Rollback (two levers, zero data loss): runtime
NOETL_EHDB_KV=primary → shadow/offrestores NATS-KV instantly; compile-timePRIMARY_SERVE_ACTIVATED = falseinsrc/ehdb/kv.rsis the structural kill switch. Documented in Roadmap → Phase 9 → Tier-3 rollback procedure.
Remaining after this: object + vector (2 tiers) + the combined tier 3-5 in-cluster kind dual-run.
The deferred in-cluster proof for both Phase 9 primary cutovers — event-log (tier 1) and projection / read-model (tier 2) — ran and passed on the local kind cluster. This closes the "kind dual-run PENDING/BATCHED" that the prior two sessions left open when the podman VM's ssh socket was wedged. LOCAL/kind only — no prod/GKE cutover; that stays gated on the user. No new ehdb/worker code — pure validation of the already-merged v5.62.0.
-
Environment recovery — the podman VM (
noetl-dev) ssh socket was wedged (handshake reset though the machine reportedrunning); apodman machine stop/startcycle reset it. Kind clusternoetlrestarted (its control-plane container had stopped). Proof it is kind, not GKE:kubectlcontextkind-noetl, API serverhttps://127.0.0.1:61866, nodenoetl-control-plane(containerd, arm64 Fedora podman VM). -
Image — worker
v5.62.0(36875e3; ships both tiers + theehdb-selfcheckbinary) built native-arm64 from therepos/workerDockerfile, saved to an archive, and loaded into the kind node's containerd (localhost/noetl-worker:v5.62.0-p9,sha256:9d222db6, linux/arm64). -
In-cluster run — a
batch/v1Job (ehdb-p9-dualrun, namespaceehdb-p9-validate,imagePullPolicy: Never) ran the shippedehdb-selfcheckverbs in a pod on nodenoetl-control-plane. Pod host kernel6.19.7-...fc43.aarch64confirms the verbs executed inside the kind container, not on the host. PodSucceeded,0restarts, no crashloop. -
Tier 1 (event-log) —
off ⇒ehdb:disabled+metrics_empty:true(exit 0);primary ⇒served_primary+served_by_ehdb:true+reversible:true+dual_run_holds:true(3×count_ok/order_ok/sequence_ok,divergence:null) +replay_matches:true+scope_ok:true+records_after_revert:4(incumbent JetStream+Postgres path restored whole on flip-back toshadow) + secret-freenoetl_ehdb_eventlog_*metrics (exit 0);serverrole⇒guard_refused(exit 4). -
Tier 2 (projection) —
off ⇒ehdb:disabled+metrics_empty:true(exit 0);primary ⇒served_primary+served_by_ehdb:true+reversible:true+dual_run_holds:true(key_ok/value_ok/checkpoint_ok,checkpoint_lag:0) +list_ok:true+read_event_ok:true+replay_idempotent:true+replay_matches:true+scope_ok:true+rows_after_revert:3(Postgres-materializer read path restored whole on flip-back) + secret-freenoetl_ehdb_projection_*metrics (exit 0);serverrole⇒guard_refused(exit 4). - Boundaries held — platform-only; event authorship unchanged (the cutover changes only the serving engine, never who appends); loose coupling; control-plane guard-refused; event log stays the source of truth (replay-is-truth proven). repos/noetl + repos/server untouched (Codex lane).
- Roadmap tier-1 + tier-2 rows flipped
PENDING/BATCHED → VALIDATED; rollback procedures unchanged (both levers still documented).
2026-07-05: Phase 9 tier 2 (projection / read-model) — primary-serve implemented + merged (kind dual-run batched)
The second per-tier primary cutover of the completion program — the
projection / read-model tier — is built, activated, reversible, and merged,
mirroring the tier-1 event-log pattern exactly. Under
NOETL_EHDB_PROJECTION=primary the EHDB projection engine now serves the
materialized read-models the control plane queries (list_executions,
per-execution read_execution_state, read_event) authoritatively in place of
the PostgreSQL materializer, dual-run parity-checked against the incumbent.
No prod/GKE cutover was performed — that is a separate later step gated on the
user. LOCAL/kind scope only.
-
ehdb — #248 (merged
d08013c):projection::exercise_primary_serve+ProjectionPrimaryInput+ProjectionPrimaryServeReport::served_by_ehdb()(served-by-EHDB proof over the full authoritative cycle: apply → the three read-model query contracts → checkpoint → idempotent re-apply → fresh-engine replay, dual-run parity-checked) +projection-primary-serveCLI verb. Reversible + additive toward the incumbent (KeepAllprojection store, replay-is-truth proven). -
worker — #162 (merged
a56583c, releasedv5.62.036875e3):PRIMARY_SERVE_ACTIVATEDfalse → true(reversible via the runtime flag);serve_primary_cycle+ProjectionServeResult+ehdb-selfcheck projection-primary-serveverb that proves served-by-EHDB and reversibility (flip back toshadow, one more execution materialized, read-models replay whole). Pins ehdbd08013c. -
Local proof (same binary the image ships): off ⇒ byte-identical no-op;
primary ⇒
served_by_ehdb:true+reversible:true+ secret-free metrics; control-plane ⇒guard_refused(exit 4). -
Kind dual-run: VALIDATED 2026-07-05 — see the
2026-07-05 in-cluster validation session
entry at the top.
ehdb-selfcheck projection-primary-serveran in a Job pod on kind nodenoetl-control-plane(imagev5.62.0sha256:9d222db6):served_by_ehdb:true+reversible:true+dual_run_holds:true+ secret-free metrics (exit 0); control-plane⇒ guard_refused(exit 4). - Boundaries held: platform-only; a projection is a derived read-model built by consuming the append-only event log — it never authors an event; loose coupling; control-plane guard-refused; event log stays the source of truth. repos/noetl + repos/server untouched (Codex lane).
- Rollback procedure documented on Roadmap.
The first per-tier primary cutover of the completion program — the event-log
tier, the bottleneck fix that motivated the whole program — is built,
activated, reversible, and merged. Under NOETL_EHDB_EVENTLOG=primary the
EHDB event-log engine now serves the platform event log authoritatively (append
- read + tail + ack + replay), dual-run parity-checked against the incumbent. No prod/GKE cutover was performed — that is a separate later step gated on the user. LOCAL/kind scope only.
-
ehdb — #247 (merged
7f014c9):exercise_primary_serve+EventLogPrimaryEvent+EventLogPrimaryServeReport(served-by-EHDB proof over the full authoritative cycle) +eventlog-primary-serveCLI verb. Reversible + additive toward the incumbent (KeepAlllog, replay-is-truth proven). -
worker — #161 (merged
ddf41de, releasedv5.61.07e98538):PRIMARY_SERVE_ACTIVATEDfalse → true(reversible via the runtime flag);serve_primary_cycle+EventLogServeResult+ehdb-selfcheck eventlog-primary-serveverb that proves served-by-EHDB and reversibility (flip back toshadow, log replays whole). Pins ehdb7f014c9. -
Local proof (same binary the image ships): off ⇒ byte-identical no-op;
primary ⇒
served_by_ehdb:true+reversible:true+ secret-free metrics; control-plane ⇒guard_refused. -
Kind dual-run: VALIDATED 2026-07-05 — see the
2026-07-05 in-cluster validation session
entry at the top.
ehdb-selfcheck eventlog-primary-serveran in a Job pod on kind nodenoetl-control-plane(imagev5.62.0sha256:9d222db6):served_by_ehdb:true+reversible:true+dual_run_holds:true(3× append parity) + secret-free metrics (exit 0); control-plane⇒ guard_refused(exit 4). - Boundaries held: platform-only; event authorship unchanged (gateway/server stay the gatekeeper — primary changes only the serving engine underneath); loose coupling; control-plane guard-refused; never authors a NoETL event.
- Rollback procedure documented on Roadmap.
Third and final Phase-8 engine slice (Roadmap Phase 8,
#241): EHDB's vector engine —
the durable vector engine underneath NoETL's internal platform vector tier
(RAG / catalog embeddings), formalizing the already-in-process Phase-E retrieval
path behind a first-class VectorDriver. Design note:
KV/State + Object/Blob + Vector Engines (Phase 8).
With this slice, all three Phase-8 engines (KV + object + vector) are
shadow-complete. No tier is cut over to primary — per-tier primary cutover is the
separately-gated Phase 9 step, so Phase 8 stays in progress overall.
-
ehdb — PR #246 (merged
f6c5a6f). Newehdb-reference::vector: theVectorDrivertrait (upsert/query(top-k) /delete) andLocalReferenceVectorDriverover one canonicalnoetl_vector_indexstream (per-point subjectnoetl.vec.<hex(collection)>.<hex(point_id)>,KeepAllretention → replay-is-truth). Query is a bounded cosine top-k filtered to the query's model + dimensionality, ranked descending with a deterministic point-id tie-break — the same scoring shape asehdb_retrieval::InMemoryRetrievalCatalog::search_similar(the Phase-E path). Bounded + secret-free (MAX_VECTOR_DIMENSIONS4096,MAX_VECTOR_QUERY_TOP_K64,MAX_VECTOR_PAYLOAD_BYTES16 KiB; over-cap → rejected, empty/non-finite/zero → invalid).compare_vector_parity= pure id-set + rank-order + score-monotonicity parity (tolerance for cross-engine float slack). 17 vector engine tests; clippy -D warnings + fmt. -
worker — PR #160 (merged
c7b8872, v5.60.00f57625).src/ehdb/vector.rsshadow behindNOETL_EHDB_VECTOR=off|shadow|primary(defaultoff):shadowdual-writes one platform vector into the EHDB engine alongside the authoritative Qdrant path, then reads back a self-retrieval top-k query and compares id-set/rank-order/monotonicity parity without serving reads or touching the authoritative Qdrant retrieval path;primaryrecognised but inert (compile-timePRIMARY_SERVE_ACTIVATED = false). Secret-freenoetl_ehdb_vector_*metrics; control-plane guard; structural no-NoETL-event-writer test.ehdb-selfcheckgainsmirror-vector/vector-suite. Pinsehdb-referenceto the merged #246 revf6c5a6f. 14 unit tests; clippy -D warnings + fmt. - Boundary preserved: the vector engine indexes derived platform embeddings — it never authors an event; platform-only (business vector collections stay external on their own Qdrant); bounded; secret-free; loose coupling (driver-selectable back to Qdrant, Phase 10).
-
Validation: engine side — 17 vector engine tests + clippy/fmt clean. Worker side
— 14
ehdb::vector+ 2ehdb::metricstests pass, clippy-D warnings+ fmt clean, and the builtehdb-selfcheckbinary proves:off→ byte-identical no-op + empty metrics (exit 0);shadow vector-suite→ 4/4 engine steps (upsert/query-parity/top-k-truncate/delete) + secret-freenoetl_ehdb_vector_*metrics (exit 0);shadow mirror-vector→mirrored, top-k parity holds (exit 0); control-plane (server) role →guard_refused(exit 4); primary →primary_unavailable, inert (exit 4); over-limit dimensionality →rejected(exit 3); zero vector →invalid(exit 4). In-container kind run deferred — env build eviction (the worker-rust image is a ~110-min cold build; same standing-evidence pattern as Phase 6/7/E). No GKE. repos/noetl + repos/server untouched (Codex-active). -
Phase 8 → Phase 9 handoff: Phase 8 is now engine-complete across all five
platform tiers (event-log, projection, KV, object, vector — the latter three from
Phase 8, the first two from Phases 6–7), each in disabled-by-default shadow. Phase 9
is the five independent per-tier primary cutovers (
shadow→primary, EHDB serves reads, incumbent retired for that tier), each gated + dual-run verified + rollback-documented + kind-before-GKE. See the Phase-9 per-tier table in Roadmap.
2026-07-05: Phase 7 started — projection / read-model engine + driver interface + disabled-by-default shadow
Second implementation slice of the completion program (Roadmap Phase 7, #241): EHDB's projection / read-model engine — the engine that builds + serves the materialized read-models the control plane queries, off the Phase-6 event-log tail, retiring the PostgreSQL materializer. Design note: Projection / Read-Model Engine (Phase 7). Design + engine slice + disabled-by-default shadow only — reads are NOT cut over off Postgres (a later gated step, Phase 9).
-
ehdb — PR #243. New
ehdb-reference::projection: theProjectionDrivertrait (apply/read_execution_state/read_event/list_executions/checkpoint) andLocalReferenceProjectionEngine, composing the append-only stream primitives over onenoetl_projection_logstore (per-execution scope vianoetl.projection.exec.<id>subject). Materializes the event read-model (keyed onevent_id, theON CONFLICT DO NOTHINGtwin), the folded execution-state read-model (theprojection_snapshottwin), and a durable consumer checkpoint (theevent_stream.positiontwin). Apply is idempotent / exactly-once keyed on the Phase-6 global sequence (skip<= checkpoint+event_iddedup); rebuild-from-log is deterministic (batch-boundary independent).ProjectionEventInput::from_event_log_recordbridges the Phase-6 tail;compare_projection_parityis the pure key/value/checkpoint parity vs the authoritative materializer.ehdb-local-referencegainsprojection-apply/projection-read-exec/projection-read-event/projection-list/projection-checkpoint/projection-from-eventlog(eventlog→projection bridge) /projection-suite(distinct exit codes 0/3/4/5). 19 unit tests (73 in the crate); clippy -D warnings + fmt. -
worker — PR #157
(merged
eadc3a5).src/ehdb/projection.rsshadow behindNOETL_EHDB_PROJECTION=off|shadow|primary(defaultoff):shadowdual-materializes the read-models from the event-log tail alongside the Postgres materializer + compares parity (compare_projection_parity) without serving reads or touching the authoritative materializer;primaryrecognised but not activated (compile-timePRIMARY_SERVE_ACTIVATED = false). Secret-freenoetl_ehdb_projection_*metrics; control-plane guard; bounded apply batch. Pinsehdb-referenceto the merged #243 reve0f1c0f. 13 unit tests; clippy -D warnings + fmt. -
Boundary preserved: the projection engine materializes read-models by
consuming the log — it never authors an event (the #103 sole-writer stays
the only
noetl.eventwriter), platform-only, bounded, secret-free. -
Validation: engine side — the built
ehdb-local-referencedrivesprojection-suite(exit 0, parity holds), idempotent re-apply (all skipped-below-checkpoint), invalid id → exit 4, and the eventlog→projection bridge folds correctly. Worker side — 63ehdb::unit tests pass (13 inprojection.rs), clippy-D warnings+ fmt clean, and the builtehdb-selfcheckbinary proves:projection-suitedisabled → byte-identical no-op + empty metrics (exit 0); shadow (worker role) →apply_1materialized 3 /apply_2_replayapplied 0, parity holds, secret-freenoetl_ehdb_projection_*metrics render (exit 0); control-plane (server) role →guard_refused(exit 4); primary →primary_unavailable, inert (exit 4). In-container kind run deferred — env build eviction (the worker-rust image is a ~110-min cold build; same standing-evidence pattern as Phase 6/E). No GKE.
First implementation slice of the completion program (Roadmap Phase 6, #241): EHDB's event-log core engine — the durable persistence + ordering + serving layer that replaces the JetStream + Postgres log-and-store path underneath the append-only producer path. Design note: Event-Log Core Engine (Phase 6). Design + disabled-by-default shadow only — the log is NOT cut over.
-
ehdb — PR #242 MERGED
squash
9a9b28d. Newehdb-reference::eventlog: theEventLogDrivertrait (append/scan_global/read_execution/tail/ack) andLocalReferenceEventLogDriver, composing the append-only stream primitives over one canonicalnoetl_event_logstream so its sequence is the global, monotonic, gapless event-log sequence. Per-execution scope vianoetl.event.exec.<execution_id>subject; durable-consumer tail/ack;compare_shadow_parity.ehdb-local-referencegainseventlog-append/eventlog-scan/eventlog-read-exec/eventlog-tail/eventlog-ack/eventlog-suite(distinct exit codes 0/3/4/5). 12 unit tests; clippy -D warnings + fmt; CI green. -
worker — PR noetl/worker#156
MERGED squash
43c8f0f. Newsrc/ehdb/eventlog.rsshadow behindNOETL_EHDB_EVENTLOG=off|shadow|primary(defaultoff):shadowdual-writes each already-authored event into the engine + compares sequence/count/order parity without serving reads or touching the authoritative path;primaryrecognised but not activated (compile-timePRIMARY_SERVE_ACTIVATED = false). Parity uses the engine's gapless invariant (global_sequence == log_record_count) — concurrency-safe, no shared bookkeeping. Secret-freenoetl_ehdb_eventlog_*metrics; control-plane guard;ehdb-selfcheck mirror-eventlog/eventlog-suite. Pinsehdb-referencec2aaad5→9a9b28d. 10 unit tests; clippy + fmt. -
Validation — local
ehdb-selfcheckdrive against the built binary: off ⇒metrics_empty:trueexit 0 (byte-identical no-op); shadow ⇒ 3 mirrors parity-holds, secret-free metrics, exit 0; control-plane (server) ⇒guard_refusedexit 4, no write; primary ⇒primary_unavailableexit 4 (never serves); parity mismatch ⇒ detected exit 5. In-containerkind-noetlrun deferred — env build eviction (~110-min cold worker-rust image; same standing evidence pattern as the Phase E first/second slices). No GKE. repos/noetl + repos/server untouched (Codex-active). -
Event authorship unchanged — gateway/server still gatekeep what is
appended; the shadow only mirrors already-authored events and never
writes
noetl.event. What changed is the engine underneath the producer path. -
Remaining Phase 6 — production segmented disk format + offset index
- compaction; sharded/multi-stream ordering; a JetStream+Postgres
EventLogDriver(Phase 10 tunable surface); the projection engine (Phase 7) attaching to the tail/ack/read-execution surface; primary-serve cutover (dual-run verify + rollback, kind before GKE).
- compaction; sharded/multi-stream ordering; a JetStream+Postgres
2026-07-05: RFC DECIDED — EHDB completion program (loose coupling + noetl self-sufficiency; EHDB is the event-log engine)
Docs/RFC session. Recorded the decided EHDB completion program on the wiki page RFC: EHDB Completion Program — Server↔EHDB Coupling + noetl Self-Sufficiency (status decided, applying — not a review gate), grounded against the current Architecture + Roadmap pages (Phase 5 integration A–E complete) and #234. Supersedes the earlier "pluggable substrate, up for review" framing.
-
Headline motivation — fix the event-sourcing log bottleneck. EHDB
IS a distributed event-log database: it writes, orders, persists, and
serves the
noetl.eventlog and builds/serves the projections, replacing the JetStream + Postgres-materializer path that is today's scaling pressure point (off-server state builder; unbounded WAL index ai-meta#166; materializer soak ai-meta#104). - What EHDB is — a family of core engines under one catalog/URN namespace: event-log, projection/read-model, KV/state, object/blob, vector — potentially separate engines per workload type.
- Critical boundary — EHDB serves NoETL platform functionality only (event log, projections, platform KV/state, catalog, system-WASM store, platform artifacts/vector). Business data is never in EHDB — tenant data stays in real business systems (Postgres, Snowflake, Cassandra, Kafka, Elasticsearch, ClickHouse, object stores) via playbook connectors. "EHDB replaces Postgres/Qdrant/NATS/object store" = NoETL's internal platform uses only.
- Decision 1 — coupling LOOSE (unchanged). Server/API/gateway control-plane-only + stateless; EHDB data-plane in-process in worker/system; seam = event log + catalog/URN, logical mapping not 1:1 pinning. Server MAY expose a thin control surface + route; MUST NOT embed a data plane. Narrow control-plane-read-cache exception.
- Decision 2 — EHDB as noetl's self-sufficient internal substrate. End-state: NoETL runs on k8s with no external infra dependency for platform functionality; EHDB default across platform tiers, every tier tunable back to the incumbent (JetStream+Postgres / NATS KV / object store / Qdrant). Default, not lock-in.
- Reconciliation (not deleted, scoped). Event authorship rules unchanged: gateway/server gatekeep what is appended through the append-only producer path; non-log data-plane roles (worker bounded step, RAG, system store) never fabricate application events. The storage/ordering/projection engine underneath that path is now EHDB (replacing JetStream + the Postgres materializer). Authorship unchanged; engine is EHDB.
- Roadmap — restructured post-integration phases into the completion trajectory Phases 6–10: 6 event-log core engine (bottleneck fix) → 7 projection/read-model engine → 8 KV/object/vector engines → 9 external-dependency retirement (k8s-only) → 10 tunable-backend config surface. Old Phase 6 (Postgres replacement) + Phase 7 (dependency collapse, #6) absorbed. Architecture + Home + Sidebar link the RFC.
-
Applying, not reviewing. Tracking issue
#241 retitled to the
completion program (decided, implementing) with the phase checklist;
pointer comment on #234.
No code repos touched (
repos/noetl/repos/serveruntouched — Codex-active). No GKE.
2026-07-05: EHDB Phase E second slice — bounded RAG retrieval wired into worker-rust (Phase E COMPLETE)
The second (and final) Phase E slice — bounded RAG retrieval — is DONE end to end, unblocking the RAG direction the first slice had deferred:
-
ehdb side: #240 MERGED
(squash
c2aaad5) → ehdbmain. Adds the bounded, read-onlyretrieve_local_reference_contexthelper toehdb-reference(runsehdb-retrieval'ssearch_textunder the hood), aningest_local_reference_retrieval_documentcompanion, andehdb-local-referenceingest-doc/retrieveCLI verbs (retrieveexit 0=hit/empty, 3=rejected, 4=invalid). Three caps enforced in the helper: top-k (ceiling 64), per-hit size (ceiling 64 KiB, char-boundary truncation), wall-clock budget (ceiling 60 s,time_capped). 9 unit/integration tests;clippy -D warnings+fmtclean. -
worker side: noetl/worker#155
MERGED (squash
d1ebaf2). New in-processsrc/ehdb/rag.rsbridges the helpers (retrieve+ingest, matching the Phase C/D/E pattern), anoetl_ehdb_rag_*metric family, andehdb-selfcheckingest-rag/retrieve-rag/rag-suite(A→E+RAG driver). Pinsehdb-reference9bb5928→c2aaad5. 8 unit tests. -
Boundaries (unit + integration tested): disabled-by-default no-op
(byte-identical
/metrics), control-plane guard (gateway/api/server refused), bounded (NOETL_EHDB_RAG_TOP_K8/64,NOETL_EHDB_RAG_MAX_CHUNK_BYTES4 KiB/64 KiB,NOETL_EHDB_RAG_TIME_BUDGET_MS5 s/60 s; over-limitRejected), stateless, event-log-authoritative (retrieve read-only; ingest writes only the private JSONL fabric, nevernoetl.event). -
Validated via the built
ehdb-selfcheckA→E+RAG drive: disabled no-op (empty metrics), enabledrag-suite(ingest 3 chunks → retrieve hit with top-k truncation → empty → over-limit rejected,ok:true,metrics_secret_free), over-limitRejected(exit 3), control-planeserverrefused (exit 4,retrieve:null), secret-freenoetl_ehdb_rag_*metrics. In-containerkind-noetlvalidation deferred (imagelocalhost/local/noetl:ehdb-rag, distinctehdb-rag-selfcheckpod): the worker-rust image is a ~110-min cold build (theehdb-referencerev bump busts the cargo-chef dependency cache) that timed out in this environment — the builtehdb-selfcheckdrive above is the standing evidence. No GKE. -
ai-meta pointers (gitlink-only, each commit only its own gitlink):
repos/ehdb→c2aaad5,repos/worker→d1ebaf2,repos/ehdb-wiki→ this commit. - Phase E is COMPLETE (both slices — system WASM store + RAG retrieval). The broad EHDB-integration umbrella #234 stays OPEN with Phase E marked done; remaining roadmap work is Phase 6 (PostgreSQL replacement path) and Phase 7 (dependency collapse).
Phase E (system-WASM store → RAG) starts, Rust-first. The first slice — the system WASM library store — is DONE end to end:
-
ehdb side already merged: #239
→ ehdb
main9bb5928— boundedpublish/bind/resolve*_system_modulehelpers inehdb-reference(immutable module manifests + environment/channel bindings). -
worker side: noetl/worker#154
MERGED (squash
3162b1d). New in-processsrc/ehdb/systemstore.rsbridges the helpers (matches the Phase C/D pattern), anoetl_ehdb_systemstore_*metric family, andehdb-selfcheckpublish-system/bind-system/resolve-system + asystem-suiteA→E driver. Bumpsehdb-reference3cefba9→9bb5928. -
Boundaries (unit + integration tested): disabled-by-default no-op
(byte-identical
/metrics), control-plane guard (gateway/api/server refused), bounded (NOETL_EHDB_SYSTEM_MAX_MODULE_BYTES16 MiB/256 MiB,NOETL_EHDB_SYSTEM_MAX_CAPABILITIES16/64; over-boundRejected), stateless, event-log-authoritative (private JSONL, nevernoetl.event). WASM execution stays host-side/sandboxed — EHDB catalogs the module ref only. -
Kind-validated against the worker-rust image
(
localhost/local/noetl:ehdb-phase-e,kind-noetl, distinctehdb-phase-e-selfcheckpod on nodenoetl-control-plane, providerkind://podman/...): disabled no-op (empty metrics) A–D + Phase E, enabled A–D drive, enabled Phase-E system-suite (absent→publish→bind→resolve rev1→publish→rebind→resolve rev2,ok:true), control-planeserver/apirefused (exit 4,resolve:null), oversized publishRejected(exit 3), unsupported targetInvalid(exit 4), 0noetl.eventwrites (only PublishLibrary/BindLibrary in the private log), secret-freenoetl_ehdb_systemstore_*metrics, Rust stack 0 restarts. No GKE. -
ai-meta pointers (gitlink-only, each commit only its own gitlink):
repos/worker→3162b1d(subsumes #153d6226a2, which never reached ai-metamain),repos/ehdb→9bb5928,repos/ehdb-wiki→ this commit. -
Next slice — bounded RAG retrieval — DEFERRED: no bounded retrieval
helper in the merged
ehdb-referenceslice; needs a follow-up ehdb slice. Umbrella #234 stays OPEN.
The Python-retire review gate cleared (approved), so the re-home is done:
-
noetl/noetl#691 MERGED (merge commit
ff3a920f) — the Python EHDB path is retired; noetlmaincarries zeronoetl/core/ehdb_*modules (guard test in place). Pairs with noetl/worker#153 (d6226a2, merged 2026-07-04), which owns the integration inworker-rustin process. -
Re-validated in kind against the worker-rust image
(
localhost/noetl-worker:ehdb-test,kind-noetl, distinctehdb-*Job name): disabled no-op (metrics empty), enabled worker full drive, cross-process durable cursor (fresh process seesacked_sequenceadvanced), control-plane guard refused (no write), secret-free metrics, event-log-authoritative. Rust stack 0 restarts. No GKE. -
ai-meta pointers:
repos/noetl→ff3a920f,repos/ehdb-wiki→ this commit (repos/worker→d6226a2bumped 2026-07-04). Each pointer-bump commit carries only its own gitlink. - noetl/ehdb#238 CLOSED; umbrella #234 stays open for Phase E (Rust-first).
Implements the Rust-first decision below (owner directive) — closes the
integration debt: the EHDB worker/playbook integration now lives in
worker-rust, in process, and the Python EHDB path is retired.
-
worker-rust in-process integration —
noetl/worker#153 MERGED
(squash
d6226a2). Newsrc/ehdbmodule:contract+guard(control-plane boundary),readiness(non-fatal bootstrap preflight oversummarize_local_reference),dataplane(bounded append/read),eventstream(bounded project/consume/ack durable-consumer drain),metrics(secret-freenoetl_ehdb_*appended to the worker/metrics). Depends on theehdb-referencecrate as a git dep pinned to3cefba9— no subprocess. Ships anehdb-selfcheckbin for in-image validation. Disabled-by-default strict no-op; control-plane guard; bounded + stateless; secret-free metrics; event-log-authoritative (no NoETL event-writer import, structurally asserted). -
Python path retired —
noetl/noetl#691 deletes the Python
EHDB modules (
noetl.core.ehdb_*), the worker bootstrap wiring, the step CLIs / smoke scripts, and stops bundling theehdb-local-referencebinary in the Python images; adds a guard test. Open, awaiting required review (protectedmain) — not merged. -
Kind validation against the worker-rust image (
kind-noetl,localhost/noetl-worker:ehdb-test): disabled no-op (byte-identical, metrics empty), enabled worker full drive (readiness ready → append → read → project → consume → ack), cross-process durable cursor (fresh process sees the advanced ack cursor), control-plane guard refused (exit 4, no log written), secret-free metrics, event-log-authoritative (only the EHDB JSONL written). No GKE. Running Rust stack undisturbed (server-rust + system-pool 0 restarts). -
Pointers:
ai-metarepos/worker→d6226a2,repos/ehdb-wiki→ this commit.repos/noetldeferred until noetl#691 merges. - Tracks noetl/ehdb#238 (closes on the Python-retire merge) under umbrella #234.
Documentation / decision session (no product code).
-
Phase D pointer finalized. noetl/noetl#690
merged (merge commit
04d3e273); theai-metarepos/noetlpointer was bumped to it, completing the Phase D pointer set (ehdb3cefba9+ ehdb-wiki064b29awere already bumped). #234 stays OPEN for Phase E. -
Architecture Decision (owner directive): Rust-first EHDB worker
integration; Python as thin wrapper. The EHDB storage engine is already
Rust (the
ehdbcrate:append/read/project/consume/ack). The Phase B–D worker/playbook integration glue (readiness hook, data-plane step, event-stream drain) was added in the legacy Python worker runtime (noetl/core), shelling out to the Rust binary as a subprocess — but prod runsworker-rust, so those disabled-by-default Python hooks do not execute in prod (they only exercise against the Python worker, as a stopgap). Target: the integration lives inworker-rustinvoking theehdbcrate in-process (no subprocess shell-out), with Python reduced to a thin binding over the Rust core (the "Polars model"). Captured on the Architecture page (Architecture Decision: Rust-first EHDB worker integration) and the Roadmap (Integration debt item + Phase E marked Rust-first). -
Tracking issue: noetl/ehdb#238
— re-home Phase B–D worker integration into
worker-rust; tracks #234.
The event-stream integration path: a bounded, disabled-by-default worker/playbook/system drain that mirrors already-emitted NoETL events into a derived EHDB stream and consumes them through a durable consumer with explicit ack-after-materialize semantics.
- EHDB adds
consume/acksubcommands +consume_local_reference_event_records/ack_local_reference_event_consumerto theehdb-local-referencehelper (composing with the Phase Cappendproject leg), and a read-onlyInMemoryStreamLog::consumercursor getter.consumecreates the durable consumer on first pull and returns pending records after its cursor without moving it;ackadvances the cursor after materialize and rejects backwards / unknown / zero sequences atomically. - NoETL adds
noetl.core.ehdb_eventstream(project/consume/ack), adapterLocalReferenceConsumeResult/LocalReferenceAckResult, worker/metrics(noetl_ehdb_eventstream_*), a worker/playbook-local step CLI (scripts/ehdb_eventstream_step.py), and a kind smoke (scripts/smoke_ehdb_eventstream.py). The Dockerfiles bumpEHDB_REFto the Phase D helper. -
Event-log-authoritative invariant: the NoETL event log stays the
append-only source of truth; EHDB is a derived, auxiliary consumer
that never writes back to it (no event-writer import; a unit test
asserts it structurally). Still disabled-by-default (strict no-op,
byte-identical
/metrics), worker/playbook/system-only with the code-level control-plane guard (assert_event_stream_access_allowed), bounded (payload + consume-limit + ack sequence ≥ 1 + time caps), stateless, secret-free metrics.
Landed:
- EHDB noetl/ehdb#237 (merged,
3cefba9; closes #236). - NoETL noetl/noetl#690
(open, green, kind-validated; awaiting required review on
noetl/noetlmain). Tracks #234.
Validation:
-
cargo test -p ehdb-reference→ 30 passed; workspace green; clippy-D warningsclean;cargo fmt --all --checkclean. -
pytest tests/core/test_ehdb_eventstream.py tests/scripts/test_ehdb_eventstream_step.py→ 27 passed (incl. a real-binary project→consume→ack cursor-restart drain); 131 EHDB tests green, no regression. - Local kind (
kind-noetl, Phase D image with the real linux binary): a Job ran all checks green — helper carries consume/ack verbs, disabled no-op + byte-identical/metrics, enabled worker drain + durable cursor restart, control-plane gateway guard refused (exit 4, no write), secret-free metrics, Phase C append/read no-regression. Nothing on GKE/prod.
Next: Phase E — system WASM store integration, then RAG retrieval.
First bounded data-plane integration (product code in noetl/noetl,
built over an EHDB append/read helper):
-
noetl.core.ehdb_dataplane(append_ehdb_domain_record/read_ehdb_domain_records) appends and reads a single domain record through the local-reference adapter — the first data-plane op, past the readiness summary of Phase B. Adapter gains typedLocalReferenceAppendResult/LocalReferenceReadResult; payload reaches the helper verbatim (byte-exact record). - Still disabled-by-default (strict no-op, byte-identical
/metrics), worker/playbook/system-only. Code-level control-plane guard (assert_data_plane_access_allowed) refuses gateway/api/server before any helper runs →guard_refused, no write. - Bounded (payload byte cap + read-limit cap + short time cap;
over-bound →
rejected) and stateless (helper opened+dropped per call). Secret-freenoetl_ehdb_dataplane_ops_total{operation,outcome}metrics on the worker/metricssurface. - Worker/playbook-local step CLI (
scripts/ehdb_dataplane_step.py) / kind smoke (scripts/smoke_ehdb_dataplane.py) — not a server endpoint.
Landed:
-
noetl/noetl#689 (merged,
merge commit
56a66da8, featurebda27200). - Over EHDB append/read subcommands
noetl/ehdb#235 (merged
3ae8950). Tracks #234.
Validation:
-
pytest tests/core/test_ehdb_dataplane.py tests/scripts/test_ehdb_dataplane_step.py→ 23 passed (incl. a real-binary append/read roundtrip); 84 existing EHDB tests still green. - Local kind (
kind-noetl, Phase C overlay image, real linux binary): a Job ran six checks green — disabled no-op + byte-identical/metrics, worker/system/playbook append→read roundtrip, control-plane gateway guard refused (exit 4, no write), secret-free metrics. Nothing on GKE/prod.
See the Architecture page's Worker/Playbook Data-Plane Step (Integration Phase C) section and the Roadmap integration phases.
NoETL-side integration milestone (product code in noetl/noetl, not the
EHDB repo):
- Bounded, stateless readiness preflight
(
noetl.core.ehdb_readiness.evaluate_ehdb_readiness) over the Phase-Aread_ehdb_local_reference_summary_from_envsummary. Wired for worker/playbook/system data-plane roles only; code-level guard (assert_data_plane_read_allowed) refuses a data-plane read for any control-plane role. Time-bounded, holds no state. - Observability
noetl_ehdb_readiness_*on the worker/metricssurface (no secret values). Disabled-by-default → strict no-op (byte-identical, no metric recorded). - Worker/playbook-local preflight command
(
scripts/ehdb_readiness_preflight.py) / kind smoke step — not a server endpoint.
Landed:
-
noetl/noetl#688 (merged,
merge commit
5960e91d, featurede23ca0b). - Tracks #234.
Validation:
-
pytest tests/core/test_ehdb_readiness.py(+ full EHDB core suite, 81 passed). - Local kind (
kind-noetl) Job on an overlay image (Phase A helper image + this code, realehdb-local-referencehelper): disabled no-op (rc=0), worker local_reference bounded read (rc=0), gateway control-plane guard (rc=4) — all passed. No GKE/prod changes.
See the Architecture page's Worker/Playbook Readiness Hook (Integration Phase B) section and the Roadmap integration phases.
Implementation branch:
kadyapam/ehdb-local-reference-summary-helper
Design target:
- Add a concrete local helper binary for NoETL-side integration tests and worker/playbook diagnostics.
- Open the local JSONL transaction log through
LocalReferenceRuntimeand report deterministic JSON counts from replayed state. - Include transaction, catalog, stream, retrieval, system-library, and storage counts so the helper reflects the current NoETL-domain storage surfaces.
- Keep this bounded to local reference inspection: no daemon, network API, gateway route, SQL planner, distributed executor, production IAM, storage mutation behavior, background worker, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmtcargo test -p ehdb-stream -p ehdb-retrieval -p ehdb-reference
Implementation branch:
kadyapam/ehdb-embedded-role-policy
Design target:
- Make EHDB's embedded distributed database direction explicit without weakening the NoETL execution model.
- Add typed NoETL embedded roles for gateway, API, worker, playbook, and system contexts.
- Add typed EHDB capabilities that distinguish control-plane embedding from catalog, transaction, stream, object, retrieval, replication, and system-library data-plane access.
- Keep gateway/API roles control-plane only by default; workers, playbooks, and system jobs carry explicit data-plane capabilities for bounded work.
- Keep this as policy/modeling only: no daemon, network API, gateway route, SQL planner, distributed executor, production IAM, storage mutation behavior, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmtcargo test -p ehdb-corecargo test --workspacecargo bench --workspace --no-run
Coverage snapshot:
- 258 Rust tests across unit, integration, and doc-test targets.
Context:
-
noetl/ehdbwas created as the future Event Horizon Database. -
ai-metaadded the code repository and wiki as submodules. - EHDB is positioned as the core database storage layer for the NoETL multitenant distributed operating-system cloud platform.
Design decisions:
- EHDB stores both operational metadata/catalog data and historical analytical data.
- The catalog lives inside EHDB as first-class transactional metadata.
- Object storage holds immutable data files: Arrow IPC, Parquet, Iceberg-compatible tables, and blobs.
- Metadata is native EHDB state, not an external PostgreSQL dependency once EHDB reaches self-hosting.
- Early implementation starts with Rust workspace boundaries and local reference adapters before distributed consensus.
Initial mission statement:
EHDB is an Arrow-native distributed database and catalog platform that stores metadata transactionally and data in multi-cloud object storage, providing a unified foundation for operational, analytical, and historical workloads.
Next actions:
- Create and track bootstrap/design issues in
noetl/ehdb. - Scaffold Rust crates and local tests.
- Keep the wiki current as the code shape changes.
Issues opened:
- #1 Bootstrap EHDB Rust workspace and CI
- #2 Design catalog-as-database metadata model
- #3 Define immutable object storage layer
- #4 Define transaction log and MVCC snapshot boundary
- #5 Plan NoETL integration path for EHDB system store
Direction update:
- EHDB should be a NoETL-domain-specific storage system, not a generic database that happens to serve NoETL.
- EHDB should grow to absorb the NoETL platform roles currently served by PostgreSQL, NATS JetStream, external object stores, Qdrant, and ClickHouse.
- RAG support is first-class: documents, chunks, embedding metadata, vector index metadata, retrieval policies, tenant context, and execution lineage belong in the EHDB model.
- NATS JetStream functionality should be represented as EHDB-native stream logs, durable consumers, replay cursors, ack state, and retention policy. A NATS bridge may exist for migration, but the target state is EHDB-owned stream durability.
Architectural implication:
NoETL should see EHDB as the storage system. Cloud object APIs, vector-index files, analytical columnar layouts, and stream persistence may exist internally, but they should not remain separate permanent NoETL platform dependencies.
Tracking issue:
Implementation branch:
kadyapam/ehdb-stream-retrieval-foundation
Implemented first local reference models:
-
ehdb-stream: typed NoETL stream records, subjects, retention policy, durable consumers, replay cursors, and ack state. -
ehdb-retrieval: documents, chunks, embedding metadata, model identity, dimension validation, and tenant/namespace-scoped text lookup fixtures. -
ehdb-transaction: replayable transaction records for catalog, stream, and retrieval mutations, with duplicate transaction ID rejection and ordered replay. -
ehdb-core: shared stream, consumer, document, chunk, and embedding model identifiers. - Integration coverage ties catalog create-table, stream publish, and retrieval registration through one transaction log.
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 25 Rust tests across unit, integration, and doc-test targets.
- Edge coverage includes duplicate catalog/stream/consumer/document/ chunk/embedding records, missing lookup behavior, unsafe object paths, invalid subjects, stream retention, cursor replay, ack monotonicity, embedding dimension validation, tenant/namespace retrieval isolation, duplicate transaction IDs, empty transactions, and replay after cursor.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~466 us |
transaction_append_replay_1000 |
1000 transaction appends + full replay | ~843 us |
This is intentionally pre-distributed and pre-service. The point is to make the NoETL-native replacement surfaces concrete before adding network protocols, consensus, or persistence adapters.
Implementation branch:
kadyapam/ehdb-local-persistence-foundation
Implemented Phase 3 local durability:
- Added serde support for EHDB durable identifiers and transaction records.
- Added
LocalJsonlTransactionLog, a fsynced append-only JSONL adapter for local developer and crash/restart tests. - Reused in-memory duplicate transaction ID and contiguous sequence checks while rebuilding state from disk.
- Added restart, duplicate-after-reopen, and corrupt-record tests.
- Added a Criterion benchmark for the durable local append/reopen path.
Tracking issue:
Validation:
cargo fmt --allcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 28 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~489 us |
transaction_append_replay_1000 |
1000 transaction appends + full replay | ~1.04 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced JSONL appends + reopen + full replay | ~456 ms |
This remains a local reference implementation. Production replicated metadata durability still belongs behind a consensus-backed transaction log boundary.
Implementation branch:
kadyapam/ehdb-local-stream-journal
Implemented Phase 3.5 local stream durability:
- Added serde support for EHDB stream configs, subjects, sequences, records, retention policies, and durable consumers.
- Added
LocalJsonlStreamLog, a fsynced append-only JSONL stream journal for local developer and restart tests. - Journaled create-stream, create-consumer, publish, and ack operations.
- Rebuilt retained records, durable consumer ack cursors, and next sequence on open.
- Added restart/cursor, retention-after-reopen, and corrupt-entry tests.
- Added a Criterion benchmark for durable local stream publish/reopen.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 31 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~629 us |
transaction_append_replay_1000 |
1000 transaction appends + full replay | ~1.04 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced JSONL appends + reopen + full replay | ~454 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~457 ms |
This remains a local reference implementation. Production stream replication and any NATS compatibility bridge should plug in behind the stream boundary.
Implementation branch:
kadyapam/ehdb-system-wasm-library-registry
Implemented Phase 3.75 system-library cataloging:
- Added
ehdb-systemfor NoETL system WASM library manifests and environment/channel bindings. - Modeled immutable module manifests with logical path, revision, digest, entry export, target, object path, byte length, capability grants, and transaction provenance.
- Modeled mutable bindings by tenant, namespace, environment, release channel, and logical path.
- Added deterministic resolution of the active module for a requested environment/channel/path.
- Added hot replacement by rebinding a stable channel to a new digest/revision while retaining prior immutable manifests.
- Added transaction-log
SystemMutationrecords for publish/bind events. - Extended the cross-domain NoETL surface test to include system library publish/bind replay alongside catalog, stream, and retrieval mutations.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 36 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~626 us |
transaction_append_replay_1000 |
1000 transaction appends + full replay | ~1.04 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced JSONL appends + reopen + full replay | ~448 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~456 ms |
This is the catalog/storage side of the NoETL WASM plug-in model. EHDB does not embed Wasmtime in this slice; worker/system-pool execution and host capability enforcement remain separate runtime concerns.
Implementation branch:
kadyapam/ehdb-system-library-journal
Implemented restartable system-library registry state:
- Added
LocalJsonlSystemLibraryCatalog, a fsynced append-only JSONL journal for system WASM library publish and bind operations. - Rebuilt immutable module manifests and mutable environment/channel bindings on open.
- Preserved hot-replacement behavior across restart by replaying stable channel rebindings to new digest/revision values.
- Added restart, hot-replacement-after-reopen, and corrupt-record tests.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 39 Rust tests across unit, integration, and doc-test targets.
This remains a local reference implementation. Production replicated system-library registry durability belongs behind the same EHDB transaction/replication boundary as catalog metadata.
Implementation branch:
kadyapam/ehdb-replay-complete-mutations
Implemented reconstructive transaction replay:
- Enabled serde support for Arrow schema datatypes and EHDB table schemas.
- Expanded transaction mutations so replay carries the facts needed to
rebuild reference state:
- catalog create-table includes schema;
- stream publish includes subject, payload, and expected sequence;
- retrieval mutations include document source metadata, chunk text and checksum, and embedding vectors;
- system library publish includes the full WASM module manifest.
- Added
ehdb-reference, a replay applier over the local catalog, stream, retrieval, and system-library reference models. - Added coverage proving transaction replay rebuilds all current reference domains from the log alone.
- Added mismatch detection for durable stream sequence values that do not match reconstructed state.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 41 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~626 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.21 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~488 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~486 ms |
This moves the reference model closer to the NoETL execution boundary: the transaction log is the durable source of truth, while local catalogs are reconstructable projections.
Implementation branch:
kadyapam/ehdb-local-reference-runtime
Implemented the first local runtime boundary over replay-complete transactions:
- Added
LocalReferenceRuntime, which opens a local JSONL transaction log and rebuilds reference state by replaying transaction records. - Added projection validation before durable append: the runtime previews the transaction record, applies it to cloned reference state, and only appends when the projection succeeds.
- Kept invalid projected commits from advancing the durable JSONL log.
- Added clone support to the in-memory catalog, stream, and retrieval reference stores so projected state can be validated without mutating the live state.
- Hardened local storage temp paths used by tests to avoid collisions in fast repeated runs.
- Added a Criterion benchmark for local runtime append/reopen.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 43 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~626 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.18 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~486 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~482 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~477 ms |
This makes the local developer runtime stricter: the transaction log is only advanced after the current reference projections agree the mutation set is valid, while restart still treats replay as the source of truth.
Implementation branch:
kadyapam/ehdb-content-addressed-objects
Implemented the first content-checked object reference boundary:
- Added
ObjectDigestas a SHA-256 digest wrapper for stored object bytes. - Extended
ObjectRefto carry path, byte length, and digest. - Added
ImmutableObjectStore::get_verified, which reads through an object reference and rejects length or digest mismatches. - Added a deterministic table/snapshot object path helper:
{tenant}/{namespace}/tables/{table}/snapshots/{snapshot}/{file}. - Added corruption detection and path-layout coverage for the local object-store adapter.
- Added a Criterion benchmark for local object put plus verified read.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 46 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~14.8 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~640 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.17 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~476 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~454 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~452 ms |
This closes an important Phase 2 gap: local object references are now content-checked and deterministic enough to be attached to future catalog snapshots and manifests.
Implementation branch:
kadyapam/ehdb-catalog-snapshots
Implemented immutable table snapshot metadata in the catalog reference:
- Added
CatalogSnapshotandCommitSnapshotwith snapshot ID, parent snapshot, content-checked object file refs, and committing transaction ID. - Added latest snapshot tracking per table.
- Rejected missing tables, empty file sets, duplicate snapshots, and parent-chain mismatches.
- Added
CatalogMutation::CommitSnapshotso snapshot metadata is replayable through the transaction log. - Extended
ehdb-referenceandLocalReferenceRuntimecoverage so replay rebuilds catalog snapshot state. - Added a Criterion benchmark for catalog snapshot commits.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 49 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~2.06 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~16.9 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~656 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.37 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~491 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~508 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~534 ms |
This connects the Phase 1 catalog model and Phase 2 object layer: content-checked object references can now become durable table snapshot metadata, and the latest snapshot projection is reconstructable from the transaction log.
Implementation branch:
kadyapam/ehdb-placement-pointers
Added explicit distributed-storage placement pointers to object references before implementing production replication:
- Added
CloudProvider,GeoLocation,DataGravityShard, andObjectPlacementtoehdb-storage. - Extended
ObjectRefso catalog snapshots carry path, byte length, digest, geo placement, and data-gravity shard metadata together. - Kept the local filesystem adapter on deterministic
local-devplacement. - Added validation coverage for geo region/zone and data-gravity shard identifiers.
- Documented that these are storage-layer routing pointers for future read/write nodes, replicators, and placement planners, not a gateway data-touch path.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 50 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~1.94 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~14.6 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~613 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.14 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~486 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~571 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~517 ms |
Implementation branch:
kadyapam/ehdb-placement-policy
Added a declarative placement policy model over geo-location and data-gravity shard metadata:
- Added
PlacementRole,PlacementTarget, andPlacementPolicytoehdb-storage. - Enforced exactly one primary placement per policy.
- Enforced a minimum copy count before a policy is accepted.
- Enforced one shared data-gravity shard across all targets.
- Rejected duplicate geo/shard targets.
- Added deterministic
local-devplacement policy support. - Added validation tests and a Criterion benchmark for policy validation.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 53 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
placement_policy_validate_1000 |
1000 three-target placement policy validations | ~1.19 ms |
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~2.12 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~14.7 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~614 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.16 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~443 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~542 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~534 ms |
This turns placement pointers into an explicit contract for future replication planners while keeping execution aligned with NoETL's gateway/worker/playbook boundary.
Implementation branch:
kadyapam/ehdb-replication-plan
Added a deterministic replication planning model over placement policy:
- Added
ObjectReplica,ReplicationAction, andReplicationPlan. - Added
plan_replication, which compares a source object and known replicas with aPlacementPolicy. - Emits
AlreadySatisfiedactions for existing policy targets. - Emits
CopyNeededactions for missing policy targets. - Rejects source/policy data-gravity shard mismatches.
- Rejects known replicas with mismatched digest, length, or shard.
- Added unit coverage and a Criterion benchmark for the planner.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 56 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
replication_plan_1000 |
1000 three-target replication plans | ~2.92 ms |
placement_policy_validate_1000 |
1000 three-target placement policy validations | ~1.18 ms |
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~1.98 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~15.7 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~626 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.16 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~516 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~554 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~540 ms |
This gives EHDB a planner-level contract for future bounded replication workers while keeping object copy execution out of the gateway and out of this reference model.
Implementation branch:
kadyapam/ehdb-replica-registry
Design target:
- Add an EHDB-owned object replica registry so replication plans are derived from durable metadata instead of caller-supplied arrays.
- Record object path, byte length, digest, placement, and data-gravity shard for each available copy.
- Reject conflicting digest, length, or shard metadata for the same object.
- Make replica registration replayable through the transaction log and local reference runtime.
- Keep copy execution out of the registry. Replication remains a future bounded worker/playbook operation, not gateway behavior.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 61 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
replication_plan_from_registry_1000 |
1000 three-target replication plans from registry state | ~3.82 ms |
replication_plan_1000 |
1000 three-target replication plans | ~2.95 ms |
replica_registry_register_1000 |
1000 object replica registrations | ~1.10 ms |
placement_policy_validate_1000 |
1000 three-target placement policy validations | ~1.17 ms |
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~1.98 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~16.0 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~635 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.16 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~509 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~507 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~516 ms |
This makes replica inventory replayable EHDB metadata while keeping object-copy execution out of gateway request handling.
Implementation branch:
kadyapam/ehdb-local-replication-executor
Design target:
- Consume deterministic replication plans from a bounded local executor.
- Verify source object bytes through the immutable object-store API before recording replica success.
- Append
StorageMutation::RegisterReplicamutations for copy-needed targets throughLocalReferenceRuntime. - Treat already-satisfied plans as no-op work.
- Keep scheduling, long-lived processes, cloud transfer adapters, and gateway data-touch behavior out of scope.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 64 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
local_replication_executor/register_25 |
25 verified source reads + fsynced replica-registration transactions + reopen | ~139 ms |
replication_plan_from_registry_1000 |
1000 three-target replication plans from registry state | ~3.75 ms |
replication_plan_1000 |
1000 three-target replication plans | ~2.81 ms |
replica_registry_register_1000 |
1000 object replica registrations | ~1.08 ms |
placement_policy_validate_1000 |
1000 three-target placement policy validations | ~1.15 ms |
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~2.00 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~20.2 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~630 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.17 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~441 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~541 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~523 ms |
This is the first local execution reference for replication work. It models an atomic worker/playbook step and keeps scheduling, cloud copy adapters, and gateway data-touch behavior outside the slice.
Implementation branch:
kadyapam/ehdb-arrow-ipc-fixture
Design target:
- Write Arrow
RecordBatchvalues into immutable object storage as IPC files. - Commit catalog snapshots over content-checked
ObjectReffile sets. - Read the latest snapshot back through verified object reads and Arrow IPC decoding.
- Preserve tenant, namespace, table, snapshot, and transaction provenance in the reference runtime.
- Keep Arrow Flight, distributed query execution, Parquet adapters, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 66 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
local_arrow_ipc_table/write_read_10 |
10 Arrow IPC write + catalog snapshot + verified read cycles | ~123 ms |
local_replication_executor/register_25 |
25 verified source reads + fsynced replica-registration transactions + reopen | ~155 ms |
replication_plan_from_registry_1000 |
1000 three-target replication plans from registry state | ~3.76 ms |
replication_plan_1000 |
1000 three-target replication plans | ~2.82 ms |
replica_registry_register_1000 |
1000 object replica registrations | ~1.08 ms |
placement_policy_validate_1000 |
1000 three-target placement policy validations | ~1.13 ms |
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~1.96 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~15.0 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~609 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.17 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~530 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~491 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~502 ms |
This proves the local catalog/object boundary for Arrow-native analytical data while leaving Arrow Flight and distributed query execution as explicit future surfaces.
Implementation branch:
kadyapam/ehdb-arrow-scan-fixture
Design target:
- Scan the latest catalog snapshot through verified Arrow IPC object reads.
- Return decoded Arrow
RecordBatchoutput. - Support optional named column projection in caller-specified order.
- Fail deterministically for missing projection columns.
- Keep predicate pushdown, SQL planning, Arrow Flight, distributed query execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 68 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
local_arrow_scan/project_latest_100 |
100 verified latest-snapshot scans with two-column projection | ~13.1 ms |
local_arrow_ipc_table/write_read_10 |
10 Arrow IPC write + catalog snapshot + verified read cycles | ~107 ms |
local_replication_executor/register_25 |
25 verified source reads + fsynced replica-registration transactions + reopen | ~133 ms |
replication_plan_from_registry_1000 |
1000 three-target replication plans from registry state | ~3.86 ms |
replication_plan_1000 |
1000 three-target replication plans | ~2.97 ms |
replica_registry_register_1000 |
1000 object replica registrations | ~1.12 ms |
placement_policy_validate_1000 |
1000 three-target placement policy validations | ~1.21 ms |
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~2.06 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~16.4 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~793 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.22 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~562 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~443 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~517 ms |
This moves the local analytical path from write/read proof to a minimal scan operation while leaving predicates, SQL planning, Flight, and distributed execution out of scope.
Implementation branch:
kadyapam/ehdb-arrow-filter-fixture
Design target:
- Extend local Arrow snapshot scanning with optional single-column equality predicates.
- Support UTF-8 and Int64 equality first.
- Apply filtering after verified object reads and Arrow IPC decode, and before optional projection.
- Fail deterministically for missing predicate columns and unsupported predicate/column type combinations.
- Keep SQL planning, predicate pushdown, object statistics, Arrow Flight, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-reference --bench local_runtimecargo bench -p ehdb-storage --bench local_storecargo bench -p ehdb-catalog --bench snapshotscargo bench -p ehdb-transaction --bench reference_models
Coverage snapshot:
- 72 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
local_arrow_scan/filter_project_latest_100 |
100 verified latest-snapshot scans with equality filter and two-column projection | ~13.0 ms |
local_arrow_ipc_table/write_read_10 |
10 Arrow IPC write + catalog snapshot + verified read cycles | ~111 ms |
local_replication_executor/register_25 |
25 verified source reads + fsynced replica-registration transactions + reopen | ~159 ms |
replication_plan_from_registry_1000 |
1000 three-target replication plans from registry state | ~3.77 ms |
replication_plan_1000 |
1000 three-target replication plans | ~2.89 ms |
replica_registry_register_1000 |
1000 object replica registrations | ~1.09 ms |
placement_policy_validate_1000 |
1000 three-target placement policy validations | ~1.36 ms |
catalog_commit_snapshots_1000 |
1000 catalog snapshot commits + latest lookup | ~2.03 ms |
local_object_store/put_get_verified_100 |
100 immutable 4 KiB local object puts + verified reads | ~15.6 ms |
stream_publish_replay_1000 |
1000 stream publishes + full replay | ~640 us |
transaction_append_replay_1000 |
1000 replay-complete transaction appends + full replay | ~1.15 ms |
local_reference_runtime/append_reopen_100 |
create stream + 100 projection-validated fsynced transaction appends + reopen + replay | ~543 ms |
local_transaction_jsonl/append_reopen_100 |
100 fsynced replay-complete JSONL appends + reopen + full replay | ~461 ms |
local_stream_jsonl/publish_reopen_100 |
100 fsynced stream publishes + reopen + full replay | ~464 ms |
This adds a first local predicate path while keeping SQL planning, predicate pushdown, Flight, distributed execution, and gateway reads out of scope.
Implementation branch:
kadyapam/ehdb-arrow-scan-service-boundary
Design target:
- Add the first Phase 4 service-facing scan API boundary.
- Create
ehdb-servicewith typed latest-table scan request/result structures. - Wrap
LocalArrowSnapshotScannerinLocalArrowScanService. - Return Arrow schema, batches, and row count from scan results.
- Preserve projection and equality-filter behavior from the local scanner.
- Keep Arrow Flight networking, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --allcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-runcargo bench -p ehdb-service --bench local_scan_service
Coverage snapshot:
- 76 Rust tests across unit, integration, and doc-test targets.
Benchmark baseline:
| Benchmark | Workload | Baseline |
|---|---|---|
local_arrow_scan_service/filter_project_latest_100 |
100 service-boundary latest-snapshot scans with equality filter and two-column projection | ~12.0 ms |
This establishes the service API shape that a future Arrow Flight server can expose while keeping data-touch work behind EHDB-owned APIs and out of the NoETL gateway.
Implementation branch:
kadyapam/ehdb-arrow-flight-ticket-codec
Design target:
- Add a versioned Arrow Flight ticket payload for latest-table scans.
- Preserve tenant, namespace, table, projection, and equality predicate fields in the encoded request.
- Round-trip through Arrow Flight
Ticketbytes. - Build command
FlightDescriptorvalues for the future Flight read API. - Reject unsupported ticket versions and malformed payloads before scan execution.
- Keep the Arrow Flight network server, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --allcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 82 Rust tests across unit, integration, and doc-test targets.
This gives Phase 4 a concrete Flight request contract without starting a long-lived service or changing the NoETL gateway boundary.
Implementation branch:
kadyapam/ehdb-arrow-flight-result-codec
Design target:
- Add a pre-network Arrow Flight result stream codec for local scan outputs.
- Encode
ArrowScanResultbatches into Arrow FlightFlightDatamessages. - Decode
FlightDatamessages back into validatedArrowScanResultvalues with schema and row count. - Preserve projected schemas and filtered rows.
- Reject empty or malformed streams deterministically.
- Keep the Arrow Flight network server/client, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --allcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 86 Rust tests across unit, integration, and doc-test targets.
This proves the local response side of the future Flight do_get path
while preserving the NoETL gateway and bounded-worker boundaries.
Implementation branch:
kadyapam/ehdb-arrow-flight-info-fixture
Design target:
- Add a pre-network Arrow Flight
FlightInfofixture for latest-table scans. - Build schema IPC bytes from the scan result schema.
- Attach the command descriptor generated from
ScanFlightTicket. - Return one ordered endpoint with the encoded scan ticket.
- Report total records and encoded FlightData byte count.
- Keep the Arrow Flight network server/client, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --allcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 88 Rust tests across unit, integration, and doc-test targets.
This proves the future get_flight_info metadata shape without starting
a network service or changing NoETL gateway data-touch boundaries.
Implementation branch:
kadyapam/ehdb-local-flight-service-facade
Design target:
- Add an in-process local facade for future Arrow Flight scan behavior.
- Provide
get_flight_infofrom typed latest-table scan requests. - Provide
do_getfrom Arrow Flight tickets to FlightData result streams. - Reuse
ScanFlightTicket,ArrowScanResult::to_flight_info, andArrowScanResult::to_flight_data. - Reject malformed tickets before scan execution and propagate missing table errors deterministically.
- Keep the Arrow Flight network server/client, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --allcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 91 Rust tests across unit, integration, and doc-test targets.
This moves Phase 4 from static protocol fixtures to local service behavior while still avoiding a long-lived network process.
Implementation branch:
kadyapam/ehdb-flight-service-trait-adapter
Design target:
- Implement the generated Arrow Flight
FlightServicetrait for the local scan path. - Accept command
FlightDescriptorrequests forget_flight_infoby decoding the existingScanFlightTicketpayload. - Accept Arrow Flight
Ticketrequests fordo_getand streamFlightDatathrough the tonic response type. - Map EHDB errors to deterministic gRPC statuses.
- Return explicit
UNIMPLEMENTEDstatuses for non-scan Flight methods. - Keep port binding, long-lived server runtime lifecycle, TLS/auth, request concurrency policy, access-log policy, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 95 Rust tests across unit, integration, and doc-test targets.
This establishes the first network-facing trait boundary while still avoiding a bound listener or persistent service process.
Implementation branch:
kadyapam/ehdb-flight-server-lifecycle-config
Design target:
- Add a validated lifecycle/config surface for the local Arrow Flight service adapter.
- Track intended bind address, max decode/encode message sizes, max concurrent requests, auth policy, and access-log policy.
- Default to loopback local-reference use with bounded message sizes, bounded concurrency, disabled local auth, and DEBUG-only access logs.
- Reject zero bounds and unauthenticated non-loopback binds.
- Build the generated
FlightServiceServerwith message limits applied. - Keep actual socket binding, long-lived runtime management, TLS/auth implementation, request scheduling, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 99 Rust tests across unit, integration, and doc-test targets.
This adds server lifecycle guardrails before EHDB exposes any bound Flight listener.
Implementation branch:
kadyapam/ehdb-loopback-flight-listener
Design target:
- Add a loopback-only listener harness behind
LocalArrowFlightServerConfig. - Bind a configured or ephemeral loopback socket and expose the actual bound local address.
- Serve the generated Arrow Flight service with configured message limits.
- Terminate cleanly through an explicit shutdown future.
- Reject non-loopback listener binds even when external auth policy is selected.
- Keep non-loopback service exposure, TLS/auth implementation, gateway integration, request scheduling, SQL planning, predicate pushdown, distributed execution, and gateway direct data access out of scope.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 101 Rust tests across unit, integration, and doc-test targets.
This proves a bounded local listener lifecycle without exposing EHDB as a networked platform service yet.
Implementation branch:
kadyapam/ehdb-loopback-flight-client-smoke
Design target:
- Start
LocalArrowFlightListeneron loopback with explicit shutdown. - Connect with the Arrow Flight client over real tonic/gRPC transport.
- Call
get_flight_infousing the existing scan command descriptor. - Follow the returned endpoint ticket with
do_get. - Decode the returned Arrow record batches and assert filtered rows.
- Keep this local-reference only: no gateway integration, non-loopback exposure, TLS/auth implementation, request scheduler, SQL planning, predicate pushdown, distributed execution, or gateway direct data access.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 102 Rust tests across unit, integration, and doc-test targets.
This verifies the first end-to-end Flight read path over local gRPC transport while preserving the NoETL execution-model boundary.
Implementation branch:
kadyapam/ehdb-flight-auth-header-policy
Design target:
- Extend
FlightAuthPolicywith a validated header-token mode for the local Arrow Flight reference service. - Reject invalid metadata header names, binary metadata headers, empty tokens, oversized tokens, and control-character tokens at config validation time.
- Enforce request metadata auth on implemented scan methods:
get_flight_infoanddo_get. - Keep default local-reference behavior unauthenticated and loopback only.
- Prove the policy directly through the generated service trait adapter and over the loopback listener with the Arrow Flight client.
- Keep this as an auth-boundary contract only: no non-loopback exposure, production TLS/identity, ACL enforcement, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 106 Rust tests across unit, integration, and doc-test targets.
This gives the local Flight harness an explicit request metadata boundary before EHDB grows a production service exposure model.
Implementation branch:
kadyapam/ehdb-flight-scope-metadata-guard
Design target:
- Add
FlightScanScopePolicyfor tenant/namespace scope metadata on implemented Arrow Flight scan calls. - Keep default local-reference behavior scope-disabled for existing in-process tests and local harnesses.
- When enabled, require
x-ehdb-tenantandx-ehdb-namespacemetadata to match the decodedScanLatestTableRequestbefore local scan execution. - Return
UNAUTHENTICATEDfor missing scan scope metadata andPERMISSION_DENIEDfor mismatched scope metadata. - Prove the guard directly through the generated service trait adapter and over the loopback listener with the Arrow Flight client.
- Keep this as a scope-boundary contract only: no ACL engine, non-loopback exposure, production TLS/identity, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 110 Rust tests across unit, integration, and doc-test targets.
This gives future catalog ACL work a concrete tenant/namespace request scope contract without changing the NoETL execution-model boundary.
Implementation branch:
kadyapam/ehdb-catalog-scan-grants
Design target:
- Add
PrincipalIdas a typed EHDB identifier. - Add
CatalogScanGrantandGrantScanto the local catalog reference. - Record tenant, namespace, table ID, principal, and granting transaction ID for scan grants.
- Reject scan grants for missing tables and duplicate table/principal grants.
- Add
InMemoryCatalog::can_scanandscan_grant_countfor future service authorization checks. - Add
CatalogMutation::GrantScanso grants are replayable through the transaction log,ehdb-reference, andLocalReferenceRuntime. - Keep this as catalog ACL metadata only: no production IAM, policy composition, revocation, service enforcement, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 113 Rust tests across unit, integration, and doc-test targets.
This adds durable catalog-side scan authorization metadata for future ACL enforcement while preserving the NoETL execution-model boundary.
Implementation branch:
kadyapam/ehdb-flight-catalog-scan-grants
Design target:
- Add
FlightScanGrantPolicyfor catalog-backed local Arrow Flight scan authorization. - Keep default local-reference behavior grant-disabled for existing harnesses.
- When enabled, require
x-ehdb-principalmetadata on implemented Arrow Flight scan calls. - Validate the metadata value as a
PrincipalId, resolve the requested table from replayed catalog state, and callInMemoryCatalog::can_scanbefore local scan execution. - Return
UNAUTHENTICATEDfor missing or invalid principal metadata andPERMISSION_DENIEDwhen the principal lacks a replayedCatalogScanGrantfor the requested table. - Prove the guard directly through the generated service trait adapter and over the loopback listener with the Arrow Flight client.
- Preserve the previous
new_with_policiesconstructor shape and add a fuller constructor for auth, scan-scope, and scan-grant policy combinations. - Keep this as local reference enforcement only: no production IAM, policy composition, revocation, non-loopback exposure, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 117 Rust tests across unit, integration, and doc-test targets.
This connects replayed catalog ACL metadata to the local Flight read path while keeping gateway as gatekeeper and EHDB scans inside explicit service/playbook boundaries.
Implementation branch:
kadyapam/ehdb-flight-access-log-policy
Design target:
- Add bounded local Arrow Flight scan access summaries for decoded scan requests.
- Keep default local-reference behavior DEBUG-only, never INFO-level.
- Add
FlightScanAccessLogEntrywith method, gRPC code, row/message counts, projection count, predicate presence, and which metadata guards were required. - Add disabled mode that emits no scan access summaries.
- Exclude auth tokens, principal values, tenant/table identifiers, object paths, predicate values, and Arrow payloads from the summary contract.
- Wire the configured policy through
LocalArrowFlightServerConfiginto implementedget_flight_infoanddo_getpaths. - Preserve existing constructor shapes and add a fuller runtime-policy constructor for auth, scope, grant, and access-log combinations.
- Keep this as local reference observability only: no non-loopback exposure, production IAM, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 120 Rust tests across unit, integration, and doc-test targets.
This completes the first log-hygiene contract for the local Flight read path without introducing high-volume INFO logs or tenant data leakage.
Implementation branch:
kadyapam/ehdb-flight-get-schema-adapter
Design target:
- Add
get_schemasupport for local Arrow Flight scan command descriptors. - Reuse the existing versioned
ScanFlightTicketcommand descriptor contract. - Add
LocalArrowFlightService::get_schemato produce Arrow FlightSchemaResultvalues from the projected latest-table scan schema. - Implement generated
FlightService::get_schemawith the same request metadata auth, tenant/namespace scan scope, catalog scan grant, and bounded access-log policies used byget_flight_infoanddo_get. - Extend the loopback client smoke path to call
FlightClient::get_schemaover tonic/gRPC before reading data. - Keep this as local reference schema discovery only: no non-loopback exposure, production IAM, gateway direct reads, SQL planning, predicate pushdown, distributed execution, request scheduler, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 120 Rust tests across unit, integration, and doc-test targets.
This completes the first schema-discovery endpoint for the local Flight read path without changing the NoETL gateway/worker/playbook boundary.
Implementation branch:
kadyapam/ehdb-flight-concurrency-guard
Design target:
- Enforce
LocalArrowFlightServerConfig::max_concurrent_requestsfor implemented local Arrow Flight scan calls. - Add a fail-fast semaphore to
LocalArrowFlightServerand pass the configured request budget when building services from config. - Preserve existing constructors by defaulting them to the local reference request budget.
- Return deterministic gRPC
RESOURCE_EXHAUSTEDwhen all local request slots are occupied. - Cover
get_flight_info,get_schema, anddo_get. - Keep this as a local reference lifecycle guard only: no request queue, scheduler, non-loopback exposure, production IAM, gateway direct reads, SQL planning, predicate pushdown, distributed execution, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 121 Rust tests across unit, integration, and doc-test targets.
This makes the validated local concurrency budget executable without introducing a long-lived scheduler or changing the NoETL execution-model boundary.
Implementation branch:
kadyapam/ehdb-retrieval-vector-search
Design target:
- Add a typed
VectorSearchrequest andVectorSearchHitresult boundary toehdb-retrieval. - Scope candidates by tenant, namespace, and embedding model.
- Compute exact cosine similarity over registered chunk embeddings.
- Validate finite non-zero embedding and query vectors.
- Apply dimension compatibility and deterministic score/tie ordering.
- Keep this as a local RAG correctness fixture only: no ANN index, retrieval daemon, gateway data path, production IAM, external Qdrant adapter, distributed query engine, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 124 Rust tests across unit, integration, and doc-test targets.
This is the first local vector lookup primitive on replayed EHDB retrieval state, moving the RAG path one step closer to replacing a permanent Qdrant dependency without changing the NoETL execution-model boundary.
Implementation branch:
kadyapam/ehdb-retrieval-service-boundary
Design target:
- Add
LocalRetrievalSearchServiceinehdb-service. - Add typed
SearchSimilarChunksRequestandSearchSimilarChunksHitservice-facing boundaries. - Read replayed
LocalReferenceRuntimeretrieval state and call the exact localVectorSearchfixture. - Return ranked chunk identity, document identity, ordinal, text, checksum, embedding model, dimensions, and score while excluding raw embedding vectors from service results.
- Cover replayed-state search, tenant/namespace/model scoping, score ordering, empty results, and validation propagation.
- Keep this as an in-process reference boundary only: no network service, gateway route, production IAM, ANN index, external Qdrant adapter, distributed query engine, retrieval daemon, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 126 Rust tests across unit, integration, and doc-test targets.
This gives NoETL a stable local RAG lookup boundary for future worker/playbook integration while preserving the gateway/worker/playbook execution model.
Implementation branch:
kadyapam/ehdb-retrieval-text-service
Design target:
- Add typed
TextSearchandTextSearchHitboundaries toehdb-retrieval. - Validate non-empty queries and positive limits.
- Scope exact case-insensitive substring matching by tenant and namespace.
- Rank by match count with deterministic document/ordinal/chunk tie ordering.
- Add
SearchTextChunksRequest,SearchTextChunksHit, andLocalRetrievalSearchService::search_textinehdb-service. - Cover local catalog search, replayed-state service search, limits, empty results, and validation propagation.
- Keep this as an in-process reference boundary only: no full-text index, BM25 engine, network service, gateway route, production IAM, external search adapter, distributed query engine, retrieval daemon, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 130 Rust tests across unit, integration, and doc-test targets.
This complements exact vector lookup with text retrieval semantics while keeping the first RAG service surface local and worker/playbook-shaped.
Implementation branch:
kadyapam/ehdb-retrieval-hybrid-search
Design target:
- Add typed
HybridSearchandHybridSearchHitboundaries toehdb-retrieval. - Validate vector query, text query, positive limit, finite non-negative weights, and at least one positive weight.
- Scope candidates by tenant, namespace, embedding model, and vector dimension.
- Combine exact cosine similarity and exact text match counts with caller-provided weights.
- Add
SearchHybridChunksRequest,SearchHybridChunksHit, andLocalRetrievalSearchService::search_hybridinehdb-service. - Cover local catalog search, replayed-state service search, limits, empty results, deterministic ordering, and validation propagation.
- Keep this as an in-process reference boundary only: no ANN index, full-text index, BM25 engine, query planner, network service, gateway route, production IAM, external Qdrant/search adapter, distributed query engine, retrieval daemon, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 135 Rust tests across unit, integration, and doc-test targets.
This gives the local RAG surface a deterministic hybrid scoring fixture without changing the NoETL execution-model boundary.
Implementation branch:
kadyapam/ehdb-retrieval-context-assembly
Design target:
- Add
AssembleRetrievalContextRequest,RetrievalContextBlock, andRetrievalContexttoehdb-service. - Build context blocks from the existing replayed local hybrid search path, preserving ordering and score metadata.
- Enforce positive hit limits through hybrid search plus positive per-block and total text budgets at context assembly.
- Return chunk id, document id, ordinal, checksum, model id, dimensions, vector score, text match count, combined score, clipped text, original text length, and truncation metadata.
- Cover ordered assembly, budget clipping, empty results, validation, and tenant/namespace scoping.
- Keep this as an in-process worker/playbook-shaped boundary only: no ANN index, BM25 engine, learned ranker, prompt template engine, LLM invocation, network service, gateway route, production IAM, external search adapter, distributed query engine, retrieval daemon, gateway direct data path, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 139 Rust tests across unit, integration, and doc-test targets.
This gives NoETL RAG tests a deterministic local context shape without moving retrieval state or data access into the gateway.
Implementation branch:
kadyapam/ehdb-retrieval-context-codec
Design target:
- Add versioned
RetrievalContextRequestPayloadandRetrievalContextResultPayloadwrappers toehdb-service. - Encode and decode local retrieval context assembly requests/results as deterministic JSON bytes.
- Reject malformed JSON and unsupported request/result payload versions before execution or handoff.
- Derive serialization for
AssembleRetrievalContextRequest,RetrievalContextBlock, andRetrievalContext. - Cover request payload round-trip, assembled result payload round-trip, malformed payloads, and unsupported versions.
- Keep this as a local worker/playbook payload boundary only: no network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 143 Rust tests across unit, integration, and doc-test targets.
This gives NoETL RAG context assembly a stable local serialization boundary without turning retrieval into a standalone service.
Implementation branch:
kadyapam/ehdb-retrieval-context-executor
Design target:
- Add
LocalRetrievalSearchService::execute_context_payload. - Decode versioned retrieval context request payload bytes.
- Assemble context from replayed
LocalReferenceRuntimeretrieval state using the existing local context assembly boundary. - Encode the assembled context as a versioned result payload.
- Propagate malformed payload, unsupported version, and invalid search/budget errors deterministically.
- Cover happy-path payload execution, malformed request payloads, unsupported request versions, invalid assembly inputs, and empty result payloads.
- Keep this as an in-process worker/playbook payload executor only: no network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 146 Rust tests across unit, integration, and doc-test targets.
This gives NoETL RAG context assembly a complete local payload-in / payload-out handoff without moving retrieval into a standalone service.
Implementation branch:
kadyapam/ehdb-retrieval-context-bounds
Design target:
- Add
RetrievalContextPayloadExecutorConfigtoehdb-service. - Validate positive max request and max result payload byte limits.
- Keep
LocalRetrievalSearchService::execute_context_payloadon a conservative default config. - Add config-aware local payload execution for worker/playbook tests.
- Reject oversized request payloads before JSON decode.
- Reject oversized encoded result payloads before returning bytes.
- Cover default config, invalid config, oversized request payloads, oversized result payloads, and happy-path configured execution.
- Keep this as an in-process worker/playbook payload guard only: no network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 150 Rust tests across unit, integration, and doc-test targets.
This bounds the local RAG payload handoff without making retrieval a service or moving data access into the gateway.
Implementation branch:
kadyapam/ehdb-retrieval-context-scope
Design target:
- Add
RetrievalContextPayloadScopetoehdb-service. - Validate decoded context assembly requests against an expected tenant and namespace before context assembly.
- Add
LocalRetrievalSearchService::execute_context_payload_with_scopethat composes existing byte bounds with the local scope check. - Keep default and config-aware payload execution unchanged.
- Cover matching scope, tenant mismatch, namespace mismatch, malformed payload propagation, and oversized request propagation.
- Keep this as an in-process worker/playbook correctness guard only: no production IAM, policy engine, ACL integration, network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 153 Rust tests across unit, integration, and doc-test targets.
This adds local tenant/namespace guardrails for RAG context payload execution without turning them into production authorization.
Implementation branch:
kadyapam/ehdb-retrieval-context-summary
Design target:
- Add
RetrievalContextPayloadExecutionSummarytoehdb-service. - Add
RetrievalContextPayloadExecutionfor summary-returning local payload execution APIs. - Keep existing byte-returning default, config-aware, and scope-aware execution APIs behavior-compatible by delegating through the summary path.
- Report request/result byte counts, context block count, total text chars, truncation status, and whether a local scope guard was required.
- Explicitly exclude tenant IDs, namespace values, query text, chunk text, tokens, vectors, payload bytes, object paths, and principals from the summary.
- Cover happy path, truncation metadata, scoped execution metadata, and request/result bound propagation.
- Keep this as redacted metrics/audit metadata for local worker/playbook tests only: no logging sink, network API, Arrow Flight retrieval endpoint, prompt template engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 157 Rust tests across unit, integration, and doc-test targets.
This gives the local RAG payload handoff a safe summary shape for future audit/metrics plumbing without exposing retrieval-sensitive content or creating a service boundary.
Implementation branch:
kadyapam/ehdb-retrieval-context-receipt
Design target:
- Add
RETRIEVAL_CONTEXT_EXECUTION_RECEIPT_VERSION. - Add
RetrievalContextPayloadExecutionReceiptPayloadto wrapRetrievalContextPayloadExecutionSummaryin a versioned JSON byte codec. - Derive serialization for the redacted summary only.
- Reject malformed receipt JSON and unsupported receipt versions deterministically.
- Prove encoded receipts exclude tenant IDs, namespace values, query text, chunk text, tokens, vectors, payload bytes, object paths, and principals.
- Keep existing retrieval context payload executor behavior unchanged.
- Keep this as a durable receipt shape for future event-log/audit plumbing only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 160 Rust tests across unit, integration, and doc-test targets.
This makes the local RAG payload summary durable and replay-friendly without turning it into logging, event publication, or a service boundary.
Implementation branch:
kadyapam/ehdb-retrieval-context-receipt-helper
Design target:
- Add
RetrievalContextPayloadExecution::encode_receipt_payload. - Produce versioned
RetrievalContextPayloadExecutionReceiptPayloadbytes directly from the redacted execution summary. - Keep existing executor and receipt codec behavior unchanged.
- Prove helper-produced receipts decode to the same summary returned by execution.
- Prove helper-produced receipt bytes exclude tenant IDs, namespace values, query text, chunk text, tokens, vectors, payload bytes, object paths, and principals.
- Keep this as local worker/playbook helper wiring only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 162 Rust tests across unit, integration, and doc-test targets.
This lets a local retrieval context execution return result bytes and receipt bytes from one execution object while preserving the redacted receipt boundary.
Implementation branch:
kadyapam/ehdb-retrieval-context-receipt-validation
Design target:
- Add
RetrievalContextPayloadExecutionSummary::validate. - Require positive request and result payload byte counts in receipt summaries.
- Reject non-zero total text chars when context block count is zero.
- Apply summary validation during
RetrievalContextPayloadExecutionReceiptPayload::encodeanddecode. - Keep execution-produced receipts and helper-produced receipts behavior-compatible.
- Cover valid empty-context receipts, invalid counter combinations, decode-time validation, and helper-produced receipt compatibility.
- Keep this as local receipt contract hardening only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 165 Rust tests across unit, integration, and doc-test targets.
This prevents malformed hand-built receipt summaries from becoming durable fixtures while preserving the redacted local worker/playbook receipt boundary.
Implementation branch:
kadyapam/ehdb-retrieval-context-artifacts
Design target:
- Add
DEFAULT_RETRIEVAL_CONTEXT_MAX_RECEIPT_PAYLOAD_BYTES. - Add
max_receipt_payload_bytestoRetrievalContextPayloadExecutorConfigwith positive validation. - Add
RetrievalContextPayloadExecutionArtifactscarrying result payload bytes and redacted receipt payload bytes. - Add artifact helpers for default, configured, and scope-aware local retrieval context payload execution.
- Keep existing byte-returning, summary-returning, scope, and receipt APIs behavior-compatible.
- Cover default config, invalid receipt limit, happy-path artifact emission, scoped artifact emission, and oversized receipt rejection.
- Keep this as a local worker/playbook handoff shape only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 168 Rust tests across unit, integration, and doc-test targets.
This gives local worker/playbook tests one bounded artifact object for result bytes and redacted receipt bytes without publishing events or opening a service boundary.
Implementation branch:
kadyapam/ehdb-retrieval-context-artifact-validation
Design target:
- Add
RetrievalContextPayloadExecutionArtifacts::receipt_summary. - Add
RetrievalContextPayloadExecutionArtifacts::validate. - Decode and validate the redacted receipt payload through the existing receipt codec.
- Reject empty result payload bytes and empty receipt payload bytes.
- Reject artifacts whose receipt summary result byte count does not match the actual result payload length.
- Ensure artifact helpers return artifacts that pass consistency validation.
- Cover valid helper artifacts, malformed receipts, empty payloads, and result-length mismatches.
- Keep this as local artifact contract hardening only: no event publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 171 Rust tests across unit, integration, and doc-test targets.
This makes the local result+receipt artifact pair self-checking before future audit/event plumbing treats it as a coherent handoff.
Implementation branch:
kadyapam/ehdb-retrieval-context-receipt-event
Design target:
- Add
RetrievalContextPayloadExecutionReceiptEventPayload. - Add the stable subject
ehdb.retrieval.context.execution.receipt. - Encode and decode a versioned JSON event envelope containing only validated redacted receipt bytes.
- Build event payloads from validated
RetrievalContextPayloadExecutionArtifacts. - Reject empty receipt payloads, malformed receipts, and unsupported event envelope versions.
- Preserve the redaction boundary by excluding result payload/context bytes from the event envelope.
- Keep this as local stream-ready payload modeling only: no automatic stream publication, stream mutation, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 174 Rust tests across unit, integration, and doc-test targets.
This gives future EHDB stream/audit plumbing a stable local payload shape while keeping publication under explicit worker/playbook control.
Implementation branch:
kadyapam/ehdb-retrieval-receipt-event-publisher
Design target:
- Add
RetrievalContextReceiptEventStreamTarget. - Add
RetrievalContextReceiptEventStreamLog. - Implement explicit local publisher support for
InMemoryStreamLogandLocalJsonlStreamLog. - Publish only caller-supplied validated receipt event payloads to the
stable subject
ehdb.retrieval.context.execution.receipt. - Require caller-owned tenant, namespace, stream name, mutable stream log, and transaction id.
- Cover in-memory publish/replay, JSONL persist/reopen/replay, missing stream errors, and malformed artifact rejection.
- Keep this as explicit local worker/playbook publication only: no automatic stream publication, background task, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 177 Rust tests across unit, integration, and doc-test targets.
This gives worker/playbook tests an explicit path to append redacted retrieval receipt events into EHDB streams without making publication implicit or gateway-driven.
Implementation branch:
kadyapam/ehdb-retrieval-receipt-event-replay
Design target:
- Add
RetrievalContextReceiptEventStreamRecord. - Add
RetrievalContextReceiptEventStreamReadLog. - Decode receipt event stream records by validating the stable subject
ehdb.retrieval.context.execution.receipt. - Validate payloads through
RetrievalContextPayloadExecutionReceiptEventPayload. - Preserve stream sequence and transaction id for local audit assertions.
- Replay from caller-supplied
InMemoryStreamLogandLocalJsonlStreamLoginstances, including cursor replay. - Cover ordered replay, cursor replay, JSONL reopen/replay, wrong subject rejection, and malformed payload rejection.
- Keep this as explicit local worker/playbook replay only: no background consumer, subscription loop, automatic processing, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, scheduler, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 180 Rust tests across unit, integration, and doc-test targets.
This gives local audit tests the matching read side for explicit receipt event publication while keeping consumption caller-controlled.
Implementation branch:
kadyapam/ehdb-retrieval-receipt-event-consumer
Design target:
- Add
RetrievalContextReceiptEventDurableConsumerLog. - Add explicit target helpers to create durable consumers, replay pending validated receipt events for a consumer, and ack receipt event sequences.
- Support caller-supplied
InMemoryStreamLogandLocalJsonlStreamLoginstances. - Keep replay validation on the stable subject
ehdb.retrieval.context.execution.receiptand receipt event payload. - Cover consumer resume/ack behavior, ack rollback rejection, missing consumer rejection, and JSONL reopen cursor behavior.
- Keep this as explicit local worker/playbook consumer control only: no background consumer, subscription loop, scheduler, automatic processing, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 183 Rust tests across unit, integration, and doc-test targets.
This gives local audit tests resumable receipt event consumption while keeping cursor advancement under explicit caller control.
Implementation branch:
kadyapam/ehdb-retrieval-receipt-stream-setup
Design target:
- Add explicit stream setup helpers to
RetrievalContextReceiptEventStreamTarget. - Build
StreamConfigfrom target tenant, namespace, stream, and caller-selected retention policy. - Create receipt event streams through caller-supplied
InMemoryStreamLogandLocalJsonlStreamLoginstances. - Keep setup explicit: publish helpers do not auto-create streams.
- Cover setup plus publish/replay, duplicate stream rejection, and JSONL stream persistence after reopen.
- Keep this as explicit local worker/playbook stream setup only: no auto-create-on-publish, scheduler, automatic processing, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 186 Rust tests across unit, integration, and doc-test targets.
This gives worker/playbook tests an explicit setup step for receipt event streams before publication.
Implementation branch:
kadyapam/ehdb-retrieval-receipt-stream-retention
Design target:
- Add
create_keep_all_streamtoRetrievalContextReceiptEventStreamTarget. - Add
create_bounded_streamtoRetrievalContextReceiptEventStreamTarget. - Reject zero bounded retention before touching the stream log.
- Cover bounded retention replay behavior, zero-bound rejection, and JSONL bounded stream reopen behavior.
- Keep this as explicit local worker/playbook stream setup only: no auto-create-on-publish, scheduler, automatic processing, logging sink, network API, Arrow Flight retrieval endpoint, prompt engine, LLM invocation, ANN index, BM25 engine, learned ranker, gateway route, production IAM, ACL integration, retrieval daemon, distributed query engine, gateway direct data path, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 189 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-stream-positive-retention
Design target:
- Reject
RetentionPolicy::MaxRecords(0)at the stream log boundary. - Keep in-memory and JSONL stream creation aligned on the same config validation.
- Reject zero max-record retention before JSONL stream setup writes a journal entry.
- Preserve keep-all retention and positive bounded retention behavior.
- Keep this as local stream log validation only: no scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 191 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-stream-subject-filtered-replay
Design target:
- Add subject matching for exact subjects, single-token
*wildcards, and terminal>tail wildcards. - Keep misplaced
>filters from matching unrelated subjects. - Add explicit subject-filtered replay for
InMemoryStreamLog. - Add matching delegated subject-filtered replay for
LocalJsonlStreamLog. - Cover exact matching, wildcard matching, cursor behavior, and JSONL reopen behavior.
- Keep this as local explicit stream-log replay only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 194 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-stream-filtered-consumer-replay
Design target:
- Add explicit subject-filtered durable consumer replay for
InMemoryStreamLog. - Add matching delegated subject-filtered durable consumer replay for
LocalJsonlStreamLog. - Filter records pending after the durable consumer ack cursor without moving that cursor.
- Preserve missing consumer
EhdbError::NotFoundbehavior. - Cover wildcard filtering, cursor behavior, missing consumers, and JSONL reopen behavior.
- Keep this as local explicit stream-log replay only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 196 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-stream-journal-subject-validation
Design target:
- Validate
StreamRecord.subjectbefore inserting replayed journal records. - Reject persisted publish entries with wildcard concrete subjects on JSONL reopen.
- Reject persisted publish entries with empty-token concrete subjects on JSONL reopen.
- Preserve valid JSONL replay behavior.
- Keep this as local stream journal replay validation only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 198 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-stream-journal-identifier-validation
Design target:
- Revalidate persisted stream config tenant, namespace, and stream identifiers during JSONL replay.
- Revalidate create-consumer, publish, and ack tenant/namespace/stream coordinates before applying journal entries.
- Revalidate persisted consumer names and stream record transaction IDs before rebuilding consumer cursor or retained-record state.
- Preserve valid JSONL replay behavior.
- Keep this as local stream journal replay validation only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 202 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-stream-sequence-serde-validation
Design target:
- Ensure
StreamSequenceserde deserialization preserves the same nonzero invariant asStreamSequence::new. - Reject zero persisted publish record sequences during JSONL stream journal replay.
- Reject zero persisted ack cursor sequences during JSONL stream journal replay.
- Preserve the existing compact numeric JSON representation for valid stream sequences.
- Keep this as local stream journal replay validation only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 241 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-stream-journal-fields
Design target:
- Make local JSONL stream journal entry replay strict.
- Reject unknown create-stream, create-consumer, publish, and ack operation fields during replay.
- Reject unknown persisted stream config fields during replay.
- Reject unknown persisted stream record fields during replay.
- Keep valid stream journal replay behavior unchanged.
- Keep this as local stream journal replay validation only: no stream publication behavior changes, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 242 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-transaction-journal-identifier-validation
Design target:
- Revalidate persisted transaction record transaction ID, tenant, and namespace identifiers during JSONL replay.
- Revalidate catalog, stream, retrieval, system-library, and storage mutation identifiers before inserting replayed records.
- Revalidate stream publish subjects in transaction mutations as concrete subjects.
- Preserve valid JSONL transaction-log replay behavior.
- Keep this as local transaction-log replay validation only: no consensus engine, background processing, network API, gateway data-touch behavior, production replication, scheduler behavior, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 208 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-transaction-journal-fields
Design target:
- Make local JSONL transaction journal replay strict.
- Reject unknown persisted transaction record fields during replay.
- Reject unknown catalog, stream, retrieval, system-library, and storage mutation fields during replay.
- Keep valid transaction journal replay behavior unchanged.
- Keep this as local transaction journal replay validation only: no network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed transaction coordinator, production replication, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 243 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-system-library-journal-fields
Design target:
- Reject unknown top-level system-library journal fields during JSONL replay.
- Reject unknown persisted publish request fields during JSONL replay.
- Reject unknown persisted bind request fields during JSONL replay.
- Preserve valid publish/bind replay behavior and hot-replacement bindings.
- Keep this as local system-library journal replay validation only: no WASM execution, background processing, network API, gateway data-touch behavior, production replication, scheduler behavior, distributed transaction coordinator, object transfer execution, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 244 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-storage-metadata-fields
Design target:
- Reject unknown JSON fields when decoding object refs.
- Reject unknown JSON fields in geo placements and placement policy targets.
- Reject unknown JSON fields in replica records, replication actions, and replication plans.
- Preserve existing placement validation, replica registry behavior, and deterministic replication planning.
- Keep this as storage metadata JSON decoding validation only: no object movement, cloud adapters, network API, gateway data-touch behavior, production replication, scheduler, background worker, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 245 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-catalog-metadata-fields
Design target:
- Reject unknown JSON fields when decoding catalog tables.
- Reject unknown JSON fields in catalog snapshots and nested object references.
- Reject unknown JSON fields in catalog scan grants.
- Reject unknown JSON fields in create-table, snapshot commit, and scan grant request shapes.
- Reject unknown JSON fields in table schema and column schema metadata.
- Preserve existing in-memory catalog behavior for table creation, snapshot parent-chain validation, and scan grant checks.
- Keep this as catalog metadata JSON decoding validation only: no network API, gateway route, production ACL/IAM engine, query planner, distributed transaction coordinator, background worker, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 246 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-retrieval-metadata-fields
Design target:
- Reject unknown JSON fields when decoding retrieval documents.
- Reject unknown JSON fields in chunks and embeddings.
- Reject unknown JSON fields in document, chunk, and embedding registration requests.
- Reject unknown JSON fields in vector, text, and hybrid search request shapes.
- Reject unknown JSON fields in local vector, text, and hybrid search hit metadata, including nested chunk and embedding payloads.
- Preserve existing local retrieval catalog behavior for registration, exact vector search, exact text search, and hybrid scoring.
- Keep this as retrieval metadata JSON decoding validation only: no ANN index, full-text index, retrieval daemon, network API, gateway route, prompt engine, LLM invocation, production IAM, query planner, background worker, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 247 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-system-library-metadata-fields
Design target:
- Reject unknown JSON fields when decoding resolved WASM system library manifests.
- Reject unknown JSON fields when decoding NoETL WASM plugin references used for worker handoff metadata.
- Reject unknown JSON fields when decoding environment/channel system library bindings.
- Preserve existing publish, bind, resolve, local JSONL replay, and hot replacement behavior.
- Keep this as system-library metadata JSON decoding validation only: no WASM execution, background processing, network API, gateway data-touch behavior, production replication, scheduler behavior, distributed transaction coordinator, object transfer execution, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 248 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-local-arrow-scan-projection-validation
Design target:
- Reject empty direct local Arrow scan projection lists before object reads.
- Reject duplicate projection columns in direct local Arrow scan requests before object reads.
- Preserve ordered valid projections and existing missing-column errors.
- Keep this as local Arrow IPC scan request validation only: no SQL planning, predicate pushdown, distributed execution, gateway direct reads, Arrow Flight protocol changes, production IAM/ACL behavior, request scheduling, object movement, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 249 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-local-arrow-scan-selector-validation
Design target:
- Validate projection-column selector identifiers in direct local Arrow scan requests before object reads.
- Validate equality-predicate column selector identifiers in direct local Arrow scan requests before object reads.
- Preserve existing
NotFoundbehavior for valid-but-missing projection and predicate columns. - Keep this as direct local Arrow IPC scan selector validation only: no SQL planning, predicate pushdown, distributed execution, gateway direct reads, Arrow Flight protocol changes, production IAM/ACL behavior, request scheduling, object movement, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 250 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-reject-duplicate-schema-columns
Design target:
- Reject duplicate
TableSchemacolumn names inehdb-core. - Preserve existing non-empty schema and per-column identifier validation.
- Keep Arrow projection and predicate selectors unambiguous before catalog state is created.
- Keep this as table schema validation only: no schema evolution, type coercion, SQL planning, predicate pushdown, distributed execution, gateway direct reads, production IAM/ACL behavior, object movement, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 251 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-revalidate-schema-column-identifiers
Design target:
- Revalidate every
TableSchemacolumn identifier inehdb-core, including preconstructed or decodedColumnSchemavalues. - Preserve existing non-empty schema and duplicate-column validation.
- Keep Arrow projection and predicate selectors unambiguous before catalog state is created.
- Keep this as table schema validation only: no schema evolution, type coercion, SQL planning, predicate pushdown, distributed execution, gateway direct reads, production IAM/ACL behavior, object movement, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 252 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-schema-json-decode-validation
Design target:
- Route
ColumnSchemaJSON decode throughColumnSchema::new. - Route
TableSchemaJSON decode throughTableSchema::new. - Preserve strict unknown-field behavior and the existing schema JSON shape.
- Reject invalid column identifiers and duplicate table schema columns during JSON decode before metadata is accepted.
- Keep this as table/column schema JSON decode validation only: no schema evolution, type coercion, SQL planning, predicate pushdown, distributed execution, gateway direct reads, production IAM/ACL behavior, object movement, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 253 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-core-identifier-json-decode-validation
Design target:
- Route core identifier newtype JSON decode through each identifier constructor.
- Preserve existing string JSON shape for identifiers.
- Reject malformed tenant, namespace, table, transaction, stream, retrieval, and related identifiers during JSON decode.
- Keep this as core identifier JSON decode validation only: no schema evolution, type coercion, SQL planning, predicate pushdown, distributed execution, gateway direct reads, production IAM/ACL behavior, object movement, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 254 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-noetl-runtime-surface-replay-fixture
Design target:
- Add a local NoETL-shaped runtime surface fixture over
LocalReferenceRuntime. - Drive catalog, stream, retrieval, system WASM, and storage replica
metadata only through appended
CommitTransactionrecords. - Reopen the runtime and verify catalog scan grants, stream consumer replay, retrieval text lookup, system library resolution, and storage replica inventory from transaction replay.
- Keep this as a local worker/playbook-shaped integration-readiness fixture only: no gateway route, direct gateway data access, persistent per-tenant service, production IAM, distributed execution, SQL planner, object movement, or external dependency replacement behavior.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 255 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-system-journal-identifier-validation
Design target:
- Revalidate persisted system-library publish entries during JSONL replay: library path, revision, digest, object path, and transaction ID.
- Revalidate persisted system-library bind entries during JSONL replay: tenant, namespace, environment, channel, path, revision, digest, and transaction ID.
- Preserve valid publish/bind replay behavior and hot-replacement bindings.
- Keep this as local system-library journal replay validation only: no WASM execution, background processing, network API, gateway data-touch behavior, production replication, scheduler behavior, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 210 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-canonical-retrieval-context-payload-encoding
Design target:
- Require local retrieval context request/result payload bytes to match
the canonical EHDB encoding produced by each payload's
encodemethod. - Reject pretty-printed or otherwise non-canonical JSON request payload bytes before local worker/playbook execution or handoff.
- Reject pretty-printed or otherwise non-canonical JSON result payload bytes before local worker/playbook execution or handoff.
- Preserve valid payloads produced by
encode. - Keep existing local payload execution paths on the stricter decode path.
- Keep this as local retrieval context worker/playbook payload byte-contract validation only: no network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 234 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-retrieval-receipt-payload-fields
Design target:
- Make local retrieval context execution receipt JSON payload decoding strict.
- Reject unknown top-level receipt payload fields.
- Reject unknown redacted execution summary fields.
- Preserve valid receipt payload round trips.
- Keep existing artifact and receipt-event helper paths on the stricter receipt decoder.
- Keep this as local retrieval context receipt payload validation only: no stream publication behavior, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 235 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-canonical-retrieval-receipt-payload-encoding
Design target:
- Require local retrieval context execution receipt payload bytes to
match the canonical EHDB encoding produced by
RetrievalContextPayloadExecutionReceiptPayload::encode. - Reject pretty-printed or otherwise non-canonical receipt JSON bytes before artifact validation, receipt decode handoff, or receipt-event helper use.
- Preserve valid receipt payloads produced by
encode. - Keep existing artifact and receipt-event helper paths on the stricter receipt decoder.
- Keep this as local retrieval context receipt payload byte-contract validation only: no stream publication behavior, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 236 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-retrieval-receipt-event-payload-fields
Design target:
- Make local retrieval context execution receipt event JSON payload decoding strict.
- Reject unknown top-level event envelope fields before replay decode, publisher helper use, or consumer handoff.
- Preserve valid event payload round trips from execution artifacts.
- Keep nested receipt payload bytes on the existing strict canonical receipt decoder.
- Keep this as local retrieval context receipt event payload validation only: no stream publication behavior changes, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 237 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-canonical-retrieval-receipt-event-payload-encoding
Design target:
- Require local retrieval context execution receipt event payload bytes
to match the canonical EHDB encoding produced by
RetrievalContextPayloadExecutionReceiptEventPayload::encode. - Reject pretty-printed or otherwise non-canonical event JSON bytes before replay decode, publisher helper use, or consumer handoff.
- Preserve valid event payload round trips from execution artifacts.
- Keep nested receipt payload bytes on the existing strict canonical receipt decoder.
- Keep this as local retrieval context receipt event payload byte-contract validation only: no stream publication behavior changes, network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 238 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-retrieval-context-payload-fields
Design target:
- Make local retrieval context request/result JSON payload decoding strict.
- Reject unknown top-level request payload fields.
- Reject unknown embedded context assembly request fields.
- Reject unknown top-level result payload fields.
- Reject unknown context object and context block fields.
- Preserve valid request/result payload round trips.
- Keep this as local retrieval context worker/playbook payload validation only: no network API, gateway route, prompt engine, LLM invocation, retrieval daemon, distributed search service, production IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 233 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-retrieval-payload-identifier-validation
Design target:
- Revalidate decoded
RetrievalContextRequestPayloadidentifiers: tenant, namespace, and embedding model. - Revalidate decoded
RetrievalContextResultPayloadcontext-block identifiers: chunk, document, and embedding model. - Validate identifiers on encode as well as decode.
- Preserve valid request/result payload round trips.
- Keep this as local RAG payload codec validation only: no ANN index, retrieval daemon, RPC protocol, Arrow Flight retrieval endpoint, gateway data-touch behavior, prompt/LLM invocation, background processing, scheduler behavior, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 212 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-ticket-identifier-validation
Design target:
- Revalidate decoded
ScanFlightTickettenant, namespace, and table-name identifiers. - Validate the same identifiers on encode, Arrow
Ticket, and command-descriptor paths. - Preserve valid scan-ticket round trips.
- Keep this as local Arrow Flight scan ticket codec validation only: no SQL planner, predicate pushdown, distributed execution, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 213 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-selector-identifier-validation
Design target:
- Revalidate decoded
ScanFlightTicketprojection-column identifiers. - Revalidate decoded
ScanFlightTicketequality-predicate column identifiers. - Validate the same selectors on encode, Arrow
Ticket, and command-descriptor paths. - Preserve valid projection and predicate scan-ticket round trips.
- Keep this as local Arrow Flight scan ticket codec validation only: no SQL planner, predicate pushdown implementation, distributed execution, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 214 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-result-stream-metadata
Design target:
- Mark locally produced
ArrowScanResultFlightDatastreams with an EHDB scan-result stream version. - Reject empty streams before Arrow decode.
- Reject missing result-stream metadata before Arrow decode.
- Reject unsupported result-stream metadata before Arrow decode.
- Preserve valid result stream round trips, including projected schemas.
- Keep this as local Arrow Flight scan result stream codec validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 215 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-result-metadata-envelope
Design target:
- Accept the EHDB scan-result stream version only on the first
FlightDatamessage. - Keep locally produced later
FlightDatamessage app metadata empty. - Reject non-empty later-message app metadata before Arrow decode.
- Preserve valid result stream round trips.
- Keep this as local Arrow Flight scan result stream codec validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 216 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-fixture-validation
Design target:
- Add a local validator for EHDB-produced scan
FlightInfofixtures. - Accept valid locally produced scan
FlightInfo. - Reject unsupported
FlightInfoapp metadata. - Reject unordered scan
FlightInfo. - Reject negative record or byte counts.
- Reject missing endpoint tickets and multiple endpoints.
- Keep this as local Arrow Flight scan
FlightInfofixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 217 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-ticket-consistency
Design target:
- Decode scan
FlightInfocommand descriptors and endpoint tickets in the local validator. - Reject a valid command descriptor paired with a valid but different endpoint ticket.
- Preserve valid locally produced
FlightInfofixtures. - Keep this as local Arrow Flight scan
FlightInfofixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 218 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-endpoint-envelope
Design target:
- Keep locally produced scan
FlightInfoendpoint locations empty. - Reject endpoint locations in the local
FlightInfovalidator. - Reject endpoint expiration timestamps in the local
FlightInfovalidator. - Reject endpoint app metadata in the local
FlightInfovalidator. - Preserve valid locally produced
FlightInfofixtures. - Keep this as local Arrow Flight scan
FlightInfofixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 219 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-schema-metadata
Design target:
- Require non-empty schema IPC bytes in the local scan
FlightInfovalidator. - Reject missing or empty schema metadata.
- Preserve valid locally produced
FlightInfofixtures. - Keep this as local Arrow Flight scan
FlightInfofixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 220 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-schema-ipc
Design target:
- Decode scan
FlightInfoschema IPC bytes during local fixture validation. - Reject non-empty malformed schema IPC metadata.
- Preserve valid locally produced
FlightInfofixtures. - Keep this as local Arrow Flight scan
FlightInfofixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 221 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-byte-count
Design target:
- Require positive
total_bytesmetadata in local scanFlightInfovalidation. - Reject zero byte-count metadata while preserving negative byte-count rejection.
- Preserve valid locally produced
FlightInfofixtures. - Keep this as local Arrow Flight scan
FlightInfofixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 222 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-result-consistency
Design target:
- Add local result-bound scan
FlightInfovalidation. - Reject
FlightInfofixtures whose schema metadata does not match the producingArrowScanResultschema. - Reject row-count and byte-count metadata that does not match the producing result.
- Reject descriptor/ticket pairs that are internally consistent but do not match the expected scan ticket.
- Keep this as local Arrow Flight scan
FlightInfofixture validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 223 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-expected-ticket
Design target:
- Add receiver-side scan
FlightInfovalidation onScanFlightTicket. - Accept returned scan
FlightInfothat matches the expected ticket. - Reject internally consistent
FlightInfodescriptor/ticket pairs that belong to a different scan request. - Validate returned
FlightInfoin the loopback client smoke path before following the endpoint ticket. - Keep this as local Arrow Flight scan
FlightInforeceiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 224 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-schema-response
Design target:
- Add receiver-side scan
FlightInfoschema validation onScanFlightTicket. - Accept returned scan
FlightInfowhose schema metadata matches the expected Arrow schema. - Reject well-formed scan
FlightInfowhose schema metadata differs from the expected schema. - Validate
get_schemaandget_flight_infoschema consistency in the loopback client smoke path before following the endpoint ticket. - Keep this as local Arrow Flight scan
FlightInforeceiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 225 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-data-info-consistency
Design target:
- Add receiver-side scan data validation against returned
FlightInfo. - Accept matching
FlightDatastreams andFlightInfometadata. - Reject row-count, byte-count, and schema mismatches.
- Validate decoded loopback client data against returned
FlightInfobefore treating batches as coherent scan output. - Keep this as local Arrow Flight scan result receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-canonical-scan-ticket-encoding
Design target:
- Require decoded local Arrow Flight scan ticket bytes to match the
canonical EHDB encoding produced by
ScanFlightTicket::encode. - Reject pretty-printed or otherwise non-canonical JSON ticket bytes before scan execution or Flight handoff.
- Preserve valid tickets produced by
ScanFlightTicket::encode,to_arrow_ticket, andcommand_descriptor. - Keep server
get_flight_info,get_schema, anddo_geton the stricter decode path. - Keep this as local Arrow Flight scan ticket byte-contract validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 231 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-strict-scan-ticket-fields
Design target:
- Make the local Arrow Flight scan ticket JSON contract strict.
- Reject unknown top-level
ScanFlightTicketfields before execution or handoff. - Reject unknown embedded latest-table scan request fields before execution or handoff.
- Reject unknown equality predicate fields before execution or handoff.
- Preserve valid ticket encode/decode, command descriptor, and local scan paths.
- Keep this as local Arrow Flight scan ticket/request payload validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 230 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-descriptor-path-validation
Design target:
- Tighten local Arrow Flight scan descriptor decoding so descriptor shape fails before scan execution.
- Reject direct scan
get_flight_infoandget_schemacommand descriptors that carry non-empty path entries. - Preserve valid command descriptors produced by
ScanFlightTicket::command_descriptor. - Keep this as local Arrow Flight scan descriptor request validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 229 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-projection-selector-validation
Design target:
- Tighten local Arrow scan request validation so invalid projection shapes fail before scan execution or Arrow Flight ticket use.
- Reject
projection: Some([])at the service and ticket request validation boundary. - Reject duplicate projection column selectors before local scan execution or Flight ticket encode/decode succeeds.
- Preserve valid projection/filter scan requests.
- Keep this as local scan request selector validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 228 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-do-get-endpoint-ticket-binding
Design target:
- Add receiver-side validation for the concrete Arrow Flight endpoint
Ticketused fordo_get. - Bind that supplied endpoint ticket back to the returned scan
FlightInfo, decoded schema, and expectedScanFlightTicket. - Accept the locally produced endpoint ticket from matching scan
FlightInfo. - Reject a well-formed endpoint ticket for a different scan request
before
do_getis treated as coherent. - Use the helper in local service/server and loopback client receiver
paths before issuing or accepting
do_getresults. - Keep this as local Arrow Flight endpoint-ticket receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-response-envelope-validation
Design target:
- Add an
ArrowScanResulthelper for the complete receiver-side scan response envelope: rawSchemaResult, returned scanFlightInfo, rawFlightData, and expectedScanFlightTicket. - Decode and validate
SchemaResult, validate returnedFlightInfo, validate rawFlightData, and return the decoded schema plus coherent result. - Accept locally produced schema/info/
FlightDataresponses for the expected scan ticket. - Reject well-formed schema responses whose decoded schema differs from
returned scan
FlightInfobefore accepting raw scan data. - Use the helper in local service and server receiver tests that consume
raw
FlightData. - Keep this as local Arrow Flight scan response receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-raw-flight-data-schema-validation
Design target:
- Add an
ArrowScanResulthelper for raw Arrow FlightFlightDataplus the decodedget_schemaschema, returned scanFlightInfo, and expectedScanFlightTicket. - Accept locally produced schema/info/
FlightDataresponses for the expected scan ticket. - Reject well-formed decoded schemas that differ from returned scan
FlightInfometadata before decoding or accepting raw scan data. - Use the helper in local service and server receiver paths before
accepting returned
FlightDatafromdo_get. - Keep this as local Arrow Flight raw scan data receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-decoded-client-response-validation
Design target:
- Add an
ArrowScanResulthelper for already-decoded Arrow batches plus the decodedget_schemaschema, returned scanFlightInfo, and expectedScanFlightTicket. - Accept locally produced schema/info/batch responses for the expected scan ticket.
- Reject well-formed decoded schemas that differ from returned scan
FlightInfometadata before treating returned batches as coherent. - Use the helper in loopback client smoke paths before accepting
returned batches from
do_get. - Keep this as local Arrow Flight decoded client response receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-schema-result-endpoint-ticket
Design target:
- Add receiver-side schema-result plus endpoint-ticket extraction on
ScanFlightTicket. - Decode returned Arrow Flight
SchemaResult, validate returned scanFlightInfoagainst the expected scan ticket and decoded schema, and return the decoded schema plus validated endpoint ticket fordo_get. - Accept locally produced schema results and scan info for the expected ticket.
- Reject well-formed schema results whose decoded schema differs from
returned scan
FlightInfometadata before returning an endpoint ticket. - Use the helper in local service and server tests before
do_get. - Keep this as local Arrow Flight schema/scan-info endpoint-ticket receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-schema-result-info-validation
Design target:
- Add receiver-side schema-result validation on
ScanFlightTicket. - Decode returned Arrow Flight
SchemaResultand validate returned scanFlightInfoagainst the expected scan ticket and decoded schema. - Accept locally produced schema results and scan info for the expected ticket.
- Reject well-formed schema results whose decoded schema differs from
returned scan
FlightInfometadata. - Use the helper in local service and server tests before treating
get_schemaandget_flight_infoas coherent; keep client paths validating already-decoded schemas against returnedFlightInfo. - Keep this as local Arrow Flight schema/scan-info receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-data-ticket-validation
Design target:
- Add receiver-side helpers on
ArrowScanResultfor validating returned scan data against both returnedFlightInfoand the expectedScanFlightTicket. - Support both raw
FlightDatastreams and decoded Arrow batches for client paths that already receive batches. - Accept locally produced scan data and returned scan info for the expected ticket.
- Reject internally valid scan info/data when the expected ticket differs.
- Use the helpers in local service, server, and loopback client smoke paths before treating returned scan data as coherent.
- Keep this as local Arrow Flight scan data receiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-schema-endpoint-ticket
Design target:
- Add schema-aware receiver-side endpoint-ticket extraction on
ScanFlightTicket. - Return the endpoint ticket only after validating returned scan
FlightInfoagainst the expected scan ticket and the expected Arrow schema. - Reject well-formed scan
FlightInfowhose schema metadata differs from the expectedget_schemaresponse before returning an endpoint ticket. - Use the schema-aware endpoint-ticket helper in local service, server,
and loopback client smoke paths before
do_get. - Keep this as local Arrow Flight scan
FlightInforeceiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-scan-info-endpoint-ticket-helper
Design target:
- Add receiver-side endpoint-ticket extraction on
ScanFlightTicket. - Return the endpoint ticket only after validating returned scan
FlightInfoagainst the expected scan ticket. - Reject internally consistent scan
FlightInfothat belongs to a different scan request before returning an endpoint ticket. - Use the validated endpoint-ticket helper in the loopback client smoke
path before
do_get. - Keep this as local Arrow Flight scan
FlightInforeceiver-side validation only: no Flight protocol expansion, distributed execution, SQL planner, predicate pushdown implementation, gateway direct reads, non-loopback exposure, production auth/IAM, background processing, or persistent per-tenant service process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 226 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-stream-subject-filters
Design target:
- Keep
Subjectconcrete for published stream record subjects. - Reject wildcard tokens
*and>inSubject::new. - Add
SubjectFilterfor replay selectors. - Accept exact filters, single-token
*wildcards, and terminal>tail wildcards inSubjectFilter. - Reject misplaced
>and partial wildcard tokens inSubjectFilter::new. - Move filtered replay APIs to
SubjectFilter. - Cover concrete wildcard rejection, valid wildcard filters, invalid filter placement, filtered replay, filtered durable consumer replay, and JSONL reopen behavior.
- Keep this as local stream log validation and replay only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 196 Rust tests across unit, integration, and doc-test targets.
Implementation branch:
kadyapam/ehdb-stream-nonempty-subject-tokens
Design target:
- Reject leading, trailing, and double-dot empty tokens in
Subject::new. - Reject leading, trailing, and double-dot empty tokens in
SubjectFilter::new. - Preserve valid concrete subjects and valid exact/wildcard filters.
- Cover concrete subject rejection, subject filter rejection, valid exact/wildcard filters, and filtered replay behavior.
- Keep this as local stream log validation only: no durable subject subscription, scheduler, background stream processing, NATS bridge, network API, gateway route, distributed stream storage, production replication, or persistent per-tenant process.
Tracking issue:
Validation:
cargo fmt --all --checkcargo test --workspacecargo clippy --workspace --all-targets -- -D warningscargo bench --workspace --no-run
Coverage snapshot:
- 196 Rust tests across unit, integration, and doc-test targets.
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)