Repository navigation
Design Performance and Load Testing
Status: Phase 1 + Layer B landed + shared-publish fix shipped & validated
(2026-07-08). The engine micro-benchmarks are merged in
ehdb-reference::benches::engine_micro
(noetl/ehdb#261 / #262), with a
committed baseline captured on a dev box (below). Layer B — the in-cluster
load harness + the EHDB-vs-incumbent head-to-head — ran in kind
(#263); results in the
Layer B section below. The
shared-tier publish O(segment) bottleneck it surfaced
(#264) is fixed (#265 + #266 —
now O(delta), flat micro-bench) and validated in kind, which corrected the
headline: the dominant deployed cost is the separate local replay-on-open
(#267), not the shared publish. See
Fix landed + a corrected attribution.
Layer B is directional (podman-VM); Layer A remains the authoritative engine
signal.
Tracks: noetl/ehdb#261 (perf/load-test workstream) · noetl/ehdb#241 (completion program) · Design: Event-Log Core Engine (Phase 6) · Design: Durable Event-Log Backend (the durability work these benches quantify) · Roadmap.
Give EHDB a repeatable performance signal so its founding claim — fixing NoETL's event-sourcing log bottleneck — is measured, not asserted. Three concrete goals:
- Quantify the event-log bottleneck fix. Put a number on the durable event-log tier against the incumbent it replaces (the PostgreSQL + NATS JetStream log-and-store path). The headline question: how much throughput and tail-latency headroom does EHDB's durable event-log/replay path buy over the incumbent on the same append + replay workload?
- Catch regressions in-engine, cheaply. Deterministic Rust micro-benchmarks per tier that run in CI-adjacent time and fail loudly when a change moves append latency, replay speed, or fold rate.
- Set SLO targets the prod-cutover runbooks can gate on — per-tier throughput / latency budgets, confirmed by the platform owner rather than pulled from thin air.
Platform data only. Every payload in these benches is a synthetic, secret-free platform event / KV entry / object / vector — never business data. EHDB serves NoETL's internal platform tiers (event log, projections, platform KV/state, platform artifacts, catalog embeddings); tenant data stays in real business DBs behind playbook connectors. See RFC.
Performance evidence comes in two layers with very different reliability.
In-process criterion benches that drive each engine tier directly, no NATS, no Postgres, no cluster. Seeded payloads, fixed workload sizes, isolated temp dirs. This is the reliable signal: same box, same seed, same numbers, run-to-run. It isolates the engine's own cost (append + fsync, replay, fold, cosine) from everything around it.
Entry point:
cargo bench -p ehdb-reference --bench engine_micro
# a single tier:
cargo bench -p ehdb-reference --bench engine_micro -- eventlog
cargo bench -p ehdb-reference --bench engine_micro -- vectorDrive real traffic through the worker's live emit_event / KV / object
paths in the local kind cluster with the NOETL_EHDB_* shadow/primary
flags armed, then read the noetl_ehdb_* runtime metrics (append counts,
outcome labels, last_ok, degraded gauges) and end-to-end latency. This
is where the EHDB-vs-incumbent head-to-head lives: the same drive
workload run with the event-log tier served by (a) the incumbent
Postgres+JetStream path and (b) EHDB durable-primary, comparing sustained
event throughput and replay/recovery time.
Layer B is directional, not authoritative. The podman-machine VM that backs kind on developer hardware is resource-constrained and flaky under sustained load (CPU-starved builds, swap pressure, occasional pod evictions — see the kind-validation history). It is good enough to show shape (EHDB keeps up / pulls ahead / falls behind) and to exercise the real metric wiring, but the absolute numbers carry a wide error bar. When Layer A and Layer B disagree, Layer A wins for engine-cost claims; Layer B is trusted for integration behavior (does the shadow mirror fire, do metrics advance, does recovery complete) rather than for precise throughput.
The head-to-head against a production-representative Postgres+JetStream (managed Cloud SQL sizing, real network) is a GKE-side exercise and is out of scope until the platform owner scopes it — this page designs it but does not run it.
What each tier's benches report, and why it matters.
| Tier | Metrics | Why |
|---|---|---|
| Event-log | append throughput (ev/s), single-append latency incl. fsync, durable-vs-local delta, segment-rotation overhead, cold-replay speed (ev/s) |
The headline. Append latency is fsync-bound and sets the single-writer-per-shard ceiling; replay speed bounds pod-restart recovery time. |
| Projection | fold/materialize rate (ev/s), checkpoint-advance/read cost | Read-model build rate off the event-log tail — the PostgreSQL-materializer replacement. |
| KV | put / get / bucket-scan / CAS-check rates | Platform NATS-KV state tier (circuit records, spool state). |
| Object | content-addressed put (small + large blob) throughput (MiB/s), get-verify, locate, list | Platform artifact / result-tier externalization. Large-blob put is bytes-moved-per-sec; small ops are registry-bound. |
| Vector | upsert latency, cosine top-k query latency at a few catalog sizes | Platform RAG / catalog-embedding tier. |
Across tiers the same three shapes recur: throughput (ops or events per second on a sustained workload), latency (p50/p95/p99 per single op — criterion reports the confidence interval; the median is the p50 proxy and the high bound approximates p95/p99 at this sample size), and cost-vs-size (how a per-op cost scales as the store grows — the property that separates a bounded-index engine from a replay-per-op reference driver).
Each LocalReference* driver is the correctness reference
implementation of its driver contract, not a production-serving engine.
Every operation reopens the pod-local JSONL log and replays it in full
to rebuild in-memory state, so every op is O(n) in the current log size
and an N-op sequence is O(N²). That is correct for shadow mode (a
derived, disposable mirror the incumbent still fronts) but it means the
reference throughput numbers degrade as the store grows — a property
these benches surface deliberately.
The durable_segment event-log backend (DurableEventLogDriver) is the
opposite: a bounded in-memory offset index over CRC-framed append
segments, fsync-per-append, O(1) amortized append. Only the
event-log tier has a durable backend today. The KV / object / vector
tiers are still reference-only — their numbers below are shadow-mode
figures, not primary-serve figures.
Hardware / environment (a dev box, not a benchmark rig): Apple M1 Max,
10 cores, 32 GiB RAM, macOS 26.3.1, APFS SSD, rustc 1.92.0, criterion
0.5, sample_size = 10. Benches ran on the macOS host directly (not the
podman VM), so fsync lands on the local APFS SSD. Treat these as an
order-of-magnitude baseline + a regression tripwire, not a datacenter
projection — a Cloud SQL / PVC-backed prod path will differ (usually
slower fsync, faster aggregate via sharding).
| Benchmark | local_reference | durable_segment | Read |
|---|---|---|---|
| sustained append, K=200 (fresh store) | 143 ev/s | 260 ev/s | durable 1.8× |
| sustained append, K=1000 (fresh store) | 96 ev/s | 255 ev/s | durable 2.7× |
| single append @ store size 100 | 9.3 ms | — | local rises with size |
| single append @ store size 1000 | 15.9 ms | 3.99 ms | durable 4× faster |
| single append @ store size 5000 | — | 3.87 ms | durable flat |
| segment rotation, K=1000 (8 MiB seg) | — | 257 ev/s | baseline |
| segment rotation, K=1000 (16 KiB seg, ~many rotations) | — | 251 ev/s | ~2% overhead |
| cold replay open, 5000 events | — | 27.1 ms | ≈185 K ev/s |
Reading the headline:
-
Durable is faster and safer at sustained append. The naive
expectation ("durable pays
fsync, so it's slower") is wrong here: thelocal_referencedriver re-reads and replays the whole JSONL twice per append, so it'sO(n)per op and collapses as the log grows (143 → 96 ev/s from K=200 → K=1000, and single-append 9.3 → 15.9 ms from size 100 → 1000). Durable holds a bounded index and stays flat (~3.9 ms/append at both size 1000 and 5000). -
Durable append is
fsync-bound at ~3.9 ms → ~256 appends/s on this SSD. Sustained (255 ev/s) and single-op (3.87–3.99 ms) agree, confirming the ceiling is the per-appendsync_data(), not CPU. This is the single-writer-per-shard limit; aggregate scales with shard count, and group-commit (batchedfsync) is the obvious lever if a single shard must exceed ~256/s. - Segment rotation is cheap — ~2% throughput cost even when forced every ~40 events (16 KiB segments) vs the default 8 MiB. Rollover is not a scaling concern.
- Cold replay is fast — ~185 K events/s to scan segments, verify CRCs, and rebuild the offset index. A 100 K-event shard recovers in ~0.5 s; this bounds pod-restart recovery time.
| Benchmark | Result |
|---|---|
| fold/materialize, 500-event batch | 1.48 K ev/s |
| fold/materialize, 2000-event batch | 1.50 K ev/s |
| checkpoint read @ 1000 applied | 4.2 ms |
Fold rate is flat per-event (~1.48 K ev/s) — one JSONL reopen amortized over the batch. This is the reference driver; a native projection engine target should be higher.
| Benchmark | Result |
|---|---|
| put sustained, K=200 | 167 put/s |
| put sustained, K=1000 | 120 put/s (degrades with size) |
| get | 5.2 ms (~192/s) |
| bucket scan (prefix, limit 128) | 5.9 ms (~168/s) |
| CAS-check (create-only conflict) | 5.2 ms (~191/s) |
All KV ops are reopen-bound at ~5 ms at size 1000; put degrades with size
(the O(n²) reference property).
| Benchmark | Result |
|---|---|
| put small (4 KiB) | 5.9 ms (~680 KiB/s) |
| put large (1 MiB) | 13.7 ms (≈73 MiB/s) |
| get-verify small @ 200 | 1.1 ms (~900/s) |
| locate @ 200 | 1.1 ms (~920/s) |
| list prefix @ 200 | 1.2 ms (~857/s) |
Large-blob put moves bytes at ~73 MiB/s (blob write + SHA-256 digest
bound) — the real throughput story for the result-tier externalization
path. get-verify reads the blob back and re-hashes it (integrity
check), so it is heavier than a raw read would be, by design.
| Benchmark | Result |
|---|---|
| upsert @ catalog 256 | 143 ms (~7/s) |
| cosine top-k=10 @ catalog 128 | 28.9 ms (~35/s) |
| cosine top-k=10 @ catalog 512 | 115 ms (~8.7/s) |
The slowest tier per op, and the clearest "measured slower than expected" finding — see below.
-
The vector reference driver is not a serving index. At dim=1536 it
re-parses the entire float-array JSONL on every op, so cost is
O(catalog × dim): a top-k query at catalog 512 already costs 115 ms, and query time scales ~linearly with catalog size (28.9 ms → 115 ms from 128 → 512). Fine as a shadow correctness fixture; not viable for primary serve. A real vector index (HNSW / IVF, memory-resident) is future work — this is the tier furthest from a production backend. -
The reference
O(n²)sustained cost is real and large — it is the entire reason the durable event-log backend exists, and these benches quantify it (durable 2.7× faster at K=1000 and the gap widens with N). The KV / object / vector tiers still carry this cost because they have no durable backend yet. -
Durable append is
fsync-ceilinged at ~256/s per writer. Not slow for a single writer, but the per-appendsync_data()is the throughput wall. If per-shard throughput must exceed a few hundred/s, group-commit (batch multiple appends behind onefsync) is the design lever — worth a follow-up if Layer B shows a single hot shard saturating.
A starting point, grounded in the baseline, framed per shard (the
durable event-log is single-writer-per-shard; aggregate = per-shard ×
shard_count). These are proposals to react to, not commitments.
| Tier / metric | Proposed target | Baseline (this box) |
|---|---|---|
| Event-log durable append, p99 latency | ≤ 10 ms | ~4 ms |
| Event-log durable append, sustained | ≥ 200 appends/s/shard | 255/s |
| Event-log cold replay | ≥ 100 K events/s | 185 K/s |
| Segment-rotation overhead | ≤ 5% of append throughput | ~2% |
| Projection fold | ≥ 1 K events/s/consumer | 1.48 K/s |
| KV / object / vector primary-serve | deferred — no durable backend yet | n/a |
Open questions for the platform owner:
- Are the event-log targets right for the prod storage medium (PVC vs
Cloud SQL vs EHDB object-tier shared segments)?
fsynclatency there will differ from local APFS. - Should projection / KV / object / vector get primary-serve SLOs now (implying native backends are on the roadmap), or stay shadow-only with the incumbent authoritative?
- Is per-shard the right SLO unit, and what shard count does prod target?
- Does Layer B need a production-representative Postgres+JetStream head-to-head before Stage-C cutover, or is the Layer-A engine delta + in-kind directional evidence sufficient?
-
Determinism. Payloads are generated from a fixed-seed xorshift64*
PRNG defined in the bench file (no
randdependency, no wall-clock, noMath.random), so a given box produces stable numbers run-to-run. -
Isolation. Every bench allocates a unique temp dir under a single
bench root (
$TMPDIR/ehdb-engine-micro-bench) and cleans it up; nothing leaks between tiers or persists after the run. -
Workload shapes. Realistic platform payloads: ~400-byte event
envelopes, circuit-state KV entries, 4 KiB / 1 MiB blobs, 1536-dim
embedding vectors. Sizes are
consts at the top of each tier's bench so they're easy to scale. -
Two measurement patterns. Sustained-batch (fresh store, append K,
report ev/s via
Throughput::Elements) exposes the growing-log cost; fixed-size single-op (pre-load size S once, measure one op) gives a clean per-op latency at a known store size. Reads use the fixed-size pattern (non-mutating); writes use whichever the metric needs. -
Reporting. criterion emits the median + 95% CI per benchmark and
writes HTML reports under
target/criterion/. Capture the consoletime:/thrpt:lines into the session log; paste the baseline table (above) into this page as the first data point and diff future runs against it. - No engine changes for measurement. All helpers are bench-local; no tier's semantics were altered to make a number look good. Where a tier measured slow, it's reported slow.
The full suite takes several minutes on the M1 Max, dominated by the
O(n²) sustained-append reference benches and the fsync-per-append
durable benches. sample_size is pinned to 10 to keep it tractable; the
vector tier uses a fixed-catalog single-op probe (rather than a sustained
loop) because a sustained O(n²) loop over 1536-float records needed
~220 s for 10 samples. Run a single tier (-- eventlog) for a fast
inner-loop check.
Ran in the local kind cluster (kind-noetl, podman) on the same M1 Max
dev box. Server v3.54.0-rc4 / worker v5.71.0-rc4, EHDB shadow,
NOETL_EHDB_EVENTLOG_BACKEND=durable_segment, durable dir on a PVC
(ehdb-durable-soak). Harness:
scripts/perf/layer_b_eventlog_load.sh (#263).
Read every number below with the VM caveat. The podman-machine VM is
CPU-constrained and its PVC fsync path (kind hostpath → qcow2 → macOS
file) is slow and noisy. These validate shape and behaviour, not
authoritative peak throughput.
Incumbent = the Postgres + NATS-JetStream log-and-store path EHDB's
event-log tier replaces, measured at the server POST /api/events write
boundary (step.enter, a pure DB+publish op) under ApacheBench; p50/p95/p99
from the per-run histogram delta:
| Path | throughput | p50 | p95 | p99 | failures |
|---|---|---|---|---|---|
| Incumbent (Postgres+NATS), sustained c=50 | ~2 150 ev/s | 12.6 ms | 33.9 ms | 52.5 ms | 0 |
| Incumbent (Postgres+NATS), burst c=200 | ~1 600 ev/s | 85.7 ms | 232 ms | 300 ms | 0 |
| EHDB durable local engine primitive (fresh segment, in-VM) | single-writer | ~3–10 ms | — | — | — |
| EHDB durable_segment as-deployed shared-tier shadow mirror | ~1–2 append/s | ~740 ms | ~975 ms | ~1.17 s | 0 |
The incumbent comfortably clears its own Phase-B target (~1 k ev/s, p99 < 20 ms was the old kind goal; here it sustains ~2 150 ev/s at p99 52 ms, degrading gracefully to p99 300 ms at 4× concurrency with zero errors).
durable_segment in the worker runtime is always the composed slice-3
shared tier (SharedTierEventLog), never the local-only DurableEventLogDriver
that Layer A benched in isolation. SharedTierEventLog::append calls
publish_shard synchronously, and publish_shard re-reads + re-writes
the entire active segment to the shared store (with sync_all) on every
append — an O(active-segment-size) cost that climbs as the segment fills
toward the 8 MiB rotation. The in-VM ehdb-selfcheck durable-eventlog
decomposition proves it:
| Append target (same VM, same binary) | latency |
|---|---|
| Fresh empty segment (local durable primitive) | ~3–10 ms |
| Live ~20 MB active segment (whole-segment re-publish) | ~0.55–1.7 s |
So the ~0.5–1.2 s the shadow mirror shows under load is not the durable
engine and not just the slow VM fsync — it is dominated by the
whole-active-segment re-publish. The fresh-segment number (~3–10 ms) lands
right on the Layer A engine baseline (3.9 ms), corroborating that the local
durable append is cheap and flat in-cluster too.
Worse, the mirror is called synchronously inside emit_event
(ControlPlaneClient::emit_event → eventlog::mirror_live_event; errors are
isolated, latency is not). So the shared-tier append cost is added to every
worker event emission, throttling emission to ~1–2 events/s and building
NOETL_COMMANDS backlog under a burst of drives (observed peak pending 11–22
that did not drain within the cap).
Yes, on the load-bearing claim; the specific 2.7× ratio is a Layer-A-only
measurement by construction. The 2.7× was durable-local vs
local_reference JSONL — neither of which the worker runs as a serving path
(it runs the shared-tier composition). What Layer B can corroborate is the
engine floor: the local durable append primitive is ~3–10 ms in-cluster,
matching Layer A's 3.9 ms. The engine is not the bottleneck — the reference
driver's O(n²) cost (the reason durable exists) and the durable engine's
flat append both hold. Layer B's new, independent finding is that the
shared-tier publish strategy re-introduces an O(segment) cost on the
hot path.
- The EHDB local durable append (~3–10 ms, a local
fsync, no Postgres/NATS/network hop) is competitive with — often faster than — the incumbent's p50 (12.6 ms). That is the promise the durable tier is built on. - The as-deployed shared-tier shadow (~0.7–1.2 s, throughput-throttling) is far slower than the incumbent. This gap is the shared-publish algorithm, not the engine — and it must be fixed before durable-shared can be a primary path. ⇒ SUPERSEDED (2026-07-11): #264/#266 (O(delta) publish) + #267 (O(1) open) fixed exactly this; the deployed durable append is now ~6 ms (see the #261 re-run).
| Strawman target | Layer A (engine) | Layer B (in-cluster) | Verdict |
|---|---|---|---|
| durable append p99 ≤ 10 ms | ~4 ms ✅ | local primitive ~4 ms ✅ / shared-tier publish flat ~12–13 ms after #266 ✅ / deployed per-op ~4–16 ms ✅ after #267 (was ~0.5 s) | Split by cost, not just backend. Both O(segment) costs are now removed: #266 (shared publish, O(delta)) + #267 (local replay-on-open, O(1) checkpoint open). Deployed per-op dropped from ~0.5 s to ~4–16 ms — see the #267 section below. |
| sustained ≥ 200 append/s/shard | 255/s ✅ | local ~256/s ✅ / deployed ~4–16 ms/op (~60–250/s engine primitive) after #267 ✅ (was ~1.3/s) | The per-op stack reconstruction + full replay was the limiter; #267 made open O(1). Sustained end-to-end rate now bounded by the emit path + fsync, not the replay. |
| cold replay ≥ 100 K ev/s | 185 K/s ✅ | not re-measured in-cluster | Layer A stands. |
| (incumbent reference, for context) | — | p99 52 ms sustained / 300 ms burst | An incumbent-parity append SLO would be ~50–100 ms p99. The engine + the fixed publish clear it; the deployed per-op path won't until #267 lands. |
SLO that needs revisiting: the "durable append" latency SLO must name which cost — engine append, shared publish, or the deployed per-op path (which additionally pays local replay-on-open). #264 (#265+#266) fixed the shared publish; the deployed path stays shadow-only until #267 removes the local replay-on-open.
The Phase-1 conditional ("group-commit if a hot shard saturates the fsync
ceiling") did not fire as the priority. The in-cluster limiter was not
the local-append fsync ceiling (~256/s) — it was the shared-tier
whole-segment re-publish. So group-commit is a secondary lever; the
priority follow-up is incremental shared publish (below).
The incremental-shared-publish fix shipped (ehdb #265 → #266), and validating it in kind corrected the Layer-B headline: the shared-tier publish was one of two O(segment) costs on the deployed per-op path, and it is not the dominant one.
-
#265 —
publish_shardsends the backend only the append-delta[published_len .. current_len](not the whole active segment) via a newSharedSegmentBackend::append_segment;FilesystemSharedBackendappends in place and re-commits the integrity marker. -
#266 — the worker builds the shared-tier stack per op (a stateless
boundary), so #265's in-memory running-digest cache was empty every append and
each publish re-seeded the digest by re-reading the whole committed prefix —
re-introducing an O(committed) read. #266 persists the resumable
XxHash64state on the segment's integrity sidecar (twox-hashserializefeature), so a freshly-constructed backend resumes the digest and folds in only the delta. Guarantees unchanged (cold-load correctness, digest keys, ordering, crash-safety); the whole-prefixdigeststring is identical. -
Engine proof (Layer A micro-bench
eventlog_shared_tier/append_at_size, host, two fsyncs/append): append is flat ~12–13 ms at pre-warmed active segment sizes 100 / 2 000 / 10 000 — no scaling with segment size. Unit tests cover the O(delta) guard, per-op-construction resume, crash-safety, and legacy reseed fallback (227ehdb-referencetests pass).
The worker rebuilds the durable stack per op (append_selected →
build_durable_stack). That opens the local DurableSegmentStore, and
DurableSegmentStore::open replays every segment to rebuild the offset index
— an O(segment) cost that runs on every mirrored event, independent of the
shared publish. Because both costs scale with segment size, the earlier
selfcheck decompose could not separate them and attributed the whole ~0.5–1.2 s
to the shared re-publish. Deploying #266 (which makes the shared publish
provably O(delta)) isolated the truth:
Probe (in-kind, ehdb-selfcheck durable-eventlog, warm cache, settled VM) |
per-op latency |
|---|---|
| Fresh empty local dir (no replay) | ~4 ms |
| Live 8 MB local (replay) + shared O(delta) | ~0.15–0.25 s |
| Live 16 MB local (replay) + shared O(delta) | ~0.6–0.7 s |
| Live 20 MB local (replay) + shared O(delta) | ~0.5–0.75 s |
The cost scales with local segment size while the shared publish is O(delta) — so the residual is the local replay-on-open, not the shared publish.
| before (as recorded) | after #266 | |
|---|---|---|
| as-deployed mirror append rate | ~1–2 append/s | ~1.3 append/s |
| mirror-append latency p50 / p95 / p99 | ~740 ms / ~975 ms / ~1.17 s | ~650 ms / ~1.10 s / ~1.65 s |
The deployed append rate did not improve — as expected once the bottleneck is understood: the shared-publish fix removed a real O(segment) cost (confirmed by the flat micro-bench and the O(delta) isolation), but the deployed per-op latency is gated by the local replay-on-open, tracked in #267. Both O(segment) costs must be removed for the deployed durable path to approach the incumbent; #264 removed the shared-publish one, #267 is the remaining (dominant) one.
Per the perf methodology (Layer A engine micro-benchmarks are the reliable signal; Layer B in-cluster is directional on the build-contended podman VM), the shared-publish fix was merged on the micro-bench evidence: the flat append-vs-segment-size proof (~12–13 ms at 100 / 2 000 / 10 000) shows the shared-tier publish SLO is now met at the engine level (O(delta), no scaling with segment size). Merging was not gated on the slow in-kind rebuild.
Deferred follow-up (not blocking): re-run Layer B in kind to confirm the as-deployed append-rate / p99 improvement once (a) #267 (local replay-on-open) lands — it is the dominant deployed cost, so the deployed rate will not move until it is fixed — and (b) the podman VM isn't build-contended (add swap, or build the worker image off-VM first). The in-kind run done this session already isolated #267 as the gate; the next run is the post-#267 confirmation.
The dominant deployed cost above — the worker rebuilding the durable stack per
op, so every mirrored append paid a fresh DurableSegmentStore::open that
replayed every segment to rebuild the offset index (O(segment)) — is fixed
in ehdb#268 (merged f6fdaea, closes
#267).
The fix. A small per-store checkpoint.json sidecar —
{event_count, active_segment_id, active_len, consumers_seen, consumer_acks} —
rewritten after each mutating op strictly after the frame fsync.
Open-for-append loads it O(1) and skips the replay; the offset index is
materialised lazily on the first read (ensure_index_loaded). It mirrors
the #266 resumable-digest sidecar. Durability is preserved — the checkpoint is a
cache, not truth:
-
Replay-is-truth stays authoritative. A missing / stale / inconsistent
checkpoint → full replay (and the checkpoint is rewritten). The strict
active-segment length anchor means the checkpoint can never name more
durable data than the segments hold, so a crash between the frame
fsyncand the checkpoint rewrite is caught → replay recovers the extra fsync'd frame(s). - CRC integrity is enforced on every read (the lazy-index replay runs full CRC + gapless checks). The append path never reads event bodies, so no corrupt data is silently served; a corrupt frame is a hard error on first read (or at open without a valid checkpoint).
- Torn-tail recovery, fsync-before-ack, single-writer/affinity routing, and byte-exactly-once are unchanged.
The prior held-open append_at_size bench reused one open store, so it never
saw the O(segment) replay. New per-op-open benches reconstruct the driver / full
shared stack per op — the deployed build_durable_stack shape:
| S (pre-warmed events) |
durable_segment/per_op_open_append_at_size (driver) |
durable_segment_shared/per_op_open_append_at_size (full stack) |
|---|---|---|
| 100 | 5.11 ms | 13.9 ms |
| 2,000 | 5.21 ms | 14.2 ms |
| 10,000 | 5.35 ms | 14.5 ms |
Flat across a 100× growth in segment size → open-for-append is O(1). Before
the fix this curve rose with S. 232 ehdb-reference tests pass (+5 checkpoint
tests); clippy -D warnings + fmt clean.
ehdb-selfcheck durable-eventlog on the user-pool worker (noetl-worker-rust),
gauge noetl_ehdb_eventlog_last_duration_seconds:
| image | per-op durable mirror | |
|---|---|---|
| before | v5.72.0-ehdb266 |
~0.509 s (local replay-on-open) |
| after | v5.72.0-ehdb267 |
~0.004–0.016 s (steady ~4–10 ms) |
A ~30–120× per-op reduction on the same ~20 MB durable state — the deployed
gate #265/#266 could not move is gone. The first-ever op after the roll paid a
one-time legacy replay (no checkpoint yet) and wrote the sidecar
(checkpoint.json, 103 B, event_count:24777); every subsequent op is O(1).
hello_world end-to-end COMPLETED on the ehdb267 pods; both user + system pools
rolled clean (0 restarts). No GKE/prod; all NOETL_EHDB_* flags stay default
off; repos/noetl + repos/server untouched.
The SLO table above updates accordingly: the deployed durable append now clears an incumbent-parity p99 (~4–16 ms vs the incumbent's ~52 ms sustained), and sustained throughput is no longer capped by the per-op stack reconstruction.
Re-measured the deployed durable append on the shipped, merged-main image
localhost/noetl-worker:ehdb178184-merged (carries #264/#265/#266 shared-publish
O(delta) + #267 O(1) open + the GC/retention work), superseding the stale
2026-07-08 Layer-B numbers that predate all of those fixes. LOCAL kind only —
no GKE/prod; noetl.event never purged.
| Metric | Stale (2026-07-08, pre-#264/#266/#267) | RE-RUN (2026-07-11, merged main) |
|---|---|---|
| Deployed durable mirror-append (in-binary op duration, shared-tier stack) | ~740 ms p50 / ~1.17 s p99 (~1–2 append/s) | ~6–7 ms typical (6.82 / 6.25 ms; one VM-contention spike 34 ms) |
A ~100–200× reduction on the as-deployed path. The ~6 ms lands on the
Layer-A engine floor (below), confirming the deployed durable append is now the
cheap local fsync — the O(active-segment) whole-segment re-publish that
dominated the stale ~0.7–1.2 s is eliminated (#266), and the per-op O(segment)
replay-on-open is gone (#267). (durable-eventlog reported
durable_replay_records 9/12/15 = the read-only reopen replayed the retained
log; the parity_mismatch/ok:false are the R7 artifact — the verb asserts a
fresh isolated 3-event sequence against a live populated shard, not a
durability failure.)
The criterion micro-bench (cargo bench -p ehdb-reference --bench engine_micro)
is the trustworthy engine number. Post-#267 per-op-open+append is flat across
local-segment size (the #267 section):
| S (events) | driver per-op open+append | full shared-tier stack per-op |
|---|---|---|
| 100 | 5.11 ms | 13.9 ms |
| 2,000 | 5.21 ms | 14.2 ms |
| 10,000 | 5.35 ms | 14.5 ms |
Flat = the O(segment) costs are gone; append no longer degrades with store size.
Same-day incumbent leg (scripts/perf/layer_b_eventlog_load.sh incumbent-sustained, POST /api/events = Postgres INSERT + NATS publish, ab
n=3000 c=50): ~90–93 req/s, p50 163 ms, p95 986 ms, p99 3.6 s, 0 failures.
This is far slower than the #263 baseline (~2 150 ev/s, p50 12.6 ms, p99
52 ms) — not a regression (the incumbent path is unchanged code); it is
pure podman-VM contention: today's VM was running all three worker pools on
fresh images plus the retention-soak append load. This is exactly why Layer B is
directional, not authoritative. Read the incumbent's engine cost from the
less-contended #263 run (p50 ~12.6 ms); read EHDB's from Layer A (~4–16 ms).
-
The deployed EHDB durable append is competitive with the incumbent. EHDB's
~6 ms in-process local
fsync(no HTTP/Postgres/NATS hop) is at/below the incumbent's engine-cost p50 (~12.6 ms on a clean VM). On a contended VM the incumbent's full HTTP write-boundary inflates (p50 163 ms today) while EHDB's local append stays ~6 ms — but that is a different measurement boundary (in-process append vs network write), so the fair claim rests on Layer A + the deployed re-measurement, not the contended incumbent HTTP number. - The stale "as-deployed is 100× slower than the incumbent" finding is obsolete. That gap was the shared-publish O(segment) algorithm (#264), now fixed. As-deployed durable append went from ~740 ms to ~6 ms.
- Authoritative engine claim: EHDB durable per-op append is ~4–16 ms, flat across store size (Layer A), clearing an incumbent-parity p99 SLO.
-
Fix local replay-on-open (#267).Done (ehdb#268,f6fdaea) — see the section above. The remaining per-op cost is the fsync(s) + the O(delta) shared publish, both flat. -
Group-commit spike (secondary) — batch multiple local appends behind one
fsyncif a single hot shard must exceed the ~256/s local ceiling. - Scope confirmation from the platform owner on the per-backend SLO split and whether a production-representative (Cloud SQL + real network) incumbent head-to-head is required before Stage-C cutover.
- Design: Durable Event-Log Backend — the segment store + fsync + replay these benches measure.
-
Design: Event-Log Core Engine (Phase 6)
— the
EventLogDrivercontract both drivers implement. -
Backend Configuration — the
NOETL_EHDB_*surface Layer B arms. - Runbook: Durable Event-Log Prod Sign-off — the §C durability gate the SLOs feed.
- Roadmap · Home.
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)