Skip to content

Design Performance and Load Testing

Kadyapam edited this page Jul 12, 2026 · 6 revisions

Design: Performance & Load Testing

Status: Phase 1 + Layer B landed + shared-publish fix shipped & validated (2026-07-08). The engine micro-benchmarks are merged in ehdb-reference::benches::engine_micro (noetl/ehdb#261 / #262), with a committed baseline captured on a dev box (below). Layer B — the in-cluster load harness + the EHDB-vs-incumbent head-to-head — ran in kind (#263); results in the Layer B section below. The shared-tier publish O(segment) bottleneck it surfaced (#264) is fixed (#265 + #266 — now O(delta), flat micro-bench) and validated in kind, which corrected the headline: the dominant deployed cost is the separate local replay-on-open (#267), not the shared publish. See Fix landed + a corrected attribution. Layer B is directional (podman-VM); Layer A remains the authoritative engine signal.

Tracks: noetl/ehdb#261 (perf/load-test workstream) · noetl/ehdb#241 (completion program) · Design: Event-Log Core Engine (Phase 6) · Design: Durable Event-Log Backend (the durability work these benches quantify) · Roadmap.

Goals

Give EHDB a repeatable performance signal so its founding claim — fixing NoETL's event-sourcing log bottleneck — is measured, not asserted. Three concrete goals:

  1. Quantify the event-log bottleneck fix. Put a number on the durable event-log tier against the incumbent it replaces (the PostgreSQL + NATS JetStream log-and-store path). The headline question: how much throughput and tail-latency headroom does EHDB's durable event-log/replay path buy over the incumbent on the same append + replay workload?
  2. Catch regressions in-engine, cheaply. Deterministic Rust micro-benchmarks per tier that run in CI-adjacent time and fail loudly when a change moves append latency, replay speed, or fold rate.
  3. Set SLO targets the prod-cutover runbooks can gate on — per-tier throughput / latency budgets, confirmed by the platform owner rather than pulled from thin air.

The boundary this respects

Platform data only. Every payload in these benches is a synthetic, secret-free platform event / KV entry / object / vector — never business data. EHDB serves NoETL's internal platform tiers (event log, projections, platform KV/state, platform artifacts, catalog embeddings); tenant data stays in real business DBs behind playbook connectors. See RFC.

The two test layers

Performance evidence comes in two layers with very different reliability.

Layer A — engine micro-benchmarks (Rust, deterministic) — this phase

In-process criterion benches that drive each engine tier directly, no NATS, no Postgres, no cluster. Seeded payloads, fixed workload sizes, isolated temp dirs. This is the reliable signal: same box, same seed, same numbers, run-to-run. It isolates the engine's own cost (append + fsync, replay, fold, cosine) from everything around it.

Entry point:

cargo bench -p ehdb-reference --bench engine_micro
# a single tier:
cargo bench -p ehdb-reference --bench engine_micro -- eventlog
cargo bench -p ehdb-reference --bench engine_micro -- vector

Layer B — in-cluster end-to-end load (kind) — next phase, design only

Drive real traffic through the worker's live emit_event / KV / object paths in the local kind cluster with the NOETL_EHDB_* shadow/primary flags armed, then read the noetl_ehdb_* runtime metrics (append counts, outcome labels, last_ok, degraded gauges) and end-to-end latency. This is where the EHDB-vs-incumbent head-to-head lives: the same drive workload run with the event-log tier served by (a) the incumbent Postgres+JetStream path and (b) EHDB durable-primary, comparing sustained event throughput and replay/recovery time.

Layer B is directional, not authoritative. The podman-machine VM that backs kind on developer hardware is resource-constrained and flaky under sustained load (CPU-starved builds, swap pressure, occasional pod evictions — see the kind-validation history). It is good enough to show shape (EHDB keeps up / pulls ahead / falls behind) and to exercise the real metric wiring, but the absolute numbers carry a wide error bar. When Layer A and Layer B disagree, Layer A wins for engine-cost claims; Layer B is trusted for integration behavior (does the shadow mirror fire, do metrics advance, does recovery complete) rather than for precise throughput.

The head-to-head against a production-representative Postgres+JetStream (managed Cloud SQL sizing, real network) is a GKE-side exercise and is out of scope until the platform owner scopes it — this page designs it but does not run it.

Metrics per tier

What each tier's benches report, and why it matters.

Tier Metrics Why
Event-log append throughput (ev/s), single-append latency incl. fsync, durable-vs-local delta, segment-rotation overhead, cold-replay speed (ev/s) The headline. Append latency is fsync-bound and sets the single-writer-per-shard ceiling; replay speed bounds pod-restart recovery time.
Projection fold/materialize rate (ev/s), checkpoint-advance/read cost Read-model build rate off the event-log tail — the PostgreSQL-materializer replacement.
KV put / get / bucket-scan / CAS-check rates Platform NATS-KV state tier (circuit records, spool state).
Object content-addressed put (small + large blob) throughput (MiB/s), get-verify, locate, list Platform artifact / result-tier externalization. Large-blob put is bytes-moved-per-sec; small ops are registry-bound.
Vector upsert latency, cosine top-k query latency at a few catalog sizes Platform RAG / catalog-embedding tier.

Across tiers the same three shapes recur: throughput (ops or events per second on a sustained workload), latency (p50/p95/p99 per single op — criterion reports the confidence interval; the median is the p50 proxy and the high bound approximates p95/p99 at this sample size), and cost-vs-size (how a per-op cost scales as the store grows — the property that separates a bounded-index engine from a replay-per-op reference driver).

A caveat that shapes every number: the reference drivers replay per op

Each LocalReference* driver is the correctness reference implementation of its driver contract, not a production-serving engine. Every operation reopens the pod-local JSONL log and replays it in full to rebuild in-memory state, so every op is O(n) in the current log size and an N-op sequence is O(N²). That is correct for shadow mode (a derived, disposable mirror the incumbent still fronts) but it means the reference throughput numbers degrade as the store grows — a property these benches surface deliberately.

The durable_segment event-log backend (DurableEventLogDriver) is the opposite: a bounded in-memory offset index over CRC-framed append segments, fsync-per-append, O(1) amortized append. Only the event-log tier has a durable backend today. The KV / object / vector tiers are still reference-only — their numbers below are shadow-mode figures, not primary-serve figures.

Baseline — first data point (2026-07-08)

Hardware / environment (a dev box, not a benchmark rig): Apple M1 Max, 10 cores, 32 GiB RAM, macOS 26.3.1, APFS SSD, rustc 1.92.0, criterion 0.5, sample_size = 10. Benches ran on the macOS host directly (not the podman VM), so fsync lands on the local APFS SSD. Treat these as an order-of-magnitude baseline + a regression tripwire, not a datacenter projection — a Cloud SQL / PVC-backed prod path will differ (usually slower fsync, faster aggregate via sharding).

Event-log — the headline (local_reference JSONL vs durable_segment)

Benchmark local_reference durable_segment Read
sustained append, K=200 (fresh store) 143 ev/s 260 ev/s durable 1.8×
sustained append, K=1000 (fresh store) 96 ev/s 255 ev/s durable 2.7×
single append @ store size 100 9.3 ms — local rises with size
single append @ store size 1000 15.9 ms 3.99 ms durable 4× faster
single append @ store size 5000 — 3.87 ms durable flat
segment rotation, K=1000 (8 MiB seg) — 257 ev/s baseline
segment rotation, K=1000 (16 KiB seg, ~many rotations) — 251 ev/s ~2% overhead
cold replay open, 5000 events — 27.1 ms ≈185 K ev/s

Reading the headline:

  • Durable is faster and safer at sustained append. The naive expectation ("durable pays fsync, so it's slower") is wrong here: the local_reference driver re-reads and replays the whole JSONL twice per append, so it's O(n) per op and collapses as the log grows (143 → 96 ev/s from K=200 → K=1000, and single-append 9.3 → 15.9 ms from size 100 → 1000). Durable holds a bounded index and stays flat (~3.9 ms/append at both size 1000 and 5000).
  • Durable append is fsync-bound at ~3.9 ms → ~256 appends/s on this SSD. Sustained (255 ev/s) and single-op (3.87–3.99 ms) agree, confirming the ceiling is the per-append sync_data(), not CPU. This is the single-writer-per-shard limit; aggregate scales with shard count, and group-commit (batched fsync) is the obvious lever if a single shard must exceed ~256/s.
  • Segment rotation is cheap — ~2% throughput cost even when forced every ~40 events (16 KiB segments) vs the default 8 MiB. Rollover is not a scaling concern.
  • Cold replay is fast — ~185 K events/s to scan segments, verify CRCs, and rebuild the offset index. A 100 K-event shard recovers in ~0.5 s; this bounds pod-restart recovery time.

Projection

Benchmark Result
fold/materialize, 500-event batch 1.48 K ev/s
fold/materialize, 2000-event batch 1.50 K ev/s
checkpoint read @ 1000 applied 4.2 ms

Fold rate is flat per-event (~1.48 K ev/s) — one JSONL reopen amortized over the batch. This is the reference driver; a native projection engine target should be higher.

KV (reference driver, at store size 1000)

Benchmark Result
put sustained, K=200 167 put/s
put sustained, K=1000 120 put/s (degrades with size)
get 5.2 ms (~192/s)
bucket scan (prefix, limit 128) 5.9 ms (~168/s)
CAS-check (create-only conflict) 5.2 ms (~191/s)

All KV ops are reopen-bound at ~5 ms at size 1000; put degrades with size (the O(n²) reference property).

Object (reference driver)

Benchmark Result
put small (4 KiB) 5.9 ms (~680 KiB/s)
put large (1 MiB) 13.7 ms (≈73 MiB/s)
get-verify small @ 200 1.1 ms (~900/s)
locate @ 200 1.1 ms (~920/s)
list prefix @ 200 1.2 ms (~857/s)

Large-blob put moves bytes at ~73 MiB/s (blob write + SHA-256 digest bound) — the real throughput story for the result-tier externalization path. get-verify reads the blob back and re-hashes it (integrity check), so it is heavier than a raw read would be, by design.

Vector (reference driver, dim=1536, model text-embedding-3-small)

Benchmark Result
upsert @ catalog 256 143 ms (~7/s)
cosine top-k=10 @ catalog 128 28.9 ms (~35/s)
cosine top-k=10 @ catalog 512 115 ms (~8.7/s)

The slowest tier per op, and the clearest "measured slower than expected" finding — see below.

Honest findings — what measured slower than expected

  • The vector reference driver is not a serving index. At dim=1536 it re-parses the entire float-array JSONL on every op, so cost is O(catalog × dim): a top-k query at catalog 512 already costs 115 ms, and query time scales ~linearly with catalog size (28.9 ms → 115 ms from 128 → 512). Fine as a shadow correctness fixture; not viable for primary serve. A real vector index (HNSW / IVF, memory-resident) is future work — this is the tier furthest from a production backend.
  • The reference O(n²) sustained cost is real and large — it is the entire reason the durable event-log backend exists, and these benches quantify it (durable 2.7× faster at K=1000 and the gap widens with N). The KV / object / vector tiers still carry this cost because they have no durable backend yet.
  • Durable append is fsync-ceilinged at ~256/s per writer. Not slow for a single writer, but the per-append sync_data() is the throughput wall. If per-shard throughput must exceed a few hundred/s, group-commit (batch multiple appends behind one fsync) is the design lever — worth a follow-up if Layer B shows a single hot shard saturating.

Proposed SLO strawman — for @alesha to confirm (NOT final)

A starting point, grounded in the baseline, framed per shard (the durable event-log is single-writer-per-shard; aggregate = per-shard × shard_count). These are proposals to react to, not commitments.

Tier / metric Proposed target Baseline (this box)
Event-log durable append, p99 latency ≤ 10 ms ~4 ms
Event-log durable append, sustained ≥ 200 appends/s/shard 255/s
Event-log cold replay ≥ 100 K events/s 185 K/s
Segment-rotation overhead ≤ 5% of append throughput ~2%
Projection fold ≥ 1 K events/s/consumer 1.48 K/s
KV / object / vector primary-serve deferred — no durable backend yet n/a

Open questions for the platform owner:

  1. Are the event-log targets right for the prod storage medium (PVC vs Cloud SQL vs EHDB object-tier shared segments)? fsync latency there will differ from local APFS.
  2. Should projection / KV / object / vector get primary-serve SLOs now (implying native backends are on the roadmap), or stay shadow-only with the incumbent authoritative?
  3. Is per-shard the right SLO unit, and what shard count does prod target?
  4. Does Layer B need a production-representative Postgres+JetStream head-to-head before Stage-C cutover, or is the Layer-A engine delta + in-kind directional evidence sufficient?

Harness & reproducibility

  • Determinism. Payloads are generated from a fixed-seed xorshift64* PRNG defined in the bench file (no rand dependency, no wall-clock, no Math.random), so a given box produces stable numbers run-to-run.
  • Isolation. Every bench allocates a unique temp dir under a single bench root ($TMPDIR/ehdb-engine-micro-bench) and cleans it up; nothing leaks between tiers or persists after the run.
  • Workload shapes. Realistic platform payloads: ~400-byte event envelopes, circuit-state KV entries, 4 KiB / 1 MiB blobs, 1536-dim embedding vectors. Sizes are consts at the top of each tier's bench so they're easy to scale.
  • Two measurement patterns. Sustained-batch (fresh store, append K, report ev/s via Throughput::Elements) exposes the growing-log cost; fixed-size single-op (pre-load size S once, measure one op) gives a clean per-op latency at a known store size. Reads use the fixed-size pattern (non-mutating); writes use whichever the metric needs.
  • Reporting. criterion emits the median + 95% CI per benchmark and writes HTML reports under target/criterion/. Capture the console time: / thrpt: lines into the session log; paste the baseline table (above) into this page as the first data point and diff future runs against it.
  • No engine changes for measurement. All helpers are bench-local; no tier's semantics were altered to make a number look good. Where a tier measured slow, it's reported slow.

Cost / runtime note

The full suite takes several minutes on the M1 Max, dominated by the O(n²) sustained-append reference benches and the fsync-per-append durable benches. sample_size is pinned to 10 to keep it tractable; the vector tier uses a fixed-catalog single-op probe (rather than a sustained loop) because a sustained O(n²) loop over 1536-float records needed ~220 s for 10 samples. Run a single tier (-- eventlog) for a fast inner-loop check.

Layer B — in-cluster results (2026-07-08)

Ran in the local kind cluster (kind-noetl, podman) on the same M1 Max dev box. Server v3.54.0-rc4 / worker v5.71.0-rc4, EHDB shadow, NOETL_EHDB_EVENTLOG_BACKEND=durable_segment, durable dir on a PVC (ehdb-durable-soak). Harness: scripts/perf/layer_b_eventlog_load.sh (#263).

Read every number below with the VM caveat. The podman-machine VM is CPU-constrained and its PVC fsync path (kind hostpath → qcow2 → macOS file) is slow and noisy. These validate shape and behaviour, not authoritative peak throughput.

Head-to-head — EHDB durable_segment vs the incumbent (same VM, same ~400 B event envelope)

Incumbent = the Postgres + NATS-JetStream log-and-store path EHDB's event-log tier replaces, measured at the server POST /api/events write boundary (step.enter, a pure DB+publish op) under ApacheBench; p50/p95/p99 from the per-run histogram delta:

Path throughput p50 p95 p99 failures
Incumbent (Postgres+NATS), sustained c=50 ~2 150 ev/s 12.6 ms 33.9 ms 52.5 ms 0
Incumbent (Postgres+NATS), burst c=200 ~1 600 ev/s 85.7 ms 232 ms 300 ms 0
EHDB durable local engine primitive (fresh segment, in-VM) single-writer ~3–10 ms — — —
EHDB durable_segment as-deployed shared-tier shadow mirror ~1–2 append/s ~740 ms ~975 ms ~1.17 s 0

The incumbent comfortably clears its own Phase-B target (~1 k ev/s, p99 < 20 ms was the old kind goal; here it sustains ~2 150 ev/s at p99 52 ms, degrading gracefully to p99 300 ms at 4× concurrency with zero errors).

The decisive decomposition — the shared-tier publish is the bottleneck, not the engine

durable_segment in the worker runtime is always the composed slice-3 shared tier (SharedTierEventLog), never the local-only DurableEventLogDriver that Layer A benched in isolation. SharedTierEventLog::append calls publish_shard synchronously, and publish_shard re-reads + re-writes the entire active segment to the shared store (with sync_all) on every append — an O(active-segment-size) cost that climbs as the segment fills toward the 8 MiB rotation. The in-VM ehdb-selfcheck durable-eventlog decomposition proves it:

Append target (same VM, same binary) latency
Fresh empty segment (local durable primitive) ~3–10 ms
Live ~20 MB active segment (whole-segment re-publish) ~0.55–1.7 s

So the ~0.5–1.2 s the shadow mirror shows under load is not the durable engine and not just the slow VM fsync — it is dominated by the whole-active-segment re-publish. The fresh-segment number (~3–10 ms) lands right on the Layer A engine baseline (3.9 ms), corroborating that the local durable append is cheap and flat in-cluster too.

Worse, the mirror is called synchronously inside emit_event (ControlPlaneClient::emit_event → eventlog::mirror_live_event; errors are isolated, latency is not). So the shared-tier append cost is added to every worker event emission, throttling emission to ~1–2 events/s and building NOETL_COMMANDS backlog under a burst of drives (observed peak pending 11–22 that did not drain within the cap).

Does it corroborate the Layer A 2.7× durable-vs-local finding?

Yes, on the load-bearing claim; the specific 2.7× ratio is a Layer-A-only measurement by construction. The 2.7× was durable-local vs local_reference JSONL — neither of which the worker runs as a serving path (it runs the shared-tier composition). What Layer B can corroborate is the engine floor: the local durable append primitive is ~3–10 ms in-cluster, matching Layer A's 3.9 ms. The engine is not the bottleneck — the reference driver's O(n²) cost (the reason durable exists) and the durable engine's flat append both hold. Layer B's new, independent finding is that the shared-tier publish strategy re-introduces an O(segment) cost on the hot path.

Engine-to-engine, EHDB is competitive with the incumbent; as-deployed, it is not

  • The EHDB local durable append (~3–10 ms, a local fsync, no Postgres/NATS/network hop) is competitive with — often faster than — the incumbent's p50 (12.6 ms). That is the promise the durable tier is built on.
  • The as-deployed shared-tier shadow (~0.7–1.2 s, throughput-throttling) is far slower than the incumbent. This gap is the shared-publish algorithm, not the engine — and it must be fixed before durable-shared can be a primary path. ⇒ SUPERSEDED (2026-07-11): #264/#266 (O(delta) publish) + #267 (O(1) open) fixed exactly this; the deployed durable append is now ~6 ms (see the #261 re-run).

SLO strawman — checked against Layer B

Strawman target Layer A (engine) Layer B (in-cluster) Verdict
durable append p99 ≤ 10 ms ~4 ms ✅ local primitive ~4 ms ✅ / shared-tier publish flat ~12–13 ms after #266 ✅ / deployed per-op ~4–16 ms ✅ after #267 (was ~0.5 s) Split by cost, not just backend. Both O(segment) costs are now removed: #266 (shared publish, O(delta)) + #267 (local replay-on-open, O(1) checkpoint open). Deployed per-op dropped from ~0.5 s to ~4–16 ms — see the #267 section below.
sustained ≥ 200 append/s/shard 255/s ✅ local ~256/s ✅ / deployed ~4–16 ms/op (~60–250/s engine primitive) after #267 ✅ (was ~1.3/s) The per-op stack reconstruction + full replay was the limiter; #267 made open O(1). Sustained end-to-end rate now bounded by the emit path + fsync, not the replay.
cold replay ≥ 100 K ev/s 185 K/s ✅ not re-measured in-cluster Layer A stands.
(incumbent reference, for context) — p99 52 ms sustained / 300 ms burst An incumbent-parity append SLO would be ~50–100 ms p99. The engine + the fixed publish clear it; the deployed per-op path won't until #267 lands.

SLO that needs revisiting: the "durable append" latency SLO must name which cost — engine append, shared publish, or the deployed per-op path (which additionally pays local replay-on-open). #264 (#265+#266) fixed the shared publish; the deployed path stays shadow-only until #267 removes the local replay-on-open.

Hot shard / fsync ceiling → group-commit?

The Phase-1 conditional ("group-commit if a hot shard saturates the fsync ceiling") did not fire as the priority. The in-cluster limiter was not the local-append fsync ceiling (~256/s) — it was the shared-tier whole-segment re-publish. So group-commit is a secondary lever; the priority follow-up is incremental shared publish (below).

Fix landed + a corrected attribution (2026-07-08)

The incremental-shared-publish fix shipped (ehdb #265 → #266), and validating it in kind corrected the Layer-B headline: the shared-tier publish was one of two O(segment) costs on the deployed per-op path, and it is not the dominant one.

What shipped — shared-tier publish is now O(delta)

  • #265 — publish_shard sends the backend only the append-delta [published_len .. current_len] (not the whole active segment) via a new SharedSegmentBackend::append_segment; FilesystemSharedBackend appends in place and re-commits the integrity marker.
  • #266 — the worker builds the shared-tier stack per op (a stateless boundary), so #265's in-memory running-digest cache was empty every append and each publish re-seeded the digest by re-reading the whole committed prefix — re-introducing an O(committed) read. #266 persists the resumable XxHash64 state on the segment's integrity sidecar (twox-hash serialize feature), so a freshly-constructed backend resumes the digest and folds in only the delta. Guarantees unchanged (cold-load correctness, digest keys, ordering, crash-safety); the whole-prefix digest string is identical.
  • Engine proof (Layer A micro-bench eventlog_shared_tier/append_at_size, host, two fsyncs/append): append is flat ~12–13 ms at pre-warmed active segment sizes 100 / 2 000 / 10 000 — no scaling with segment size. Unit tests cover the O(delta) guard, per-op-construction resume, crash-safety, and legacy reseed fallback (227 ehdb-reference tests pass).

The corrected finding — local-replay-on-open dominates the deployed path

The worker rebuilds the durable stack per op (append_selected → build_durable_stack). That opens the local DurableSegmentStore, and DurableSegmentStore::open replays every segment to rebuild the offset index — an O(segment) cost that runs on every mirrored event, independent of the shared publish. Because both costs scale with segment size, the earlier selfcheck decompose could not separate them and attributed the whole ~0.5–1.2 s to the shared re-publish. Deploying #266 (which makes the shared publish provably O(delta)) isolated the truth:

Probe (in-kind, ehdb-selfcheck durable-eventlog, warm cache, settled VM) per-op latency
Fresh empty local dir (no replay) ~4 ms
Live 8 MB local (replay) + shared O(delta) ~0.15–0.25 s
Live 16 MB local (replay) + shared O(delta) ~0.6–0.7 s
Live 20 MB local (replay) + shared O(delta) ~0.5–0.75 s

The cost scales with local segment size while the shared publish is O(delta) — so the residual is the local replay-on-open, not the shared publish.

Deployed before→after (harness ehdb-drive, same ~20 MB durable state)

before (as recorded) after #266
as-deployed mirror append rate ~1–2 append/s ~1.3 append/s
mirror-append latency p50 / p95 / p99 ~740 ms / ~975 ms / ~1.17 s ~650 ms / ~1.10 s / ~1.65 s

The deployed append rate did not improve — as expected once the bottleneck is understood: the shared-publish fix removed a real O(segment) cost (confirmed by the flat micro-bench and the O(delta) isolation), but the deployed per-op latency is gated by the local replay-on-open, tracked in #267. Both O(segment) costs must be removed for the deployed durable path to approach the incumbent; #264 removed the shared-publish one, #267 is the remaining (dominant) one.

Merge decision — micro-bench is the authoritative signal (2026-07-08)

Per the perf methodology (Layer A engine micro-benchmarks are the reliable signal; Layer B in-cluster is directional on the build-contended podman VM), the shared-publish fix was merged on the micro-bench evidence: the flat append-vs-segment-size proof (~12–13 ms at 100 / 2 000 / 10 000) shows the shared-tier publish SLO is now met at the engine level (O(delta), no scaling with segment size). Merging was not gated on the slow in-kind rebuild.

Deferred follow-up (not blocking): re-run Layer B in kind to confirm the as-deployed append-rate / p99 improvement once (a) #267 (local replay-on-open) lands — it is the dominant deployed cost, so the deployed rate will not move until it is fixed — and (b) the podman VM isn't build-contended (add swap, or build the worker image off-VM first). The in-kind run done this session already isolated #267 as the gate; the next run is the post-#267 confirmation.

#267 fixed — local replay-on-open is now O(1) (2026-07-09)

The dominant deployed cost above — the worker rebuilding the durable stack per op, so every mirrored append paid a fresh DurableSegmentStore::open that replayed every segment to rebuild the offset index (O(segment)) — is fixed in ehdb#268 (merged f6fdaea, closes #267).

The fix. A small per-store checkpoint.json sidecar — {event_count, active_segment_id, active_len, consumers_seen, consumer_acks} — rewritten after each mutating op strictly after the frame fsync. Open-for-append loads it O(1) and skips the replay; the offset index is materialised lazily on the first read (ensure_index_loaded). It mirrors the #266 resumable-digest sidecar. Durability is preserved — the checkpoint is a cache, not truth:

  • Replay-is-truth stays authoritative. A missing / stale / inconsistent checkpoint → full replay (and the checkpoint is rewritten). The strict active-segment length anchor means the checkpoint can never name more durable data than the segments hold, so a crash between the frame fsync and the checkpoint rewrite is caught → replay recovers the extra fsync'd frame(s).
  • CRC integrity is enforced on every read (the lazy-index replay runs full CRC + gapless checks). The append path never reads event bodies, so no corrupt data is silently served; a corrupt frame is a hard error on first read (or at open without a valid checkpoint).
  • Torn-tail recovery, fsync-before-ack, single-writer/affinity routing, and byte-exactly-once are unchanged.

Micro-bench — per-op-open+append is now FLAT across local-segment size

The prior held-open append_at_size bench reused one open store, so it never saw the O(segment) replay. New per-op-open benches reconstruct the driver / full shared stack per op — the deployed build_durable_stack shape:

S (pre-warmed events) durable_segment/per_op_open_append_at_size (driver) durable_segment_shared/per_op_open_append_at_size (full stack)
100 5.11 ms 13.9 ms
2,000 5.21 ms 14.2 ms
10,000 5.35 ms 14.5 ms

Flat across a 100× growth in segment size → open-for-append is O(1). Before the fix this curve rose with S. 232 ehdb-reference tests pass (+5 checkpoint tests); clippy -D warnings + fmt clean.

Deployed before→after — the payoff (kind user pool, ~20 MB / 24 765-record store)

ehdb-selfcheck durable-eventlog on the user-pool worker (noetl-worker-rust), gauge noetl_ehdb_eventlog_last_duration_seconds:

image per-op durable mirror
before v5.72.0-ehdb266 ~0.509 s (local replay-on-open)
after v5.72.0-ehdb267 ~0.004–0.016 s (steady ~4–10 ms)

A ~30–120× per-op reduction on the same ~20 MB durable state — the deployed gate #265/#266 could not move is gone. The first-ever op after the roll paid a one-time legacy replay (no checkpoint yet) and wrote the sidecar (checkpoint.json, 103 B, event_count:24777); every subsequent op is O(1). hello_world end-to-end COMPLETED on the ehdb267 pods; both user + system pools rolled clean (0 restarts). No GKE/prod; all NOETL_EHDB_* flags stay default off; repos/noetl + repos/server untouched.

The SLO table above updates accordingly: the deployed durable append now clears an incumbent-parity p99 (~4–16 ms vs the incumbent's ~52 ms sustained), and sustained throughput is no longer capped by the per-op stack reconstruction.

#261 head-to-head RE-RUN on merged main (2026-07-11)

Re-measured the deployed durable append on the shipped, merged-main image localhost/noetl-worker:ehdb178184-merged (carries #264/#265/#266 shared-publish O(delta) + #267 O(1) open + the GC/retention work), superseding the stale 2026-07-08 Layer-B numbers that predate all of those fixes. LOCAL kind only — no GKE/prod; noetl.event never purged.

The before → after — the shared-publish bottleneck is gone

Metric Stale (2026-07-08, pre-#264/#266/#267) RE-RUN (2026-07-11, merged main)
Deployed durable mirror-append (in-binary op duration, shared-tier stack) ~740 ms p50 / ~1.17 s p99 (~1–2 append/s) ~6–7 ms typical (6.82 / 6.25 ms; one VM-contention spike 34 ms)

A ~100–200× reduction on the as-deployed path. The ~6 ms lands on the Layer-A engine floor (below), confirming the deployed durable append is now the cheap local fsync — the O(active-segment) whole-segment re-publish that dominated the stale ~0.7–1.2 s is eliminated (#266), and the per-op O(segment) replay-on-open is gone (#267). (durable-eventlog reported durable_replay_records 9/12/15 = the read-only reopen replayed the retained log; the parity_mismatch/ok:false are the R7 artifact — the verb asserts a fresh isolated 3-event sequence against a live populated shard, not a durability failure.)

Layer A — authoritative (deterministic host bench, VM-independent)

The criterion micro-bench (cargo bench -p ehdb-reference --bench engine_micro) is the trustworthy engine number. Post-#267 per-op-open+append is flat across local-segment size (the #267 section):

S (events) driver per-op open+append full shared-tier stack per-op
100 5.11 ms 13.9 ms
2,000 5.21 ms 14.2 ms
10,000 5.35 ms 14.5 ms

Flat = the O(segment) costs are gone; append no longer degrades with store size.

Layer B — directional incumbent baseline (VM-contention caveat in force)

Same-day incumbent leg (scripts/perf/layer_b_eventlog_load.sh incumbent-sustained, POST /api/events = Postgres INSERT + NATS publish, ab n=3000 c=50): ~90–93 req/s, p50 163 ms, p95 986 ms, p99 3.6 s, 0 failures. This is far slower than the #263 baseline (~2 150 ev/s, p50 12.6 ms, p99 52 ms) — not a regression (the incumbent path is unchanged code); it is pure podman-VM contention: today's VM was running all three worker pools on fresh images plus the retention-soak append load. This is exactly why Layer B is directional, not authoritative. Read the incumbent's engine cost from the less-contended #263 run (p50 ~12.6 ms); read EHDB's from Layer A (~4–16 ms).

Head-to-head verdict (honest)

  • The deployed EHDB durable append is competitive with the incumbent. EHDB's ~6 ms in-process local fsync (no HTTP/Postgres/NATS hop) is at/below the incumbent's engine-cost p50 (~12.6 ms on a clean VM). On a contended VM the incumbent's full HTTP write-boundary inflates (p50 163 ms today) while EHDB's local append stays ~6 ms — but that is a different measurement boundary (in-process append vs network write), so the fair claim rests on Layer A + the deployed re-measurement, not the contended incumbent HTTP number.
  • The stale "as-deployed is 100× slower than the incumbent" finding is obsolete. That gap was the shared-publish O(segment) algorithm (#264), now fixed. As-deployed durable append went from ~740 ms to ~6 ms.
  • Authoritative engine claim: EHDB durable per-op append is ~4–16 ms, flat across store size (Layer A), clearing an incumbent-parity p99 SLO.

Next phase

  1. Fix local replay-on-open (#267). Done (ehdb#268, f6fdaea) — see the section above. The remaining per-op cost is the fsync(s) + the O(delta) shared publish, both flat.
  2. Group-commit spike (secondary) — batch multiple local appends behind one fsync if a single hot shard must exceed the ~256/s local ceiling.
  3. Scope confirmation from the platform owner on the per-backend SLO split and whether a production-representative (Cloud SQL + real network) incumbent head-to-head is required before Stage-C cutover.

Related

Clone this wiki locally