Repository navigation
Runbook L1 Command Bus Cutover
STATUS: ✅ T4 COMPLETE — EHDB is the production command bus on shastaratech prod, and since 2026-07-30 it is ~2.4× FASTER than the NATS envelope it replaced. The staged Option-2 (single command shard) cutover shipped 2026-07-27 on server v3.58.1 / worker v5.81.1 (80/80, 0 dup, 0 loss). Round 4 on 2026-07-30 upgraded prod to server v3.58.3 / worker v5.81.3 / ehdb
86a24f9and cleared both remaining gates:
- Dispatch latency: p50 138.5 ms / p95 156.4 / p99 181.0 (n=60, unsaturated) against the 338 ms NATS baseline — ~2.4× faster than NATS, down from 285–557 ms at round 3. → noetl/ai-meta#205 CLOSED.
- Writer-restart survival: writer pod deleted mid-load → 38/38 COMPLETED, 114 = 38×3 commands, 0 dup / 0 loss, ~2.7 s redial, cursor-resume rather than shard replay. → noetl/ai-meta#208 CLOSED.
⚠️ Round 4 also found that round 3's sign-off had already stopped holding. The writer restarted on 2026-07-28 and user-playbook dispatch was silently dead for ~2.4 days — every user-pool claimer wedged on a half-open socket, masked by system-pool traffic that kept flowing. Read the round-4 section before trusting any "cutover complete" claim that has not restarted the writer.NATS is still installed as the rollback path (one command:
NOETL_COMMAND_BUS=nats). T5 (delete NATS) is HELD on the human.The current T5 go/no-go item is autoscaling, not latency:
noetl-worker-ruststill triggers onnats-jetstreamlag, so deleting NATS would strand the user pool at 8 concurrent slots. The EHDB-lag scaler is built (noetl/ops#242, kind-proven 3 → 6 → 12 → 20 → 6 → 2) and applied to prod paused — and paused means nothing is watching, because KEDA 2.15 deletes the HPA and stops the scaler loop. The user pool has had no autoscaling at all since 2026-07-26. → noetl/ai-meta#210, worth landing independently of T5. Also still open: #209 (crash loses the unsealed tail — a loss class) and #206 (out_of_order_appendsnot exposed on writer:9102).The full flip required fixing a real delivery-loss bug first — noetl/ai-meta#203, root-caused and fixed in noetl/ehdb#300. Detail in the Prod execution log below, which records all four rounds.
Prior — T4 KIND RE-VALIDATED 2026-07-22 with subject-based routing; both findings RESOLVED; prod-go-READY pending human sign-off. The T0–T3 shadow arc and the T4 wiring (
NOETL_COMMAND_BUSflag; shadow KEDA scaler) are merged, default-off, kind/local only. T4 end-to-end kind validation PASSED all correctness dimensions (real playbook over EHDB; exactly-once 108/108 0-dup across 4 competing consumers; crashed-replica redelivery 12/12 0-loss; KEDA lag live; one-command rollback to NATS clean) at NATS latency parity (the ~800ms floor is a server-side pipeline artifact common to both buses; pure transport = 137µs).Finding #1 (pool isolation) — RESOLVED via a general subject-based routing/subscription primitive (the honest NATS-subject equivalent, the RFC's G4 gap): every command routes by a
Subject(commands.<pool>.shard.<n>); workers subscribe with aSubjectFilter(NATS*/>wildcards) and the coordinator only ever hands a member a matching command. Pool isolation and #166 shard routing are subject dimensions of one mechanism. Proven on the deploy: a user-pool worker claimed 0 commands of asystem/execution; the__orchestrate__drive (a system command even inside a shared execution) went only to the system pool. ehdbf5dd76b(noetl/ehdb#299) + worker noetl/worker#186; server unchanged (subject derived worker-side from theexecution_poolit already stamps). Finding #2 (DNS) — RESOLVED + proven: the server + workers connect tonoetl-cmdbus-writer.noetl.svc.cluster.local(a DNS name, resolved at connect time; no hardcoded IP).Next: the prod canary / full-flip (T4) and delete-NATS (T5) — human-gated. No agent flips the live bus or touches prod. Detail: ai-meta#194 (comment 2026-07-22 round 2). This page is the plan + env catalogue + one-command rollback for the operator.
All in the ehdb-feed crate over the L0 Watch(shard,cursor) change-feed.
Additive, NATS still authoritative, nothing in prod.
| Phase | What | Evidence |
|---|---|---|
| T0 | shadow feed — ChangeFeed primitive + networked FeedWriter/serve/FeedSubscription
|
latency gate: bus append→subscriber p50 57µs / p99 137µs (NATS-parity; NATS intra-node p99 ≈ 0.1–1ms). Per-hop: feed-read 13µs, transport ~29µs one-way. The ~4ms first-observed was posture-A fsync durability, a separate tunable dimension (sub-ms on NVMe; group-commit amortizes; NATS-JetStream shares it), not the bus. Parity: 0 missed / 0 spurious. |
| T1 | consumer groups + ack / ack_wait redelivery + shard routing (ShardConsumerGroup) |
competing consumers (each record to exactly one member), ack_wait redelivery, crashed-member full redelivery (at-least-once, 0 loss), committed-cursor resume, per-shard isolation, replica-kill fallback |
| T2 | KEDA lag signal — per-shard backlog gauge + Prometheus /metrics
|
lag tracks backlog; exposition well-formed; /metrics scrapeable. Ready before T4 (hard-ordering rule). |
| T3 | gateway/SPA live feed over SSE (serve_sse) |
text/event-stream; SSE id:/Last-Event-ID maps onto the cursor; reconnect resumes 0-missed/0-dup |
A single environment variable selects the command transport, defaulting to NATS so the cutover is opt-in and reversible. The code exists (path A, merged — see the env catalogue below):
NOETL_COMMAND_BUS = nats # (default) publish/subscribe over NATS JetStream — today's path
| ehdb # server publishes to the per-shard writer; workers claim from it
| shadow # BOTH: NATS authoritative + EHDB mirrored for parity comparison
-
Server (
noetl-server/repos/server, PR #281): whenehdb/shadow, the command dispatch also/instead routes the notification to the owning shard's writer viaPublishRouter.append(the networked publish path). The server is not in the delivery path either way. -
Worker (
noetl-worker/repos/worker, PR #184): the host pool (NOETL_COMMAND_BUS_HOST=true, the system-pool shard owner) opens the shard's durable command-log writer and serves its ingest / claim //metricsfaces; the consumer (ehdbmode) claims over the network from its shard coordinator viaclaim_next/ack/nack— a shared competing consumer across replicas (path A), so each command goes to exactly one worker, withack_waitredelivery on crash. No NATS subscriber is created inehdbmode. -
KEDA (
repos/ops, PR #240): aScaledObjectonsum(ehdb_feed_total_lag)via Prometheus, shipped paused (autoscaling.keda.sh/paused: "true") so the signal is validated before it drives the real pool.
| Var | Where | Meaning |
|---|---|---|
NOETL_COMMAND_BUS |
server + worker |
nats (default) / ehdb / shadow
|
NOETL_COMMAND_SHARD_COUNT |
server + worker | shard modulus (must match across both) |
NOETL_COMMAND_BUS_WRITER_ADDRS |
server |
shard@host:port,… — the writers' ingest ports |
NOETL_COMMAND_BUS_HOST |
worker |
true on the system-pool shard owner → host the writer |
NOETL_COMMAND_BUS_SHARD |
worker (host) | this worker's shard |
NOETL_COMMAND_BUS_WRITER_DIR |
worker (host) | durable command-log dir — must be a PVC |
NOETL_COMMAND_BUS_INGEST_BIND |
worker (host) | ingest listen addr (server publishes here) |
NOETL_COMMAND_BUS_CLAIM_BIND |
worker (host) | claim listen addr (replicas compete here) |
NOETL_COMMAND_BUS_METRICS_BIND |
worker (host) |
/metrics (ehdb_feed_total_lag) listen addr |
NOETL_COMMAND_BUS_CLAIM_ADDR |
worker (consumer) | the shard coordinator to claim from |
NOETL_COMMAND_BUS_ACK_WAIT_SECS |
worker (host) | redelivery window (default 30) |
The flag is read once at process start. No dual-write of commands in ehdb
mode: a given process is on exactly one bus (shadow publishes to both but
workers still consume NATS). Cutover is a rolling restart with the flag
flipped, not a data migration.
-
Kind validation (
repos/ops/automation/development/noetl.yaml,--runtime local): bring up the stack withNOETL_COMMAND_BUS=ehdbon server + worker. Run the e2e core+integration suite. Confirm: playbooks execute end-to-end, ack/redelivery works under an induced worker kill, lag gauge drives the shadowScaledObject, no orphaned commands. - Shadow-parity confirmation (already the T0–T3 posture): with NATS still authoritative, the EHDB feed observed the same command stream with 0 missed / 0 spurious and NATS-parity latency. Re-confirm on the target image.
-
Canary shard (GKE): flip
NOETL_COMMAND_BUS=ehdbfor one shard's worker replica set + the server's publish path for that shard only (per-shard subject routing already exists, #166). Watch: command completion rate, redelivery counter, lag gauge, p99 dispatch latency vs the NATS baseline. Hold for a soak window. - Full flip: roll the remaining shards. NATS is now carrying no commands but is left running (not deleted) — this is the reversible watermark.
- Bake: run at full flip for the agreed soak with NATS still installed and the rollback one command away.
# Server + worker back onto NATS, rolling restart:
kubectl -n noetl set env deploy/noetl-server NOETL_COMMAND_BUS=nats
kubectl -n noetl set env deploy/noetl-worker NOETL_COMMAND_BUS=nats
kubectl -n noetl set env deploy/noetl-worker-system-pool NOETL_COMMAND_BUS=nats
kubectl -n noetl rollout restart deploy/noetl-server deploy/noetl-worker deploy/noetl-worker-system-pool
Because NATS was never deleted and commands are not dual-written, this is a clean revert: in-flight EHDB-fed commands finish (ack/redelivery bounded by ack_wait), new commands publish to NATS again. No data reconciliation.
Only after T4 has baked and the operator explicitly approves. Removing the NATS Deployment/StatefulSet + its Secret (see also noetl/ai-meta#188 — the plaintext NATS credential to retire at the same time) is irreversible: after it, rollback means re-provisioning NATS. Do not couple T5 to T4 — they are distinct approvals.
- Target server + worker images built with the
NOETL_COMMAND_BUSflag, kind-validated (step 1) — full-functionality e2e green. (server v3.58.1 / worker v5.81.1.) - Latency on the target image re-measured on representative hardware
(not the sandbox FS) and within the NATS envelope, including the
durability posture chosen for the command log.
⚠️ Cleared with a caveat — correctness passed and dispatch stays sub-second, but p50 285–557 ms vs the ~200 ms NATS baseline is a real elevation. Accepted for T4; open for T5 (noetl/ai-meta#205). -
ScaledObjectreadingehdb_feed_total_lagapplied in shadow and observed to track real backlog. - Redelivery + at-least-once verified under an induced worker kill on the target image. (Canary graceful pod-kill 12/12 0-loss on prod; 25/25 in kind under load.)
- Per-shard writer durability posture decided (fsync-per-append vs group-commit) with the crash-window trade-off signed off.
- Rollback command rehearsed in kind — and exercised for real on prod in round 2, which rolled back cleanly.
- Operator sign-off recorded: 2026-07-27, explicit human GO, headless synthetic load, no organic traffic.
Two of the original blockers cleared on the 2026-07-30 prod deploy (round 4 below). One new blocker replaced them, and it is the one that now gates T5.
-
Dispatch latency accepted or optimized (noetl/ai-meta#205) — CLEARED 2026-07-30, and inverted. Prod p50 138.5 ms / p99 181.0 ms against the 338 ms NATS baseline: the EHDB bus is now ~2.4× faster than the envelope it replaces, not 1.5–2.8× slower. Issue closed.
-
The bus survives a writer restart (noetl/ai-meta#208) — CLEARED 2026-07-30. This was not on the original checklist because nobody had tested it; it turned out to be broken, and silently broken in prod for ~2.4 days. Fixed and prod-proven: writer pod deleted mid-load → 38/38 COMPLETED, 114 = 38×3 commands, 0 dup / 0 loss, ~2.7 s redial, cursor-resume rather than shard replay. T5 removes the NATS fallback, so this had to be true before T5 — it is the clearest argument for testing restart behaviour before deleting the escape hatch.
-
The user pool scales on EHDB lag (noetl/ai-meta#210) — the new go/no-go item.
noetl-worker-rust's ScaledObject triggers onnats-jetstreamconsumer lag; deleting NATS removes the user pool's only scaler signal and freezes it at 2 replicas ×WORKER_MAX_CONCURRENT4 = 8 concurrent slots. Under a saturating burst that already shows as p50 2123 ms of pure queueing (measured 2026-07-30 — the bus was fine; the pool was not). Sequence: 1. merge noetl/ehdb#303 —ehdb_feed_subject_lag{subject}; 2. land the worker adoption + deploy; 3. merge noetl/ops#242 — the ScaledObject + thePodMonitoringthat has been live-but-uncommitted since T4; 4. swapvalueLocationtoehdb_feed_subject_lag{subject="commands.shared.shard.0"}; 5. unpause (human-gated — it changes live user-pool behaviour); 6. soak.⚠️ **The pool has had no autoscaling at all since 2026-07-26.** A paused ScaledObject does not merely withhold scaling: KEDA 2.15 **deletes the HPA and stops running the scaler loop**, and `.status.externalMetricNames` never refreshes — so the object looks configured while nothing watches. Unpausing is therefore an improvement over today's posture **independent of T5**; do not let it wait on the T5 decision. -
The writer's SIGTERM seal does not race in-flight ingest, and a crash does not lose the unsealed tail (noetl/ai-meta#209). Graceful restarts are covered by the #208 seal +
terminationGracePeriodSeconds90 s. An ungraceful stop (SIGKILL, node loss) can still lose the unsealed tail of the active part. On a command bus that is a loss class, and T5 removes the fallback that makes it recoverable. -
out_of_order_appendsexposed on writer:9102(noetl/ai-meta#206) so a future feed-ordering regression is directly observable. -
A bake window on organic (not synthetic) traffic.
-
Plaintext NATS credential retired in the same change set (noetl/ai-meta#188).
-
Explicit, separate human approval — T5 is irreversible.
Four rounds. Round 1 held at canary on a topology blocker; round 2 chose the topology and failed the full flip on a real bug; round 3 shipped it; round 4 made it fast and made it survive a restart — and found that round 3's sign-off had already stopped holding.
| Round | Date | Outcome |
|---|---|---|
| 1 | 2026-07-26 | Step 0 + shadow PASS (2-shard). Held at canary — per-shard user-pool blocker → #202. Rolled back to NATS. |
| 2 | 2026-07-27 | Option 2 chosen + kind-validated + prod reconfigured to 1 command shard. Shadow + canary PASS. Full flip FAILED — ~10% of commands ingested but never claimed, feed lag=0 → #203. Rolled back to NATS; prod clean. |
| 3 | 2026-07-27 | With the #203 fix (server v3.58.1 / worker v5.81.1): FULL FLIP SUCCEEDED — 80/80, 0 dup, 0 loss. Bus = EHDB end-to-end. |
| 4 | 2026-07-30 | #205 + #208 bundle (server v3.58.3 / worker v5.81.3 / ehdb 86a24f9). Latency p50 285–557 → 138.5 ms — now ~2.4× faster than NATS. Writer-restart survival PASS. |
T4 was run on the new shastaratech prod cluster
(gke_shastaratech-noetl-prod_us-central1_noetl-prod-autopilot, project
shastaratech-noetl-prod, ns noetl) — headless, synthetic load, explicit
human GO. Reproducible IaC + the full executed-run record are in
noetl/ai-meta under playbooks/194-l1-t4-prod-iac/ (README +
step0-rollforward.sh, step2-shadow.sh, rollback.sh) and
playbooks/194-l1-t4-prod-cutover.md. Soak windows were compressed to
synthetic-load correctness gates (no organic-traffic window in one session).
-
2-shard cluster —
NOETL_COMMAND_SHARD_COUNT=2,NOETL_SHARD_SUBJECT_ROUTE=true; system pool split intonoetl-worker-system-pool(shard0) +noetl-worker-system-pool-shard1. The command bus therefore needs one writer per shard (PublishRouterkeeps one client per shard and errors on a missing shard;spawn_writer_hostbinds oneClaimCoordinatortoNOETL_COMMAND_BUS_SHARD). Deployednoetl-cmdbus-writer-0+-1, each apremium-rwo20Gi PVC + ClusterIP. -
GMP, not VictoriaMetrics — a GMP
PodMonitoringcovers both writers; live gates read the writer:9102directly via port-forward. -
Autopilot — storageclass
premium-rwo(PD-SSD). -
Images pulled from ghcr — server rolled to v3.58.0
(
ghcr.io/noetl/server@sha256:99b842…), all worker deploys to v5.81.0 (ghcr.io/noetl/worker@sha256:27807d…).
-
Step 0 (roll forward + writers, bus NATS) — PASS. server → v3.58.0, all 3
worker deploys → v5.81.0; both writers up, PVCs Bound; writers inert in
natsmode. Gate: 15/15hello_worldCOMPLETED over NATS, 0 restarts. - Step 1 (NATS baseline) — captured. command-path issued→claimed (n=90): p50 338 / p95 511 / p99 520 ms.
-
Step 2 (SHADOW) — PASS, 0 divergence. server + both writers →
shadow, workers stayed on NATS. 20 controlled execs: server published 121 → NATS == 121 → EHDB; the EHDB feed grew by exactly 121 (shard0 +61, shard1 +60); 0 shadow errors ⇒ 0 missed / 0 spurious. Server connected to writers by DNS (no IP literal). Latency p50 273 / p95 526 / p99 579 ms — inside the NATS envelope. Writer RSS/CPU flat (3m / 4Mi), feed lag climbs monotonically (expected — nothing claims in shadow), 0 restarts. - Step 3 (CANARY) — HELD (not attempted). Blocker below.
On the EHDB bus every command is physically partitioned by execution_id
across the shards — including the shared pool (proven in shadow: 61/60).
On NATS today only the system pool is sharded (sharding.rs:
ROUTABLE_POOL = "system"; the shared subject is unsharded). A worker claims
from exactly one writer (NOETL_COMMAND_BUS_CLAIM_ADDR is a single
host:port; one ClaimClient). The system pool is already split shard0/shard1
→ each maps cleanly to its writer. The user pool (noetl-worker-rust) is a
single deployment → flipping it to ehdb with one claim addr would strand
~half the shared commands (the ones on the other shard). This split does not
exist in ops#241 and was never integration-tested — kind ran single-shard.
-
Split the user pool per shard (mirror the system pool): two deployments
noetl-worker-rust-shard0(claimwriter-0:9101) +noetl-worker-rust-shard1(claimwriter-1:9101), each ~half the replicas, plus per-shard KEDA. Then canary one shard's user pool, then the other, then the system pool + server. Keeps the deployed #166 Phase 5 sharded command routing. -
Collapse the command-bus axis to a single shard — set
NOETL_COMMAND_SHARD_COUNT=1for the command-bus axis only (distinct from the server's event/state sharding), so all commands land on writer-0 and this runbook's validated single-shard / single-writer / single-user-pool design applies unchanged. Changes system-pool command routing from partitioned to competing (still exactly-once). Simpler; diverges from #166 Phase 5.
Tracked in noetl/ai-meta#202.
Images server v3.58.0 / workers v5.81.0 (validated); writers deployed
but inert (NOETL_COMMAND_BUS=nats); bus on NATS (shadow rolled back
after the parity gate); everything staged for canary once the #202 decision
landed.
Decision taken: Option 2 — collapse the command-bus axis to a single
shard. The two sharding axes were proven independent in code (sharding.rs:172
"correctness never depends on affinity"; the command bus reads only
NOETL_COMMAND_SHARD_COUNT, the state materializer only
NOETL_RESULT_SHARD_COUNT plus its own event stream) — so collapsing the
command bus does not touch #166 state sharding.
Kind re-validation first, on a rig mirroring prod-after-Option-2
(dedicated single writer + 2 state-shard system pools +
NOETL_COMMAND_SHARD_COUNT=1): single writer holds all commands
(ehdb_feed_total_lag→0); user pool + both state-shard system pools compete
exactly-once (0 dup across 5 consumers); pod-kill redelivery 8/8 0-loss;
rollback clean. Three config bugs caught here rather than on prod:
- the two system pools need a shared NATS consumer, or NATS
double-delivers while the bus is still
nats; - the writer needs its
NATS_STREAMset; -
a writer in
ehdbmode also needsNOETL_COMMAND_BUS_CLAIM_ADDR— it is a consumer of its own shard too.
Prod reconfigure to 1 command shard (bus stayed NATS) — clean. Both system
pools unified onto a shared broad consumer (noetl_worker_system_rust,
noetl.commands.system.>), server NOETL_COMMAND_SHARD_COUNT=1, writer-1
scaled to 0. #166 state sharding INTACT — both pools keep
NOETL_SHARD_COUNT=2 / index, STATE_BUILDER=offserver,
STATE_SHARD_WRITE=true; state builder healthy. Gate 12/12.
Then shadow PASS and canary PASS — but the full flip FAILED. Under the
full flip roughly 10% of commands were ingested into the writer feed and
never delivered to any claimer, and feed lag read 0 the whole time, so
the loss was silent. Rolled back to NATS; prod left clean. Filed as
noetl/ai-meta#203.
Root cause — an out-of-order append versus the feed cursor. The command
feed's producer (noetl-server) assigned each record's sort key itself — the
command's snowflake event_id — and the single writer trusted it. But the
Dataset contract the feed cursor and range pruning depend on is "records
are appended in ascending sort_key order within a partition". Snowflake ids
are sparse, and under concurrent publish a lower id can reach the writer
after a higher one. Then:
- the follower cursor has already advanced to the higher key
(
ChangeFeed::pollsetscursor = max sort_key read); -
refillreadsread_partition_after(shard, cursor)→ only> cursor→ the late lower key is filtered out, never enterspending, never delivered; -
lag()is cursor-relative too, so the lost record — below the cursor — is never counted. Silent loss,lag=0.
Fix — the writer assigns the ordering key (append_writer_assigned), so
the ascending contract is guaranteed by the only component that can guarantee
it. noetl/ehdb#300 → ehdb d4b6235
→ noetl-worker v5.81.1 (the writer host lives in the worker, so this is
the image that carries the runtime fix) + noetl-server v3.58.1. A new
out_of_order_appends counter records the condition — not yet exposed on
:9102, tracked as
noetl/ai-meta#206.
Kind re-validation under sustained load, identical 40-concurrent burst:
| delivered | stuck (command.issued, no command.claimed) |
feed lag | |
|---|---|---|---|
| pre-fix | 17/40 | 23 permanent | 0 (silent) |
| post-fix (v5.81.1) | 41/41 | 0 | 0 |
Plus exactly-once 335=335 (0 dup), graceful pod-kill redelivery 25/25, #166
state sharding intact, rollback to nats clean.
Fixed images deployed by digest (server v3.58.1
@sha256:5a73f4a5…, worker v5.81.1 @sha256:db156fa6…, copied
ghcr→project AR with crane), then the same staged Option-2 sequence re-run:
- Step 0 (bus NATS, roll forward): 20/20.
- Shadow: 0 divergence — 123 published == 123 in the feed on the single writer; writer flat.
-
Canary (user pool →
ehdb): 30/30; pool isolation — 0 system claims by the user pool (the hard requirement); graceful pod-kill redelivery 12/12, 0 loss. -
Full flip (both system pools + server →
ehdb): 30/30 + 50/50 soak = 80/80, 0 dup, 0 loss; serverNATS_pub=0; feed lag → 0; writer 0 restarts; #166 state sharding intact throughout.
The ~10% silent loss that failed round 2 was gone at the exact same stage. #203 closed as prod-validated.
The production command bus is EHDB. Server + all worker pools run
NOETL_COMMAND_BUS=ehdb on server v3.58.1 / worker v5.81.1, single
command shard, one writer on a premium-rwo PVC.
NATS is still installed and is the rollback path — one command:
kubectl -n noetl set env deploy/noetl-server-rust \
deploy/noetl-worker-rust deploy/noetl-worker-system-pool \
deploy/noetl-worker-system-pool-shard1 NOETL_COMMAND_BUS=nats
bash rollback.sh in playbooks/194-l1-t4-prod-iac/ reverts the rest.
T5 (delete NATS) NOT done — separate, irreversible approval; see the T5 checklist above. It couples noetl/ai-meta#188 (the plaintext NATS credential retires with NATS) and is gated on noetl/ai-meta#205 (dispatch latency).
Measured on the same cluster and the same synthetic load, command
issued → claimed:
| Bus | p50 | p99 | min |
|---|---|---|---|
| NATS (baseline) | ~200 ms | — | — |
| EHDB (post-flip) | 285–557 ms (load-variable) | ~1040 ms | ~145 ms |
Roughly 1.5–2.8× the NATS baseline. Sub-second, and it affected neither
correctness nor completion — but it is a real regression on the hot dispatch
path, and it is the open go/no-go item for T5. Two plausible contributors,
not yet isolated: the 250 ms poll-claim interval versus NATS's push
delivery (a push variant is already possible — the FeedWriter::tip_receiver()
seam exists), and the writer-side ordering-key append the #203 fix added.
Tracked in noetl/ai-meta#205.
Superseded by round 4 (2026-07-30). Both hypotheses above were wrong. The cost was neither the poll interval (claim delivery measured 16–33 µs, flat) nor the #203 ordering-key append: it was an
fsyncheld inside the engine lock (~4 ms, capping the bus at ~230 cmd/s) multiplied by the server's publish mutex held across the round trip. Group commit + off-lock fsync + pipelined publish took prod p50 to 138.5 ms, below the NATS baseline. See round 4.
One rollout carrying two fixes, so prod deploys once and both gates are
measured in the same window. NOETL_COMMAND_BUS=ehdb before and after — a
rolling image update on the live bus, not a bus change.
| Component | Version | Carries |
|---|---|---|
| ehdb engine | 86a24f9 |
#301 group commit + off-lock fsync + pipelined publish; #302 writer-restart survival |
noetl-server |
v3.58.3 | server#289 (pipelined publish), server#290 (publish retry) |
noetl-worker (all pools + writer) |
v5.81.3 | worker#195 (group-commit adoption), worker#196 (resume + seal) |
Post-deploy: 6/6 pods Ready, 0 restarts; #166 state sharding intact
(NOETL_SHARD_COUNT=2, index 0/1, offserver, STATE_SHARD_WRITE=true);
NOETL_COMMAND_SHARD_COUNT=1, single writer.
Before any deploy, user-playbook dispatch was dead and had been for ~2.4
days. noetl-cmdbus-writer restarted 2026-07-28T07:49Z; both
noetl-worker-rust pods' last log line was from that instant —
EHDB claim connect failed; retrying claim_addr=…:9101 error=Connection refused (os error 111)
— then silence. Synthetic executions issued and were never claimed
(ehdb_feed_total_lag climbing, executions stuck at command.issued).
noetl-worker-system-pool-shard1 happened to restart later (07-29 16:17), so
system commands kept completing and masked the outage completely.
Two things to carry forward:
- A cutover sign-off is only as strong as its restart story. Round 3 proved 80/80 with zero loss and was correct about that — it just never restarted the writer. T4's "complete" did not survive the first one.
- Silent single-pool dispatch death is not observable today. System-pool traffic keeps every cluster-level signal green. Per-pool lag (ehdb#303) is the series that would have caught it — which makes #303 an observability fix as much as a scaler input.
Gate 1 — dispatch latency (#205): PASS
command.issued → command.claimed, read off noetl.event:
| Regime | n | p50 | p95 | p99 |
|---|---|---|---|---|
| NATS baseline (round 3, 07-27) | 90 | 338 ms | 511 ms | 520 ms |
| EHDB at round 3 (07-27) | — | 285–557 ms | — | ~1040 ms |
| EHDB, round 4, unsaturated | 60 | 138.5 ms | 156.4 ms | 181.0 ms |
| EHDB, round 4, post writer-restart | 45 | 140.5 ms | 191.6 ms | 208.0 ms |
p50 138.5 ms against the 338 ms NATS baseline — ~2.4× faster than NATS, and 2–4× faster than EHDB was at round 3. The loopback attribution (281 → 6.7 ms p50 at 64 publishers / 24 claimers; 225 → 6983 cmd/s) holds on real infrastructure. Issue closed.
targetValue: "2" (a replica per 2 queued commands) rather than waiting for
backlog to exceed capacity.
Gate 2 — writer-restart survival (#208): PASS
Writer pod deleted mid-stream under continuous load:
- 38 executions accepted → 38 COMPLETED, 0 stuck.
- Commands: 40 before + 74 after = 114 = 38 × 3 exactly, zero duplicates, zero loss.
- Redial ~2.7 s: writer died 18:18:22; the worker logged
EHDB claim_next failed; reconnecting to the claim coordinator … early eofat 18:18:22.056 and retried every ~250 ms; coordinator up 18:18:24.765. The pre-fix code emitted no such line — it wedged. - Cursor-resume, not replay, proven by the exact 114 = 38 × 3 accounting.
- Steady-state latency unchanged across the restart.
-
terminationGracePeriodSecondsraised 30 → 90 s for the SIGTERM seal.
-
The resume signal is not readable on its own. The resume took the
durable path (
origin="persisted"), but the logged cursor reads0andehdb_feed_shard_committedreset 408 → 165 across the restart — the counter is segment-relative, not a global monotonic offset. No replay happened, but the log line alone cannot distinguish "resumed" from "replayed from 0" — exactly the signal this runbook would lean on during an incident. → ehdb#304, held. -
2 × HTTP 500 on
POST /api/executeduring the writer's absence — the #290 publish retry (3 × 250 ms) does not span a ~2.7 s pod restart. Fail-closed (the caller is told; every accepted execution completed), but not transparent. → server#291, held, widening the retry to a 10 s wall-clock deadline.
Prod monitoring is Google Managed Prometheus, not VictoriaMetrics — the
vmservicescrape CRD is not installed on this cluster and there is no
vmstack namespace. The writer scrape is PodMonitoring/noetl-cmdbus-writer
(10 s, cmdbus-lag port), created during the T4 cutover and never committed
to noetl/ops until ops#242.
Consequence for the T5 autoscaler: there is no in-cluster PromQL endpoint,
so the KEDA trigger is metrics-api scraping the writer's :9102 directly
rather than a prometheus trigger. That also keeps autoscaling independent of
the monitoring pipeline's health, which is the right coupling for a T5
prerequisite.
Confirmed live on prod: ehdb_feed_total_lag 0,
ehdb_feed_shard_lag{shard="0"} 0,
ehdb_feed_shard_committed{shard="0"} 5853,
up{job="noetl-cmdbus-writer"} 1.
NATS untouched — ns nats and ns nats-supercluster still installed.
T5 NOT performed. Playbook + rollback recipe:
playbooks/205-208-prod-rollout/README.md in noetl/ai-meta.
- Program: EHDB Layered Platform (L0→L3) — L1 is the NATS takeover.
- Umbrella: noetl/ai-meta#194.
- The event-log tier cutover runbook: Runbook — Prod Cutover: Event-Log Tier.
- noetl/ai-meta#188 — retire the plaintext NATS credential as part of T5.
- noetl/ai-meta#202 — the 1-shard-vs-split decision. Closed: Option 2 chosen, executed, prod-validated.
- noetl/ai-meta#203 — the feed delivery-loss bug that failed round 2. Closed, fixed in noetl/ehdb#300.
- noetl/ai-meta#205 — dispatch latency vs NATS. Closed 2026-07-30 — p50 138.5 ms vs the 338 ms NATS baseline (~2.4× faster), fixed in noetl/ehdb#301.
- noetl/ai-meta#208 — the bus did not survive a writer restart. Closed 2026-07-30, fixed in noetl/ehdb#302.
- noetl/ai-meta#210 — the user pool has had no autoscaling since 2026-07-26; landing the EHDB-lag scaler is the current T5 go/no-go item. Open.
- noetl/ai-meta#209 — writer SIGTERM seal race / unsealed tail on a crash. Open — a loss class that should land before T5 removes the NATS fallback.
-
noetl/ai-meta#206 — expose
out_of_order_appendson the writer:9102. Open. - Held PRs from round 4: noetl/ehdb#303 (per-subject lag), noetl/ehdb#304 (resume signal), noetl/server#291 (publish retry deadline), noetl/ops#242 (ScaledObject + PodMonitoring).
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)