Skip to content

Runbook L1 Command Bus Cutover

Kadyapam edited this page Jul 31, 2026 · 7 revisions

Runbook — L1 Command-Bus Cutover (EHDB feed ⟶ off NATS)

STATUS: ✅ T4 COMPLETE — EHDB is the production command bus on shastaratech prod, and since 2026-07-30 it is ~2.4× FASTER than the NATS envelope it replaced. The staged Option-2 (single command shard) cutover shipped 2026-07-27 on server v3.58.1 / worker v5.81.1 (80/80, 0 dup, 0 loss). Round 4 on 2026-07-30 upgraded prod to server v3.58.3 / worker v5.81.3 / ehdb 86a24f9 and cleared both remaining gates:

  • Dispatch latency: p50 138.5 ms / p95 156.4 / p99 181.0 (n=60, unsaturated) against the 338 ms NATS baseline — ~2.4× faster than NATS, down from 285–557 ms at round 3. → noetl/ai-meta#205 CLOSED.
  • Writer-restart survival: writer pod deleted mid-load → 38/38 COMPLETED, 114 = 38×3 commands, 0 dup / 0 loss, ~2.7 s redial, cursor-resume rather than shard replay. → noetl/ai-meta#208 CLOSED.

⚠️ Round 4 also found that round 3's sign-off had already stopped holding. The writer restarted on 2026-07-28 and user-playbook dispatch was silently dead for ~2.4 days — every user-pool claimer wedged on a half-open socket, masked by system-pool traffic that kept flowing. Read the round-4 section before trusting any "cutover complete" claim that has not restarted the writer.

NATS is still installed as the rollback path (one command: NOETL_COMMAND_BUS=nats). T5 (delete NATS) is HELD on the human.

The current T5 go/no-go item is autoscaling, not latency: noetl-worker-rust still triggers on nats-jetstream lag, so deleting NATS would strand the user pool at 8 concurrent slots. The EHDB-lag scaler is built (noetl/ops#242, kind-proven 3 → 6 → 12 → 20 → 6 → 2) and applied to prod paused — and paused means nothing is watching, because KEDA 2.15 deletes the HPA and stops the scaler loop. The user pool has had no autoscaling at all since 2026-07-26. → noetl/ai-meta#210, worth landing independently of T5. Also still open: #209 (crash loses the unsealed tail — a loss class) and #206 (out_of_order_appends not exposed on writer :9102).

The full flip required fixing a real delivery-loss bug first — noetl/ai-meta#203, root-caused and fixed in noetl/ehdb#300. Detail in the Prod execution log below, which records all four rounds.

Prior — T4 KIND RE-VALIDATED 2026-07-22 with subject-based routing; both findings RESOLVED; prod-go-READY pending human sign-off. The T0–T3 shadow arc and the T4 wiring (NOETL_COMMAND_BUS flag; shadow KEDA scaler) are merged, default-off, kind/local only. T4 end-to-end kind validation PASSED all correctness dimensions (real playbook over EHDB; exactly-once 108/108 0-dup across 4 competing consumers; crashed-replica redelivery 12/12 0-loss; KEDA lag live; one-command rollback to NATS clean) at NATS latency parity (the ~800ms floor is a server-side pipeline artifact common to both buses; pure transport = 137µs).

Finding #1 (pool isolation) — RESOLVED via a general subject-based routing/subscription primitive (the honest NATS-subject equivalent, the RFC's G4 gap): every command routes by a Subject (commands.<pool>.shard.<n>); workers subscribe with a SubjectFilter (NATS */> wildcards) and the coordinator only ever hands a member a matching command. Pool isolation and #166 shard routing are subject dimensions of one mechanism. Proven on the deploy: a user-pool worker claimed 0 commands of a system/ execution; the __orchestrate__ drive (a system command even inside a shared execution) went only to the system pool. ehdb f5dd76b (noetl/ehdb#299) + worker noetl/worker#186; server unchanged (subject derived worker-side from the execution_pool it already stamps). Finding #2 (DNS) — RESOLVED + proven: the server + workers connect to noetl-cmdbus-writer.noetl.svc.cluster.local (a DNS name, resolved at connect time; no hardcoded IP).

Next: the prod canary / full-flip (T4) and delete-NATS (T5) — human-gated. No agent flips the live bus or touches prod. Detail: ai-meta#194 (comment 2026-07-22 round 2). This page is the plan + env catalogue + one-command rollback for the operator.

What is already built (T0–T3, shadow, merged, kind/local only)

All in the ehdb-feed crate over the L0 Watch(shard,cursor) change-feed. Additive, NATS still authoritative, nothing in prod.

Phase What Evidence
T0 shadow feed — ChangeFeed primitive + networked FeedWriter/serve/FeedSubscription latency gate: bus append→subscriber p50 57µs / p99 137µs (NATS-parity; NATS intra-node p99 ≈ 0.1–1ms). Per-hop: feed-read 13µs, transport ~29µs one-way. The ~4ms first-observed was posture-A fsync durability, a separate tunable dimension (sub-ms on NVMe; group-commit amortizes; NATS-JetStream shares it), not the bus. Parity: 0 missed / 0 spurious.
T1 consumer groups + ack / ack_wait redelivery + shard routing (ShardConsumerGroup) competing consumers (each record to exactly one member), ack_wait redelivery, crashed-member full redelivery (at-least-once, 0 loss), committed-cursor resume, per-shard isolation, replica-kill fallback
T2 KEDA lag signal — per-shard backlog gauge + Prometheus /metrics lag tracks backlog; exposition well-formed; /metrics scrapeable. Ready before T4 (hard-ordering rule).
T3 gateway/SPA live feed over SSE (serve_sse) text/event-stream; SSE id:/Last-Event-ID maps onto the cursor; reconnect resumes 0-missed/0-dup

The cutover design (T4)

The flag

A single environment variable selects the command transport, defaulting to NATS so the cutover is opt-in and reversible. The code exists (path A, merged — see the env catalogue below):

NOETL_COMMAND_BUS = nats    # (default) publish/subscribe over NATS JetStream — today's path
                   | ehdb    # server publishes to the per-shard writer; workers claim from it
                   | shadow   # BOTH: NATS authoritative + EHDB mirrored for parity comparison
  • Server (noetl-server / repos/server, PR #281): when ehdb/shadow, the command dispatch also/instead routes the notification to the owning shard's writer via PublishRouter.append (the networked publish path). The server is not in the delivery path either way.
  • Worker (noetl-worker / repos/worker, PR #184): the host pool (NOETL_COMMAND_BUS_HOST=true, the system-pool shard owner) opens the shard's durable command-log writer and serves its ingest / claim / /metrics faces; the consumer (ehdb mode) claims over the network from its shard coordinator via claim_next/ack/nack — a shared competing consumer across replicas (path A), so each command goes to exactly one worker, with ack_wait redelivery on crash. No NATS subscriber is created in ehdb mode.
  • KEDA (repos/ops, PR #240): a ScaledObject on sum(ehdb_feed_total_lag) via Prometheus, shipped paused (autoscaling.keda.sh/paused: "true") so the signal is validated before it drives the real pool.

Environment catalogue (as implemented)

Var Where Meaning
NOETL_COMMAND_BUS server + worker nats (default) / ehdb / shadow
NOETL_COMMAND_SHARD_COUNT server + worker shard modulus (must match across both)
NOETL_COMMAND_BUS_WRITER_ADDRS server shard@host:port,… — the writers' ingest ports
NOETL_COMMAND_BUS_HOST worker true on the system-pool shard owner → host the writer
NOETL_COMMAND_BUS_SHARD worker (host) this worker's shard
NOETL_COMMAND_BUS_WRITER_DIR worker (host) durable command-log dir — must be a PVC
NOETL_COMMAND_BUS_INGEST_BIND worker (host) ingest listen addr (server publishes here)
NOETL_COMMAND_BUS_CLAIM_BIND worker (host) claim listen addr (replicas compete here)
NOETL_COMMAND_BUS_METRICS_BIND worker (host) /metrics (ehdb_feed_total_lag) listen addr
NOETL_COMMAND_BUS_CLAIM_ADDR worker (consumer) the shard coordinator to claim from
NOETL_COMMAND_BUS_ACK_WAIT_SECS worker (host) redelivery window (default 30)

The flag is read once at process start. No dual-write of commands in ehdb mode: a given process is on exactly one bus (shadow publishes to both but workers still consume NATS). Cutover is a rolling restart with the flag flipped, not a data migration.

Sequence (operator, kind first, then GKE — each step gated)

  1. Kind validation (repos/ops/automation/development/noetl.yaml, --runtime local): bring up the stack with NOETL_COMMAND_BUS=ehdb on server + worker. Run the e2e core+integration suite. Confirm: playbooks execute end-to-end, ack/redelivery works under an induced worker kill, lag gauge drives the shadow ScaledObject, no orphaned commands.
  2. Shadow-parity confirmation (already the T0–T3 posture): with NATS still authoritative, the EHDB feed observed the same command stream with 0 missed / 0 spurious and NATS-parity latency. Re-confirm on the target image.
  3. Canary shard (GKE): flip NOETL_COMMAND_BUS=ehdb for one shard's worker replica set + the server's publish path for that shard only (per-shard subject routing already exists, #166). Watch: command completion rate, redelivery counter, lag gauge, p99 dispatch latency vs the NATS baseline. Hold for a soak window.
  4. Full flip: roll the remaining shards. NATS is now carrying no commands but is left running (not deleted) — this is the reversible watermark.
  5. Bake: run at full flip for the agreed soak with NATS still installed and the rollback one command away.

One-command rollback (any step, before T5)

# Server + worker back onto NATS, rolling restart:
kubectl -n noetl set env deploy/noetl-server  NOETL_COMMAND_BUS=nats
kubectl -n noetl set env deploy/noetl-worker  NOETL_COMMAND_BUS=nats
kubectl -n noetl set env deploy/noetl-worker-system-pool NOETL_COMMAND_BUS=nats
kubectl -n noetl rollout restart deploy/noetl-server deploy/noetl-worker deploy/noetl-worker-system-pool

Because NATS was never deleted and commands are not dual-written, this is a clean revert: in-flight EHDB-fed commands finish (ack/redelivery bounded by ack_wait), new commands publish to NATS again. No data reconciliation.

T5 — delete NATS (the point of no return, separately gated)

Only after T4 has baked and the operator explicitly approves. Removing the NATS Deployment/StatefulSet + its Secret (see also noetl/ai-meta#188 — the plaintext NATS credential to retire at the same time) is irreversible: after it, rollback means re-provisioning NATS. Do not couple T5 to T4 — they are distinct approvals.

Go / no-go checklist (T4) — CLEARED 2026-07-27

  • Target server + worker images built with the NOETL_COMMAND_BUS flag, kind-validated (step 1) — full-functionality e2e green. (server v3.58.1 / worker v5.81.1.)
  • Latency on the target image re-measured on representative hardware (not the sandbox FS) and within the NATS envelope, including the durability posture chosen for the command log. ⚠️ Cleared with a caveat — correctness passed and dispatch stays sub-second, but p50 285–557 ms vs the ~200 ms NATS baseline is a real elevation. Accepted for T4; open for T5 (noetl/ai-meta#205).
  • ScaledObject reading ehdb_feed_total_lag applied in shadow and observed to track real backlog.
  • Redelivery + at-least-once verified under an induced worker kill on the target image. (Canary graceful pod-kill 12/12 0-loss on prod; 25/25 in kind under load.)
  • Per-shard writer durability posture decided (fsync-per-append vs group-commit) with the crash-window trade-off signed off.
  • Rollback command rehearsed in kind — and exercised for real on prod in round 2, which rolled back cleanly.
  • Operator sign-off recorded: 2026-07-27, explicit human GO, headless synthetic load, no organic traffic.

Go / no-go checklist (T5 — delete NATS) — NOT cleared

Two of the original blockers cleared on the 2026-07-30 prod deploy (round 4 below). One new blocker replaced them, and it is the one that now gates T5.

  • Dispatch latency accepted or optimized (noetl/ai-meta#205) — CLEARED 2026-07-30, and inverted. Prod p50 138.5 ms / p99 181.0 ms against the 338 ms NATS baseline: the EHDB bus is now ~2.4× faster than the envelope it replaces, not 1.5–2.8× slower. Issue closed.

  • The bus survives a writer restart (noetl/ai-meta#208) — CLEARED 2026-07-30. This was not on the original checklist because nobody had tested it; it turned out to be broken, and silently broken in prod for ~2.4 days. Fixed and prod-proven: writer pod deleted mid-load → 38/38 COMPLETED, 114 = 38×3 commands, 0 dup / 0 loss, ~2.7 s redial, cursor-resume rather than shard replay. T5 removes the NATS fallback, so this had to be true before T5 — it is the clearest argument for testing restart behaviour before deleting the escape hatch.

  • The user pool scales on EHDB lag (noetl/ai-meta#210) — the new go/no-go item. noetl-worker-rust's ScaledObject triggers on nats-jetstream consumer lag; deleting NATS removes the user pool's only scaler signal and freezes it at 2 replicas × WORKER_MAX_CONCURRENT 4 = 8 concurrent slots. Under a saturating burst that already shows as p50 2123 ms of pure queueing (measured 2026-07-30 — the bus was fine; the pool was not). Sequence: 1. merge noetl/ehdb#303 — ehdb_feed_subject_lag{subject}; 2. land the worker adoption + deploy; 3. merge noetl/ops#242 — the ScaledObject + the PodMonitoring that has been live-but-uncommitted since T4; 4. swap valueLocation to ehdb_feed_subject_lag{subject="commands.shared.shard.0"}; 5. unpause (human-gated — it changes live user-pool behaviour); 6. soak.

    ⚠️ **The pool has had no autoscaling at all since 2026-07-26.** A paused
    ScaledObject does not merely withhold scaling: KEDA 2.15 **deletes the
    HPA and stops running the scaler loop**, and `.status.externalMetricNames`
    never refreshes — so the object looks configured while nothing watches.
    Unpausing is therefore an improvement over today's posture **independent
    of T5**; do not let it wait on the T5 decision.
    
  • The writer's SIGTERM seal does not race in-flight ingest, and a crash does not lose the unsealed tail (noetl/ai-meta#209). Graceful restarts are covered by the #208 seal + terminationGracePeriodSeconds 90 s. An ungraceful stop (SIGKILL, node loss) can still lose the unsealed tail of the active part. On a command bus that is a loss class, and T5 removes the fallback that makes it recoverable.

  • out_of_order_appends exposed on writer :9102 (noetl/ai-meta#206) so a future feed-ordering regression is directly observable.

  • A bake window on organic (not synthetic) traffic.

  • Plaintext NATS credential retired in the same change set (noetl/ai-meta#188).

  • Explicit, separate human approval — T5 is irreversible.

Prod execution log — shastaratech prod

Four rounds. Round 1 held at canary on a topology blocker; round 2 chose the topology and failed the full flip on a real bug; round 3 shipped it; round 4 made it fast and made it survive a restart — and found that round 3's sign-off had already stopped holding.

Round Date Outcome
1 2026-07-26 Step 0 + shadow PASS (2-shard). Held at canary — per-shard user-pool blocker → #202. Rolled back to NATS.
2 2026-07-27 Option 2 chosen + kind-validated + prod reconfigured to 1 command shard. Shadow + canary PASS. Full flip FAILED — ~10% of commands ingested but never claimed, feed lag=0 → #203. Rolled back to NATS; prod clean.
3 2026-07-27 With the #203 fix (server v3.58.1 / worker v5.81.1): FULL FLIP SUCCEEDED — 80/80, 0 dup, 0 loss. Bus = EHDB end-to-end.
4 2026-07-30 #205 + #208 bundle (server v3.58.3 / worker v5.81.3 / ehdb 86a24f9). Latency p50 285–557 → 138.5 ms — now ~2.4× faster than NATS. Writer-restart survival PASS. ⚠️ On arrival, round 3's state was already broken — user dispatch silently dead ~2.4 days.

Round 1 — 2026-07-26 (Step 0 + shadow, held at canary)

T4 was run on the new shastaratech prod cluster (gke_shastaratech-noetl-prod_us-central1_noetl-prod-autopilot, project shastaratech-noetl-prod, ns noetl) — headless, synthetic load, explicit human GO. Reproducible IaC + the full executed-run record are in noetl/ai-meta under playbooks/194-l1-t4-prod-iac/ (README + step0-rollforward.sh, step2-shadow.sh, rollback.sh) and playbooks/194-l1-t4-prod-cutover.md. Soak windows were compressed to synthetic-load correctness gates (no organic-traffic window in one session).

What was different from this runbook (which was written single-shard)

  • 2-shard cluster — NOETL_COMMAND_SHARD_COUNT=2, NOETL_SHARD_SUBJECT_ROUTE=true; system pool split into noetl-worker-system-pool (shard0) + noetl-worker-system-pool-shard1. The command bus therefore needs one writer per shard (PublishRouter keeps one client per shard and errors on a missing shard; spawn_writer_host binds one ClaimCoordinator to NOETL_COMMAND_BUS_SHARD). Deployed noetl-cmdbus-writer-0 + -1, each a premium-rwo 20Gi PVC + ClusterIP.
  • GMP, not VictoriaMetrics — a GMP PodMonitoring covers both writers; live gates read the writer :9102 directly via port-forward.
  • Autopilot — storageclass premium-rwo (PD-SSD).
  • Images pulled from ghcr — server rolled to v3.58.0 (ghcr.io/noetl/server@sha256:99b842…), all worker deploys to v5.81.0 (ghcr.io/noetl/worker@sha256:27807d…).

Results

  • Step 0 (roll forward + writers, bus NATS) — PASS. server → v3.58.0, all 3 worker deploys → v5.81.0; both writers up, PVCs Bound; writers inert in nats mode. Gate: 15/15 hello_world COMPLETED over NATS, 0 restarts.
  • Step 1 (NATS baseline) — captured. command-path issued→claimed (n=90): p50 338 / p95 511 / p99 520 ms.
  • Step 2 (SHADOW) — PASS, 0 divergence. server + both writers → shadow, workers stayed on NATS. 20 controlled execs: server published 121 → NATS == 121 → EHDB; the EHDB feed grew by exactly 121 (shard0 +61, shard1 +60); 0 shadow errors ⇒ 0 missed / 0 spurious. Server connected to writers by DNS (no IP literal). Latency p50 273 / p95 526 / p99 579 ms — inside the NATS envelope. Writer RSS/CPU flat (3m / 4Mi), feed lag climbs monotonically (expected — nothing claims in shadow), 0 restarts.
  • Step 3 (CANARY) — HELD (not attempted). Blocker below.

Blocker — per-shard user-pool split (2-shard only)

On the EHDB bus every command is physically partitioned by execution_id across the shards — including the shared pool (proven in shadow: 61/60). On NATS today only the system pool is sharded (sharding.rs: ROUTABLE_POOL = "system"; the shared subject is unsharded). A worker claims from exactly one writer (NOETL_COMMAND_BUS_CLAIM_ADDR is a single host:port; one ClaimClient). The system pool is already split shard0/shard1 → each maps cleanly to its writer. The user pool (noetl-worker-rust) is a single deployment → flipping it to ehdb with one claim addr would strand ~half the shared commands (the ones on the other shard). This split does not exist in ops#241 and was never integration-tested — kind ran single-shard.

Two forward paths (human design decision → re-validate on 2-shard kind → prod)

  1. Split the user pool per shard (mirror the system pool): two deployments noetl-worker-rust-shard0 (claim writer-0:9101) + noetl-worker-rust-shard1 (claim writer-1:9101), each ~half the replicas, plus per-shard KEDA. Then canary one shard's user pool, then the other, then the system pool + server. Keeps the deployed #166 Phase 5 sharded command routing.
  2. Collapse the command-bus axis to a single shard — set NOETL_COMMAND_SHARD_COUNT=1 for the command-bus axis only (distinct from the server's event/state sharding), so all commands land on writer-0 and this runbook's validated single-shard / single-writer / single-user-pool design applies unchanged. Changes system-pool command routing from partitioned to competing (still exactly-once). Simpler; diverges from #166 Phase 5.

Tracked in noetl/ai-meta#202.

Hold state left on prod after round 1 (clean, reversible)

Images server v3.58.0 / workers v5.81.0 (validated); writers deployed but inert (NOETL_COMMAND_BUS=nats); bus on NATS (shadow rolled back after the parity gate); everything staged for canary once the #202 decision landed.

Round 2 — 2026-07-27 (Option 2 executed; full flip FAILED on a real bug)

Decision taken: Option 2 — collapse the command-bus axis to a single shard. The two sharding axes were proven independent in code (sharding.rs:172 "correctness never depends on affinity"; the command bus reads only NOETL_COMMAND_SHARD_COUNT, the state materializer only NOETL_RESULT_SHARD_COUNT plus its own event stream) — so collapsing the command bus does not touch #166 state sharding.

Kind re-validation first, on a rig mirroring prod-after-Option-2 (dedicated single writer + 2 state-shard system pools + NOETL_COMMAND_SHARD_COUNT=1): single writer holds all commands (ehdb_feed_total_lag→0); user pool + both state-shard system pools compete exactly-once (0 dup across 5 consumers); pod-kill redelivery 8/8 0-loss; rollback clean. Three config bugs caught here rather than on prod:

  1. the two system pools need a shared NATS consumer, or NATS double-delivers while the bus is still nats;
  2. the writer needs its NATS_STREAM set;
  3. a writer in ehdb mode also needs NOETL_COMMAND_BUS_CLAIM_ADDR — it is a consumer of its own shard too.

Prod reconfigure to 1 command shard (bus stayed NATS) — clean. Both system pools unified onto a shared broad consumer (noetl_worker_system_rust, noetl.commands.system.>), server NOETL_COMMAND_SHARD_COUNT=1, writer-1 scaled to 0. #166 state sharding INTACT — both pools keep NOETL_SHARD_COUNT=2 / index, STATE_BUILDER=offserver, STATE_SHARD_WRITE=true; state builder healthy. Gate 12/12.

Then shadow PASS and canary PASS — but the full flip FAILED. Under the full flip roughly 10% of commands were ingested into the writer feed and never delivered to any claimer, and feed lag read 0 the whole time, so the loss was silent. Rolled back to NATS; prod left clean. Filed as noetl/ai-meta#203.

The #203 bug and its fix (read this before running a cutover)

Root cause — an out-of-order append versus the feed cursor. The command feed's producer (noetl-server) assigned each record's sort key itself — the command's snowflake event_id — and the single writer trusted it. But the Dataset contract the feed cursor and range pruning depend on is "records are appended in ascending sort_key order within a partition". Snowflake ids are sparse, and under concurrent publish a lower id can reach the writer after a higher one. Then:

  1. the follower cursor has already advanced to the higher key (ChangeFeed::poll sets cursor = max sort_key read);
  2. refill reads read_partition_after(shard, cursor) → only > cursor → the late lower key is filtered out, never enters pending, never delivered;
  3. lag() is cursor-relative too, so the lost record — below the cursor — is never counted. Silent loss, lag=0.

Fix — the writer assigns the ordering key (append_writer_assigned), so the ascending contract is guaranteed by the only component that can guarantee it. noetl/ehdb#300 → ehdb d4b6235 → noetl-worker v5.81.1 (the writer host lives in the worker, so this is the image that carries the runtime fix) + noetl-server v3.58.1. A new out_of_order_appends counter records the condition — not yet exposed on :9102, tracked as noetl/ai-meta#206.

Kind re-validation under sustained load, identical 40-concurrent burst:

delivered stuck (command.issued, no command.claimed) feed lag
pre-fix 17/40 23 permanent 0 (silent)
post-fix (v5.81.1) 41/41 0 0

Plus exactly-once 335=335 (0 dup), graceful pod-kill redelivery 25/25, #166 state sharding intact, rollback to nats clean.

Round 3 — 2026-07-27 (FULL FLIP SUCCEEDED — bus = EHDB end-to-end)

Fixed images deployed by digest (server v3.58.1 @sha256:5a73f4a5…, worker v5.81.1 @sha256:db156fa6…, copied ghcr→project AR with crane), then the same staged Option-2 sequence re-run:

  • Step 0 (bus NATS, roll forward): 20/20.
  • Shadow: 0 divergence — 123 published == 123 in the feed on the single writer; writer flat.
  • Canary (user pool → ehdb): 30/30; pool isolation — 0 system claims by the user pool (the hard requirement); graceful pod-kill redelivery 12/12, 0 loss.
  • Full flip (both system pools + server → ehdb): 30/30 + 50/50 soak = 80/80, 0 dup, 0 loss; server NATS_pub=0; feed lag → 0; writer 0 restarts; #166 state sharding intact throughout.

The ~10% silent loss that failed round 2 was gone at the exact same stage. #203 closed as prod-validated.

State left on prod (current)

The production command bus is EHDB. Server + all worker pools run NOETL_COMMAND_BUS=ehdb on server v3.58.1 / worker v5.81.1, single command shard, one writer on a premium-rwo PVC.

NATS is still installed and is the rollback path — one command:

kubectl -n noetl set env deploy/noetl-server-rust \
  deploy/noetl-worker-rust deploy/noetl-worker-system-pool \
  deploy/noetl-worker-system-pool-shard1 NOETL_COMMAND_BUS=nats

bash rollback.sh in playbooks/194-l1-t4-prod-iac/ reverts the rest.

T5 (delete NATS) NOT done — separate, irreversible approval; see the T5 checklist above. It couples noetl/ai-meta#188 (the plaintext NATS credential retires with NATS) and is gated on noetl/ai-meta#205 (dispatch latency).

Latency — the one number that got worse

Measured on the same cluster and the same synthetic load, command issued → claimed:

Bus p50 p99 min
NATS (baseline) ~200 ms — —
EHDB (post-flip) 285–557 ms (load-variable) ~1040 ms ~145 ms

Roughly 1.5–2.8× the NATS baseline. Sub-second, and it affected neither correctness nor completion — but it is a real regression on the hot dispatch path, and it is the open go/no-go item for T5. Two plausible contributors, not yet isolated: the 250 ms poll-claim interval versus NATS's push delivery (a push variant is already possible — the FeedWriter::tip_receiver() seam exists), and the writer-side ordering-key append the #203 fix added. Tracked in noetl/ai-meta#205.

Superseded by round 4 (2026-07-30). Both hypotheses above were wrong. The cost was neither the poll interval (claim delivery measured 16–33 µs, flat) nor the #203 ordering-key append: it was an fsync held inside the engine lock (~4 ms, capping the bus at ~230 cmd/s) multiplied by the server's publish mutex held across the round trip. Group commit + off-lock fsync + pipelined publish took prod p50 to 138.5 ms, below the NATS baseline. See round 4.

Round 4 — 2026-07-30 (latency fixed, restart survival fixed; the bus is now faster than NATS)

One rollout carrying two fixes, so prod deploys once and both gates are measured in the same window. NOETL_COMMAND_BUS=ehdb before and after — a rolling image update on the live bus, not a bus change.

Component Version Carries
ehdb engine 86a24f9 #301 group commit + off-lock fsync + pipelined publish; #302 writer-restart survival
noetl-server v3.58.3 server#289 (pipelined publish), server#290 (publish retry)
noetl-worker (all pools + writer) v5.81.3 worker#195 (group-commit adoption), worker#196 (resume + seal)

Post-deploy: 6/6 pods Ready, 0 restarts; #166 state sharding intact (NOETL_SHARD_COUNT=2, index 0/1, offserver, STATE_SHARD_WRITE=true); NOETL_COMMAND_SHARD_COUNT=1, single writer.

⚠️ Found on arrival: round 3's sign-off had already stopped holding

Before any deploy, user-playbook dispatch was dead and had been for ~2.4 days. noetl-cmdbus-writer restarted 2026-07-28T07:49Z; both noetl-worker-rust pods' last log line was from that instant —

EHDB claim connect failed; retrying claim_addr=…:9101 error=Connection refused (os error 111)

— then silence. Synthetic executions issued and were never claimed (ehdb_feed_total_lag climbing, executions stuck at command.issued). noetl-worker-system-pool-shard1 happened to restart later (07-29 16:17), so system commands kept completing and masked the outage completely.

Two things to carry forward:

  1. A cutover sign-off is only as strong as its restart story. Round 3 proved 80/80 with zero loss and was correct about that — it just never restarted the writer. T4's "complete" did not survive the first one.
  2. Silent single-pool dispatch death is not observable today. System-pool traffic keeps every cluster-level signal green. Per-pool lag (ehdb#303) is the series that would have caught it — which makes #303 an observability fix as much as a scaler input.

Gate 1 — dispatch latency (#205): PASS

command.issued → command.claimed, read off noetl.event:

Regime n p50 p95 p99
NATS baseline (round 3, 07-27) 90 338 ms 511 ms 520 ms
EHDB at round 3 (07-27) — 285–557 ms — ~1040 ms
EHDB, round 4, unsaturated 60 138.5 ms 156.4 ms 181.0 ms
EHDB, round 4, post writer-restart 45 140.5 ms 191.6 ms 208.0 ms

p50 138.5 ms against the 338 ms NATS baseline — ~2.4× faster than NATS, and 2–4× faster than EHDB was at round 3. The loopback attribution (281 → 6.7 ms p50 at 64 publishers / 24 claimers; 225 → 6983 cmd/s) holds on real infrastructure. Issue closed.

⚠️ Measure unsaturated, or you measure the wrong thing. 90 commands fired in 1.2 s gave p50 2123 / p99 3735 ms — that is queueing behind the user pool's 8 slots, not bus latency. This number is why the T5 autoscaler picked targetValue: "2" (a replica per 2 queued commands) rather than waiting for backlog to exceed capacity.

⚠️ kind cannot reproduce this measurement. Its worker pool peaks at 28–131 cmd/s against a ~1000 cmd/s bus ceiling, so the publish queue the fix removes never forms; run-to-run variance there exceeded any before/after signal. Prod, unsaturated, is the only rig that resolves it.

Gate 2 — writer-restart survival (#208): PASS

Writer pod deleted mid-stream under continuous load:

  • 38 executions accepted → 38 COMPLETED, 0 stuck.
  • Commands: 40 before + 74 after = 114 = 38 × 3 exactly, zero duplicates, zero loss.
  • Redial ~2.7 s: writer died 18:18:22; the worker logged EHDB claim_next failed; reconnecting to the claim coordinator … early eof at 18:18:22.056 and retried every ~250 ms; coordinator up 18:18:24.765. The pre-fix code emitted no such line — it wedged.
  • Cursor-resume, not replay, proven by the exact 114 = 38 × 3 accounting.
  • Steady-state latency unchanged across the restart.
  • terminationGracePeriodSeconds raised 30 → 90 s for the SIGTERM seal.

Two rough edges left open

  1. The resume signal is not readable on its own. The resume took the durable path (origin="persisted"), but the logged cursor reads 0 and ehdb_feed_shard_committed reset 408 → 165 across the restart — the counter is segment-relative, not a global monotonic offset. No replay happened, but the log line alone cannot distinguish "resumed" from "replayed from 0" — exactly the signal this runbook would lean on during an incident. → ehdb#304, held.
  2. 2 × HTTP 500 on POST /api/execute during the writer's absence — the #290 publish retry (3 × 250 ms) does not span a ~2.7 s pod restart. Fail-closed (the caller is told; every accepted execution completed), but not transparent. → server#291, held, widening the retry to a 10 s wall-clock deadline.

Monitoring note (correcting a common assumption)

Prod monitoring is Google Managed Prometheus, not VictoriaMetrics — the vmservicescrape CRD is not installed on this cluster and there is no vmstack namespace. The writer scrape is PodMonitoring/noetl-cmdbus-writer (10 s, cmdbus-lag port), created during the T4 cutover and never committed to noetl/ops until ops#242. Consequence for the T5 autoscaler: there is no in-cluster PromQL endpoint, so the KEDA trigger is metrics-api scraping the writer's :9102 directly rather than a prometheus trigger. That also keeps autoscaling independent of the monitoring pipeline's health, which is the right coupling for a T5 prerequisite.

Confirmed live on prod: ehdb_feed_total_lag 0, ehdb_feed_shard_lag{shard="0"} 0, ehdb_feed_shard_committed{shard="0"} 5853, up{job="noetl-cmdbus-writer"} 1.

NATS untouched — ns nats and ns nats-supercluster still installed. T5 NOT performed. Playbook + rollback recipe: playbooks/205-208-prod-rollout/README.md in noetl/ai-meta.

Related

Clone this wiki locally