Repository navigation
Plan Tier Failure Domain
Resilient-core Phase 3, ops half. Overview: Architecture — resilient KV core · coordination ai-meta#339
⚠ Stopped before applying, deliberately. This needs a data migration and is not cleanly reversible once the switch is made. The instruction was to bring the plan rather than apply in that case.
NOETL_EVENT_BUS_WRITER_DIR = /data/eventbus ← volume "eventbus-data"
NOETL_EHDB_TIER_SERVICE_DIR = /data/eventbus/ehdb-tier ← the SAME volume
The tier's "replica" is a subdirectory of the volume holding the thing it is a
copy of, so replication buys no independent failure domain.
failure_domain.rs names the consequence: an RF of N over one domain is an RF
of 1 wearing a larger number.
/data/eventbus 2.4 G
/data/eventbus/ehdb-tier 1.8 G ← the tier is most of the volume
files under ehdb-tier 3 ← few, large parts
volume 48.9 G, 5% used, 46.5 G available
Good news for the mechanics: 3 files, 1.8 GB, and the destination has room.
- ⚠⚠ The writer hosts BOTH buses.
noetl-cmdbus-writer-0serves the command bus (:9100-02) and the event bus (:9103-08). Any restart of it stops dispatch. This is the pod whose unschedulability caused the 55-minute outage in ai-meta#323. - ⚠⚠ That outage is this exact shape. #323 added a volume for a PVC that had
never been created; the pod could not schedule. The PVC must exist and be
Boundbefore any manifest references it. -
The copy is not consistent while the writer runs. The tier is being
appended to, so a live
cpcan capture a torn part. The writer must be quiesced, or the copy re-synced after a pause. -
Reverting after the switch loses writes. Once
NOETL_EHDB_TIER_SERVICE_DIRpoints at the new volume, anything written there is not on the old path. A revert must copy back, so "revert" is a second migration, not a flag flip.
✅ One thing that is not a blocker: the writer uses volumes referencing
pre-created PVCs, not volumeClaimTemplates (which is empty). Adding a volume
to the pod template is therefore a normal, mutable update — no
delete-and-recreate of the StatefulSet.
Each step is separately revertible up to step 4; step 4 is the one-way door.
Step 1 — create the PVC (additive, zero impact).
kubectl --context "$PROD" -n noetl apply -f - <<'YAML'
apiVersion: v1
kind: PersistentVolumeClaim
metadata: { name: noetl-ehdb-tier-0-data, namespace: noetl }
spec:
accessModes: [ReadWriteOnce]
storageClassName: premium-rwo
resources: { requests: { storage: 20Gi } }
YAML
kubectl --context "$PROD" -n noetl get pvc noetl-ehdb-tier-0-data # MUST be BoundNothing references it yet, so this cannot affect the running system. Revert: delete the PVC.
Step 2 — mount it (rolls the writer; ~30-60 s of dispatch downtime).
Add a 4th volume + mount at /data/ehdb-tier. Do not change the env var yet —
the tier keeps writing to the old path, so the new volume is inert and the
mount is provably harmless before anything depends on it.
Full-spec kubectl diff --server-side must show exactly the added volume and
volumeMount, nothing else.
Revert: remove the volume + mount (rolls again).
Step 3 — quiesce and copy. With the writer paused (or during a known-idle window), copy 3 files / 1.8 GB, then verify by digest, not size:
kubectl -n noetl exec noetl-cmdbus-writer-0 -- sh -c \
'cp -a /data/eventbus/ehdb-tier/. /data/ehdb-tier/ && \
find /data/eventbus/ehdb-tier -type f | sort | xargs md5sum | md5sum && \
find /data/ehdb-tier -type f | sort | xargs md5sum | md5sum'The two digests must match. Revert: delete the copy; the original is untouched.
Step 4 — switch the env var (the one-way door).
NOETL_EHDB_TIER_SERVICE_DIR=/data/ehdb-tier, rolls the writer.
Revert: switch back and copy any new writes back, because the two paths
diverge from this moment.
Step 5 — verify, then reclaim.
Confirm tier reads still answer and parity is no worse; only then delete the old
/data/eventbus/ehdb-tier — a separate, later decision.
Worth putting to the owner, because it changes the risk profile completely:
The tier is a mirror, not an authority. Postgres holds every event. So an alternative is to start the new volume empty and let the tier refill — skipping steps 3 and 4's consistency hazard entirely.
⚠ The cost is that historical tier content is gone until repaired, and tier coverage is already poor: 48% of recent executions have missing events (ai-meta#342), and per earlier measurement ~95% of the log predates the tier's enablement. So the data being migrated is of limited and unknown value.
Recommendation: decide #342 (the mirror drops) first. Migrating 1.8 GB of a mirror that is currently losing 39.6% of what it is handed moves a known-bad copy onto better storage. Fixing the mirror, then letting a clean tier refill on the new volume, is both safer and less work.
The owner chose refill, not migrate, and deferred the change until the mirror is clean.
Rationale, in one line: migrating 1.8 GB of a mirror that is currently losing 39.6% of what it is handed relocates a known-bad copy.
So the sequence is:
- Land and arm the durable mirror repair (noetl/ai-meta#342, noetl/server#426) and confirm the tier converges to Postgres.
- Then create the new PVC and mount it empty, and let the tier refill from the authoritative log.
This drops the two hazards that made the original plan a one-way door:
- No consistency window. Nothing is copied, so there is no torn-part risk and no need to quiesce the writer for a copy.
- No copy-back on revert. A fresh volume that refills is discardable; the revert is "point the env var back and delete the PVC", with Postgres still holding every event.
Steps 1 and 2 of the sequence above (create the PVC, mount it inert) stay valid
and stay safe in that order — the PVC must be Bound before anything
references it, which is the #323
outage shape.
⚠ Still needs a window for a writer restart, because the writer hosts both buses. That remains an owner call.
- Home
- Architecture
- Architecture — the four engines
- Architecture — resilient KV core
- Consistency Invariants (per tier)
- Roadmap
- Sessions Log
- Claude Handoff
- RFC: Completion Program
- RFC: External EHDB Driver
- L1 Command-Bus Cutover (T4/T5 — prepared, human-gated)
- Prod Cutover — Event-Log Tier (Phase 9, Tier 1)
- Runbook: Async Event-Log Mirror
- Durable Event-Log — Prod Durability Sign-off (§C, slice 6)