docs: B3 — the unsealed tail replicates off-box, and the lock boundary
New page Design-Tail-Replication: the measured gap (prod seal interval 900s
with oldest_unsealed_age 719 and nothing waiting to upload), why per-batch
objects rather than re-uploading the part, the three invariants and what
each prevents, the RED-control-first evidence, and the six series.
Also records the v0.8.1 lock-boundary fix, because it is the kind of thing
someone adding another timer-driven substrate method needs to read first:
the remote put must not run under the engine lock the live append path
takes.
Framing kept honest throughout: the tick IS the loss window, so B3 takes it
from 900s to 15s — bounded, not zero.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: replica backfill — attaching a replica replicates nothing existing
New page Design-Replica-Backfill: the one enqueue site (on seal), the prod
measurement (40 parts in the remote manifest, 39 at replica_count 1 while
survives_node_loss read 1), why the remote was worse than empty, the
local_path-is-None trap under the fix, and why backfills carry their own
counters. Cross-linked from Home Core Pages and the sidebar.
Releases table was stale — v0.5.1 and v0.5.2 were missing; added those
plus v0.6.0 and v0.7.0.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs(home): auto-release is ON — and is not auto-deploy
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs(home): v0.5.0 — the first automated ehdb release, and where it is consumed
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
wiki: the north-star distributed-registry direction, and P1 landed
Records EHDB as NoETL's distributed registry and control plane, with the inventory that
corrected three of my own prior claims: D8 is in the production crate rather than unwired,
vectors already exist rather than being a greenfield hard part, and four of the seven
capabilities are substantially built. The real gaps are watch, time-based leases, ephemeral
registration, secret-reference types and unsealed-tail replication.
States the distributed lower bound honestly: a coordination-free data path gives CALM-style
availability and NOT linearizable cross-shard reads, so "etcd semantics, distributed better"
is wrong. P7 (tail replication, still RF=1) and P8 (election + fencing) are named as the
hard parts rather than vectors or the registry.
P1 landed: wall-clock TTL liveness, with the upgrade hazard pinned -- a timestamp-less
record is unknown, not dead, because the alternative evicts the whole fleet on upgrade.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
wiki: the first measurements of the production L0 engine
ehdb-l0 had 88 .rs files and ZERO benches while ehdb-reference had 22 and two, so every
number on the existing performance page is about the reference model rather than the engine
on the path. These are about the engine.
⭐ The two agree where they overlap: that page measured durable append as fsync-bound at
~3.9 ms / ~256 appends/s, and the production engine measures 3.749 ms / 267 appends/s --
within 4%, from a different crate and harness.
Two new results. The fsync is ~95% of posture-A write cost: 267 rec/s fsync-per-append
against 5 450 group-committed at batch 80, a 20.4x difference, and group commit needs
FlushPolicy::CallerDriven or batching costs MORE. And a read steps 15.6x per record at
seal_max_records = 1024, so a chain over 1024 events costs ~13x what extrapolating from the
sub-1024 figures predicts -- filed as noetl/ai-meta#453, where whether the step is
acceptable is an open judgement call rather than an assumed defect.
Also records the method: three of three append instruments were wrong before they were
right, each caught by its own planted-effect control, including one where the instrument was
wrong and not the engine.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
docs: CORRECTION — v0.4.4 is rolled back, and a core assumption of the design is false
noetl/ai-meta#362.
populate_from_log recomputes chain edges from log order and treats that order as
a stable prefix that only grows at the tail. An event can commit into the MIDDLE
of an ORDER BY event_id read, because a snowflake event_id is minted before the
insert — so commit order is not id order. Every position after the interior
insert then reads as a content conflict.
Proven by position arithmetic: the store held at position 72 what the log holds at
position 73 — exactly one missing interior row.
⚠ No available column is commit-ordered. created_at is stamped at MINT time, so
it always agrees with event_id order and cannot detect the reordering; I ran that
comparison as a refutation and it refuted nothing.
v0.4.4 stands — both defects it fixed are real — but it was not what was breaking
prod, and the earlier 5-in-79 was most likely this same mechanism.
docs: v0.4.4 — a behind snapshot is staleness, and the chain store is live in prod
noetl/ai-meta#360. Design-Execution-Chain-Store + Sessions-Log + Home.
New section "Staleness is not a conflict — and the same shape is a second
writer", which is the design point of this release: the crate REPORTS
`FromLog::StaleLog` as unresolved rather than deciding, because from a single
snapshot benign staleness and a foreign writer are indistinguishable, and only
the consumer can re-read the log to tell them apart. My first version decided
it, and the noetl/ai-meta#358 regression test failed on the first run.
Also recorded: a content conflict now revokes serve-trust (previously this
function wrote nothing, so the still-matching coverage handed a
`chain_if_authoritative` caller a chain just proved wrong); coverage is
monotonic; and the monotonic-coverage mutant SURVIVED its first test, which
populated 5 then read 2 and returned StaleLog before the guard was ever reached.
⚠ Corrections: the status line said "flag-gated, default OFF, not enabled
anywhere" and the Home link said the store "cannot serve yet". It is ARMED IN
PRODUCTION since 2026-09-29 17:30Z on server v3.117.2 — 772 comparisons, 0
failures. Version floor raised v0.4.3 -> v0.4.4, because v0.4.3 carries the
false-divergence classification that rolled back the first ramp.
Design: the execution-partitioned chain store, and why it cannot serve yet
New page for ehdb-l0's DurableChainStore + ChainPopulator (RFC
noetl/ai-meta#355). Written from the diff and from the kind measurements,
not from the RFC's summary.
The load-bearing parts, recorded because each is a guard whose purpose
someone would otherwise delete as noise:
* the three states the watermark separates — `None` never means "no
events", `Some(vec![])` does;
* the ASYMMETRIC append/watermark ordering, and the defect that produced
it (marking before every append left Authoritative{1,1} over zero
events on the common mid-flight-arming path). v0.4.1 carries that bug;
v0.4.2 is the first tag without it;
* `apply_replicated` deliberately does NOT enforce the head, because async
replication delivers out of order — found by a mutation test that could
not construct a gap at all;
* `chain_is_complete` compares against the partition SPAN, so it catches a
hole in the middle and NOT a missing beginning — which is exactly why
the open defect is undetectable from inside the store.
⚠⚠ Both measured defects are stated with their numbers and the shared
cause (a chain edge stamped from an in-memory map that does not survive a
restart), so nobody reads this page and concludes the store is ready to
serve.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
wiki: link the resilient-KV-core architecture overview from Home + sidebar
Embedded-state foundations: format gate, stored cursor, D3/D8 exercised, append validation, foca
Records what noetl/ai-meta#332 added to THIS repo, and for each guard what it
exists to prevent -- so it is not later deleted as noise. Includes the bugs the
work's own tests found: the format guard's blanket Err(_) => Ok(None) that turned
every read failure into 'absent', and the dot-run check that is load-bearing for
exactly one shape because dots are legal in an id.
wiki: record the F1-F5 remediation and refresh the invariants
The Consistency-Invariants page said 'detected by nothing' for two invariants
that now have detectors, which is exactly the drift that page exists to catch.
I-EL-2 (single writer) is now DETECTABLE but still not enforced: fencing runs in
shadow, counting stale-epoch writes while letting them through, and the election
issues tokens without being authoritative. The ordering hazard is recorded on
the page itself -- enforcing before the election issues real tokens is an outage
rather than a degradation, because with no election every writer's epoch is 0.
I-EL-4 (reaching the substrate) now names ehdb_l0_unreplicated_age_seconds as
its detector, live on the writer's /metrics since worker v5.125.0, and keeps the
two misleading readings explicitly labelled: upload_lag_micros_total is
seal-relative and blind to the dominant term, and ehdb_feed_shard_lag is consumer
backlog wearing an adjacent name.
Also records that prod as it stands would FAIL the new replica-domain check --
the tier dir is nested inside the writer dir on one PVC -- so validating at open
before fixing the layout is a startup outage by construction.
Home carries the summary above the fold; Sessions-Log gets the dated entry.
Refs noetl/ehdb#324, noetl/ehdb#328, noetl/ehdb#329, noetl/ehdb#330, noetl/ehdb#331, noetl/ehdb#332
wiki: the four-engine architecture + per-tier consistency invariants
ehdb#323, under #324. Two new pages, plus Home and the sidebar reconciled to
the narrowed scope.
Architecture-Four-Engines documents the four owned engines with a diagram, and
records WHY vector and OLAP collapse into projections -- matching what is built
rather than what was promised: ehdb-reference/src/vector.rs is already a bounded
cosine search over the collection's live points, so there is NO ANN index to
remove. The gap was in the promise. It also separates three vocabularies that
are routinely conflated: engine vs tier vs StoreTier/QueryTier.
⚠⚠ Both pages surface a discrepancy the Backend-Configuration page invites:
it presents five tiers each with an off/shadow/primary mode, which reads as
though setting primary makes any of them serve. SERVE_WIRED_TIERS is
["eventlog", "projection"] -- KV and object accept primary as CONFIGURATION and
cannot serve as CAPABILITY. Home now says so above the fold.
Consistency-Invariants states, per tier, what each guarantee is, WHAT FORCES IT
to hold, and HOW YOU WOULD FIND OUT if it stopped -- because a guarantee with no
enforcement and no detector is a description, not an invariant. Three come out
badly and are labelled rather than smoothed over: I-EL-2 (single writer) is
enforced by nothing stronger than replicas: 1 and detected by nothing;
I-EL-4 (reaching the substrate) is bounded by volume not time and detected by
nothing that works, since upload lag is measured from seal and feed lag is
consumer backlog; and in prod the substrate is the SAME PVC as the primary, so
it buys no independent failure domain.
Home's mission no longer claims EHDB will absorb Qdrant and ClickHouse.
Refs noetl/ehdb#323, noetl/ehdb#324, noetl/ehdb#320, noetl/ehdb#321, noetl/ehdb#322
wiki(#155): the async event-log mirror, and correct the stale "primary not activated" claims
Adds Runbook-Async-Event-Log-Mirror — the operational surface for the
mirror that went asynchronous on prod 2026-08-19, and which this wiki
did not mention at all. Leads with the one-flag rollback and with the
pairing that makes it dangerous: async on with the tolerance at 0 makes
the comparator judge a healthy tier on its own liveness.
Records the parts that are easy to get wrong rather than only the
settings. That the window must be MEASURED and then set above the
observed p99 — prod came back p99 699ms, so the window is 30s. That it
sits deliberately below SETTLE_SECS so the background sampler still
compares every execution. That the two lag controls are a
discrimination rather than a clean-parity check, because a comparator
taught to ignore everything passes the naive version perfectly. And
that reserve rather than send is load-bearing, because the first
implementation lost 32 of 60 records while every counter still read
correctly.
Corrects two current-state claims that had gone stale:
* Home said the reference tiers are O-of-n per op and "not
primary-serve". Still true for KV/object/vector; false for the event
log since the runtime cache landed and since the tier went primary on
2026-08-13.
* Roadmap said the primary cutover "stays a later step" and that primary
is "recognised but not activated". Both annotated as superseded
rather than deleted, since they remain accurate as the plan of record.
Sessions-Log is left untouched — it is history and correct as written.
Home gains a dated current-state banner so a reader meets the live state
before the older planning prose.
Also records that Option 2's batch flag is still OFF on prod, batch=0 on
both replicas, so the substrate is landed and measurable but inert.
Tracked at noetl/ai-meta#284.
wiki(L1 round 4): prod deploy — EHDB bus ~2.4x faster than NATS (#205 closed) + writer-restart survival (#208 closed); T5 gate is now the user-pool autoscaler (#210); GMP-not-VM correction
wiki(L1 #205): dispatch latency root-caused + fixed; T4/T5 status corrected
- Program page: T4 row corrected (full flip SUCCEEDED on shastaratech
prod 2026-07-27 round 3, was still 'PREPARED, NOT EXECUTED'); T5 row
now names its two blockers (#205 latency, #208 writer-restart survival);
new section on the #205 attribution + fix.
- Sessions log: 2026-07-28 entry with the per-hop attribution, the
three-part fix, the loopback before/after, the kind correctness PASS,
and why kind cannot resolve the latency delta.
- Home: implementation note for the fix.
Refs noetl/ai-meta#205, noetl/ai-meta#208
L1 T4 COMPLETE — EHDB is the production command bus (full flip 80/80 on shastaratech prod)
Runbook: status banner rewritten to SUCCESS; T4 go/no-go checklist cleared
with the latency caveat + a new T5 checklist; prod execution log restructured
into three rounds (held-at-canary / failed-on-#203 / succeeded) with the #203
root cause + fix + before-after table, the post-flip prod state, the rollback
command, and the latency regression.
Program page: L1 status T0-T4 done / T5 held; T4+T5 table rows rewritten.
Home: program bullet leads with prod-command-bus state.
Sessions-Log: new 2026-07-27 entry.
Refs noetl/ai-meta#194, #202 (closed), #203 (closed), #205, #206
Refine boundary to write-behind-cache model; re-scope D5/D6 issues (+#198 sink)
EHDB is never the durable system of record for business data (customer's
connector-backed store is), but MAY hold business processing context transiently
as a write-behind cache: two roles — (a) permanent control-plane log stays lean/
reference; (b) transient processing cache bounded + sunk to the customer store +
evictable. Supersedes the 'never holds business payload' absolute. Re-scopes the
fixes: #195 reframed (keep PERMANENT noetl.event lean; transient cache keeps full
context for the drive — dissolves the earlier drive-breaking risk since the drive
reads the transient cache not the permanent log); NEW #198 the write-behind sink
path (transient context -> customer system of record + sink-confirmation-gated
eviction; builds on connector tools + result-tier byte-source + GC); #196
reframed (permanence in noetl.command); #197 unchanged.
Add noetl-only scoping invariant + ClickHouse meta-catalog to L0 design
Program invariant (recorded prominently): EHDB is noetl-internal-only, a FIXED
set of predefined datasets (D1-D10: event log, commands, projections, KV,
object/blob, vector, catalog, runtime, system-WASM, provider-facts), NOT a
general DB — no arbitrary schemas/DDL, no query planner, no hosted-DB surface;
business data never in EHDB. It SHRINKS every layer. L0 design now = VM/VictoriaLogs
write engine + a ClickHouse-MergeTree-style meta-catalog (manifest + sparse index
-> object-store pointer; fixed-schema table-format-over-object-storage). Scope
cuts: no general inverted index (few fixed per-dataset indexes only); L3 cut from
a SQL engine to fixed reads + read-only export (~2-3q -> ~1q). Umbrella:
https://github.com/noetl/ai-meta/issues/194
Reframe program page to layered platform (L0-first); NATS takeover = L1
Major reframe: EHDB becomes a layered platform built L0-first — L0 replicated
object store (modeled on VictoriaMetrics/VictoriaLogs: buffered flush ->
immutable parts -> background merge -> inverted-index-first -> columnar) -> L1
streaming (the NATS takeover) -> L2 KV -> L3 append-log SQL. L0 resolves the HA
debate: durability+replication from the object store => fungible writers,
per-shard-Raft (T-RF) retired, only a light L1 ordering lease remains. Records
the honest departure (VM keeps the hot path on local disk; object storage is
backup-only, so L0's live object-store durability tier is net-new beyond VM),
the reuse assessment (#254 segments approx immutable parts; gaps = inverted
index + merge engine + object-store tier), per-layer cost, and the first L0
slice (L0.1). Master RFC: docs/rfc/ehdb-layered-platform.md. Umbrella:
https://github.com/noetl/ai-meta/issues/194
Program page: lock HA posture — shards-only now, RF deferred (replication-ready)
Verified parity (prod NATS replicas:1, streams R1). Ships (c) single-writer-
per-shard; per-shard replication factor deferred to a named non-blocking phase
(T-RF). Adds the replication-ready seam (the #254 durable_eventlog_affinity
single-writer + open_read_only followers + durable_eventlog_shared log-shipping
+ durable external cursors => T-RF additive, no log/cursor rewrite), the
failover sequence + seconds-scale loss-free stall estimate, and the worker
reconnect-on-failover story (unchanged StatefulSet DNS). Umbrella:
https://github.com/noetl/ai-meta/issues/194
Program page: revise to (c) per-shard-writer-as-broker topology
Supersedes (b) co-locate on latency (two delivery hops). (c): the stateful
per-shard writer owns its shard's durable log AND delivery; workers subscribe
directly to the writer; stateless server publishes to the writer and is out of
the delivery path -> one delivery hop, matching NATS. Records the honest premise
correction from the code trace (#166 sharding dormant; system pool is a
single-replica Deployment, no StatefulSet/PVC/identity; drive hands back to the
server), so (c) is a build not a relocation (~2q, more than (b)). Adds the
alternatives ledger, honest downsides (connection fan-out, per-shard HA,
concentration on the #163 OOM pod), revised cost, and the (c)-shaped T0 spec
with a latency-vs-NATS comparison. Umbrella:
https://github.com/noetl/ai-meta/issues/194
Program page: lock (b) co-locate topology for EHDB->NATS takeover
Records the locked topology (2026-07-15): durable log + change-feed stay in
the per-shard writer (#166, stateful); the noetl-server stays stateless (#115)
and owns delivery only (tails the change-feed, pushes over the SSE hub); option
(a) server-embeds-EHDB rejected. Adds the three-component split, re-derived
Track T phases, revised cost (~1.5 quarters), and the (b)-shaped T0 spec.
Program umbrella: https://github.com/noetl/ai-meta/issues/194
Add EHDB→NATS takeover program page (decision + plan, nothing built)
Records the locked decision (2026-07-15): remove NATS via noetl-server-owned
push reusing the gateway SSE ConnectionHub over EHDB's durable log +
change-feed; standalone ehdb-server broker rejected. Track S (storage cutover,
in flight) + Track T (transport, server-owned-push, T0->T5, unbuilt), revised
cost, three open sub-decisions, T0 slice spec. Cross-linked from Home +
_Sidebar. Program umbrella: https://github.com/noetl/ai-meta/issues/194
docs(rfc): external EHDB driver — outward-facing read access (design)
Design-first RFC for an external driver/client so third-party apps and
BI tools can read EHDB's engine tiers. Recommends Arrow Flight + Flight
SQL over a dedicated data-plane endpoint fronting the worker (loose
coupling preserved), scoped read-only API tokens, committed-only reads.
Five decision forks surfaced. Grounds on the existing ehdb-service
Flight scaffold + the #178 internal read contract. Registered in
_Sidebar + Home.
Tracks noetl/ai-meta#184.
docs(perf): add Design-Performance-and-Load-Testing + baseline; Sessions-Log, Home, Roadmap, _Sidebar
docs(durable-eventlog): slice 6 prod-durability sign-off package (ehdb#254)
New Runbook-Durable-EventLog-Prod-Signoff page — go/no-go checklist (D1-D11),
evidence bundle (slices 1-5), residual-risk register (R1-R7), and the extended
durable-shadow -> durable-primary prod rollout sequence. Cross-linked from the
tier-1 cutover runbook (§C update + Related), the design page, and _Sidebar;
Sessions-Log + Home updated. Docs/planning only — no prod action; ehdb#254
slice-6 box stays unchecked.
docs(subject-digest): extend the subject-length fix to KV + vector tiers (ehdb#259 / worker#172)
KV + vector registry subjects now use a fixed-width SHA-256 digest token
(noetl.kv.<bucket>.<sha256hex(key)>, noetl.vec.<sha256hex(col)>.<sha256hex(pt)>)
instead of hex-of-full-id, matching the object tier — bounding every per-id
subject under the 256-char Subject cap. Adds a 'Subject digest tokens (all
three tiers)' section to the Phase-8 design page, updates the KV/object/vector
sections + object callout, prepends a Sessions-Log entry, and notes the
hardening on Home. Forward-safety (no live breakage today); in-kind re-proof
deferred. Refs noetl/ai-meta#241.
docs(live-in-kind): KV mirror PROVEN on live spool-circuit runtime — all four wired tiers now fire on live drives (worker v5.69.0)
Stood up a dedicated subscription runtime (WORKER_MODE=subscription, v5.69.0 +
shadow) activating subscriptions/spool_outage_stream (spool + http-probed
downstream). SpoolRuntime::persist_circuit → kv::mirror_live_put fires every ~2s
probe tick + on circuit open/close: noetl_ehdb_kv_ops_total{operation="mirror",
outcome="mirrored"} climbed 10→50+, kv_last_ok=1, zero invalid. Ref-store circuit
record subject noetl.kv.noetl_subscription_circuit.<hex> = circuit.<sub_id>
(short key, ~90-char subject). event-log + object + projection + KV all proven;
vector deferred. LOCAL/kind only; prod untouched.
docs(live-in-kind): worker v5.69.0 deploy proof — OBJECT invalid→mirrored, PROJECTION proven, event-log regression holds; KV live circuit pending
Deployed worker v5.69.0 to both kind data-plane pools (shadow), drove a real
tests/large_tabular_result_test (exec 333041450319089664, COMPLETED). Object
mirror flipped from v5.68.0 outcome="invalid" (2) to outcome="mirrored"
(sys 2 / usr 1) on real state-shard/result-tier keys — subject now
noetl.obj.<sha256hex> (74 chars), the ehdb#256 bbc5047 fix. Projection windowed
drain hook advanced (materialized 4, no false divergence). Event-log regression
held (mirrored 6). KV still pending: idle subscription pool has no spool harness
so persist_circuit never fires (mechanism selfcheck-proven; sub-pool roll safe).
LOCAL/kind only; prod untouched.