Skip to content

History / Home

Revisions

  • docs: B3 — the unsealed tail replicates off-box, and the lock boundary New page Design-Tail-Replication: the measured gap (prod seal interval 900s with oldest_unsealed_age 719 and nothing waiting to upload), why per-batch objects rather than re-uploading the part, the three invariants and what each prevents, the RED-control-first evidence, and the six series. Also records the v0.8.1 lock-boundary fix, because it is the kind of thing someone adding another timer-driven substrate method needs to read first: the remote put must not run under the engine lock the live append path takes. Framing kept honest throughout: the tick IS the loss window, so B3 takes it from 900s to 15s — bounded, not zero. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 10, 2026
  • docs: replica backfill — attaching a replica replicates nothing existing New page Design-Replica-Backfill: the one enqueue site (on seal), the prod measurement (40 parts in the remote manifest, 39 at replica_count 1 while survives_node_loss read 1), why the remote was worse than empty, the local_path-is-None trap under the fix, and why backfills carry their own counters. Cross-linked from Home Core Pages and the sidebar. Releases table was stale — v0.5.1 and v0.5.2 were missing; added those plus v0.6.0 and v0.7.0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 10, 2026
  • docs(home): auto-release is ON — and is not auto-deploy Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
  • docs(home): v0.5.0 — the first automated ehdb release, and where it is consumed Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
  • wiki: the north-star distributed-registry direction, and P1 landed Records EHDB as NoETL's distributed registry and control plane, with the inventory that corrected three of my own prior claims: D8 is in the production crate rather than unwired, vectors already exist rather than being a greenfield hard part, and four of the seven capabilities are substantially built. The real gaps are watch, time-based leases, ephemeral registration, secret-reference types and unsealed-tail replication. States the distributed lower bound honestly: a coordination-free data path gives CALM-style availability and NOT linearizable cross-shard reads, so "etcd semantics, distributed better" is wrong. P7 (tail replication, still RF=1) and P8 (election + fencing) are named as the hard parts rather than vectors or the registry. P1 landed: wall-clock TTL liveness, with the upgrade hazard pinned -- a timestamp-less record is unknown, not dead, because the alternative evicts the whole fleet on upgrade. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
  • wiki: the first measurements of the production L0 engine ehdb-l0 had 88 .rs files and ZERO benches while ehdb-reference had 22 and two, so every number on the existing performance page is about the reference model rather than the engine on the path. These are about the engine. ⭐ The two agree where they overlap: that page measured durable append as fsync-bound at ~3.9 ms / ~256 appends/s, and the production engine measures 3.749 ms / 267 appends/s -- within 4%, from a different crate and harness. Two new results. The fsync is ~95% of posture-A write cost: 267 rec/s fsync-per-append against 5 450 group-committed at batch 80, a 20.4x difference, and group commit needs FlushPolicy::CallerDriven or batching costs MORE. And a read steps 15.6x per record at seal_max_records = 1024, so a chain over 1024 events costs ~13x what extrapolating from the sub-1024 figures predicts -- filed as noetl/ai-meta#453, where whether the step is acceptable is an open judgement call rather than an assumed defect. Also records the method: three of three append instruments were wrong before they were right, each caught by its own planted-effect control, including one where the instrument was wrong and not the engine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
  • docs: CORRECTION — v0.4.4 is rolled back, and a core assumption of the design is false noetl/ai-meta#362. populate_from_log recomputes chain edges from log order and treats that order as a stable prefix that only grows at the tail. An event can commit into the MIDDLE of an ORDER BY event_id read, because a snowflake event_id is minted before the insert — so commit order is not id order. Every position after the interior insert then reads as a content conflict. Proven by position arithmetic: the store held at position 72 what the log holds at position 73 — exactly one missing interior row. ⚠ No available column is commit-ordered. created_at is stamped at MINT time, so it always agrees with event_id order and cannot detect the reordering; I ran that comparison as a refutation and it refuted nothing. v0.4.4 stands — both defects it fixed are real — but it was not what was breaking prod, and the earlier 5-in-79 was most likely this same mechanism.

    @claude claude committed Sep 30, 2026
  • docs: v0.4.4 — a behind snapshot is staleness, and the chain store is live in prod noetl/ai-meta#360. Design-Execution-Chain-Store + Sessions-Log + Home. New section "Staleness is not a conflict — and the same shape is a second writer", which is the design point of this release: the crate REPORTS `FromLog::StaleLog` as unresolved rather than deciding, because from a single snapshot benign staleness and a foreign writer are indistinguishable, and only the consumer can re-read the log to tell them apart. My first version decided it, and the noetl/ai-meta#358 regression test failed on the first run. Also recorded: a content conflict now revokes serve-trust (previously this function wrote nothing, so the still-matching coverage handed a `chain_if_authoritative` caller a chain just proved wrong); coverage is monotonic; and the monotonic-coverage mutant SURVIVED its first test, which populated 5 then read 2 and returned StaleLog before the guard was ever reached. ⚠ Corrections: the status line said "flag-gated, default OFF, not enabled anywhere" and the Home link said the store "cannot serve yet". It is ARMED IN PRODUCTION since 2026-09-29 17:30Z on server v3.117.2 — 772 comparisons, 0 failures. Version floor raised v0.4.3 -> v0.4.4, because v0.4.3 carries the false-divergence classification that rolled back the first ramp.

    @claude claude committed Sep 29, 2026
  • Design: the execution-partitioned chain store, and why it cannot serve yet New page for ehdb-l0's DurableChainStore + ChainPopulator (RFC noetl/ai-meta#355). Written from the diff and from the kind measurements, not from the RFC's summary. The load-bearing parts, recorded because each is a guard whose purpose someone would otherwise delete as noise: * the three states the watermark separates — `None` never means "no events", `Some(vec![])` does; * the ASYMMETRIC append/watermark ordering, and the defect that produced it (marking before every append left Authoritative{1,1} over zero events on the common mid-flight-arming path). v0.4.1 carries that bug; v0.4.2 is the first tag without it; * `apply_replicated` deliberately does NOT enforce the head, because async replication delivers out of order — found by a mutation test that could not construct a gap at all; * `chain_is_complete` compares against the partition SPAN, so it catches a hole in the middle and NOT a missing beginning — which is exactly why the open defect is undetectable from inside the store. ⚠⚠ Both measured defects are stated with their numbers and the shared cause (a chain edge stamped from an in-memory map that does not survive a restart), so nobody reads this page and concludes the store is ready to serve. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Sep 28, 2026
  • wiki: link the resilient-KV-core architecture overview from Home + sidebar

    @akuksin akuksin committed Sep 11, 2026
  • Embedded-state foundations: format gate, stored cursor, D3/D8 exercised, append validation, foca Records what noetl/ai-meta#332 added to THIS repo, and for each guard what it exists to prevent -- so it is not later deleted as noise. Includes the bugs the work's own tests found: the format guard's blanket Err(_) => Ok(None) that turned every read failure into 'absent', and the dot-run check that is load-bearing for exactly one shape because dots are legal in an id.

    @claude claude committed Sep 8, 2026
  • wiki: record the F1-F5 remediation and refresh the invariants The Consistency-Invariants page said 'detected by nothing' for two invariants that now have detectors, which is exactly the drift that page exists to catch. I-EL-2 (single writer) is now DETECTABLE but still not enforced: fencing runs in shadow, counting stale-epoch writes while letting them through, and the election issues tokens without being authoritative. The ordering hazard is recorded on the page itself -- enforcing before the election issues real tokens is an outage rather than a degradation, because with no election every writer's epoch is 0. I-EL-4 (reaching the substrate) now names ehdb_l0_unreplicated_age_seconds as its detector, live on the writer's /metrics since worker v5.125.0, and keeps the two misleading readings explicitly labelled: upload_lag_micros_total is seal-relative and blind to the dominant term, and ehdb_feed_shard_lag is consumer backlog wearing an adjacent name. Also records that prod as it stands would FAIL the new replica-domain check -- the tier dir is nested inside the writer dir on one PVC -- so validating at open before fixing the layout is a startup outage by construction. Home carries the summary above the fold; Sessions-Log gets the dated entry. Refs noetl/ehdb#324, noetl/ehdb#328, noetl/ehdb#329, noetl/ehdb#330, noetl/ehdb#331, noetl/ehdb#332

    @claude claude committed Aug 30, 2026
  • wiki: the four-engine architecture + per-tier consistency invariants ehdb#323, under #324. Two new pages, plus Home and the sidebar reconciled to the narrowed scope. Architecture-Four-Engines documents the four owned engines with a diagram, and records WHY vector and OLAP collapse into projections -- matching what is built rather than what was promised: ehdb-reference/src/vector.rs is already a bounded cosine search over the collection's live points, so there is NO ANN index to remove. The gap was in the promise. It also separates three vocabularies that are routinely conflated: engine vs tier vs StoreTier/QueryTier. ⚠⚠ Both pages surface a discrepancy the Backend-Configuration page invites: it presents five tiers each with an off/shadow/primary mode, which reads as though setting primary makes any of them serve. SERVE_WIRED_TIERS is ["eventlog", "projection"] -- KV and object accept primary as CONFIGURATION and cannot serve as CAPABILITY. Home now says so above the fold. Consistency-Invariants states, per tier, what each guarantee is, WHAT FORCES IT to hold, and HOW YOU WOULD FIND OUT if it stopped -- because a guarantee with no enforcement and no detector is a description, not an invariant. Three come out badly and are labelled rather than smoothed over: I-EL-2 (single writer) is enforced by nothing stronger than replicas: 1 and detected by nothing; I-EL-4 (reaching the substrate) is bounded by volume not time and detected by nothing that works, since upload lag is measured from seal and feed lag is consumer backlog; and in prod the substrate is the SAME PVC as the primary, so it buys no independent failure domain. Home's mission no longer claims EHDB will absorb Qdrant and ClickHouse. Refs noetl/ehdb#323, noetl/ehdb#324, noetl/ehdb#320, noetl/ehdb#321, noetl/ehdb#322

    @claude claude committed Aug 30, 2026
  • wiki(#155): the async event-log mirror, and correct the stale "primary not activated" claims Adds Runbook-Async-Event-Log-Mirror — the operational surface for the mirror that went asynchronous on prod 2026-08-19, and which this wiki did not mention at all. Leads with the one-flag rollback and with the pairing that makes it dangerous: async on with the tolerance at 0 makes the comparator judge a healthy tier on its own liveness. Records the parts that are easy to get wrong rather than only the settings. That the window must be MEASURED and then set above the observed p99 — prod came back p99 699ms, so the window is 30s. That it sits deliberately below SETTLE_SECS so the background sampler still compares every execution. That the two lag controls are a discrimination rather than a clean-parity check, because a comparator taught to ignore everything passes the naive version perfectly. And that reserve rather than send is load-bearing, because the first implementation lost 32 of 60 records while every counter still read correctly. Corrects two current-state claims that had gone stale: * Home said the reference tiers are O-of-n per op and "not primary-serve". Still true for KV/object/vector; false for the event log since the runtime cache landed and since the tier went primary on 2026-08-13. * Roadmap said the primary cutover "stays a later step" and that primary is "recognised but not activated". Both annotated as superseded rather than deleted, since they remain accurate as the plan of record. Sessions-Log is left untouched — it is history and correct as written. Home gains a dated current-state banner so a reader meets the live state before the older planning prose. Also records that Option 2's batch flag is still OFF on prod, batch=0 on both replicas, so the substrate is landed and measurable but inert. Tracked at noetl/ai-meta#284.

    @claude claude committed Aug 20, 2026
  • wiki(L1 round 4): prod deploy — EHDB bus ~2.4x faster than NATS (#205 closed) + writer-restart survival (#208 closed); T5 gate is now the user-pool autoscaler (#210); GMP-not-VM correction

    @akuksin akuksin committed Jul 31, 2026
  • wiki(L1 #205): dispatch latency root-caused + fixed; T4/T5 status corrected - Program page: T4 row corrected (full flip SUCCEEDED on shastaratech prod 2026-07-27 round 3, was still 'PREPARED, NOT EXECUTED'); T5 row now names its two blockers (#205 latency, #208 writer-restart survival); new section on the #205 attribution + fix. - Sessions log: 2026-07-28 entry with the per-hop attribution, the three-part fix, the loopback before/after, the kind correctness PASS, and why kind cannot resolve the latency delta. - Home: implementation note for the fix. Refs noetl/ai-meta#205, noetl/ai-meta#208

    @claude claude committed Jul 29, 2026
  • L1 T4 COMPLETE — EHDB is the production command bus (full flip 80/80 on shastaratech prod) Runbook: status banner rewritten to SUCCESS; T4 go/no-go checklist cleared with the latency caveat + a new T5 checklist; prod execution log restructured into three rounds (held-at-canary / failed-on-#203 / succeeded) with the #203 root cause + fix + before-after table, the post-flip prod state, the rollback command, and the latency regression. Program page: L1 status T0-T4 done / T5 held; T4+T5 table rows rewritten. Home: program bullet leads with prod-command-bus state. Sessions-Log: new 2026-07-27 entry. Refs noetl/ai-meta#194, #202 (closed), #203 (closed), #205, #206

    @akuksin akuksin committed Jul 28, 2026
  • Refine boundary to write-behind-cache model; re-scope D5/D6 issues (+#198 sink) EHDB is never the durable system of record for business data (customer's connector-backed store is), but MAY hold business processing context transiently as a write-behind cache: two roles — (a) permanent control-plane log stays lean/ reference; (b) transient processing cache bounded + sunk to the customer store + evictable. Supersedes the 'never holds business payload' absolute. Re-scopes the fixes: #195 reframed (keep PERMANENT noetl.event lean; transient cache keeps full context for the drive — dissolves the earlier drive-breaking risk since the drive reads the transient cache not the permanent log); NEW #198 the write-behind sink path (transient context -> customer system of record + sink-confirmation-gated eviction; builds on connector tools + result-tier byte-source + GC); #196 reframed (permanence in noetl.command); #197 unchanged.

    @claude claude committed Jul 23, 2026
  • Add noetl-only scoping invariant + ClickHouse meta-catalog to L0 design Program invariant (recorded prominently): EHDB is noetl-internal-only, a FIXED set of predefined datasets (D1-D10: event log, commands, projections, KV, object/blob, vector, catalog, runtime, system-WASM, provider-facts), NOT a general DB — no arbitrary schemas/DDL, no query planner, no hosted-DB surface; business data never in EHDB. It SHRINKS every layer. L0 design now = VM/VictoriaLogs write engine + a ClickHouse-MergeTree-style meta-catalog (manifest + sparse index -> object-store pointer; fixed-schema table-format-over-object-storage). Scope cuts: no general inverted index (few fixed per-dataset indexes only); L3 cut from a SQL engine to fixed reads + read-only export (~2-3q -> ~1q). Umbrella: https://github.com/noetl/ai-meta/issues/194

    @claude claude committed Jul 16, 2026
  • Reframe program page to layered platform (L0-first); NATS takeover = L1 Major reframe: EHDB becomes a layered platform built L0-first — L0 replicated object store (modeled on VictoriaMetrics/VictoriaLogs: buffered flush -> immutable parts -> background merge -> inverted-index-first -> columnar) -> L1 streaming (the NATS takeover) -> L2 KV -> L3 append-log SQL. L0 resolves the HA debate: durability+replication from the object store => fungible writers, per-shard-Raft (T-RF) retired, only a light L1 ordering lease remains. Records the honest departure (VM keeps the hot path on local disk; object storage is backup-only, so L0's live object-store durability tier is net-new beyond VM), the reuse assessment (#254 segments approx immutable parts; gaps = inverted index + merge engine + object-store tier), per-layer cost, and the first L0 slice (L0.1). Master RFC: docs/rfc/ehdb-layered-platform.md. Umbrella: https://github.com/noetl/ai-meta/issues/194

    @claude claude committed Jul 16, 2026
  • Program page: lock HA posture — shards-only now, RF deferred (replication-ready) Verified parity (prod NATS replicas:1, streams R1). Ships (c) single-writer- per-shard; per-shard replication factor deferred to a named non-blocking phase (T-RF). Adds the replication-ready seam (the #254 durable_eventlog_affinity single-writer + open_read_only followers + durable_eventlog_shared log-shipping + durable external cursors => T-RF additive, no log/cursor rewrite), the failover sequence + seconds-scale loss-free stall estimate, and the worker reconnect-on-failover story (unchanged StatefulSet DNS). Umbrella: https://github.com/noetl/ai-meta/issues/194

    @claude claude committed Jul 16, 2026
  • Program page: revise to (c) per-shard-writer-as-broker topology Supersedes (b) co-locate on latency (two delivery hops). (c): the stateful per-shard writer owns its shard's durable log AND delivery; workers subscribe directly to the writer; stateless server publishes to the writer and is out of the delivery path -> one delivery hop, matching NATS. Records the honest premise correction from the code trace (#166 sharding dormant; system pool is a single-replica Deployment, no StatefulSet/PVC/identity; drive hands back to the server), so (c) is a build not a relocation (~2q, more than (b)). Adds the alternatives ledger, honest downsides (connection fan-out, per-shard HA, concentration on the #163 OOM pod), revised cost, and the (c)-shaped T0 spec with a latency-vs-NATS comparison. Umbrella: https://github.com/noetl/ai-meta/issues/194

    @claude claude committed Jul 16, 2026
  • Program page: lock (b) co-locate topology for EHDB->NATS takeover Records the locked topology (2026-07-15): durable log + change-feed stay in the per-shard writer (#166, stateful); the noetl-server stays stateless (#115) and owns delivery only (tails the change-feed, pushes over the SSE hub); option (a) server-embeds-EHDB rejected. Adds the three-component split, re-derived Track T phases, revised cost (~1.5 quarters), and the (b)-shaped T0 spec. Program umbrella: https://github.com/noetl/ai-meta/issues/194

    @claude claude committed Jul 16, 2026
  • Add EHDB→NATS takeover program page (decision + plan, nothing built) Records the locked decision (2026-07-15): remove NATS via noetl-server-owned push reusing the gateway SSE ConnectionHub over EHDB's durable log + change-feed; standalone ehdb-server broker rejected. Track S (storage cutover, in flight) + Track T (transport, server-owned-push, T0->T5, unbuilt), revised cost, three open sub-decisions, T0 slice spec. Cross-linked from Home + _Sidebar. Program umbrella: https://github.com/noetl/ai-meta/issues/194

    @claude claude committed Jul 16, 2026
  • docs(rfc): external EHDB driver — outward-facing read access (design) Design-first RFC for an external driver/client so third-party apps and BI tools can read EHDB's engine tiers. Recommends Arrow Flight + Flight SQL over a dedicated data-plane endpoint fronting the worker (loose coupling preserved), scoped read-only API tokens, committed-only reads. Five decision forks surfaced. Grounds on the existing ehdb-service Flight scaffold + the #178 internal read contract. Registered in _Sidebar + Home. Tracks noetl/ai-meta#184.

    @claude claude committed Jul 10, 2026
  • docs(perf): add Design-Performance-and-Load-Testing + baseline; Sessions-Log, Home, Roadmap, _Sidebar

    @claude claude committed Jul 8, 2026
  • docs(durable-eventlog): slice 6 prod-durability sign-off package (ehdb#254) New Runbook-Durable-EventLog-Prod-Signoff page — go/no-go checklist (D1-D11), evidence bundle (slices 1-5), residual-risk register (R1-R7), and the extended durable-shadow -> durable-primary prod rollout sequence. Cross-linked from the tier-1 cutover runbook (§C update + Related), the design page, and _Sidebar; Sessions-Log + Home updated. Docs/planning only — no prod action; ehdb#254 slice-6 box stays unchecked.

    @kadyapam kadyapam committed Jul 8, 2026
  • docs(subject-digest): extend the subject-length fix to KV + vector tiers (ehdb#259 / worker#172) KV + vector registry subjects now use a fixed-width SHA-256 digest token (noetl.kv.<bucket>.<sha256hex(key)>, noetl.vec.<sha256hex(col)>.<sha256hex(pt)>) instead of hex-of-full-id, matching the object tier — bounding every per-id subject under the 256-char Subject cap. Adds a 'Subject digest tokens (all three tiers)' section to the Phase-8 design page, updates the KV/object/vector sections + object callout, prepends a Sessions-Log entry, and notes the hardening on Home. Forward-safety (no live breakage today); in-kind re-proof deferred. Refs noetl/ai-meta#241.

    @kadyapam kadyapam committed Jul 8, 2026
  • docs(live-in-kind): KV mirror PROVEN on live spool-circuit runtime — all four wired tiers now fire on live drives (worker v5.69.0) Stood up a dedicated subscription runtime (WORKER_MODE=subscription, v5.69.0 + shadow) activating subscriptions/spool_outage_stream (spool + http-probed downstream). SpoolRuntime::persist_circuit → kv::mirror_live_put fires every ~2s probe tick + on circuit open/close: noetl_ehdb_kv_ops_total{operation="mirror", outcome="mirrored"} climbed 10→50+, kv_last_ok=1, zero invalid. Ref-store circuit record subject noetl.kv.noetl_subscription_circuit.<hex> = circuit.<sub_id> (short key, ~90-char subject). event-log + object + projection + KV all proven; vector deferred. LOCAL/kind only; prod untouched.

    @kadyapam kadyapam committed Jul 8, 2026
  • docs(live-in-kind): worker v5.69.0 deploy proof — OBJECT invalid→mirrored, PROJECTION proven, event-log regression holds; KV live circuit pending Deployed worker v5.69.0 to both kind data-plane pools (shadow), drove a real tests/large_tabular_result_test (exec 333041450319089664, COMPLETED). Object mirror flipped from v5.68.0 outcome="invalid" (2) to outcome="mirrored" (sys 2 / usr 1) on real state-shard/result-tier keys — subject now noetl.obj.<sha256hex> (74 chars), the ehdb#256 bbc5047 fix. Projection windowed drain hook advanced (materialized 4, no false divergence). Event-log regression held (mirrored 6). KV still pending: idle subscription pool has no spool harness so persist_circuit never fires (mechanism selfcheck-proven; sub-pool roll safe). LOCAL/kind only; prod untouched.

    @kadyapam kadyapam committed Jul 8, 2026