Skip to content

History

Revisions

  • docs: D2 tail-object cleanup, and why the refusal is the design Adds the D2 section to Design-Tail-Replication plus the v0.9.0 release row. The part worth reading is not the deletion but the refusal: the watermark stops at the first non-durable part, because max-over-durable skips a local-only part in the middle and would authorise deleting a tail object whose records exist nowhere off-box. Records both mutants that the safety test catches, and why tail_objects_retained is a gauge. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 10, 2026
    4184fdb
  • docs: B3 — the unsealed tail replicates off-box, and the lock boundary New page Design-Tail-Replication: the measured gap (prod seal interval 900s with oldest_unsealed_age 719 and nothing waiting to upload), why per-batch objects rather than re-uploading the part, the three invariants and what each prevents, the RED-control-first evidence, and the six series. Also records the v0.8.1 lock-boundary fix, because it is the kind of thing someone adding another timer-driven substrate method needs to read first: the remote put must not run under the engine lock the live append path takes. Framing kept honest throughout: the tick IS the loss window, so B3 takes it from 900s to 15s — bounded, not zero. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 10, 2026
    a9b3604
  • docs: the pre-backfill remote fails LOUDLY, not silently I had written that the bucket "reads as a complete store and is not a recoverable one", which can be read as implying a silent partial recovery. Measured against the emulator: the open succeeds (the manifest IS complete) and replay_all then errors with "no reachable replica (1 listed)". It refuses the whole read rather than returning a short log. The trap is real but detectable, and the distinction changes the operational severity — so it should not be left ambiguous. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 10, 2026
    576b729
  • docs: replica backfill — attaching a replica replicates nothing existing New page Design-Replica-Backfill: the one enqueue site (on seal), the prod measurement (40 parts in the remote manifest, 39 at replica_count 1 while survives_node_loss read 1), why the remote was worse than empty, the local_path-is-None trap under the fix, and why backfills carry their own counters. Cross-linked from Home Core Pages and the sidebar. Releases table was stale — v0.5.1 and v0.5.2 were missing; added those plus v0.6.0 and v0.7.0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 10, 2026
    a93d8e3
  • docs(home): auto-release is ON — and is not auto-deploy Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    8566408
  • docs(home): v0.5.0 — the first automated ehdb release, and where it is consumed Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    0c116c3
  • docs(measures): the supersede hook — 9.2x fewer records, and D8 is disqualified by watch_since Plus the release automation: gated off, verified skipped, tag still v0.4.5. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    5d10f78
  • docs(sessions): P6 — recall is 1.0 by construction, cost is O(ops), compaction cannot help Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    dd43707
  • docs(measures): P6 — recall is 1.0 by construction, query cost is O(ops), compaction cannot help Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    6b8e6f0
  • docs(measures): memory O(parts), the mint-order trap, drain 530x, state gauges, benches in CI Plus north-star P2-P5 landed and a 2026-10-08 session entry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    916e947
  • wiki: the north-star distributed-registry direction, and P1 landed Records EHDB as NoETL's distributed registry and control plane, with the inventory that corrected three of my own prior claims: D8 is in the production crate rather than unwired, vectors already exist rather than being a greenfield hard part, and four of the seven capabilities are substantially built. The real gaps are watch, time-based leases, ephemeral registration, secret-reference types and unsealed-tail replication. States the distributed lower bound honestly: a coordination-free data path gives CALM-style availability and NOT linearizable cross-shard reads, so "etcd semantics, distributed better" is wrong. P7 (tail replication, still RF=1) and P8 (election + fencing) are named as the hard parts rather than vectors or the registry. P1 landed: wall-clock TTL liveness, with the upgrade hazard pinned -- a timestamp-less record is unknown, not dead, because the alternative evicts the whole fleet on upgrade. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    6c5f0e1
  • wiki: the manifest quadratic is fixed, measured by a control that still sees it retain=0 disables pruning and IS the pre-fix behaviour, so the planted defect is a configuration the engine still accepts rather than a mutation. It still measures exponent 1.94 while the default measures 1.23, and the file count is held flat at ~39 across 4x the appends against 79 -> 320 unbounded. Residual recorded: bounded BYTES still grow at 1.23 because retention bounds snapshot count, not snapshot size. Adds tail latency, which criterion cannot give -- append p99 7550us against p50 4014us, and group commit 187x cheaper at p50 rather than the 20.4x the throughput figure shows. Adds the acceptance-criteria standard and its honest scoring: column C (observability) is barely started, because metrics.rs exports 7 public functions for an 88-file engine. Filed as noetl/ai-meta#454. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    a6ea5be
  • wiki: the first measurements of the production L0 engine ehdb-l0 had 88 .rs files and ZERO benches while ehdb-reference had 22 and two, so every number on the existing performance page is about the reference model rather than the engine on the path. These are about the engine. ⭐ The two agree where they overlap: that page measured durable append as fsync-bound at ~3.9 ms / ~256 appends/s, and the production engine measures 3.749 ms / 267 appends/s -- within 4%, from a different crate and harness. Two new results. The fsync is ~95% of posture-A write cost: 267 rec/s fsync-per-append against 5 450 group-committed at batch 80, a 20.4x difference, and group commit needs FlushPolicy::CallerDriven or batching costs MORE. And a read steps 15.6x per record at seal_max_records = 1024, so a chain over 1024 events costs ~13x what extrapolating from the sub-1024 figures predicts -- filed as noetl/ai-meta#453, where whether the step is acceptable is an open judgement call rather than an assumed defect. Also records the method: three of three append instruments were wrong before they were right, each caught by its own planted-effect control, including one where the instrument was wrong and not the engine. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
    3b98c70
  • wiki: make the guard.rs source link absolute `worker/src/ehdb/guard.rs` was written as a relative path, which resolves to `./worker/src/ehdb/guard.rs` inside the wiki and does not exist. Source links are absolute GitHub URLs per agents/rules/wiki-maintenance.md; the file is verified present at src/ehdb/guard.rs on noetl/worker main. Found by running ai-meta's check-docs.sh across all 15 wiki submodules -- the wikis cannot host Actions themselves, so nothing had ever checked them. Refs noetl/ai-meta#375 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 3, 2026
    8f14afd
  • docs: v0.4.5 — link-defined ordering, and the capability it removes noetl/ai-meta#362. The chain is ordered by following prev-links; event_id order is never consulted, so an event committing into the middle of an id-ordered read is a link rather than a shifted position. ⚠ Records the trade explicitly: v0.4.3/v0.4.4 could reconstruct a forked chain from log order and v0.4.5 refuses it instead. That reconstruction was the same id-order assumption that caused the false divergence, so removing it is the fix — but the store now covers fewer executions, and a source that refuses everything also reports zero divergence.

    @claude claude committed Sep 30, 2026
    2b58f64
  • docs: CORRECTION — v0.4.4 is rolled back, and a core assumption of the design is false noetl/ai-meta#362. populate_from_log recomputes chain edges from log order and treats that order as a stable prefix that only grows at the tail. An event can commit into the MIDDLE of an ORDER BY event_id read, because a snowflake event_id is minted before the insert — so commit order is not id order. Every position after the interior insert then reads as a content conflict. Proven by position arithmetic: the store held at position 72 what the log holds at position 73 — exactly one missing interior row. ⚠ No available column is commit-ordered. created_at is stamped at MINT time, so it always agrees with event_id order and cannot detect the reordering; I ran that comparison as a refutation and it refuted nothing. v0.4.4 stands — both defects it fixed are real — but it was not what was breaking prod, and the earlier 5-in-79 was most likely this same mechanism.

    @claude claude committed Sep 30, 2026
    0003290
  • docs: v0.4.4 — a behind snapshot is staleness, and the chain store is live in prod noetl/ai-meta#360. Design-Execution-Chain-Store + Sessions-Log + Home. New section "Staleness is not a conflict — and the same shape is a second writer", which is the design point of this release: the crate REPORTS `FromLog::StaleLog` as unresolved rather than deciding, because from a single snapshot benign staleness and a foreign writer are indistinguishable, and only the consumer can re-read the log to tell them apart. My first version decided it, and the noetl/ai-meta#358 regression test failed on the first run. Also recorded: a content conflict now revokes serve-trust (previously this function wrote nothing, so the still-matching coverage handed a `chain_if_authoritative` caller a chain just proved wrong); coverage is monotonic; and the monotonic-coverage mutant SURVIVED its first test, which populated 5 then read 2 and returned StaleLog before the guard was ever reached. ⚠ Corrections: the status line said "flag-gated, default OFF, not enabled anywhere" and the Home link said the store "cannot serve yet". It is ARMED IN PRODUCTION since 2026-09-29 17:30Z on server v3.117.2 — 772 comparisons, 0 failures. Version floor raised v0.4.3 -> v0.4.4, because v0.4.3 carries the false-divergence classification that rolled back the first ramp.

    @claude claude committed Sep 29, 2026
    fb752b4
  • Design: the chain store shipped in v0.4.3 — record the version FLOOR The page said 'built, not enabled'. It is now on main and tagged, so the thing a reader needs is the VERSION FLOOR, not the build status: v0.4.3 or later, because v0.4.2 has the store and the watermark but neither populate_from_log nor the coverage record — and a reader on v0.4.2 therefore cannot tell a truncated partition from a whole one, which is the defect the page documents. Also records that FORMAT_VERSION is 1 across v0.3.0/v0.4.2/v0.4.3, so moving between them does not trip the on-disk layout gate — the question anyone bumping the pin will have. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Sep 29, 2026
    addf49b
  • Design: populate_from_log + the coverage record; both defects closed The chain store's page said it could not serve, and named two defects. Both are now fixed and kind-proven, so the page has to say how rather than leave a reader with a stale blocker. Adds the section a reader needs most — populate_from_log and the COVERAGE record — and states plainly what the watermark cannot express: its first_seq is the STORE's sequence, 1 for any fresh partition regardless of where the execution began, which is exactly how a chain starting at an execution's third event read as authoritative. ⚠ Records that the root signal is deliberately NOT an event-type name. Only 62 of 595 executions start with playbook_started; 533 start with playbook.initialized, so a check built on the first guess would have been inert for 90% of them. The root is position 0 in the log. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Sep 29, 2026
    c4175e3
  • Design: the execution-partitioned chain store, and why it cannot serve yet New page for ehdb-l0's DurableChainStore + ChainPopulator (RFC noetl/ai-meta#355). Written from the diff and from the kind measurements, not from the RFC's summary. The load-bearing parts, recorded because each is a guard whose purpose someone would otherwise delete as noise: * the three states the watermark separates — `None` never means "no events", `Some(vec![])` does; * the ASYMMETRIC append/watermark ordering, and the defect that produced it (marking before every append left Authoritative{1,1} over zero events on the common mid-flight-arming path). v0.4.1 carries that bug; v0.4.2 is the first tag without it; * `apply_replicated` deliberately does NOT enforce the head, because async replication delivers out of order — found by a mutation test that could not construct a gap at all; * `chain_is_complete` compares against the partition SPAN, so it catches a hole in the middle and NOT a missing beginning — which is exactly why the open defect is undetectable from inside the store. ⚠⚠ Both measured defects are stated with their numbers and the shared cause (a chain edge stamped from an in-memory map that does not survive a restart), so nobody reads this page and concludes the store is ready to serve. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Sep 28, 2026
    ff6653c
  • Architecture: the writer mounts three STATIC PVCs, not volumeClaimTemplates Correcting my own text in the previous commit, which said the write role 'carries a PVC' — singular, and implying a StatefulSet-managed claim. Verified 2026-09-16: volumeClaimTemplates is NONE and three separately-created PVCs are mounted by name, so the claims outlive the StatefulSet. The mirror runbook landed the same finding upstream while this was being written. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Sep 16, 2026
    08e45b9
  • Invariants + architecture: the transition semantics and the layer/role axes noetl/ehdb#323. Both pages existed; both were missing the half the issue asks for, and it is the half the cutover questions turn on. Consistency-Invariants gains §5 — what a reader observes during a shadow -> primary transition: * the three states, and why the middle one misleads — a shadow tier is fully written and fully compared and still serves nothing, so append rate and store size say the mirror is alive and say nothing about whether a flip is correct * what is dual-written, re-measured on prod 2026-09-16: MINT_AUTHORITATIVE=true, STORE_DUAL_WRITE=false, PROJECTION_READ_SOURCE=wal, SERVE_ON_BEHIND=true (so projection reads are NOT read-your-writes), OBJECT_STORE_BACKEND=gcs * that the NOETL_EHDB_<TIER> mode variables are not set on the prod pods at all, so live modes come from code defaults rather than any manifest * what parity checks and what it deliberately does not: three DIVERGENCE_KINDS; superseded/unmirrored are observations, not divergences; arrival order is counted rather than judged (#346 — the first definition of "different" was wrong and produced 8-of-74 false divergence) * the recovery ladder spine -> tier -> postgres, and why the 2026-09-16 audit finding (every event-log read still has Postgres authority behind it) is exactly what a flip would spend * a four-question checklist to run before any flip Architecture-Four-Engines gains the two axes it was missing: * the L0..L3 layer stack, marked for what is CODE vs what is PLAN — only L0 is a crate, L1 is the feed, L2/L3 are plan * node roles as DEPLOYMENT facts, stated as such because there is no NodeRole type in the codebase, with the write role's PVC + pinned digest called out as the reason writer-ordered rollouts are gated Documentation only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Sep 16, 2026
    61e5765
  • Mirror runbook: batch path is live; surge-free rolls; the writer has no volumeClaimTemplates ai-meta#343/#344 closed out. Records what proved the transport fix (the path=batch counter moving 0 -> 460, after 0 of 80,264 records ever), the repair going from partial/1-of-39 to repaired/64-of-64, and parity 14/14 agree. Adds two operational traps that cost an outage: a maxSurge=25% roll on a 1-replica pool rounds to +1 pod at 3Gi and cannot schedule at 99% node memory; and the writer StatefulSet mounts its PVCs by NAME with no volumeClaimTemplates, so deleting them leaves the pod unschedulable forever. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @kadyapam kadyapam committed Sep 13, 2026
    b146b14
  • Mirror runbook: the timeout is an un-batched fan-out; tier records cannot be corrected in place ai-meta#344 and #345. Two things an operator needs before reaching for the obvious fix in either direction. The mirror timeouts are not a slow network: the relay appends one record at a time (batch path 0 of 80,264) against an fsync-per-append writer at ~118ms each. Raising APPEND_TIMEOUT buys time for a loop that should not exist. Arming the batch path needs chunking first, because the request cap stays at 1 MiB by design. And a record already in the tier cannot be corrected: dedup ignores rather than replaces, and the event log has no replace/upsert/delete op. A re-mirror to fix content is a no-op that reports success. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @kadyapam kadyapam committed Sep 13, 2026
    dfa383a
  • Mirror runbook: a repair closed the count gap and opened a content gap ai-meta#343. Three findings from the 2026-09-12 session, all reachable from the mirror runbook because that is where an operator looks when a repair reports success and parity still disagrees. 1. The #342 repair re-mirrored the PERSISTED row, whose result carries a reference, while the live path mirrors the in-memory row, whose result is inlined. Repairing again rewrote the same reference, so it could never converge. Every event still differing was one the repair had re-mirrored: 3 of 39 and 3 of 26, zero outside. Fixed in server#430; repair now reports 'hydrated'. 2. not_comparable is not a divergence. Three of a fixed 40 were unreadable, not divergent, and a boolean could not say so. The sweep now publishes accounted_for so the denominator is stated rather than subtracted. 3. A mirror send failure now carries its cause as a label. Do not change APPEND_TIMEOUT unless kind="timeout" is what moves. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @kadyapam kadyapam committed Sep 12, 2026
    982e032
  • Architecture: record the mirror-coverage root cause, the parity proof, and what remains Two causes at 30/48: a verification artifact (raw vs hydrated Postgres fold) and real mirror loss. Parity proof on a fixed 40: agree 10 -> 26, 0 regressed. Records that the first proof attempt moved 0 of 40 because the measuring endpoint was never on the path fixed -- an instrument that cannot see a fix cannot validate one.

    @akuksin akuksin committed Sep 12, 2026
    d66437a
  • Plan: decision recorded — refill after #342, do not migrate Choosing refill over migration removes both hazards that made the original plan a one-way door: no copy consistency window, and no copy-back on revert. Deferred until the mirror is clean, because migrating 1.8 GB of a mirror losing 39.6% of what it is handed relocates a known-bad copy.

    @akuksin akuksin committed Sep 12, 2026
    59673bf
  • Plan: give the event-log tier its own failure domain (not applied) Needs a data migration (1.8 GB / 3 files) and the env-var switch is a one-way door, so this stops at the plan per the owner gate. Notes that the writer uses pre-created PVCs rather than volumeClaimTemplates, so adding a volume is a normal mutable update -- but that the PVC must be Bound before anything references it, which is the exact #323 outage shape. Recommends deciding the mirror-drop issue (#342) first: migrating a mirror that is currently losing 39.6% of what it is handed moves a known-bad copy.

    @akuksin akuksin committed Sep 12, 2026
    c270287
  • wiki: link the resilient-KV-core architecture overview from Home + sidebar

    @akuksin akuksin committed Sep 11, 2026
    767f6c5
  • Architecture: EHDB as a resilient KV core — CockroachDB's layers minus SQL The overview for the resilient-core programme (ai-meta#339). Maps Cockroach's five layers onto EHDB as implemented at v0.2.0, with prod readings. Key finding: the L4 gap is narrower than 'no replication'. Sealed-part N-way copy is implemented; the unsealed tail is RF=1. The window is already measured AND bounded on prod (SEAL_MAX_AGE_MS=5000, both engines export it) -- but open_replicated and validate_replica_domains have NO production callers, and the tier dir is nested inside the writer dir on one PVC. Keeps the existing no-consensus decision for sealed parts (immutable copies cannot conflict) and adds Raft only over the unsealed tail, where no immutable object exists yet. No SQL layer, permanently.

    @akuksin akuksin committed Sep 11, 2026
    3bd16a3