Skip to content

History / Sessions Log

Revisions

  • docs: B3 — the unsealed tail replicates off-box, and the lock boundary New page Design-Tail-Replication: the measured gap (prod seal interval 900s with oldest_unsealed_age 719 and nothing waiting to upload), why per-batch objects rather than re-uploading the part, the three invariants and what each prevents, the RED-control-first evidence, and the six series. Also records the v0.8.1 lock-boundary fix, because it is the kind of thing someone adding another timer-driven substrate method needs to read first: the remote put must not run under the engine lock the live append path takes. Framing kept honest throughout: the tick IS the loss window, so B3 takes it from 900s to 15s — bounded, not zero. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 10, 2026
  • docs: replica backfill — attaching a replica replicates nothing existing New page Design-Replica-Backfill: the one enqueue site (on seal), the prod measurement (40 parts in the remote manifest, 39 at replica_count 1 while survives_node_loss read 1), why the remote was worse than empty, the local_path-is-None trap under the fix, and why backfills carry their own counters. Cross-linked from Home Core Pages and the sidebar. Releases table was stale — v0.5.1 and v0.5.2 were missing; added those plus v0.6.0 and v0.7.0. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 10, 2026
  • docs(measures): the supersede hook — 9.2x fewer records, and D8 is disqualified by watch_since Plus the release automation: gated off, verified skipped, tag still v0.4.5. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
  • docs(sessions): P6 — recall is 1.0 by construction, cost is O(ops), compaction cannot help Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
  • docs(measures): memory O(parts), the mint-order trap, drain 530x, state gauges, benches in CI Plus north-star P2-P5 landed and a 2026-10-08 session entry. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

    @claude claude committed Oct 8, 2026
  • docs: CORRECTION — v0.4.4 is rolled back, and a core assumption of the design is false noetl/ai-meta#362. populate_from_log recomputes chain edges from log order and treats that order as a stable prefix that only grows at the tail. An event can commit into the MIDDLE of an ORDER BY event_id read, because a snowflake event_id is minted before the insert — so commit order is not id order. Every position after the interior insert then reads as a content conflict. Proven by position arithmetic: the store held at position 72 what the log holds at position 73 — exactly one missing interior row. ⚠ No available column is commit-ordered. created_at is stamped at MINT time, so it always agrees with event_id order and cannot detect the reordering; I ran that comparison as a refutation and it refuted nothing. v0.4.4 stands — both defects it fixed are real — but it was not what was breaking prod, and the earlier 5-in-79 was most likely this same mechanism.

    @claude claude committed Sep 30, 2026
  • docs: v0.4.4 — a behind snapshot is staleness, and the chain store is live in prod noetl/ai-meta#360. Design-Execution-Chain-Store + Sessions-Log + Home. New section "Staleness is not a conflict — and the same shape is a second writer", which is the design point of this release: the crate REPORTS `FromLog::StaleLog` as unresolved rather than deciding, because from a single snapshot benign staleness and a foreign writer are indistinguishable, and only the consumer can re-read the log to tell them apart. My first version decided it, and the noetl/ai-meta#358 regression test failed on the first run. Also recorded: a content conflict now revokes serve-trust (previously this function wrote nothing, so the still-matching coverage handed a `chain_if_authoritative` caller a chain just proved wrong); coverage is monotonic; and the monotonic-coverage mutant SURVIVED its first test, which populated 5 then read 2 and returned StaleLog before the guard was ever reached. ⚠ Corrections: the status line said "flag-gated, default OFF, not enabled anywhere" and the Home link said the store "cannot serve yet". It is ARMED IN PRODUCTION since 2026-09-29 17:30Z on server v3.117.2 — 772 comparisons, 0 failures. Version floor raised v0.4.3 -> v0.4.4, because v0.4.3 carries the false-divergence classification that rolled back the first ramp.

    @claude claude committed Sep 29, 2026
  • wiki: record the F1-F5 remediation and refresh the invariants The Consistency-Invariants page said 'detected by nothing' for two invariants that now have detectors, which is exactly the drift that page exists to catch. I-EL-2 (single writer) is now DETECTABLE but still not enforced: fencing runs in shadow, counting stale-epoch writes while letting them through, and the election issues tokens without being authoritative. The ordering hazard is recorded on the page itself -- enforcing before the election issues real tokens is an outage rather than a degradation, because with no election every writer's epoch is 0. I-EL-4 (reaching the substrate) now names ehdb_l0_unreplicated_age_seconds as its detector, live on the writer's /metrics since worker v5.125.0, and keeps the two misleading readings explicitly labelled: upload_lag_micros_total is seal-relative and blind to the dominant term, and ehdb_feed_shard_lag is consumer backlog wearing an adjacent name. Also records that prod as it stands would FAIL the new replica-domain check -- the tier dir is nested inside the writer dir on one PVC -- so validating at open before fixing the layout is a startup outage by construction. Home carries the summary above the fold; Sessions-Log gets the dated entry. Refs noetl/ehdb#324, noetl/ehdb#328, noetl/ehdb#329, noetl/ehdb#330, noetl/ehdb#331, noetl/ehdb#332

    @claude claude committed Aug 30, 2026
  • wiki(L1 round 4): prod deploy — EHDB bus ~2.4x faster than NATS (#205 closed) + writer-restart survival (#208 closed); T5 gate is now the user-pool autoscaler (#210); GMP-not-VM correction

    @akuksin akuksin committed Jul 31, 2026
  • wiki(L1 #208): writer-restart survival — both defects fixed + a third in the publish path; kind before/after; 3 PRs held

    @akuksin akuksin committed Jul 30, 2026
  • wiki(L1 #205): dispatch latency root-caused + fixed; T4/T5 status corrected - Program page: T4 row corrected (full flip SUCCEEDED on shastaratech prod 2026-07-27 round 3, was still 'PREPARED, NOT EXECUTED'); T5 row now names its two blockers (#205 latency, #208 writer-restart survival); new section on the #205 attribution + fix. - Sessions log: 2026-07-28 entry with the per-hop attribution, the three-part fix, the loopback before/after, the kind correctness PASS, and why kind cannot resolve the latency delta. - Home: implementation note for the fix. Refs noetl/ai-meta#205, noetl/ai-meta#208

    @claude claude committed Jul 29, 2026
  • L1 T4 COMPLETE — EHDB is the production command bus (full flip 80/80 on shastaratech prod) Runbook: status banner rewritten to SUCCESS; T4 go/no-go checklist cleared with the latency caveat + a new T5 checklist; prod execution log restructured into three rounds (held-at-canary / failed-on-#203 / succeeded) with the #203 root cause + fix + before-after table, the post-flip prod state, the rollback command, and the latency regression. Program page: L1 status T0-T4 done / T5 held; T4+T5 table rows rewritten. Home: program bullet leads with prod-command-bus state. Sessions-Log: new 2026-07-27 entry. Refs noetl/ai-meta#194, #202 (closed), #203 (closed), #205, #206

    @akuksin akuksin committed Jul 28, 2026
  • docs(ehdb): L1 T0-T3 shadow arc complete + T4 cutover runbook (prepared, gated) Program page: L1 T0-T3 status table (all merged; latency gate PASS, bus p99 137us). New Runbook-L1-Command-Bus-Cutover: the NOETL_COMMAND_BUS flag, cutover sequence + validation gates, go/no-go checklist, and the one-command rollback — PREPARED, NOT EXECUTED (T4/T5 human-gated). Sessions-Log + sidebar updated. Refs noetl/ai-meta#194

    @claude claude committed Jul 20, 2026
  • docs(ehdb): L0 object-store foundation complete — datasets D1-D10 merged Program page: status → L0 COMPLETE (engine L0.1-L0.7 + all predefined datasets D1-D10 merged, main 0af99fc); added the D1-D10 dataset table + the op-log/fold shape note. Sessions-Log: D6-D10 landing entry. Refs noetl/ai-meta#194

    @claude claude committed Jul 18, 2026
  • docs: record Alesha's APPROVED GO to prod durable-SHADOW (sign-off complete; shadow runbook refreshed; no prod action)

    @claude claude committed Jul 12, 2026
  • docs: prod-durability decision packet — D11 GAP->MET + #261 head-to-head re-run (deployed ~740ms->~6ms)

    @claude claude committed Jul 12, 2026
  • docs(durable-eventlog): limits-based retention design + organic in-kind soak (shadow self-bounds 19->2 segments)

    @claude claude committed Jul 10, 2026
  • docs(durable-eventlog): segment-GC worker periodic wiring + in-kind soak result (D11 deployed exercise)

    @claude claude committed Jul 10, 2026
  • docs(durable-eventlog): shared-tier segment GC — coherent reclamation across the shared medium (ehdb#270)

    @claude claude committed Jul 10, 2026
  • docs(durable-eventlog): segment GC design + Sessions-Log (the R1/D11 gap) Document the interest-based segment GC (ehdb#269) on the durable event-log backend design page: consumer-ack watermark retention rule, the base offset (retained log gapless-from-reclaimed+1), the durable reclaim.json manifest as the crash-atomic commit point, write-forward of consumer state, the NOETL_EHDB_EVENTLOG_GC selection surface (off by default), CLI + cost, the shared-tier scope note (follow-up), and the D11 verdict (GAP -> delivered for single-writer local). Sessions-Log entry + remaining-slices checklist updated. Refs noetl/ehdb#254, noetl/ehdb#269

    @claude claude committed Jul 10, 2026
  • perf(#267): local replay-on-open fixed — O(1) checkpoint open; deployed ~0.5s→~4-16ms/op Design-Performance-and-Load-Testing (#267-fixed section + micro-bench flat + deployed before→after + SLO table), Design-Durable-EventLog-Backend (checkpoint sidecar subsection), Sessions-Log (2026-07-09 entry).

    @claude claude committed Jul 9, 2026
  • perf: record merge-on-micro-bench decision (#264 fix merged, worker#175) + explicit deferred in-kind follow-up (post-#267)

    @claude claude committed Jul 9, 2026
  • perf: shared-tier publish fix (#264/#265/#266) + corrected headline — deployed bottleneck is local replay-on-open (#267)

    @claude claude committed Jul 9, 2026
  • docs(perf): Layer B in-cluster results + EHDB-vs-incumbent head-to-head Records the kind-side load-harness numbers (ehdb#263): incumbent Postgres+NATS append ~2150 ev/s p99 52ms; EHDB durable local engine primitive ~3-10ms (corroborates Layer A); as-deployed shared-tier shadow mirror ~0.55-1.7s/append. Headline: SharedTierEventLog::append re-publishes the whole active segment synchronously in emit_event (O(segment)) — the bottleneck is the shared-publish strategy, not the engine. SLO split-by- backend flagged; follow-up = incremental shared publish + off-hot-path.

    @claude claude committed Jul 8, 2026
  • docs(perf): add Design-Performance-and-Load-Testing + baseline; Sessions-Log, Home, Roadmap, _Sidebar

    @claude claude committed Jul 8, 2026
  • docs(kind): Phase 3.5 — #151 PRs merged, event-log leak fixed, Auth0 green, OpenAI+IBKR scoped out

    @claude claude committed Jul 8, 2026
  • docs(kind): Phase 3 — #151 keychain fix + durable GSM bridge; residual = external account limits Refs noetl/ai-meta#151 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @claude claude committed Jul 8, 2026
  • docs(kind-validation): Phase 2 — GSM external integration + Muno end-to-end kind CAN reach Google Secret Manager (Alesha). Deploy-time GSM metadata bridge (host-ADC shim + in-cluster socat relay + worker hostAliases + server NOETL_GCP_METADATA_TOKEN_URL) unblocks both GSM resolution paths. - 9 external providers LIVE with real GSM creds: Duffel, Google Places, HotelBeds hotels/activities/transfers, Firestore, Snowflake, OpenAI, Anthropic. - Muno itinerary planner full flow GREEN in kind (flight->book real Duffel TEST order->hotels->activities->transfers->summary->map, all live data). - Kafka + Pub/Sub subscription drains PASS. - Residual: #151 keychain-template gap (cred reachable, fixture pattern broken), IBKR live-gateway dependency, fixture drift. No GKE/prod touched. No secret values printed. No repos/noetl or repos/server source change. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @kadyapam kadyapam committed Jul 8, 2026
  • kind full-functionality validation Phase 1 — inventory, baseline, first-run matrix Living test matrix for Alesha's kind-first directive (validate all playbooks in LOCAL kind before GKE). Phase 1: inventory (268 files), standardised baseline (all worker pools v5.70.0, EHDB shadow, durable non-primary), first runs — core 61 PASS/4 FAIL, extended 9 PASS/11 non-green, EHDB probes 2/2. Two platform bugs (pagination continuation wedge; large-result output_select resolve 404); rest = cred setup + fixture DSL drift. New page + _Sidebar + Sessions-Log. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

    @kadyapam kadyapam committed Jul 8, 2026
  • docs(durable-eventlog): slice 6 prod-durability sign-off package (ehdb#254) New Runbook-Durable-EventLog-Prod-Signoff page — go/no-go checklist (D1-D11), evidence bundle (slices 1-5), residual-risk register (R1-R7), and the extended durable-shadow -> durable-primary prod rollout sequence. Cross-linked from the tier-1 cutover runbook (§C update + Related), the design page, and _Sidebar; Sessions-Log + Home updated. Docs/planning only — no prod action; ehdb#254 slice-6 box stays unchecked.

    @kadyapam kadyapam committed Jul 8, 2026