henyey mainnet daily — 2026-05-18 #2800
tomerweller
announced in
Announcements
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Validator
be066514(uptime 11h 55m; 1 deploy in last 24h after 2 quarantined builds; ~5m avg build)agree=21/missing=0,lag_ms=445)first_to_self_externalize_seconds): 51msexternalized_seconds): 4.77sDeploys (1 effective)
be066514Convert INV-H2 panic to corrective branch in assert_lcl_consistency (INV-H2 panic at herder.rs:940 during catchup replay (separate from #2789) #2791) (Convert INV-H2 panic to corrective branch in assert_lcl_consistency (#2791) #2792)Validating → Catching Up → Validatingwith WARN + tracking advance; INV-H2 corrective WARN count steady at 5 from the initial replay burst)Note: also landed but rolled back (quarantined) before deploy:
52ac1665Tighten henyey-db API-narrowing doc (Tighten henyey-db API-narrowing doc to acknowledge rusqlite exposure #2786) — quarantined after INV-H2 crash loop0df88603Narrow metrics: recovery-stalled burst (delta=240) + 155k overlay backpressure on post-deploy catchup-behind #2713 hard-reset suppression to at-or-near-tip (metrics: mid-run recovery cycle (lost-sync + hard-reset + recovery-stalled) on 52ac1665 at 17h uptime #2789) (Narrow #2713 hard-reset suppression to at-or-near-tip (#2789) #2790) — quarantined after second INV-H2 crash on same line 940Incidents (1 major resolved)
crates/herder/src/herder.rs:940withLCL ≥ tracking_consensus_slotduring fast-tracking catchup burst-closebe066514(Convert INV-H2 panic to corrective branch in assert_lcl_consistency (#2791) #2792 — converted panic to WARN + metric + idempotent self-correct that advances tracking to LCL+1)Issues activity
Filed today (4):
0df88603(Narrow #2713 hard-reset suppression to at-or-near-tip (#2789) #2790)be066514(Convert INV-H2 panic to corrective branch in assert_lcl_consistency (#2791) #2792)b3e4ec83Closed today (9):
b3e4ec83/review-pr PR-detection bugbe066514INV-H2 panic (catchup replay)0df88603mid-run recovery cycle039c7b5a/review-pr Step 2 paraphrasingc33f3b19zsh regression test for monitor-decisions74f4b1a2Archive Done Project Items PROJECT_BOARD_TOKEN diagnose7380c388ScpPersistenceManager purge transactionalbb431484ScpPersistenceManager persist wired22c5a88d(Owner, TempDir) drop-order auditStill open (3 long-running):
Watch items
next_checkpointnot yet published by archive (5 min wait by design). metrics: mid-run recovery cycle (lost-sync + hard-reset + recovery-stalled) on 52ac1665 at 17h uptime #2789 fix narrows hard-reset suppression correctly butspawn_catchupthen bounces off "archive hasn't published checkpoint yet" gate. Validator survives every time (INV-H2 panic at herder.rs:940 during catchup replay (separate from #2789) #2791 fix prevents the resulting INV-H2 crash). Inter-wedge interval has lengthened from 1h7m → 4h11m → 4h13m. Flagged on metrics: mid-run recovery cycle (lost-sync + hard-reset + recovery-stalled) on 52ac1665 at 17h uptime #2789 as possible follow-up.52ac1665,0df88603in/home/tomer/data/deploy_quarantine.txt. Monitor will not auto-redeploy them.Tick aggregates (last 24h)
6c74937d(daily-summary at 13:07 UTC) — 1 expected, 0 fired today (manual post; cron fires consistently 10-30 min late, see open question below). The 20m monitor-tick is run by /loop self-pacing, not cron..claude/skills/directly today;b3e4ec83/039c7b5a/57438a74/c33f3b19are skill or script changes already listed in deploysOpen questions
Daily-summary cron reliability:
6c74937d("0 13 * * *") has fired 10-28 min late every day this week (13:17 on May 17, 13:32 on May 16, 13:35 on May 15) and didn't fire by 13:22 today (this post is manual). The current cron is "session-only" so it only persists as long as the host Claude session does. Migrate to GH Actions cron or accept the lateness pattern?Wedge follow-up: 4 wedges/12h on
be066514(metrics: mid-run recovery cycle (lost-sync + hard-reset + recovery-stalled) on 52ac1665 at 17h uptime #2789 fix landed). Each costs ~5min of RPC unhealth + Catching Up state. The fix narrowed one suppression butspawn_catchupstill blocks on "archive hasn't published checkpoint yet". Worth a follow-up issue to either (a) skip that gate when peer_gap > N, or (b) accept the 5-min wedge as a "wait for next checkpoint" by design?All reactions