Collector 0.7.25 — token-free managed Grok hook (installed-host curl + 0600 header file), doctor hook checks, Codex OTLP record-shape diagnostics
Plimsoll collector 0.7.25 — release notes (eco-6hoxj)
Base: collector main 9892437b (PR #322 merged 2026-09-13 06:49:21Z on the released 0.7.24 main). This release carries eco-6hoxj.67 alone: it rewrites the delivery scheduler's failure handling and introduces the first durable server-directed retry floor, so it gets its own revert boundary and its own observation window (review r1 judgement).
Change
- eco-6hoxj.67 — preserve upload progress and honour server retry cooldowns (PR #322, head b13a0a2 = Astra's fe4ebb0 + lead delta d1c262d + lead correction b13a0a2). A sync cycle that acknowledges rows and then hits a network-class failure no longer escalates:
sync_failedlogsfailureStreak: 0,backoffMs: 0,uploadedEvents: <count>, a timestamp and a concrete code (ECONNRESET,ETIMEDOUT,UND_ERR_SOCKET) and the failed batch retries on the next ordinary cadence tick instead of 10–60 minutes later. Genuine outages still ladder 60 s → 120 s → 240 s → … to the 1 h ceiling and recover on their own. A hosted503withRetry-Afteris obeyed instead of ignored, on both paths and bounded on both: the process-local scheduler cooldown ismin(max(0, retryAfterMs), MAX_DELAY_MS)(1 h) on the failure and the success path (sync-backoff.tsserverDelayMs), and the durable per-row floor written intoupload_outboxis raised to the server'snotBeforebut never pastat + delivery.maxBackoffSeconds(outbox.tsflooredAttemptAt, insideretry()so no caller can talk the ledger past its own ceiling). HTTP-dateRetry-Afteris measured against the responseDateheader, not the local clock. Session follow-ups are carried during a cooldown.GET /statusgainssync.failureStreak,sync.nextAttemptAt,sync.notBefore,sync.lastError(null on a successful cycle) andsync.lastCycleUploadedEvents;plimsoll statusshows the same block from one/statusrequest andsync_failedkeeps itsmessage. The dead maintenance clause in the local-pressure test is gone.proof:sync-backoff(34 checks, with negative controls that reproduce the unclamped 2033/2094 retry dates) runs as its own CI step. - Review: r1 BLOCK on F1 (the new proof was hidden behind the sync-storage-retry step under
bash -e) and F2 (unbounded, durableRetry-Aftercooldown) —lanes/eco-6hoxj.67-review-pr322/harvest/REVIEW-67.mdsha256 2d973d56…; resolved by the lead correction lane (F1–F6,CHANGE-67-fix.patchsha256 0e66ef62…) and checked by the lead with the reviewer's own harness on studio4 (lead-delta-b13a0a29-full.log, exit 0): reviewer-prescribed battery green (sync-backoff 34/34, sync-storage-retry, outbox, delivery, hook-spool 96/96, http-boundary 23/23, upload-timer-bounds, typecheck), clean merge onto main with the suites green there, probe p6 (Retry-After: 200000000→backoffMs 3600000,notBefore+1 h, ledger floor +30 s) and p7 (Retry-After: 86400+ restart → 600 events uploaded at 95 s,remainingDelivery 0) PASS, real homes identical, collector pid unchanged. CONFIRM at b13a0a2 (REVIEW-67-CONFIRM.txt). CI on b13a0a2: attempt 1 failed on a hosted-runner Chrome devtools race in the dashboard browser proof (unrelated; bead eco-6hoxj.75), attempt 2 success (06:49:10Z, run 34742541538, proof job 11m25s).
Cost
- No request-rate increase: a cycle is still bounded by
maxBatchesPerCycle(20) andmaxProbesPerCycle(31), and the partial-success path returns the cycle rate to the configured cadence — the pre-defect normal, not above it. The .64 hosted route budget is unchanged. - One bounded steady-state difference: while the maintenance circuit is open (5 min on a first maintenance failure, 15 min on repeats) escalation is suppressed for every network-class failure, including a genuine hosted outage — worst case one request per
syncIntervalSeconds(≈12 cycles per hour at the 300 s cadence instead of ≈4).
Residuals (tracked / documented)
- A restart drops the process-local cooldown (durable row floors survive), so a restarting collector can poll an endpoint that asked for a pause — bounded to one request per interval per restart (README).
- Witness-only probe failures have no leased row, so their cooldown is process-local only (README).
- A clamped cooldown means a long hosted maintenance window (
Retry-After: 86400) is honoured for at mostdelivery.maxBackoffSeconds(3600 s by default, coinciding with the scheduler's 1 h ceiling), after which the collector retries at its cadence and a repeated 503 simply re-issues itsRetry-After. An operator who lowersmaxBackoffSecondsgets the tighter durable floor and the unchanged 1 h scheduler bound. retryAfterMillisecondsstill accepts any target before year 9999 by design; the bound lives in the two named consumers (serverDelayMs,flooredAttemptAt) and a third consumer must clamp too.- The partial-lease 503 branch is proof-verified against a real 503 and the real SQLite ledger, not forced from a live daemon.
- 0.7.26 candidates under review, not in this release: eco-6hoxj.73/.63 capture-health truth (PR #324) and eco-6hoxj.74 hook-spool durability notes (PR #325).
Rollout
rollout-0725.sh <merge> <pr>: prepare (waits for the push-to-main run) → publish → host facts → assemble (fit sources from the 0.7.24 windows) → stage studio4 / stream studio5+6 / MacBook if reachable → windows (fit re-checked immediately before each) → Studio0 loop (two consecutive fit admits + drain) → native acceptance (native-acceptance-0725.py: sync contract present on /status, sync.notBefore ≤ 3720 s ahead, sync.nextAttemptAt ≤ 90120 s ahead, coherent status). Studio0 acceptance: partial-success cycles log failureStreak: 0 and retry on the next tick; no 10–60 min upload stalls after a network-class failure; remainingDelivery keeps draining. Rollback: the previous release stays installed under lifecycle/versions/0.7.24; the managed window rolls back on a failed restore.