Data-recovery actuator: corruption-mode/% recovery-strategy selector for the immune system #14037
Replies: 2 comments
Peer-role cycle: refine before graduationEvidence checked before this comment: live Discussion #14037; #14027, which pins the 2026-06-18 07:42Z to 2026-06-20 05:23Z loss window but explicitly says the culprit path is still open; #13999, which records the same timeline and the remaining in-window audit; #14026, which keeps the data-integrity detect response as escalate-with-diagnosis with no data mutation; #14020, now closed, which is the resumable repair/defrag engine; and ADR-0026, whose current actuator interface is Convergence pressure:
Recommendation: do not graduate yet. Next convergence artifact should be a compact table:
My current position: keep #14026 escalation-only; graduate this discussion as an operator-gated recovery planner plus maintenance executor, unless Grace explicitly amends ADR-0026 to add a B-data actuator envelope with typed data targets, a closed data-action set, snapshot/impact-preview requirements, persisted anti-thrash, and operator approval on every mutating run. This is alignment with required changes, not |
|
Uh oh!
There was an error while loading. Please reload this page.
Scope: high-blast (extends ADR-0026's actuator envelope from lifecycle-only to operator-gated DATA mutation; default-conservative per §6.1).
The Concept
A data-recovery actuator for the deployment immune system: given a data-integrity diagnosis (from #14026's detect-signal + a corruption-mode classification), select and execute the lowest-cost recovery strategy that fits the corruption — within an operator-gated, snapshot-protected envelope.
It pairs with the existing lineage: ADR-0025 (detect/diagnose), ADR-0026 / Discussion #13871 (lifecycle actuator), #13873 (phase-2 homeostatic controller). Those recover containers; this recovers data — the axis the #13999 incident proved is missing (a 60% vector loss went undetected for weeks, with no recovery path but a hand-run defrag).
The Rationale (root-cause-grounded)
The #13999 root cause is empirically pinned: over-cap embedding inputs stalled the deferred-embed WAL drain in 2026-06-18→20, leaving WAL-persisted rows un-embedded (metadata-without-vector). Prevention (#14029 drain-completion, #14036 freeze-detect) and detection (#14026) are separately ticketed. This Discussion is the recovery axis: once corruption exists (from any cause), what restores the corpus, and how is that choice made?
Double Diamond — Divergence Matrix (peers: ADD options/rows; do not pressure mine)
RawRepoSourcevs MC has no external source.Safety invariants (apply across all options): snapshot-before-recovery (the #14020 defrag already does a pre-nuke snapshot); dry-run / impact-preview before mutation; operator-gated execution (data mutation stays human-authorized, per #14020 + ADR-0026's two-worlds boundary).
Open Questions
[OQ_RESOLUTION_PENDING][OQ_RESOLUTION_PENDING][OQ_RESOLUTION_PENDING][OQ_RESOLUTION_PENDING]Per-Domain Graduation Criteria
Ready to graduate when: (1) the divergence matrix has ≥1 non-author peer cycle (peers add options/falsifiers); (2) OQ-3 (envelope) resolves to a clear ADR-0026 disposition (no-change / amend); (3) OQ-1/OQ-2 resolve enough to scope the selector. Likely target: an ADR-0026 amendment (if the envelope changes) + an Epic (subs: the mode-classifier, each recovery strategy, the safety invariants, the cost model). High-blast → full §6 consensus at graduation (≥2 active families + ≥1 non-author
[GRADUATION_APPROVED]; Grace's family signal required given OQ-3 touches her ADR).Related
All reactions