Releases: anotherpanacea-eng/apodictic
Release list
v2.11.0
v2.11.0 - 2026-08-26
Specs — Approval-Gated Reconstruction
New companion-module spec docs/approval-gated-reconstruction.md (v0.2.0, unbuilt):
approval-gated reconstruction over a provenance-linked claim graph — normalize
Argument_State.md + manuscript into a content-addressed approval graph, author
adjudicates every node and edge, a drafter authors a fresh document from the approved
subgraph only, and a two-layer conformance gate (mechanical argument-reconstruction
validator + semantic exclusion/coverage/novelty-quarantine checks) proves it, emitting a
reconstruction receipt that /ready validates. Five top-level invariants, including
rejection-stickiness (no reconciliation or revision path may shrink the exclusion set).
Spec folded three independent review passes before freeze; builds in five increments,
starting with scripts/approval_graph.py.
Approval-Gated Reconstruction — ledger-authority redesign (Rev 0.3.2)
Adopt ADR 0002: Approval_Events.jsonl becomes the sole authoritative state, one decision
bundle is one ledger record (carrying its MINTED content payload), and Approval_Graph.md
plus Adjudication_Session.json become deterministic projections rebuilt from the ledger.
Crash recovery collapses to verify/truncate-uncommitted-tail/replay that fails closed, and the
three-artifact cross-consistency checks reduce to deterministic projection rebuild. This supersedes the
multi-artifact recovery/reconciliation machinery of Rev 0.2.1–0.2.3, whose crash-prefix cases
move to an Increment-1 truncate-and-replay fixture family. Still an unbuilt spec; the
built-when marker does not fire.
- Close the authority gaps found after the redesign: stored terminal-verifiable bundle hashes,
ledger-resident source context, a legal QUARANTINE bundle, carried-typing edge identity,
newline-delimited torn-tail recovery, ledger-derived notes, cache rebuild semantics, and an
ephemeral process lock for live-session state. - Pin canonical nested JSON key order, novel-origin mint payloads, current-draft version identity,
QUARANTINE Origin/provenance/context agreement, atomic anchors/flags refresh replay, and
distinct machinereasonversus projectednoteroles.
Approval-Gated Reconstruction Phase 0
Harden the unbuilt companion-module specification before Increment 1: require explicit
graph/draft-ready/acceptance stage selection, define separate node and edge field shapes,
make History entries resolve to a canonical hash-chained event ledger, specify reconstructible
session cache state plus an ephemeral live-process lock, state the limit of singleton
rejection-history detection, and
fail Stage C closed until Increment 4 defines a structured monotonic gate-config order.
Bind event labels to transitions, ledger Inclusion changes and revision pairs, preserve
human-readable History summaries, and pin Argument State 0.3 metadata handling. Define
multi-event decision-bundle recovery, node-only Inclusion and per-axis ledger replay,
canonical provenance entries, and the literal legacy-untyped edge representation.
Argument Register
- Add writer-confirmed AT5 generative/lens routing with GN0–GN4 diagnostics, explicit
high-stakes precedence, and per-cash-out asserted-burden records. - Version Argument State Schema as v0.3.0 and extend argument-spine/Argument_State
contracts so register calibration remains inspectable and never silently bypasses the
Deficit Lock. - Close the E7 hybrid gate with two independent strict blind records whose canonical cash-out
joins preserve generative treatment for the journey and full asserted burden for the policy
landing; both vendored ledger/state pairs passstance-calibrationand are exercised by
validate.sh --check-all.
Fiction calibration and editor-panel evidence
- Pinned and executed the fiction M2a slice with cross-vendor Terra + Opus,
including the 11-member reconstructable corpus, strict blind runner, and
committed audit package (22 hash-bound outputs, tidy scores, manifest,
scorecard, and verifier). The first run converged on 3/4 matched pairs and
correctly remains red on continuity clean-member specificity pending a
locked-key corrective retest; Lane-2 results remain report-only until M2b. - Regenerated the access-controlled blind-editor packet as one shared argument
M2 / fiction M2b closed-response panel: ≥3 sealed independent editors, GT8 and
Reliability-ledger coverage, exact α transforms, seven independent fiction
reliability units, frozen key projections, and judgment-free tidy-CSV
compilation. The packet is ready; recruitment and measured licensing remain
human-gated.
SETEC contract safety
- Gate the vendored
voice_distancefixture on SETEC'sregister_families/v2register-family contract: both taxonomy markers, astrengthdrawn from the closedstrong/moderate/weak/mismatchvocabulary, and no legacyverdict. The check is unconditional, so a sync that re-vendored a pre-v2 fixture fails rather than passing on a grandfather clause. Rejection cases run against a hand-built payload, never the vendored golden — a negative arm fed by a producer artifact goes vacuous the moment that artifact legitimately changes.
Faster M5 override-hygiene scan
Anchor validator-conventions' M5 Python scan on literal <!-- occurrences instead of running the four-alternative override-marker regex from every position of every sibling script. Same verdicts (self-test cases guard the windowed semantics, including whitespace-gap and deep-offset forms); the canonical CI gate drops from ~28s to ~15s locally.
Faster CI without weaker gates
Run validator self-tests, canonical-framework checks, and static/build checks in parallel behind the existing required validate result. The aggregate validator also avoids executing the shared gate/gate-state self-test suite twice.
Docs — SETEC supplement spec reconciliation
Corrects docs/pass3-pass7-setec-supplement-spec.md: its 2026-06-19 reconciliation box asserted that the SETEC surfaces sliding_window_heatmap and voice_drift_tracker "do not exist," but both shipped in setec-voiceprint before that note was written (902 and 1,421 lines, with test suites and registry fragments). A dated correction box records the capability ids, shipping commits, and a verify command; the stale bullet is kept as provenance and flagged. The §3 rows and §8 cross-references now carry the capability ids plus the real reason neither is reachable from APODICTIC — both fragments carry consumers: [] and no json_delivery, so they miss both the sync_setec.py vendoring filter and the setec_run.py dispatcher promotion. Pinning them is producer-side work, recorded as follow-up. Also refreshes the stale v1.117.0 vendored-contract reference to the pinned v1.126.0.
Validation portability
- Make the refutation hostile-test assertions robust under Bash
pipefail, so a correctly rejected fabricated-budget fixture cannot be misreported as missing its diagnostic evidence.
Release safety and portability
- Build Codex and Antigravity ZIP assets with a fail-closed, stdlib-only Node packager instead of requiring external Unix
zipandunzipbinaries. - Force validator-launched Python processes to emit UTF-8, keeping hostile-output assertions and release verification deterministic on Windows consoles.
- Preflight both host package builders before mutating release files; the staging pipeline never creates tags, refuses when Git state cannot be inspected, and the separate owner-only tag helper only cuts from a clean
mainat freshly fetchedorigin/main.
Rhetorical Stance Triage
- Add pre-lock earned/unearned stance triage for overstatement-family argument findings,
with explicit writer authority and Firewall boundaries. - Add a mechanical shape/join validator that proves recorded calibration effects and
blocks demotion at high-stakes or prescriptive cash-out spans without judging rhetoric. - Require every recorded earned verdict to carry its register-floor or stance-demotion
effect, reject stance metadata on GT8 premise flags, and keep blocked unearned or
divergent findings representable at full severity. - Document the validator's completeness boundary: supplied cash-out joins are checked,
while omitted joins and undeclared stance decisions remain auditor-owned. - Recognize realistic
GT8 …mechanism identifiers, pin register-floor precedence over
instance stance, and state why post-triageCould-Fixcannot prove a blocked demotion. - Reject finding-level generative floors under an asserted document, restrict actual
demotions to the declared WR/SM/BP and overstatement families, and parseNONEcash-out
and high-stakes source records exactly rather than by substring.
Roadmap reconciliation
Reconciles the strategic board to the released v2.8.0–v2.10.0 state: records the completed argument-regrounding descent and visualization chart set; preserves the genuinely open benchmark work; and corrects the editor-panel prerequisite to the current shared, regenerated ≥3-editor protocol.
SETEC consumer client contract (Increment 1)
- Fixed the version parser's silent-partial-parse bug (
_parse_version("garbage") == (), and a prerelease like1.129.0-rc.1could satisfy a numerically-equal stable floor).setec_discovery/setec_capabilitiesnow use a SemVer-subset parser +meets_floorthat correctly rejects1.129.0-rc.1 < 1.129.0; an unparseable version always raises rather than silently truncating. - Changed the warning classifier's unmatched-prose fallback from
cosmetictoreliability(fail-upward default) — the classifier itself is unchanged and permanent, only its default tier for text that matches none of the 11 reliability patterns. - Added a named drift tri-state (
offline_unresolved/resolved_match/resolved_drift) totools/check_setec_contract.py's live-drift check, with a hermetic self-test proving a...
v2.10.0
v2.10.0 - 2026-07-13
Argument engine — AGD scan/audit agreement benchmark (R3B §4 Phase-2 scoring)
R3B §4 of the argument re-grounding roadmap — the offline, deterministic scoring half of the AGD producer/consumer behavioral benchmark. Phase 1 (acquisition) is committed as 20 run-manifests under evals/benchmark/agd-scan/manifests/ (5 R3A fixtures × {fable, codex} × rep{1,2}); the new evals/benchmark/agd-scan/score.py (stdlib-only) replays each one offline through the consumer channel and scores scan↔audit agreement.
Replay. Each manifest is replayed through the same shim an audit runs — ai_prose_agd_move_scan.py --judge manifest --judge-manifest <manifest> --expect-fingerprint <the manifest's own fingerprint> --json (SETEC_VOICEPRINT_DIR → the producer worktree; taken from the environment, a scoring run errors loudly if unset) — taking results.observations from the envelope. A non-available envelope or a fingerprint-drift refusal is a scoring ERROR for that cell, loudly reported and never silently skipped (and it fails the hard gate if it hits a gate cell).
Coordinate + match rule (mechanical §1b). An M-record's coordinate is its resolved Source anchor: the anchor's verbatim quote is resolved in the source under argument_agd's normalization (whitespace folded, quote/dash chars normalized, case-sensitive), the source is split into the producer's blank-line paragraphs, and the one normalized paragraph containing the normalized anchor is the record's derived paragraph (zero or >1 = a loud error, never a guess). A scan observation matches iff same derived paragraph AND same family AND normalized-span interval overlap (each span at its first occurrence; adjacent-but-disjoint spans do not overlap). Per-M-record outcome (§4 taxonomy): CORRECT = ≥1 same-family overlapping observation; INCORRECT = ≥1 overlapping observation, all a different family; SOFT = no overlapping observation. The prose locus label is never a comparison operand. The §10.9 parsing and Source-anchor normalization are imported from scripts/argument_agd.py (extract_section, parse_block, _norm, anchor_resolves), not forked.
Reporting + ownership. Per cell the harness enumerates misses (records with no same-family overlapping observation) and extras (observations overlapping no M-record coordinate — calibration data for the audit's decline path, never an error). The one hard gate (§4 wedge claim): fixture cue-free-structural-discounting, M1 (family DISCOUNTING) reaches ≥1 CORRECT rep per vendor and 0 INCORRECT reps across all reps — score.py exits 1 if it fails. --self-test covers the matcher/taxonomy/gate on synthetic in-memory arms (no file IO); --report / --write-report emit the §M-AGD section into docs/argument-benchmark-calibration-round.md, generated by running the scorer (result cells are never hand-written). The scorer measures agreement only (R4A ADR D5: the producer observes, the consumer alone adjudicates) — no code, no verdict, no observation-count-as-quality, no aggregate beyond the agreement bookkeeping. A purely additive harness: manifests, fixtures, argument_agd.py, and validate.sh are untouched.
Sources: Sinnott-Armstrong & Fogelin, Understanding Arguments 9e, ch. 3 (assuring / guarding / discounting); R3A AGD Move Audit lineage (argument-agd-audit.md); R4A ADR D5 (producer observes, consumer adjudicates); the R1 convergence-record convention (docs/argument-benchmark-calibration-round.md).
Argument engine — R4B AIF export + ADR review folds (non-blocking)
Folds five non-blocking review findings from the R4B pass. ADR hygiene: restores the
accepted docs/adr/0001-argument-layer-boundary.md Consequences sentence ("those unmapped
rows are the visible R4B/R2/R3 work-list") that the R4B PR had silently rewritten, and moves
the reinterpretation ("R4B reviewed those rationales into present-tense conclusions rather than
treating them as a quota for speculative mappings") into the dated R4B implementation note,
so the amendment history is explicit rather than overwriting the accepted decision.
Exporter (scripts/argument_aif.py). The _populated placeholder heuristic no longer
treats any underscore-led line as an empty stub: only the reserved machine-seeded shapes
(_seeded…_, _pending…_, or a bracketed […] stub) count as unpopulated, so authored italic
content smuggled into an out-of-profile section (e.g. a §7 body of _smuggled real out-of-profile content_) now discloses its OUT-OF-PROFILE-SECTION loss instead of being dropped silently,
while genuinely-seeded stubs still disclose nothing. Under --source, check now binds the
export's recorded source.artifact basename to the supplied source filename, so a tampered
basename can no longer ride through a source-closure check. Sourceless check now enforces the
canonical key order (it re-emits the parsed object in the exporter's emission order and
compares bytes), so a key-reordered-but-otherwise-canonical artifact fails without a source;
docs/argument-aif-export.md states that loss-set completeness still genuinely requires
--source. The self-test banner now derives its arm count from a counter (17 arms) instead of a
stale literal, adds a dedicated final-support-enum-invalid build-failure arm, and asserts both
the out-of-profile disclosure and the reserved-placeholder non-disclosure. The committed goldens
(evals/fixtures/argument-aif/final-audit/export.json and the reference example) are unchanged
and regenerate byte-identically; the scripts/ ↔ plugins/apodictic/scripts/ mirror stays
byte-identical.
Argument engine — R4B vocabulary migration + optional AIF-Core export
Renames Step 9's third decision test from Warrant-recoverability to
Bridge-recoverability without changing its question, ordering, verdicts,
codes, or scoring; Dialectical Clarity advances to v2.1 while Argument_State
stays v0.2.0. The taxonomy crosswalk advances to v0.2.0 with a separately
drift-bound eight-row concept layer for R2 relations and R3 AGD moves; the closed
82-code layer remains unchanged and all 48 unmapped rationales are now stable
present-tense conclusions.
Adds the optional apodictic.aif-core-export.v1 one-way adapter: deterministic,
atomic JSON with source-owned claim/support IDs, verbatim warrant scheme_text,
typed I/RA/CA incidence, explicit loss accounting, final-vs-pre-draft precedence,
and byte-level --source closure. It never invents claim-ladder, PA, alternative-
conflict, or AGD topology and never feeds external graphs back into diagnosis.
Mechanical gates certify shape and closure only; mappings remain reviewed
scholarly assertions. Source: ARG-tech, The AIF Specification, Definition 1.1
(AIF graph); Toulmin, The Uses of Argument (1958); Sinnott-Armstrong & Fogelin,
Understanding Arguments, 9th ed.
ClaimLicense ↔ Toulmin crosswalk (R5)
Documents the fleet's warrant/license resemblance as a cross-level analogy rather than an identity. A new field-by-field crosswalk distinguishes Toulmin's object-level argument roles from SETEC Voiceprint's meta-level ClaimLicense, fixes the current whole-object→warrant mapping as unmapped, and supplies hostile anti-equivalence cases plus a promotion test for any future explicit entitlement_basis / inference_rule. The snapshot is grounded in setec-voiceprint v1.124.0 and its schema_version: 1.0 producer envelopes; this is documentation and ADR clarification only, with no SETEC edit, vend, dependency-floor change, schema, validator, or runtime behavior change.
v2.9.0
v2.9.0 - 2026-07-12
Validators — Krippendorff's α (agreement-alpha)
A new pure-stdlib validator, scripts/agreement_alpha.py, computes Krippendorff's α inter-rater reliability over a tidy rater,unit,value CSV — the mechanical half of the benchmarks' agreement-as-license promotion workflow, shipped before any human panel data exists so the M2 round is analysis-ready and α is computed by audited, self-tested code rather than an ad-hoc notebook (and shared with the fiction benchmark's M2b, which specifies no calculator of its own). Run via validate.sh agreement-alpha <ratings.csv> [--metric nominal|ordinal]. It builds the coincidence matrix (each within-unit ordered value pair contributing 1/(m_u−1); units with <2 values contribute nothing), then α = 1 − D_o/D_e under the nominal (δ²=0/1) or ordinal (Krippendorff's cumulative-marginal metric) difference function. Contract, pinned: a required rater,unit,value header (missing/wrong → ERROR), stdlib-csv quoting, blank value = missing, a malformed row (wrong column count or empty rater/unit) → ERROR naming the line number (never silently skipped); D_e=0 (a constant column) → alpha=UNDEFINED (D_e=0) + WARN, treated as not clearing any threshold. The bootstrap 95% CI resamples UNITS with replacement — ≥1000 resamples (default 1000; --resamples may only raise the floor), the Hayes & Krippendorff (2007) unit-level scheme — from a fixed default seed (--seed override) so two auditors reproduce the same interval byte-for-byte; licensing is on the CI lower bound, not the point estimate (the small-n false-promotion guard). The panel floor is enforced mechanically: panel-licensed requires a ≥3-editor panel (docs/argument-benchmark-spec.md §GT schema), so a run with fewer than three participating editors (those carrying ≥1 non-missing rating — an all-blank column is a listed-but-absent editor, not a rating one) still computes and prints α for information but returns an explicit non-clearing FAILED (exit 2) with no OK: line, which the promotion path rejects; a two-rater high-agreement dataset can no longer back into a license. Its hermetic --self-test locks the arithmetic to hand-derived values — Krippendorff's published worked example (nominal α = 0.691, n = 26), a fully worked tiny nominal case (8/15), an ordinal-beats-nominal ordered example (0.700 vs 0.4545), perfect agreement (1.0), systematic disagreement (−0.5), the D_e=0 UNDEFINED path, a deterministic seeded-CI regression lock, the missing-header / malformed-row / single-rater / single-unit ERROR arms, and the panel-floor arms (a two-rater run prints α but is non-clearing; a ≥3-editor panel licenses; the reviewer's 100-unit two-disagreement high-α repro is rejected; an all-blank third editor does not count toward the floor). Registered on all five surfaces (AGG token, dispatcher case with --self-test + python3-degrade path, help line, byte-identical scripts/ ↔ plugins/apodictic/scripts/ mirror, this fragment); no --check-all corpus block — panel ratings live outside git. No hard-coded validator count is introduced (AGG_COUNT stays derived). Method sources: Krippendorff, K. (2011/2013), Computing Krippendorff's Alpha-Reliability; Hayes, A. F., & Krippendorff, K. (2007), Answering the Call for a Standard Reliability Measure for Coding Data, Communication Methods and Measures 1(1), 77–89.
Nonfiction Argument Engine — Reliability Ledger (Agreement-as-License, GT schema v0.3.0)
The Argument Benchmark's ground truth now carries its trust distinction as machine-checked data rather than prose. Each fixture's Provenance block gains a required Reliability ledger line assigning every GT anchor a status (authoritative / provisional / panel-licensed / low-agreement) and a decision-use (gate / confirm / report), and the GT schema bumps to v0.3.0. scripts/argument_groundtruth.py gains Check 6: the ledger group grammar (GT<a>(–GT<b>)?: <status>, <use>, (?![0-9]) boundary-guarded so GT10 cannot truncate-parse, |//-in-value rejected as copied guidance), exact GT1–GT8 coverage (no gaps/overlaps), the enforcement matrix (gate requires a licensed status — authoritative/panel-licensed; provisional may only confirm/report; low-agreement may only report), and a stale-heading cross-check that finally consumes the formerly-dead (PROVISIONAL) heading bool — a licensed status under a PROVISIONAL-marked heading is an error (the M2-promotion tripwire). A sibling round-record conformance mode (--round-record <record.md> --fixtures-dir <dir>, wired into --check-all over docs/argument-benchmark-calibration-round.md) is the one mechanical guard at the run-side attribution seam: every booked - BOOKED: ENGINE-FAULT <fixture-slug> GT<n>[ OVER-FIRE] must cite an anchor whose ledger licenses it (gate: any booking; confirm: only with an explicit OVER-FIRE tag — the asymmetric ruling; report: none). All 16 registered groundtruth.md fixtures migrate (one added ledger line each; the GT8 (provisional migration default) parentheticals preserved byte-for-byte; two PD-control Scope sentences corrected to GT4–GT6 and GT8 are provisional / advisory), and Checks 1–5 plus the strict GT8 contract are byte-stable. The convergence criteria are unchanged; only attribution becomes reliability-aware and asymmetric — a confirm-anchor over-fire stays ENGINE-FAULT (the specificity gate does not soften), while a confirm-anchor false-negative downgrades to key-suspect (ground-truth ambiguity) and routes to run-blind re-registration or the blind M2 panel, never an engine regression. Promotion α thresholds are pre-registered in the spec (Krippendorff's canonical cutoffs: α ≥ .800 → panel-licensed, .667 ≤ α < .800 → provisional, α < .667 → low-agreement; ordinal-weighted for the GT4 sub-scales, nominal for GT2/GT6/GT7; licensed on the bootstrap-CI lower bound), computed by the audited agreement-alpha CLI shipped separately. This is the argument-benchmark instantiation of the psychometric frame the fiction benchmark already ships (scripts/fiction_groundtruth.py) — inter-rater agreement licenses whether a label may gate; it never scores the engine — with a deliberately domain-honest token set (a documentation-only fiction↔argument cross-map, no drift guard). 76 parser self-test arms total (40 inherited from the warrant-split base), 36 new — Check-6, round-record, plus a Codex-#193-review hardening pass: the ledger field is anchored to the exact - **Reliability:** label within the ## Provenance block (a near-label like - **Not Reliability:** or a misplaced ledger under ## Notes is rejected, not silently accepted); the provisional heading bool is sticky-true across every heading covering a GT anchor (a later duplicate heading cannot clear a (PROVISIONAL) marker and evade the stale-heading tripwire); and the round-record BOOKED matcher is claim-broad/validate-strict — house-bold - **BOOKED:** is legal, while near-miss dialects (- **BOOKED** — …, - BOOKED:: …, - BOOKED* …, * BOOKED: …) are loud errors, never silent skips.
Argument engine — AGD Move Audit companion (argument-agd; Argument_State §10.9)
R3A of the argument re-grounding roadmap — the frontier build: a companion audit that inventories the text's performative argument moves — ASSURING (authority/certainty in place of support), GUARDING (a claim weakened to shrink its commitment), DISCOUNTING (an objection anticipated and set aside) — identified functionally at a transition, not by cue words (cue lexicons under-determine function: most hedge-cue lemmas also occur as non-cues — Velldal et al. 2012, Comp. Ling. 38(2), BioScope, domain-specific; the identification-by-transition posture is analogically scaffolded on Inference Anchoring Theory, Budzynska et al. 2016, a dialogue result whose monological transfer is this audit's own claim, validated by its fixtures). All three moves are legitimate; the audit never codes a move for being a move. Each inventoried move is challenged with its family's protocol — ASSURING × STRIP (purely subtractive: delete the assurance span; does independent support remain? — S&F 9e ch. 5), GUARDING × COMMITMENT (de-hedge-only: force the text's own claim to its unguarded form, adding no specificity the text lacks; ch. 16's self-sealer test), DISCOUNTING × ENGAGEMENT (constructs no text: decoy-vs-strongest reuses Step 6a Test A/B + the R2 Basis field by cross-reference, plus a downstream-consequence prong) — under a total family × challenge × result matrix (SURVIVES / COLLAPSES [/-DECOY/-COSTLESS] / SELF-SEALS (GUARDING-only) / INDETERMINATE (run-but-unresolved, documented) / NOT-CHALLENGED (inventory-only)). GUARDING records carry a Trajectory (S&F's disappearing guard). Candidate diagnoses are flag-only and licensed exclusively by failed function: Candidates: NONE is required on SURVIVES/NOT-CHALLENGED/INDETERMINATE (sole exception: the Trajectory: DISAPPEARING whitelist {FM-A16, WR3, BP4}); candidates come from a derived 62-code namespace (the Dialectical Clarity codes minus the AT1–AT4 type labels, via argument_crosswalk's scoped registry parse — no fourth hardcoded copy — plus FM-A1–20 from argument_groundtruth._FM_A_MAX), carry a PENDING → CONFIRMED/DECLINED reconciliation contract (the human editor or a full engine re-run adjudicates; the companion never writes codes outside its §10.9 annotation block, honoring the annotation protocol), and never enter severity propagation directly. The audit writes Argument_State.md §10.9 behind a machine-readable coverage manifest (declared spans cross-checked against the §2 claim ladder; exclusions with reasons; Completion: PARTIAL = WA...
v2.8.0
v2.8.0 - 2026-07-07
Nonfiction Argument Engine — Reviewer-Anticipation (genre layer, Increment 2)
The genre layer's reviewer-anticipation surface (W5) now rides the argument-spine validator. When a genre_profile block is present and its reviewer_objections array is empty or absent, W5 surfaces the empty required class as a WARN (ERROR under --strict), directing the writer to pre-list the evaluator's likely objections — which seed §6, The Strongest Case Against — before drafting. The Firewall holds: W5 checks only that the writer's list is non-empty; it never authors, suggests, or validates the content of an objection. Overridable via <!-- override: argument-spine-reviewer <genre> — <rationale> --> (a W5-specific slug, decoupled from W4's argument-spine-genre). W5 adds no validator — the derived self-testable count is unchanged. The three canonical genre worked examples (grant / academic / pitch) each carry a reviewer_objections pre-list seeded into a §6 Reviewer-anticipation block, so they pass --check-all under --strict; a per-genre reviewer-anticipation calibration table (evaluator stance · top objection classes · W5 pre-list → §6 mapping) lands in Dialectical Clarity's Genre & Audience Calibration.
Validators — cross-surface uncalibrated-band calibration-honesty guard + decision-audit doc fixes
Both consumer decision-audit surfaces — narrative_decision_audit (StoryScope) and argument_decision_audit (ArgScope) — ship an uncalibrated verdict band with null thresholds; the discipline that the band is surfaced as provenance only, never as a calibrated verdict was enforced entirely in prose (the audit docs' anti-verdict sections + four pass-dependencies.md rows), with no mechanical check that a reader-facing rendering doesn't present the uncalibrated band as calibrated. New AGG_VALIDATORS arm calibration-honesty closes that gap: a whole-letter, per-paragraph scan (region-scoping was a false foundation — the only per-audit sectioner scopes appendix bodies only, decision-audit findings land in the unscoped synthesis body, and the canonical scaffolded letter carries no decision-audit region) that flags a paragraph only when a claim shape matches and no qualifier is co-present — CS1 band-placement (… scores in the … band), CS2 calibrated-verdict, CS3 threshold-claim on a thresholds-null surface, CS4 aggregate-as-verdict; a co-present uncalibrated / provenance-only / advisory / not a verdict qualifier quenches the flag, so the mandated clean-path boilerplate, the seven vendored StoryScope bundle labels ("AI-elevated: Structural streamlining" …), ArgScope's two bundle labels, and the doc's own "fair" summary sentence all PASS. WARN by default, ERROR under --strict (the honest class — an open natural-language co-presence check, content_advisory W1 / stance-consistency F4, failure direction toward NOT firing): the guard closes the enumerated claim shapes; residual semantic leakage no vocabulary can close is presented to a human who owns the reading. Disjoint from severity-floor (D5): the submission-readiness band vocabulary (Strong Fit / Highest Band / …) is excised from CS1, so a legal readiness verdict does not fire here, and calibration-honesty never reads flag counts. Per-paragraph override via the shared override_marker SSoT (<!-- override: calibration-honesty — <why> -->, code-span-stripped, boundary-matched). Cross-surface + surface-agnostic — a future uncalibrated handoff: experimental consumer inherits it by construction. A dedicated canonical example-decision-audit-letter.md (a clean narrative-decision finding + a violating argument-decision finding coexisting) exercises the arm non-vacuously in --check-all. The CS shapes / qualifier set / allowlist are this guard's own new vocabulary (not the Must/Should/Could-Fix severity token, so severity_vocab/M8 does not apply). Firewall-hardening chore; no paper cite.
Also two docs-only decision-audit fixes: narrative-decision-audit.md's Version-floor line still described a consumer-side resolve_floor pre-check (pre-R2 language, contradicting the same file's corrected line and the shim it documents) — rewritten to the post-R2 runtime-dispatcher enforcement matching the ArgScope sibling; and AUDIT_SELECTION_MATRIX.md's Narrative-Decision cell described a non-existent choice-and-consequence propagation lens — rewritten to describe StoryScope feature scoring.
Coaching Deepening — Coaching History & Pattern Recognition (the ethics-sensitive surface)
Over multiple revision cycles the coach can now surface a cross-session process pattern — the same finding deferred across an unbroken run of sessions (deferral-recurrence, floor ≥3), or a revision-arc phase left open across consecutive sessions (phase-incompletion, floor ≥2) — as a rolling, opt-in, local-only [Project]_Coaching_History_[runlabel].md of apodictic.coaching_observation.v1 blocks, each mechanically derived from the recorded finding-disposition / revision-arc records (a count, never a vibe) and carrying no editorial severity (no Must/Should/Could token, no apodictic:finding block). This is APODICTIC's ONE ethically-sensitive surface; the two Fable conditions (2026-07-05) are mechanized as the load-bearing gates.
New coaching-history validator (scripts/coaching_history.py, dual-mirror) + the apodictic.coaching_observation.v1 schema: H1 schema + unique CH-NN + per-pattern count floor, H2 provenance/anti-fabrication (every <F-id> deferred @ session <n> resolves to a recorded deferred disposition; len(evidence) ≥ count; the cited sessions are actually consecutive — a gap fails), H3 descriptive-not-judgmental (reuses author_fingerprint._PRESCRIPTIVE_RE + a trait-blame lexicon), H4 no-severity-leak (severity_vocab.SEVERITY_TOKEN_RE), H7 tentative-framing (no trait verdict / no bare-scoreboard rendering / an invitation present — the transference-health floor), W1 local-only. The two ethics gates: H5 single-home / no coach-only shadow and H6 + the delete <project_root> subcommand deletion-honored (RECOMPUTE-not-trust per the PR #161 recorded-field rule / disposition_check DP2.6 — under the <!-- coaching-history: deleted --> tombstone the full H5 scan must be empty + no surviving artifact + no seq; residue = ERROR with no override; opted-in+deleted = ERROR). delete removes every artifact (root + runs/*), drops the sidecar seq, and flips consent to the tombstone. Opt-in gate (home: Diagnostic_State.md; no marker → no-op exit 2). revision-coach/SKILL.md §9 carries the human-terminus contract (opt-in gate, single-home / no-projection, read-by-id, tentative-noticing presentation, surface-the-delete). Spec: docs/coaching-history.md.
H5 scan coverage (Opus review F1, P1). The projection-ban is only as good as its coverage: an early draft scanned a 6-glob authored-artifact allowlist, so a coaching observation projected into an unlisted type (the editorial letter, *_Core_DE_Synthesis_*) escaped H5 and survived delete — a false "deletion honored" that broke both Fable gates (H6 rests on H5's projection ban). Fixed: the self-identifying (i) parsed coaching_observation block + (ii) schema-id signatures are now scanned over every *.md in scope (minus the one Coaching_History artifact — zero false-positive risk, no legitimate file carries either); the (iii) evidence-grammar + bare-CH-NN scans (which do carry manuscript false-positive risk) stay manuscript-excluded, but the manuscript is now identified positively via the persisted *_Manuscript_Snapshot_*.md (annotation_manifest._SNAPSHOT_GLOB), not an allowlist. Regression fixtures (self-test + --check-all): an editorial-letter projection FAILs H5 and, after delete, is CAUGHT by the H6 recompute; a manuscript snapshot carrying "CH-12" / prose "@ session" does not false-trip.
H2 anti-fabrication honestly scoped (Opus review F2). The guarantee is fabrication-resistant ONLY on governed projects (the independent per-session gate_events[].disposition_deltas records). On non-governed projects the multi-session history is self-reported by /coach (the folded record's sessions list), so H2 now requires a non-governed deferral-recurrence observation to carry a visible honesty caveat that its streak is from the coach's own notes, not an independently verified record (WARN, ERROR --strict, per-id override — the writer is never shown an unverified pattern as if it were checked; the H7 principle). SKILL §9 now instructs /coach to persist the sessions list at disposition time and to state the caveat; the canonical fixture teaches both.
Completeness sweep — the incomplete-verification class, closed across ALL patterns and ALL depths (Codex round-1 P1 + P2 + the class sweep). The same "checked the known locations, not all of them" class that produced F1/F2 recurred twice more; swept whole:
- P1 (phase-incompletion verification, BLOCKING).
phase-incompletionevidence was shape-checked only — a fabricated streak at non-existent sessions (101/102) passed--strict. Grepped exhaustively: no per-session revision-arc-phase completion is recorded anywhere (finding_statesis a rolling last-write map;finding_deltascarry no session ordinal; the arc is a stateless overwriting re-plan). So phase-incompletion has no governed verification path on any project — it is inherently self-reported. H2 now (a) does not claim a verification it cannot perform, (b) requires the honesty caveat on every phase-incompletion observation, governed or not. Pattern-by-pattern status is documented:deferral-recurrence= verified-on-governed / caveated-on-non-governed;phase-incompletion= caveated-always. No pattern is shape-checked-only. - **P2 (recursive se...
v2.7.0
v2.7.0 - 2026-07-02
Disposition supersedence is recomputed, never trusted (disposition-check DP2.6)
Sibling-sweep fix from the PR #161 Codex P1 class (an exemption gate trusting a recorded field):
disposition-check's active() exempted a declined/deferred Must-Fix from the DP1 readiness
caveat on the sidecar's self-reported finding_states[id] == "revised" alone — one JSON field
edit on a non-governed sidecar waived the caveat, silenced the DP2.5 sync audit for the id, and
no validator on the enforcement path corroborated (finding-trace E5 is deliberately scoped to
report-mentioned ids; /ready runs only disposition-check at verdict time). Supersedence now
requires a corroborating <!-- resolved: F-id --> marker in a reachable completion artifact
(run folder, project root, or runs/* archives — evidence-only, so DP2.1's same-run scoping is
untouched); an uncorroborated revised leaves the disposition ACTIVE (DP1 caveat still owed)
and is surfaced as the new DP2.6 WARN (ERROR under --strict). Adds 10 self-tests (the
disposition-check suite → 46): the keystone exploit repro, archived-evidence exemption,
evidence-not-leaked negative, DP2.5-no-longer-suppressed, non-UTF8 evidence fails closed, and a
run-folder + runs/* end-to-end both directions; --check-all gains hostile arm 4 (fabricated
supersedence must fail closed). Build-review folds: _read gains the UnicodeDecodeError guard
(the adjacent-exception class this PR later seeded a repo-wide sweep for), and the trusting-rule
sentences in state-lifecycle.md / submission-readiness.md are amended to the corroborated
form. Spec: docs/disposition-supersedence-recompute.md.
Adversarial finding disconfirmation: HIGH means "survived"
New synthesis-time Finding Disconfirmation Pass — Step 6b in run-synthesis.md
§Processing Protocol (lettered, no renumbering; after the Deficit Lock and the Adversarial
Self-Check, before the stress test). Per eligible finding (Must-Fix ∪ HIGH Should-Fix, cap
15 — Must-Fix always processed), the pass hunts counter-evidence spans, generates rival
readings, judges survived | weakened | refuted, and records the attempt in a new
[Project]_Refutation_Record_[runlabel].md artifact (apodictic.refutation.v1 +
apodictic.refutation_budget.v1 blocks; hybrid/swarm runs dispatch a dedicated
disconfirmation subagent whose input excludes the letter draft and confidence tokens —
anti-anchoring). The synthesis agent then transcribes only the confidence consequences into
the locked ledger blocks: survived = unchanged (never confidence-raising — HIGH now
requires a survived attempt, output-policy.md §Confidence Calibration), weakened capped
at MEDIUM, refuted = LOW/UNCERTAIN, severity never remapped. Cap-bound HIGHs are disclosed
(<!-- refutation: not-attempted-budget F-… --> + Appendix B), never silently skipped; a
refuted finding ships labeled with its counter-evidence quoted, never dropped.
Three new validators in one scripts/refutation_check.py module (registered in
AGG_VALIDATORS, the run_spot_check gate, and --check-all against an extended
example-run-folder fixture with hostile arms): refutation-coverage (no HIGH without
survived refutation; dangling ids; marker abuse), refutation-evidence (verbatim
single-line snapshot quotes — fabricated quotes void the attempt; snapshot_sha256 binding;
missing-snapshot split — ERROR on core-de/full-de runs under --require-snapshot/run-shape
detection, WARN-with-demotions-void elsewhere; budget arithmetic), and
refutation-write-scope (no severity channel in the record; exact confidence transcription
per the outcome caps). Behavioral eval fixtures land under
evals/fixtures/finding-disconfirmation/ (rubber-stamp, demotion-abuse, budget).
Finding Dispositions — engine-level declined/deferred set-asides
A deliberately set-aside finding is now a mechanical record, not a prose note a later session must
remember to read (docs/finding-dispositions.md). Dispositions are an overlay, never a fourth
lifecycle state: execution.finding_dispositions maps F-id → apodictic.finding_disposition.v1
(new closed-key schema, no severity field by construction) beside an untouched
finding_states — the fold stays rank-monotonic, finding-trace E3 unchanged. Pinned durable
markers (<!-- declined: F-… — reason --> / <!-- deferred: F-… until: trigger — reason -->) are
the canonical source, with ONE shared grammar helper (apodictic_artifacts._DISPOSITION_RE +
parse_disposition_markers(), code-span-stripped via the override_marker SSoT) imported by both
run_gate.py and the new validator. Dual writer mirrors revised exactly: governed projects
freeze full records into gate_events[].disposition_deltas at revision_round clear (fold-derived
map, pointer == fold, same-event revise+disclaim launder rejected by gate-state); non-governed
projects write the validated record directly. New disposition-check validator (self-testable
count is derived): DP0 record shape (trigger-iff-deferred), DP1 the /ready teeth — an active
declined/deferred Must-Fix must be named on the assessment's pinned **Declined/Deferred Must-Fixes:** caveat lines, run BEFORE the verdict is delivered — and DP2 no-laundering
(same-run resolved+declined contradiction, phantom keys, ledger↔calibration severity mismatch,
triage-tally decrement via the shared structured_findings.severity_tally, bidirectional
marker/sidecar sync with governed-lag exemption). Canonical fixture
references/example-run-folder-dispositions/ wired into --check-all with three hostile arms
(caveat stripped → DP1 fails; trigger stripped → DP0 fails; record dropped → DP2.5 warns, --strict
fails). Consumers: the revision-coach Loop Dispatch skips declined ids, trigger-reviews deferrals
(ISO dates fire mechanically), and records set-asides at the stalled off-ramp; feedback-triage
gains W3 (a declined item citing a ledger F-… id with no recorded disposition) plus the
decline-reconciliation offer; /ready gains the caveat block + a CONDITIONALLY VIABLE ceiling
for undisclosed declined Must-Fixes (fired deferrals count as open); the roundtrip resume displays
active dispositions and never machine-proposes declined/deferred. State gardening preserves
active disposition markers and compresses superseded ones to the archive line form. The honesty
gates (softness-check, deficit-lock, severity-floor, structured-findings, finding-trace,
regression-diff) read dispositions nowhere — a disposition grants no severity relief and never
decrements a count.
The Firewall — claim honesty
Reframed the Firewall's user-facing claim to match what the audit actually does. The absolute "it never invents content — no new plot events, characters, dialogue, or imagery" overstated the boundary: the audit routinely (and usefully) names that a structural element is missing or mis-weighted — "the climax needs a cost paid on-page," "this character carries no climactic agency," "this argument never re-qualifies its scope" — which is a class of solution, not content invention. The honest line is never drafts your prose and never scripts the specific content (events, characters, dialogue, evidence) — not never names a needed element. Updated the canonical references/firewall.md (with an explicit naming-a-need-vs-inventing-content distinction), the README ("How It Works" + Key Terms), and the /apodictic command (whose canonical pointer now targets references/firewall.md).
Workflows — Multi-Session Revision Arc Planning
New revision-coaching capability: a phased multi-week revision arc (Phase 1 structural root causes → Phase 2 downstream consequences → Phase 3 polish) that sequences the full Findings Ledger — the layer above the per-session Loop Dispatch, and the generalization of Retcon Planning's single-decision arc to the whole Ledger. The coach SEQUENCES findings; it never prescribes execution (the Coaching Firewall). Adds the apodictic.revision_arc.v1 schema (one block per manuscript; phases an ordered list, ≥1, no fixed 3-enum; stateless re-plan with no round/version field), the revision-arc validator, the multi-session-arc-planning.md skill (revision-coach mode 8), docs/multi-session-arc-planning.md, and the canonical example-revision-arc.md gate.
Honest posture (the Retcon pattern — load-bearing). The Root-Cause mapping is not machine-readable (finding.v1 carries only structured severity; the diagnostic-state root-cause map is prose), so the coach's dependency reasoning is trusted, not gated. The revision-arc validator gates the PLAN's provenance + self-consistency + firewall only — A1 schema + nested phase shape, A2 provenance closure (every finding_ref resolves to a real Ledger finding), A3 self-consistency ONLY (each finding in exactly one phase; a Must-Fix finding the arc itself labels a structural root cause is not parked in the polish phase — not a true-causal-graph check), A4 non-empty sequencing rationale; advisory W1 firewall-drift (a rationale that prescribes execution) and W2 orphan (a Must-Fix Ledger finding absent from the arc).
Nonfiction Argument Engine — Modularization (Workstream A)
Promoted the nonfiction argument path from implicit to named and bounded:
Named skill. skills/nonfiction-argument-engine/SKILL.md is now a first-class skill alongside core-editor, specialized-audits, revision-coach, and plot-architecture. The skill defines the delegation contract (triggered from core-editor §Delegation Rules when intake resolves constraint=nonfiction + persuasive-argument form), scope boundary, owned references, argument engine workflow, and QA guardrails. core-editor/SKILL.md gained an explicit § Nonfiction Argument Engine delegation rule replacing the previous implicit routing through specialized-audits.
Firewall single-sourced. The canon...
v2.6.2
v2.6.2 - 2026-06-26
Validators — override-marker hardening
Closed the code-span-decoy override bypass fleet-wide. Eighteen validators
(content-advisory, persona-divergence, intake-interview, author-fingerprint,
world-bible, continuity-bible, legal-risk, scene-ethics, retcon-plan,
state-card-diff, annotation-manifest, crosslink, argument-spine,
reader-instrument, promise-contract, regression-diff, style-explanation,
honesty-check) carried their own re.compile(r"<!--\s*override: …") — seventeen
of them matching on raw text, so an override marker quoted inside a ``` fence or
an inline code span (a documentation example) was honored as a live directive.
They now route through the shared `override_marker` SSoT: new `override_targets`
(id- / pair- / presence-scoped) and `override_payloads` (free-text) helpers strip
code spans first and boundary-match the slug, so one module owns both stripping and
marker-matching. The meta-linter gains two gates so the class cannot re-enter: M5
now flags the compiled / inline override-marker regex form (not just the bare
substring), and M6 flags a local code-span / fence stripper (delegate to
`override_marker.strip_code_spans`). The shared helpers match `override:` case-
sensitively (the old per-validator regexes were `re.IGNORECASE`); every shipped and
documented marker is lowercase, so this is a deliberate, inert tightening in the
fail-closed direction — an off-spec mixed-case marker no longer silences a finding.
v2.6.1
v2.6.1 - 2026-06-23
Audit Inventory — Content Advisory + Reader-Persona Simulation now registered
Two specialized audits that shipped in v2.6.0 — Content Advisory (content-advisory, a sensitivity-surface audit sibling to Reception Risk) and Reader-Persona Simulation (persona-divergence, a Pass-1 reader-dynamics overlay) — were never carded in the user-facing audit surfaces, so they were absent from the downstream Gemini website (generated from release-registry.json). Both are now registered across every audit-inventory surface: release-registry.json (two new Craft items; availableAudits 35 → 37, specializedAudits 32 → 34), the canonical audit-routing-table.md signal-emitting-audits inventory, AUDIT_SELECTION_MATRIX.md, and overview-dashboard.html (inventory-synced markers re-synced to audits=45). The inventory-parity gate (scripts/check-inventory-parity.mjs) now also guards release-registry.json: every shipped reference under specialized-audits/references/ must be carded in the registry's categories[].items[].files or explicitly listed in a new NOT_CARDED allowlist — so a future specialized audit can no longer ship uncarded the way these two did. A hermetic --self-test case proves the new check is non-vacuous.
v2.6.0
v2.6.0 - 2026-06-23
Annotated-Manuscript Deliverable — letter ↔ margin cross-links (Increment 3)
Navigation between the editorial letter and the marked-up copy is now bidirectional. The margin→letter direction already existed (every comment ends "(See letter §F-…)"); this adds letter→margin. A crosslink render treats the letter as a "second snapshot" and injects a CriticMarkup back-link span — {>>→ marked-up copy: <id> @ <kind>:<value><<} — immediately after each letter <!-- finding: F-… --> marker whose finding has a manifest annotation, copying the anchor verbatim from the gated manifest. The same reverse transform (delete every {>> … <<} span) proves the letter is untouched, behind the same two-sided sigil precondition (the letter, being authored prose, must not already contain a CriticMarkup sigil — without this the transform would silently delete an author's own span). A new crosslink validator gates bidirectional integrity: X1 the margin comment carries the forward link, X2 each back-link's anchor equals the manifest's (no drift), X3 no dangling either way (no phantom back-link; no missing reverse link), X4 no letter mutation; W1 annotated-but-uncited is advisory. Firewall-clean: the back-links carry only finding IDs + anchor tokens drawn verbatim from the manifest, never authored prose. Validators 42 → 43. Consumer-only — the shipped synthesis does not yet emit letter markers matching the annotation manifest, so the feature is inert on the real corpus until that producer lands; the canonical --check-all fixture's worked letter + crosslinked letter are hand-constructed.
Annotated-Manuscript export — DOCX with anchored comments → Google Docs (Increment 4)
The marked-up manuscript can now be exported as a .docx that imports into Google Docs with each finding as a native anchored comment — the professional editorial deliverable (Word/Google-Docs comments are the industry standard), and the DOCX target in one. annotation_export.py docx <run_folder> projects the gated manifest + snapshot into a .docx (Office Open XML — a ZIP of XML parts, stdlib-only: zipfile + hand-written XML, no python-docx): the manuscript text fills word/document.xml as one <w:p> per snapshot line, each finding's anchored span is wrapped <w:commentRangeStart/End w:id="N"/> + <w:commentReference w:id="N"/>, and word/comments.xml carries the verbatim comment — so Word and Google Docs render an anchored comment on the exact span. Firewall-clean by construction: a pure projection — the document text is the verbatim snapshot (XML-escaped via the same exact &/</> 3-entity pair as the HTML path, runs split only at comment boundaries), comments are verbatim, and the model assembles fixed OOXML boilerplate (7 parts incl. minimal styles.xml/settings.xml to avoid Word's repair-prompt) around copied bytes. The ZIP is byte-deterministic (ZIP_STORED; explicit ZipInfo per part with pinned date_time/create_system=3/external_attr; fixed order; seekable buffer; a w:date derived from the run date (the manifest runlabel at noon UTC — deterministic, not wall-clock, and no prior-day rollback in western timezones)) so a committed .docx fixture is byte-stable across machines. The new docx-export validator gates the on-disk artifact: D1 artifact integrity (the .docx equals a fresh deterministic build byte-for-byte — the authoritative lock for a binary, the HTML-H1 discipline), D2 text round-trip (the document.xml <w:t> text, exact-unescaped, one <w:p>/line + one trailing newline, reproduces the snapshot — the comment markup carries no body text), D3 comment resolution + fidelity (the commentRangeStart/End/Reference ids ↔ comments.xml ids form a bijection equal to the manifest finding set, and each comment is verbatim, keyed via the deterministic sorted(finding_id)→N map). Ships with the canonical example-annotated-manuscript/docx/ fixture wired into --check-all (byte-identical to a fresh export) and *.docx binary in .gitattributes. Validators 47 → 48. One-time maintainer acceptance: open the fixture in Word (no repair prompt) and import into Google Docs (comments land anchored) — OOXML on-paper validity doesn't guarantee either. PDF remains on the horizon; the manifest stays the one canonical artifact, every format (CriticMarkup, Obsidian, HTML, DOCX/GDocs) a projection.
Annotated-Manuscript export — self-contained read-only HTML (Increment 3)
The marked-up manuscript can now be exported as a self-contained .html openable in any browser — no Obsidian, no plugin, no Markdown tool. annotation_export.py html <run_folder> projects the gated manifest + snapshot into one HTML file: the manuscript in a faithful <pre class="manuscript"> (CSS white-space: pre-wrap + a serif face reads as prose while preserving the snapshot's exact bytes), with a footnote-style <sup id="ref-F-…"><a href="#fn-F-…">[F-…]</a></sup> marker at each anchor, and a <section class="findings"> listing each finding (<li id="fn-F-…">{verbatim comment} <a href="#ref-F-…">↩</a></li>) — so the browser gives native bidirectional anchor navigation, with embedded CSS and zero network refs. Firewall-clean by construction: a pure projection — the snapshot is HTML-escaped (&→& first, then </>), footnote markers are spliced between escaped prose segments at raw-snapshot offsets (escaping never touches a marker, so offsets never drift), and the new html-export validator gates it by identity: H1 round-trip (delete the manifest-keyed <sup id="ref-…"> markers + the exact 3-entity inverse unescape — not a general decoder — reproduces the snapshot byte-for-byte), H2 anchor resolution (<sup>↔<li> bijection equal to the manifest finding-id set, catching an un-manifested marker/finding), H3 comment fidelity (each <li> equals the HTML-escaped verbatim comment + the exact back-ref). The validator reads and gates the on-disk html/ artifact. A <pre>-faithful view (not reflowed <p>/<h>, which would defeat the byte-exact round-trip) keeps the firewall proof exact; reflowed-prose HTML is a future increment. Ships with the canonical example-annotated-manuscript/html/ fixture wired into --check-all (byte-identical to a fresh export) and a hostile self-test (HTML metachars upstream of an anchor + literal &/< in a comment) that the metachar-free canonical fixture can't exercise. Validators 46 → 47. Google Docs is the next render target (DOCX/PDF on the horizon); the manifest stays the one canonical artifact, every format a projection.
Annotated-Manuscript Obsidian export — bidirectional letter cross-links (Increment 2)
The native-Obsidian export is now clickable both ways, with no plugin. Building on Increment 1's footnoted copy, annotation_export.py now also projects the Obsidian letter and wires the navigation: (a) each copy footnote definition gains a forward [[<letter>#^<finding-id>|→ letter]] wikilink (web-verified that wikilinks render clickable inside footnote definitions — the key contrast with CriticMarkup); (b) the Obsidian letter appends an Obsidian ^<finding-id> block id to each finding's line so those forward links resolve; (c) the gated crosslinked letter's CriticMarkup back-link spans convert in place to reverse [[<copy>#<heading>]] wikilinks, using the resolved snapshot heading text (Chapter 9) — never the manifest's normalized Ch 9 token — with a file-level [[<copy>]] link for line-range/quote/document anchors that have no addressable heading (W1). Heading-level reverse-nav is reliable because Obsidian's heading-anchor slug strips the footnote ref Increment 1 appended to the heading line (# Chapter 9[^F-RR-01] → resolves as Chapter 9, web-verified). Firewall-preserved: still a pure projection — the comment is carried verbatim (O3 now checks comment + the exact forward wikilink, no authored-text gap), and the letter's editorial prose is untouched: O5 proves stripping the Obsidian letter's additions (wikilinks + block ids) reproduces the same bytes as stripping the crosslinked letter's CriticMarkup spans (the X4 analog, two-sided precondition). O4 gates link resolution (every forward link → a real letter block id; every reverse link → a real copy heading, footnote refs stripped) — so a manifest-token-vs-heading-text mismatch fails at build. Extends obsidian-export (no new validator — still 46); the validator reads and gates the on-disk copy and letter. The canonical obsidian/ fixtures (copy + new letter) are wired into --check-all (both byte-identical to a fresh export). Scope stays Obsidian-only; read-only HTML → Google Docs are the next render targets, DOCX/PDF on the horizon. The manifest remains the one canonical artifact; every format (CriticMarkup, Obsidian, future HTML/GDocs/DOCX/PDF) is a projection.
Annotated-Manuscript Export — Obsidian native footnotes (no plugin)
The marked-up manuscript can now be opened in vanilla Obsidian, no plugin, with every finding as a clickable footnote. The new annotation_export.py projects the gated annotation manifest + snapshot into Obsidian-native Markdown: each finding becomes a footnote reference [^<finding_id>] at its anchor locus (quote → after the sentence; chapter/section → on the heading line; line-range → end of line; document → file-level), and its definition carries the verbatim manifest comment (which already includes the (See letter §id.) pointer). Obsidian renders footnotes natively — clickable superscript, hover preview, the core Footnotes View pane — whereas CriticMarkup {>> <<} shows as literal brace clutter without a community plugin; that (not anchor links, which are native) was the real obstacle. Firewall-clean by construction: a pure projection of the gated manifest — the reverse transform (st...
v2.4.0
v2.4.0 - 2026-06-14
Nonfiction Argument Engine — uncompared-recommendation carve-out
dialectical-clarity.md gains classification rule 2a: for an argument whose controlling claim is a recommendation to act (AT3 — "X should do Y"), the comparative dimension is constitutive of the claim, so an AT3 recommendation that discharges none of its comparative burden (BP5 primary + OB3, no funding mechanism) is not evaluable as a recommendation — a defeat under decision test two (Evidence-evaluability) — and the verdict is Structurally Unsound (FM-A10, The Uncompared Proposal), with a matching note at the Step-9 Final Diagnostic Question. This is a bounded exception to rule 2's default-to-SOUND discipline, scoped to AT3 recommendations only (descriptive/explanatory/interpretive theses are untouched) and guarded by a wholly-absent vs. partially-discharged line — a recommendation that engages even one alternative thinly stays a Should-Fix soft spot in a sound argument. Calibration update (post-benchmark): the guard is tightened so that naming any alternative — even a weak or strawmanned foil — counts as partially-discharged (soft spot); only the total absence of any comparison triggers Unsound. (A benchmark run had andreessen-techno-optimist-manifesto regress SOUND→UNSOUND because a strawman "the only alternative is Communism" framing was misread as zero comparison.) Two independent cross-vendor reviewers (Gemini + GPT-5.5) ratified the narrowing and both flagged a token-foil gaming risk, so an anti-gaming clause was added: a named foil disables only the automatic FM-A10 defeat; a merely decorative foil (no mechanism/criteria/costs/tradeoff) can still be Unsound via the general evaluability test (rule 2), not the AT3 auto-trigger. It brings the engine into line with the policy-brief-uncompared ground-truth key (GT7 = UNSOUND), which the engine previously read SOUND. Verdict-behavior change for argument-shaped runs; gated on a benchmark convergence run (no --check-all gate covers behavioral calibration). No validator/schema change. See docs/argument-benchmark-calibration-round.md.
Command Surface — trimmed to 13
Retired three redundant command entry points, all reachable through /start: /revision-plan (a compatibility alias for /coach), /develop-edit (the default full_draft + repair router path), and /diagnose (a targeted repair). A writer with a draft lands on the /start router first anyway, so these added surface without adding capability. The distinct doors stay first-class (/ready, /pre-writing, /coach, /audit, /research, /plot-coach, /legal-risk, /triage-feedback, /reader-questions, /new-project, /projects). The registry command taxonomy (category / status / routerEquivalent / writerQuestion) is the single source the grouped README lists generate from. Routing references that pointed at the retired commands now point at /start.
Validators — finding-trace completion glob narrowed
finding-trace's _COMPLETION_GLOBS narrowed from *_Revision_*.md to *_Revision_Report_*.md, so a deadline-coaching *_Revision_Calendar_*.md is no longer mis-classified as a completion artifact (which would let its mentions advance a finding toward revised). Aligns finding-trace with the Increment-4a revision_round gate, which already narrowed its revision_report key. Negative-test guarded (calendar_not_completion); the revision-stage glob (_REVISION_GLOBS, for plan-coverage) stays broad.
Workflows — Legal Risk Register detection layer
The Legal Risk Register gains a detection layer (core-editor/references/legal-risk-register.md §Detection guidance / §Escalation-trigger taxonomy): per-class textual signals for what to flag under each risk_class (defamation / privacy / rights-clearance / other, with the finer categories — intrusion, false light, trade secrets, incitement, … — as sub-signals), a severity model (base tier raised by documented modifiers: +identifiable-living-private-person, +serious-allegation, +weak-or-no-documentation, +international-distribution, +author-signed-agreement, +marketing/cover/merchandise-use, +minor-or-vulnerable-subject), route-to-counsel bright lines, a flag-don't-resolve posture for jurisdiction divergence, and a compact controlled-vocabulary escalation-trigger taxonomy (~20 codes → default tiers) for the escalation_trigger field. Lean by design — the runtime module carries the heuristics; the cross-model research + citations live in docs/legal-risk-detection-level-setting.md. Firewall unchanged: detection only; a qualified lawyer is the final gate. No schema or validator change.
Workflows — Legal Risk Register router wiring
The built Legal Risk Register module is now reachable from the router and a command, not just internally. constraint:risk ("sensitive or legally risky content") offers the register and, on accept, attaches [Project]_Legal_Risk_Register_[runlabel].md as a companion artifact (synthesis constraint hook in run-synthesis.md — the first constraint:-keyed presentation overlay, and the first offer-then-attach one, since the not-a-lawyer framing warrants a confirm). New /legal-risk command as a direct entry point. The route map flips to Built (§3 option D, §6 Table B, §4a). The firewall is unchanged — flag, don't practice law. (Still future: auto-recommending the register for memoir/autofiction with identifiable real people without the explicit flag.)
Onboarding — install decision-aid and glossary
The README install section now opens with a Which install do I need? table that maps each host (Antigravity, Codex, Claude Code CLI, Cowork) to its path and fastest route, so a newcomer doesn't have to read all five install flows to find theirs. A new Key Terms section defines the load-bearing vocabulary a first-timer meets cold — contract, controlling idea, the Firewall, pass, macro block, audit, genre module, editorial letter, the Must/Should/Could severity tiers (and the Deficit Lock), spine, and reverse outline.
Onboarding — visual surfaces brought current
The README now opens with a See It in Action section linking the rendered sample editorial letters, the targeted-audit and pre-writing samples, and the two interactive maps (overview dashboard, route explorer) so newcomers can see real output before installing, plus a Your First Five Minutes walkthrough.
The overview-dashboard.html and route-explorer.html visuals (and their .codex.html twins) were stale at a v1.0.2 snapshot and have been brought current: the route explorer no longer reports shipped workflows (Fragment Synthesis, partial-manuscript mode, Submission Readiness, Submission Triage, Feedback Triage, the Nonfiction Argument Engine, editor scaffolding, diagnostic vocabulary, Series Continuity) as "not yet built"; the Legal Risk Register is shown as built-but-not-yet-routed; only multi-party/team intake remains a true gap. The overview dashboard now shows the canonical 8-block macro map (Emotional Dynamics restored as its own block), the full 50 spines / 12 families (adding Kishōtenketsu and Jo-ha-kyū), and the current version. The dashboard's click-to-expand cards are now keyboard-operable (role="button", tabindex, Enter/Space handlers).
Onboarding — overview dashboard front door
The overview dashboard header now opens with a "New here?" getting-started callout — a one-line plain-language orientation ("this page is a map; you don't need to memorize it") plus the first action (/start) — so a brand-new user landing on the page sees what to do before scrolling into the technical sections. Additive only (one <div> + two CSS rules; no redesign, no new sections, no network/deps). Applied to both the canonical overview-dashboard.html and its authored .codex.html twin (which keeps the apodictic-start wrapper naming per the codex override convention).
Validators — post-merge review nits
Three small fixes from a review of the merged batch (no new validators; count unchanged at 38→40 baseline):
reader-instrumentB3 (fabrication smell-test). The "unsourced question" advisory was gated on the Ledger having### Unresolved Questionsbullets, so anunresolved-questionreader-question with an inventedsource_notepassed unflagged when the Ledger had none (the more suspicious case — citing a UQ that can't exist). The advisory now also fires when the Ledger has zero UQ bullets. Stays a WARN: UQ provenance is non-referential by design, so this is a fabrication smell-test, not a hard gate.manuscript-vizrender gate (false pacing curve).W2 scene orderis advisory, but a reordered manifest draws a false pacing curve — the one warning that corrupts the render's core output. Therendersubcommand now refuses on a scene-order divergence too (not just ERROR-level gate failures), overridable with--force;W1 coveragestays advisory so a legitimate partial map still renders.- Swarm cost copy. The intake-router execution-mode menu rows (B/C/I) said a bare "roughly 5x" while
run-core.mdnotes the 2026-06 re-test measured ~8.5×+ on long fiction — understating cost at the decision point. The menu rows now carry the measured figure.
Validators — manuscript-viz E5 + check-mirror hardening
Follow-ups from an independent post-merge review of the Horizon-Tier-1 validator train. manuscript-viz gains E5 (no duplicate entry): a scenes[].scene_id or findings[].id repeated in the manifest is now an ERROR — a duplicate double-draws a pacing bar / double-counts a chapter's severity bar (a chart element showing a value the sources did not contain), which the per-id E2/E4 checks pass on. check-mirror now flags a CM_ROOT_ONLY utility (e.g. sync_setec.py) that has strayed to the plugin side or diverged across both copies, instead of skipping it by nam...
v2.3.1
v2.3.1 - 2026-06-07
Tooling — decoupled web-app UI generation
release-generate.mjs no longer reaches into the private APODICTIC-Gemini sibling
to write its App.tsx / LandingPage.tsx. That generation now lives in the app,
which pulls this repo's release-registry.json (vendored alongside the plugin)
and runs its own generator. apodictic's generator produces only its own docs;
removed the now-dead TS-emit helpers (−175 lines). Fixes the silent drift that
occurred whenever the sibling wasn't checked out during a release.
Distribution & changelog tooling
The generated codex/ and antigravity/ workspaces are no longer committed to
the repo. They are built from the canonical plugins/ source and published as
release assets by a new .github/workflows/release.yml (triggered on v*
tags): apodictic-codex-marketplace.zip, the new apodictic-antigravity.zip, and
apodictic.plugin. Install is now download-and-open — no clone, no Node — with
build-from-source still available. This removes the ×3 parity-churn multiplier
that every release-touching change otherwise paid. (Decision: GitHub #52, Option B.)
release-verify.mjs and CI now --self-check the host builds (regenerate in a temp
dir and validate internal consistency) instead of diffing against committed trees,
and CI gained generator/parity gates (release-generate --check, both build
--self-checks, assemble-changelog --check).
The changelog moves to a changelog.d/ fragment directory at the repo root:
each change drops one ### -headed fragment instead of editing the shared
changelog, and scripts/assemble-changelog.mjs cuts them into a dated section at
release time (wired into release.sh after the version bump). (GitHub #51.)
Registry — research modes 4 → 6
release-registry.json's researchModes array was missing two of the six modes it
claims in counts.researchModes and lists in commands/research.md: Citation
Verification (citation-verifier) and Field Reconnaissance (field-recon).
Added both (their backing reference docs already shipped), so the array matches the
count and the registry-derived web-app UI surfaces all six research modes.