Skip to content

v2.9.0

Choose a tag to compare

@github-actions github-actions released this 12 Jul 19:12
· 109 commits to main since this release

v2.9.0 - 2026-07-12

Validators — Krippendorff's α (agreement-alpha)

A new pure-stdlib validator, scripts/agreement_alpha.py, computes Krippendorff's α inter-rater reliability over a tidy rater,unit,value CSV — the mechanical half of the benchmarks' agreement-as-license promotion workflow, shipped before any human panel data exists so the M2 round is analysis-ready and α is computed by audited, self-tested code rather than an ad-hoc notebook (and shared with the fiction benchmark's M2b, which specifies no calculator of its own). Run via validate.sh agreement-alpha <ratings.csv> [--metric nominal|ordinal]. It builds the coincidence matrix (each within-unit ordered value pair contributing 1/(m_u−1); units with <2 values contribute nothing), then α = 1 − D_o/D_e under the nominal (δ²=0/1) or ordinal (Krippendorff's cumulative-marginal metric) difference function. Contract, pinned: a required rater,unit,value header (missing/wrong → ERROR), stdlib-csv quoting, blank value = missing, a malformed row (wrong column count or empty rater/unit) → ERROR naming the line number (never silently skipped); D_e=0 (a constant column) → alpha=UNDEFINED (D_e=0) + WARN, treated as not clearing any threshold. The bootstrap 95% CI resamples UNITS with replacement — ≥1000 resamples (default 1000; --resamples may only raise the floor), the Hayes & Krippendorff (2007) unit-level scheme — from a fixed default seed (--seed override) so two auditors reproduce the same interval byte-for-byte; licensing is on the CI lower bound, not the point estimate (the small-n false-promotion guard). The panel floor is enforced mechanically: panel-licensed requires a ≥3-editor panel (docs/argument-benchmark-spec.md §GT schema), so a run with fewer than three participating editors (those carrying ≥1 non-missing rating — an all-blank column is a listed-but-absent editor, not a rating one) still computes and prints α for information but returns an explicit non-clearing FAILED (exit 2) with no OK: line, which the promotion path rejects; a two-rater high-agreement dataset can no longer back into a license. Its hermetic --self-test locks the arithmetic to hand-derived values — Krippendorff's published worked example (nominal α = 0.691, n = 26), a fully worked tiny nominal case (8/15), an ordinal-beats-nominal ordered example (0.700 vs 0.4545), perfect agreement (1.0), systematic disagreement (−0.5), the D_e=0 UNDEFINED path, a deterministic seeded-CI regression lock, the missing-header / malformed-row / single-rater / single-unit ERROR arms, and the panel-floor arms (a two-rater run prints α but is non-clearing; a ≥3-editor panel licenses; the reviewer's 100-unit two-disagreement high-α repro is rejected; an all-blank third editor does not count toward the floor). Registered on all five surfaces (AGG token, dispatcher case with --self-test + python3-degrade path, help line, byte-identical scripts/plugins/apodictic/scripts/ mirror, this fragment); no --check-all corpus block — panel ratings live outside git. No hard-coded validator count is introduced (AGG_COUNT stays derived). Method sources: Krippendorff, K. (2011/2013), Computing Krippendorff's Alpha-Reliability; Hayes, A. F., & Krippendorff, K. (2007), Answering the Call for a Standard Reliability Measure for Coding Data, Communication Methods and Measures 1(1), 77–89.

Nonfiction Argument Engine — Reliability Ledger (Agreement-as-License, GT schema v0.3.0)

The Argument Benchmark's ground truth now carries its trust distinction as machine-checked data rather than prose. Each fixture's Provenance block gains a required Reliability ledger line assigning every GT anchor a status (authoritative / provisional / panel-licensed / low-agreement) and a decision-use (gate / confirm / report), and the GT schema bumps to v0.3.0. scripts/argument_groundtruth.py gains Check 6: the ledger group grammar (GT<a>(–GT<b>)?: <status>, <use>, (?![0-9]) boundary-guarded so GT10 cannot truncate-parse, |//-in-value rejected as copied guidance), exact GT1–GT8 coverage (no gaps/overlaps), the enforcement matrix (gate requires a licensed status — authoritative/panel-licensed; provisional may only confirm/report; low-agreement may only report), and a stale-heading cross-check that finally consumes the formerly-dead (PROVISIONAL) heading bool — a licensed status under a PROVISIONAL-marked heading is an error (the M2-promotion tripwire). A sibling round-record conformance mode (--round-record <record.md> --fixtures-dir <dir>, wired into --check-all over docs/argument-benchmark-calibration-round.md) is the one mechanical guard at the run-side attribution seam: every booked - BOOKED: ENGINE-FAULT <fixture-slug> GT<n>[ OVER-FIRE] must cite an anchor whose ledger licenses it (gate: any booking; confirm: only with an explicit OVER-FIRE tag — the asymmetric ruling; report: none). All 16 registered groundtruth.md fixtures migrate (one added ledger line each; the GT8 (provisional migration default) parentheticals preserved byte-for-byte; two PD-control Scope sentences corrected to GT4–GT6 and GT8 are provisional / advisory), and Checks 1–5 plus the strict GT8 contract are byte-stable. The convergence criteria are unchanged; only attribution becomes reliability-aware and asymmetric — a confirm-anchor over-fire stays ENGINE-FAULT (the specificity gate does not soften), while a confirm-anchor false-negative downgrades to key-suspect (ground-truth ambiguity) and routes to run-blind re-registration or the blind M2 panel, never an engine regression. Promotion α thresholds are pre-registered in the spec (Krippendorff's canonical cutoffs: α ≥ .800 → panel-licensed, .667 ≤ α < .800 → provisional, α < .667 → low-agreement; ordinal-weighted for the GT4 sub-scales, nominal for GT2/GT6/GT7; licensed on the bootstrap-CI lower bound), computed by the audited agreement-alpha CLI shipped separately. This is the argument-benchmark instantiation of the psychometric frame the fiction benchmark already ships (scripts/fiction_groundtruth.py) — inter-rater agreement licenses whether a label may gate; it never scores the engine — with a deliberately domain-honest token set (a documentation-only fiction↔argument cross-map, no drift guard). 76 parser self-test arms total (40 inherited from the warrant-split base), 36 new — Check-6, round-record, plus a Codex-#193-review hardening pass: the ledger field is anchored to the exact - **Reliability:** label within the ## Provenance block (a near-label like - **Not Reliability:** or a misplaced ledger under ## Notes is rejected, not silently accepted); the provisional heading bool is sticky-true across every heading covering a GT anchor (a later duplicate heading cannot clear a (PROVISIONAL) marker and evade the stale-heading tripwire); and the round-record BOOKED matcher is claim-broad/validate-strict — house-bold - **BOOKED:** is legal, while near-miss dialects (- **BOOKED** — …, - BOOKED:: …, - BOOKED* …, * BOOKED: …) are loud errors, never silent skips.

Argument engine — AGD Move Audit companion (argument-agd; Argument_State §10.9)

R3A of the argument re-grounding roadmap — the frontier build: a companion audit that inventories the text's performative argument movesASSURING (authority/certainty in place of support), GUARDING (a claim weakened to shrink its commitment), DISCOUNTING (an objection anticipated and set aside) — identified functionally at a transition, not by cue words (cue lexicons under-determine function: most hedge-cue lemmas also occur as non-cues — Velldal et al. 2012, Comp. Ling. 38(2), BioScope, domain-specific; the identification-by-transition posture is analogically scaffolded on Inference Anchoring Theory, Budzynska et al. 2016, a dialogue result whose monological transfer is this audit's own claim, validated by its fixtures). All three moves are legitimate; the audit never codes a move for being a move. Each inventoried move is challenged with its family's protocol — ASSURING × STRIP (purely subtractive: delete the assurance span; does independent support remain? — S&F 9e ch. 5), GUARDING × COMMITMENT (de-hedge-only: force the text's own claim to its unguarded form, adding no specificity the text lacks; ch. 16's self-sealer test), DISCOUNTING × ENGAGEMENT (constructs no text: decoy-vs-strongest reuses Step 6a Test A/B + the R2 Basis field by cross-reference, plus a downstream-consequence prong) — under a total family × challenge × result matrix (SURVIVES / COLLAPSES [/-DECOY/-COSTLESS] / SELF-SEALS (GUARDING-only) / INDETERMINATE (run-but-unresolved, documented) / NOT-CHALLENGED (inventory-only)). GUARDING records carry a Trajectory (S&F's disappearing guard). Candidate diagnoses are flag-only and licensed exclusively by failed function: Candidates: NONE is required on SURVIVES/NOT-CHALLENGED/INDETERMINATE (sole exception: the Trajectory: DISAPPEARING whitelist {FM-A16, WR3, BP4}); candidates come from a derived 62-code namespace (the Dialectical Clarity codes minus the AT1–AT4 type labels, via argument_crosswalk's scoped registry parse — no fourth hardcoded copy — plus FM-A1–20 from argument_groundtruth._FM_A_MAX), carry a PENDING → CONFIRMED/DECLINED reconciliation contract (the human editor or a full engine re-run adjudicates; the companion never writes codes outside its §10.9 annotation block, honoring the annotation protocol), and never enter severity propagation directly. The audit writes Argument_State.md §10.9 behind a machine-readable coverage manifest (declared spans cross-checked against the §2 claim ladder; exclusions with reasons; Completion: PARTIAL = WARN, ERROR under --strict).

The mechanical gate, scripts/argument_agd.py (stdlib, + byte-identical mirror; validate.sh argument-agd), enforces: matrix totality; the neutrality firewall both directions (incl. NOT-CHALLENGED-carries-no-construction and ENGAGEMENT-constructs-no-text); candidate namespace + reconciliation grammar (CONFIRMED requires adjudicator + target ref); the DISCOUNTING cross-ref contract (COLLAPSES-DECOY ⇒ a resolving Displaced strongest: pointer to §6's STRONGEST OBJECTION; COLLAPSES-COSTLESS ⇒ a resolving Discounted: grading target; NOT-INVENTORIED legal only on SURVIVES/INDETERMINATE); coverage (records' loci ∈ declared spans; COMPLETE ⇒ spans ⊇ C0 + the §2 ladder, degrade-not-fail when §2 is absent); and — with --sourceSource-anchor resolution by normalized substring (whitespace folded, quotes/dashes normalized; paraphrase and elision near-misses FAIL; greenfield — R2's anchors are convention-only and the GT lineage deliberately retired naive substring matching). Registered on all five surfaces (AGG token, dispatcher with --self-test + python3-degrade, help line, mirror, this fragment) plus a --check-all resolve-and-skip arm over five worked fixtures (evals/fixtures/argument-agd/: legitimate / abusive-cued / cue-only / cue-free structural discounting (the decisive functional-not-lexical case) / ambiguous-INDETERMINATE — original authored prose with stored source, anchors resolved under --strict). Argument_State schema stays v0.2.0 (companion subsections do not bump, per the schema's own rule); the §10 initializer enumerations (schema doc + the Dialectical Clarity §8 artifact text, which was already stale at "10.1–10.5") now read 10.1–10.9. No new failure code; no crosswalk change (AGD↔S&F crosswalk entries are R4B's). Sources: Sinnott-Armstrong & Fogelin, Understanding Arguments 9e, ch. 3/5/16; Budzynska, Janier, Reed & Saint-Dizier (2016), Arg. & Computation 7(1); Velldal, Øvrelid, Read & Oepen (2012), Comp. Ling. 38(2) — ACL J12-2005.

Argument engine — AGD Move Scan consumer (R3B seam; agd_move_scan → AGD Move Audit §10.9)

R3B A-1 of the argument re-grounding roadmap — the consumer half of the AGD producer/consumer seam. SETEC's agd_move_scan surface (shipped in the tagged v1.124.0 release, handoff: experimental, calibration_status: heuristic) reports LOCATED, verbatim-anchored candidate AGD move observations (family + verbatim span + paragraph_index + cue; cue: null = cue-free) for an argument-shaped nonfiction passage. APODICTIC now consumes them under the load-bearing ownership rule (R4A ADR D5: the producer observes, the consumer alone assigns codes): the scan is a POINTER, never a finding — every challenge, result, and diagnosis stays the AGD Move Audit's.

Consumer surface. A canonical stdout-forwarding shim scripts/ai_prose_agd_move_scan.py (SURFACE = "agd_move_scan", so discover_shim_surfaces counts it) routes through SETEC's normalized dispatcher (R2), with the per-surface floor (1.124.0) enforced runtime-side. sync_setec re-vendored the projected capabilities manifest (now carrying agd_move_scan with its 1.124.0 floor), re-vendored the contract fixtures (adds the agd_move_scan.json golden), and repinned setec-plugin.lock to the v1.124.0 release tag. The offline drift gate's CHECK-1 self-consistency now covers the surface, and tests/setec-contract/test_setec_contract.py bumps the shim-count floor (11 → 12), maps the surface to its 1.124.0 floor, and adds a positive runner-path assertion (routes the vendored golden through run_supplement, asserting results.observations survives — not mere fixture presence).

Consumption contract (Layer-1 scan pre-pass). The AGD Move Audit's Layer 1 gains a bounded pre-pass: if agd_move_scan.json is present in the run folder, read results.observations before building the move inventory; absent or malformed ⇒ proceed and record Scan: not consulted (<absent|error>) — loud, never silent (the artifact write is an orchestration stdout-redirect, not shim code). The §10.9 coverage manifest gains an optional mechanical Scan: line — Scan: consulted (<n> observations; <k> inventoried, <m> declined) + one indented Declined: "<span fragment>" — <one-clause reason> per declined observation, OR Scan: not consulted (absent|error) (that two-value reason enum only; an absent line is the valid pre-R3B case). Each scan observation is either inventoried (an M-record at that locus — identification authority stays with the audit) or declined with a one-clause reason; the comparison coordinate is the M-record's resolved Source anchor → derived paragraph index → normalized span overlap within that paragraph (the prose locus label is never a comparison operand). scripts/argument_agd.py (+ byte-identical mirror) learns and shape-checks the line — integers; k + m = n; m = the count of Declined: lines; and, with --scan agd_move_scan.json (as --check-all supplies), n = len(results.observations) — with hostile self-test arms (bad counts, k+m≠n, m≠declined-lines, bad reason value, malformed Declined: line). Structural line parse throughout; bounded leaf regexes on the count/reason tokens only.

Worked fixture + source realism. The five R3A fixture sources are split into their natural 2–3 blank-line paragraphs (a whitespace-only edit; every §10.9 Source anchor is intra-sentence and keeps resolving — the --check-all arm re-verified). The cue-free structural-discounting fixture ships a committed agd_move_scan.json (generated by the producer's manifest judge against the split source) plus a §10.9 Scan: block that exercises the cross-check with one inventoried observation (the cue-free DISCOUNTING move whose anchor resolves in a different paragraph than its prose locus label — matching on the label would flip the verdict) and one declined (an ASSURING data-report observation with no strippable assurance span), validated by the extended argument-agd under --check-all. The stale B5 "deferred dynamic signals — not present" note in argument-decision-audit.md is refreshed: the disappearing-guard and discounting/straw-man signals now ship via agd_move_scan as observed-not-pinned, unanchored, aggregate-excluded located observations (the §131 B3/B4 posture).

No new apodictic diagnostic code, no GT anchor, no crosswalk change (the AGD↔S&F crosswalk is R4B's); the Phase-1/2 behavioral benchmark is a separate later track. Sources: Sinnott-Armstrong & Fogelin, Understanding Arguments 9e, ch. 3 (assuring / guarding / discounting); R3A AGD Move Audit lineage (argument-agd-audit.md); R4A ADR D5 (producer observes, consumer adjudicates).

Nonfiction Argument Engine — Matched Clean/Broken Pairs (Argument Benchmark specificity retrofit)

The Argument Benchmark's two synthetic planted-defect fixtures (op-ed-warrant-leap, policy-brief-uncompared) are retrofitted into matched clean/broken pairs so specificity is measured within-work, not across works: today the broken fixtures and the positive controls are different prose, so "the engine flagged the broken fixture" is confounded with "found the fixture's own authored roughness" and "correctly specific" is confounded with "the control is just different prose." Each broken member moves (history-preserving git mv) to <pair>/broken/, and a derived clean twin is authored at <pair>/clean/ by an enumerated, minimal repair-edit set that discharges exactly the broken key's registered GT2 defect and nothing else (every other byte identical) — for op-ed-warrant-leap the double warrant leap (a causal-warrant insertion supplying a per-ride rate + confounder disposal, and a remedy-warrant insertion showing the lighter remedies were tried and failed); for policy-brief-uncompared the AT3 comparative + feasibility burden (a genuine same-goal service-investment comparison with criteria and a conceded tradeoff, plus a cost figure and funding mechanism discharging the planted FM-A18). Both twins are born on GT schema v0.3.0 with the Reliability ledger at the authoritative tier and are author-deterministic (the mutation diff is the answer key), so they add no new second-editor ask. scripts/argument_groundtruth.py gains Check 7 — matched-pair provenance (Matched-pair member / Paired-with, fiction's field names and <pair>/<member> value shape) implemented as a deliberate stricter superset of fiction's pairing grammar: a new lowercase leading-token member regex _PAIR_MEMBER_RE (hostile boundary lookahead so cleanX cannot truncate-parse — NOT fiction's substring parse, and the uppercase _GT7_VERDICT_RE/_GT8_FLAGS_RE are neither touched nor reused), both-fields-together enforcement, complement pairing (a clean names its broken twin and vice versa), slug self-consistency (the wrong-twin gate — the key's own Fixture slug pair must equal the Paired-with pair), and clean-side derivation gates (a clean member requires a non-N/A Base text + repair record and a GT2 marked N/A — positive control — the argument-side inversion of fiction's broken-plant-record gate). Absence of both pairing fields is a zero-behavior no-op, so all 14 unpaired legacy keys pass byte-unmodified. validate.sh --check-all's argument corpus loop gains the two-depth glob (<pair>/{clean,broken}/groundtruth.md) so nested members are validated at all, plus an orphan-twin completeness pass (a member with no sibling is a loud FAIL) that closes the missing-twin hole mechanically alongside Check 7 rule 5's wrong-twin hole. Two mechanical guards make the "deterministic twin" claim real: Check 7 rule 6 and a scripted repair-diff acceptance gate (each pair's diff broken/fixture.md clean/fixture.md is pure-additive with hunks mapping 1:1 to the enumerated repair-record loci — 2 each, zero unrecorded drift). Protocol side: RUN-PROTOCOL.md §Step 4 gains a fourth convergence class (broken converges failure-bearing, clean converges Must-Fix-qualified-positive-control, and the pair delta — the planted discriminator fires as a structural failure in the broken run and NOT in the clean run, on the same prose, per config), with a §Step 1/2 pair-blindness rule (members are separate blind submissions in separate runner sessions; a runner is never told a submission has a twin). Scoring is unchanged (protocol-level only — no new metric, no rubric dimension). Eleven new Check-7 self-test arms (none weakened to fiction's looser behavior); byte-identical scripts/plugins/apodictic/scripts/ mirror; no hard-coded fixture/validator counts introduced. Out of scope (M2): a PD broken twin of federalist-10 (broken derived from the real clean base, carrying Base text + plant record — the fiction direction); douglass-fourth-of-july is excluded (a planted defect in an unconventional-but-warranted form-control would score sensitivity by pathologizing the form). Pattern: BUILD-PREFLIGHT.md §"CONSTRUCTIVE PATTERN — ground truth for a partly-subjective domain" (matched counterfactual pairs — score the delta), adapted from the fiction benchmark (fiction-GT #187).

Codex #196 diff-review fold (refreshed onto main after #197's structural GT-validator refactor, whose _provenance_block/_parse_booked helpers this branch adopts): the clean-member GT2 positive-control gate is now a structural leading-line marker match rather than a body substring (a substantive GT2 that merely mentioned "positive control" no longer passes); the scripted repair-diff acceptance gate is wired into --check-all as a real --repair-diff mode (pure-additive, hunks 1:1 with the enumerated repair loci parsed by a structural line walk — not prose-only); the nested opt-out is closed (a file under <pair>/{clean,broken}/ may not drop its pairing fields — the directory is the source of truth via a path-derived member hint); and the residual regex holes are closed structurally (_FIXTURE_SLUG_PAIR_RE end-anchored so suffix garbage cannot truncate-parse; the orphan-twin pass now requires both fixture.md and groundtruth.md per member and twin). Ten new self-test arms (5 Check-7 hardening + 5 repair-diff) join the original eleven; the byte-identical scripts/plugins/apodictic/scripts/ mirror holds; no hard-coded counts introduced. The behavioral pair-delta convergence (8/8 cross-vendor, Fable + Codex 5.6) is recorded in docs/argument-benchmark-calibration-round.md §Matched clean/broken pairs.

Second Codex #196 fold (fresh-head review): the orphan-twin completeness pass is now driven off the member directories (<pair>/{clean,broken}), not a groundtruth.md glob — a fixture-only member directory (a fixture.md with no groundtruth.md) was invisible to every corpus loop, all of which key on groundtruth.md, and so evaded completeness validation entirely (P1); it is now a loud FAIL. And the repair-diff gate now rejects duplicate repair-locus identifiers (a repeated Locus <n> inflated the loci count and could spuriously match the hunk total, purporting a 1:1 map that isn't one — P2); each enumerated locus must carry a distinct id. Two more self-test arms; mirror parity and full --check-all (all 18 GT files ok) hold.

Third Codex #196 fold (fresh-head re-review): Check 7 rule 1 is now bidirectional — a file that declares a Matched-pair member but lives at a flat path (its parent directory is a slug, not clean/broken) is rejected, closing the inverse of the nested opt-out (a flat member declaration escaped the directory-enumerated completeness pass, so its missing twin was never caught); the CLI passes the raw parent-dir basename as member_hint so this is distinguishable (None is reserved for the in-memory self-test). And the repair-loci parser adopts the claim-broad / validate-strict pattern: a locus-shaped bullet that is not a well-formed - **Locus <n> — … (a case variant LOCUS 1, a number-less Locus, Locus1) is now a loud malformed-locus error instead of being silently dropped, and a second Base text + repair record field (whose loci would be invisible to the map) is rejected. Five more self-test arms; mirror parity and full --check-all hold.

Fourth Codex #196 fold (fresh-head residuals): Check 7 now binds the in-file pair slug to the actual parent pair directory (two internally consistent wrong-pair/* declarations can no longer pass at another pair's path), parses pairing declarations only from the exact ## Provenance block, rejects pair-shaped fields misplaced under Notes (including on flat fixtures), and rejects contradictory duplicate fields instead of trusting the first match. The repair-locus strict leaf now enforces the documented bold-title + dash form rather than accepting any line that merely begins Locus <n>. Five hostile self-test arms pin the three bypass classes; mirror parity and the full corpus gate hold.

Nonfiction Argument Engine — GT-validator Check 3 exemption hardened to a structural marker

scripts/argument_groundtruth.py (and its byte-identical plugins/apodictic/scripts/ mirror) closes the last instance of the substring-exemption class that Codex #196 fixed for Check 7 rule 6ii. The Check 3 GT2 locus↔code-family consistency check previously exempted "N/A" GT2s via a body substring test (re.search("N/A", gt2) OR "positive control" in gt2.lower()), so a real WARRANT-locus GT2 that carried only SUPPORT codes — the spec's "WARRANT diagnosed as SUPPORT" error — was wrongly exempted whenever it merely mentioned the words "N/A" or "positive control" anywhere in its prose, and the locus↔code mismatch slipped through. The exemption is now structural: it fires only when the GT2 section's first non-blank line (extracted by _parse_gt_sections) OPENS with an N/A marker (new _GT2_NA_MARKER_RE, the same leading-anchored posture as _GT2_POSCTRL_MARKER_RE). The marker is deliberately the broader N/A family — a bare - **N/A.** (the modest-proposal-satire fixture, whose literal-support/warrant locus is legitimately N/A, not a positive control) as well as the N/A — positive control controls (a subset) — so every currently-exempt fixture stays exempt; a real Primary failure layer GT2 is checked regardless of any downstream N/A / positive-control mention. Verified non-regressive: the exempt/run classification is byte-for-byte identical to the prior behavior across all 18 registered GT files (0 flips), all stay ok under validate.sh --check-all, and two new self-test arms pin it — a WARRANT-locus/SUPPORT-codes GT2 that mentions both phrases in prose now FAILs, and a bare-N/A satire GT2 stays exempt. Byte-identical scripts/plugins/apodictic/scripts/ mirror; no schema, spec, or contract change. Stacked on #196 (which introduced _GT2_POSCTRL_MARKER_RE).

Argument engine — warrant-defeaters as typed §6 objections (R2; Argument_State v0.2.0)

R2 of the argument re-grounding roadmap: complete the Toulmin model's rebuttal node — the "this holds unless …" exception conditions on a warrant — as a typed kind of §6 objection, not a parallel field. The §6 objection record (docs/argument-state-schema.md §6; Dialectical Clarity Step 6 template) gains a typing vocabulary: Target (C0/Cn/Cn.warrant/Cn.support/argument-wide), Relation (WARRANT-DEFEATER / CLAIM-CHALLENGE / EVIDENCE-CHALLENGE / VALUE-CONFLICT / ALTERNATIVE), and Basis (TEXT-INTERNAL / IMPORTED) — two decoupled axes: Relation is what the objection targets (a WARRANT-DEFEATER is an "unless…" exception on a specific Cn.warrant's connecting principle, holding that warrant fixed — Toulmin's rebuttal), Basis is derivation provenance (the existing Step 6a Test-B/Test-A distinction, mechanized). A Test-B-derived self-undermining remedy is TEXT-INTERNAL + CLAIM-CHALLENGE/ALTERNATIVE, not a warrant-defeater; a generic warrant exception is WARRANT-DEFEATER + IMPORTED. Every warrant-defeater records its Condition; a TEXT-INTERNAL one must cite Derivation anchors (the warrant/grounds locations it is derived from — engine-surfaced, author-adjudicated diagnosis, never an asserted validity constraint; the Firewall's no-invented-warrants discipline). A new bounded per-warrant sweep (Step 6, "6a-sweep") runs the Test-B question scoped to each §4 warrant and records material results as typed records; the existing Engaged×Quality axes grade them once (no new engagement vocabulary), and an ignored material defeater reaches the author through the ordinary objection pathway (Engaged: N + OB codes — OB3 when central), so severity propagation rides the existing table unchanged. §4 gains a projection line onlyDefeater refs: (within-run pointers, populated by Step 6, never authored at Step 4) or NO-MATERIAL-TEXT-INTERNAL-DEFEATER-IDENTIFIED — <review_basis stating the materiality bar> (an open-world search result, never a warrant-soundness assurance; an absent line = not swept, i.e. unknown, never clean — the silent-skip guard). No new failure code (WR stays 4, OB stays 6 — the R4A crosswalk's exact-82-set pin is untouched; Relation/Basis are field enums, and their crosswalk mapping to Toulmin/AIF-CA is deferred to R4B with the AIF adapter).

Pre-draft, the opposition surfaces stay writer-owned: reviewer_objections items may now be a legacy bare string (unchanged; blank-ness remains W5's advisory concern — behavior-preserved against the documented whitespace-only arm) or an object {objection: required non-whitespace, target?, relation?}; anti_thesis stays a required string with optional flat scalar siblings anti_thesis_target/anti_thesis_relation (schema pattern/enum-checked; W1's echo consumer untouched). The genre-profile schema's reviewer_objections.items is deliberately untyped — the stdlib subset validator has no union support and a typed items would reject the object form before hand-rolled code ran — so ALL per-item validation lives in argument_spine.py's new _reviewer_objection_errs (the _genre_section_errs pattern): object shape (required objection, target/relation value-spaces, post-draft audit keys rejected as wrong-surface, closed key set), plus a new B5 dangling-objection-target arm (a typed target's Cn must resolve against the spine ladder via the A5/A8 machinery; skip-not-fail when no valid spine — the B3 precedent; C0/argument-wide need no resolution) and a W5 widening so an all-object-form list counts as real content (previously it would false-fire "declares no reviewer_objections" — ERROR under --strict). The review fold also rejects explicit JSON null for either optional typed object field: omission remains legal, but a present target or relation must satisfy its declared value space. 19 new self-test arms (101 total) cover the full truth table (valid variants, missing/blank objection, numeric item, bad or null target/relation, stray post-draft key, unknown key, B5 dangling + skip-without-spine + C0/argument-wide, all-object W5-green under --strict, sibling-field pattern/enum rejects, blank-bare-string-stays-advisory). The canonical genre fixtures now exercise all three item forms under --check-all --strict (grant = mixed, pitch = all-object, academic = legacy bare). Both docs (docs/nonfiction-pre-draft.md, the pre-writing-pathway reference) updated; check-mirror byte-identity held; no new AGG_VALIDATORS entry (the derived count is unchanged). Post-draft §6 typed fields and the §4 projection are convention-governed (parity with every §6/§4 field today). Design source: Toulmin, S. (1958), The Uses of Argument, Cambridge University Press (the warrant/rebuttal model); the typed-objection single-home frame follows a cross-vendor design review that rejected a parallel-field design as duplicating the semantic object and adding a third engagement vocabulary.

Nonfiction Argument Engine — Warrant-Leap Primacy Over the Uncompared-Proposal Defeat (rule 2a)

Dialectical Clarity's Step-9 rule 2a (the AT3 uncompared-recommendation FM-A10 defeat) gains a primacy-override for the case where an AT3 recommendation is defeated both by the zero-comparison FM-A10 rule and by a causal/diagnostic warrant leap (WR0 MISSING on C0's causal bridge — the recommendation's problem is not warranted, e.g. "the harm rose when X appeared" ⇒ "X caused the harm" with confounders unaddressed). Previously the engine's "name the first defeated test" ordering ranked the FM-A10 comparative defeat — which fires at decision test two (Evidence-evaluability) — ahead of the warrant leap at test three (Warrant-recoverability), so a piece with both defeaters was reported as an Uncompared-Proposal (FM-A10 / BURDEN) primary even though the causal warrant leap is the deeper break. The override makes the causal warrant leap the primary structural break (FM-A6, The Warrant Leap) and co-reports FM-A10 as a subordinate defeat, on a logical-dependency rationale rather than the table order: a recommendation to "do Y about problem P" rests on L1 (P is real and attributable to Y's target) and L2 (Y beats the alternatives), and L2 presupposes L1 — the same dependency the repair order already encodes (warrant before remedy-defense). The verdict is unchanged (UNWARRANTED under either failure); only the primary locus / pattern changes, and only for this co-defeat case. A negative test preserves the ordinary case: an AT3 recommendation whose problem is adequately warranted (or whose only WR0 is incidental / RECOVERABLE) and whose sole defeat is the missing comparison keeps FM-A10 primary (e.g. a fare-free-transit brief whose benefits are evidenced but whose alternatives are never weighed). Cross-pointers were added at the Step-9 Final Diagnostic Question ("name the first defeated test" / "name which step breaks first") and the operative-test FM-A10 note so the override is discoverable from every place the ordering rule is stated. Resolves the M1 convergence-run caveat #1 (docs/argument-benchmark-calibration-round.md), where two independent blind engines (Fable + Codex 5.6) both returned the correct UNWARRANTED verdict on op-ed-warrant-leap but ranked FM-A10 / BURDEN over the registered FM-A6 / WARRANT primary. The fixture's groundtruth.md GT2 is sharpened to match: FM-A10 → BP5 is registered as a subordinate second defeater (uncoded objection to avoid colliding with GT3's confounding-zone central-objection code), with a split Q2 scoring boundary (surfaced-but-subordinated → cap Q2 = 2; warrant-leap-omitted → Q2 = 0 miss) and an explicit rejection of a co-equal GT-as-a-set reading under GT-set guards #1 and #4. Behavioral change → gated on a cross-vendor blind convergence re-run before merge.

Argument taxonomy — layer-boundary ADR + machine-readable crosswalk (argument-crosswalk-check)

R4A of the argument re-grounding roadmap: fix the layer boundaries once so the R2–R5 coherence work builds on stable IDs. Ships an Architecture Decision Record (docs/adr/0001-argument-layer-boundary.md) and a machine-readable crosswalk (evals/argument-crosswalk/crosswalk.json, v0.1.0) mapping the 82 internal argumentation code-layer entries (46 Dialectical Clarity codes + 20 FM-A patterns + 8 scheme hints + 3 warrant-verdict values + 5 premise-plausibility flag types) to four external reference vocabularies — AIF (node ontology; an optional serialization adapter, ADR D2), Walton schemes + critical questions, S&F fallacy families + the p.333 refutation split, and Wachsmuth/GAQCorpus logic/rhetoric/dialectic dimensions — each row carrying an explicit cardinality tag (exact/broader/narrower/related/unmapped). The ADR records: Argument_State stays the internal SSOT (D1); AIF is an adapter, not the vocabulary (D2); external vocabularies are reference targets crosswalked-to, never authorities that rename internal codes (D3); mapping cardinality is first-class and visible, and a mechanical global non-injectivity floor forbids a silently all-1:1 crosswalk, so the historical "1:1" overclaim cannot silently recur (D4); apodictic alone assigns codes, voiceprint emits located observations (D5, mirroring the SETEC seam); the verdict-axis→S&F mapping is R4A/R1-owned (D7). No taxonomy is renamed; R4A is purely additive.

The gate, scripts/argument_crosswalk.py (stdlib-only, validate.sh argument-crosswalk-check), certifies shape, not scholarly correctness — the crosswalk's targets/cardinality/provenance/rationale are authored assertions about external theory that no structural check can verify, so the ADR draws that boundary explicitly and the mechanical additions raise the floor without overclaiming: (1) drift-binding — the authoritative 82-ID set is derived at check time, not a fourth hardcoded copy (the verdict/flag/FM-A slices import the live owners _GT7_CLASSES / _GT8_ROW_FLAG_TYPES / _FM_A_MAX from argument_groundtruth.py; the 46 DC codes + 8 scheme hints, which have no machine-readable owner, are parsed structurally from dialectical-clarity.md by heading/table walk — the structural discipline, not a doc-wide regex), so a 47th code, a renamed flag type, or a stale registry self-count desyncs loudly; (2) count tripwires freeze the enumerated 46 (naming OB5) and 20 (naming FM-A20) against the registry's own stale "45"/"nineteen"; (3) a provenance locator ({work, loc, id}) is required on every populated target, mirroring the existing argument_groundtruth.py per-fault provenance precedent; (4) a global non-injectivity floor makes a silently all-1:1 crosswalk impossible; (5) closed external-ref value-spaces (an unregistered Walton scheme / Wachsmuth sub-criterion → ERROR); (6) unmapped ⇔ no-targets, with a required non-empty rationale on every unmapped row (the visible R2/R3/R4B work-list — 48 of the 82 at v0.1.0). Registered on all five surfaces (AGG token, dispatcher case with --self-test + python3-degrade path, help line, byte-identical scripts/plugins/apodictic/scripts/ mirror, this fragment) plus a --check-all real-file arm over the shipped crosswalk.json (resolve-and-skip when evals/ is absent from host workspaces). Behaviorally inert: no engine reading or scoring changes, no Argument_State/dialectical-clarity.md step-logic edits. Reference sources: the AIF specification (arg-tech.org); Walton, D., Reed, C., & Macagno, F. (2008), Argumentation Schemes, Cambridge University Press; Sinnott-Armstrong, W., & Fogelin, R. (2015), Understanding Arguments, 9th ed., Cengage; Wachsmuth, H., et al. (2017), Argumentation Quality Assessment: Theory vs. Practice, ACL — ACL Anthology P17-2039 (the citation the repo's argquality_judge.py already uses; the fuller 15-dimension taxonomy is Wachsmuth et al. 2017 EACL E17-1017), as the taxonomy behind GAQCorpus (Lauscher et al. 2020, arXiv:2006.00843).

Nonfiction Argument Engine — Warrant Verdict / Premise-Plausibility Split (GT schema v0.2.0)

Dialectical Clarity's Step-9 verdict enum is re-grounded from SOUND / UNCONVENTIONAL-BUT-EFFECTIVE / UNSOUND to WARRANTED / UNCONVENTIONAL-BUT-WARRANTED / UNWARRANTED, splitting the conflated "soundness" judgment into two axes: a warrant verdict (whether the reasoning warrants the conclusion on the text's own terms — the inference axis, Wachsmuth cogency's Local Relevance + Local Sufficiency) and separate, non-adjudicative premise-plausibility flags (whether a load-bearing premise is contestable — the acceptability axis, Local Acceptability). The token stays WARRANTED rather than COGENT precisely because full cogency folds in premise acceptability, which the Firewall forbids the engine from adjudicating in a verdict; premise flags surface a premise for the writer to earn but never rule it true or false and never change the warrant verdict by themselves. The FM-A10 (Uncompared Proposal) rule-2a evaluability defeat is preserved exactly — only the token changed (UNSOUND → UNWARRANTED). The Argument Benchmark GT schema bumps to v0.2.0: GT7 is renamed Warrant verdict, a required flag-only GT8 — Premise-plausibility flags section is added, and all 16 registered groundtruth.md fixtures migrate (10 take NONE_REGISTERED (provisional migration default)). scripts/argument_groundtruth.py gains GT1–GT8 coverage, warrant-verdict enum validation with active rejection of the retired label/tokens (present-but-unparseable GT7 is now an ERROR, not a silent skip), and GT8 leading-token parsing with a field-scoped case-sensitive truth-token Firewall check (flag-type cell only; the Why flagged / Firewall boundary prose is exempt). GT8 is an M1 contract/firewall check, not a scored dimension — a scored premise-flag dimension is deferred to M2. A Fable-5 three-lens review (anchor/parity, conceptual/firewall, adversarial/edge-case) was folded pre-merge: the GT8 registered path is strict (field↔row id agreement both directions as a multiset — a doubled P1, P1 field id or two conflicting P1 detail rows are rejected before the coverage compare rather than collapsed away by set equality, bolded-id rows matched, exactly-5-cell rows, full-match flag enum, token/heading-number boundaries, all three retired GT7 encodings rejected as residue — 40 self-test arms), the Hard-Gate-vs-Must-Fix conflation in the Step-9 decision rule was fixed behavior-preservingly, ground joined the load-bearing-role enum, and the GT8 ownership boundary (a flag must not duplicate a phenomenon a code already carries) landed in the ground-truth template. Move 1 of a multi-increment re-grounding of the argument taxonomy in Understanding Arguments (Sinnott-Armstrong & Fogelin, 9th ed.) and the argument-mining literature; grounded in Wachsmuth et al., "Computational Argumentation Quality Assessment in Natural Language" (EACL 2017) and Inference Anchoring Theory (Budzynska & Reed).

Execution Modes — Cost-Floor Dispatch Override

New cost-floor validator + run-core path: a cost-constrained user can now cap a run below the token-fit floor at the cheapest load-viable mode, with honest tradeoff disclosure — the missing counterpart to quality_risk_override (which only declines an upward escalation). A cap is recorded as BOTH a <!-- override: cost-floor-{single-agent,sequential,hybrid} — … --> marker and a cost_floor_override: <mode>[; context_tier: standard|large] — … metadata token (via the override_marker SSoT). Execution Protocol step 3 gains (d) final = min( max(token-fit-floor, RAW-quality-risk-target), cost_cap ). scripts/validate.sh cost-floor gates record integrity — CF1 marker integrity incl. the bidirectional orphan-token check, CF2 marker↔token sync, CF3 single-agent context_tier: large + mandatory-packet load-under-600K bound (with a below-floor WARN for a sequential/hybrid cap under the token-fit floor), CF4 no silent quality-risk demotion (a fired Q-trigger above the cap needs its own quality-risk-Q[n] marker) — without selecting the mode; the orchestrator owns selection. preflight.sh now prints a ## Per-Mode Token Cost Estimate section (single-agent · sequential ×3 · hybrid ×4 · swarm ×7, not ratio-comparable), the intake-router §2b below-floor pick is routed through the same gated front door, and dispatch_record.py --report counts cost_floor_override records (CR-6 detectability). The shared Q1–Q5 detection was extracted into a behavior-preserving _fired_triggers enumerator that both quality-risk-triggers and cost-floor call.

Nonfiction Argument Engine — GT-validator markdown matchers refactored to structural parsing

scripts/argument_groundtruth.py (and its byte-identical plugins/apodictic/scripts/ mirror) replaces the two markdown-structure matchers behind the recurring Codex-review edge bugs on the reliability-ledger PR (#193) — three of the four edge bugs there were in these two (two in the Provenance scope, one in the BOOKED matcher) — with structural line/heading parsing (pure stdlib; the fleet no-deps rule holds). The ## Provenance-block scope, formerly a multi-line ^##[ \t]+Provenance…(.*?)(?=^##\s) block regex that took three review rounds to stop matching ## Provenance Notes (suffix lookalike) and ##\nProvenance (newline split), is now _provenance_block — a heading walk mirroring _parse_gt_sections that matches the heading by exact title equality and closes the block at the next ##-level heading, with no \b/\s-across-newline boundary games. The round-record BOOKED matcher, formerly a claim-broad / validate-strict / body regex trio that reached correctness only after review rounds hand-tuned it to reject the near-miss dialects (- **BOOKED** — ENGINE-FAULT, - BOOKED::, - BOOKED*, * BOOKED:), is now _parse_booked: a deliberately broad structural CLAIM (any bullet, optional bold, the word BOOKED) so no near-miss dialect silently escapes, a strict canonical form (- BOOKED: / house-bold - **BOOKED:**) by string slicing, and a token-split body (ENGINE-FAULT <slug> GT<n>[ OVER-FIRE], zero-padding rejected). The fourth #193 edge bug's guard — the - **Reliability:** field-label leaf regex — is deliberately kept as a bounded single-line token match (regex is the right tool for a leaf token; the trap was multi-line/heading scope). The change is behavior-preserving: all 76 parser self-test arms pass unchanged (including the reliability_provenance_lookalike / _newline_split / _near_label_rejected / _misplaced_ledger and every roundrec_dialect_* regression arm), and a differential sweep of the retired regexes vs the new helpers over a hostile input battery confirms they agree — the sole delta is a functionally-inert error-message category (a whitespace-only-body - BOOKED: now reads "unrecognized dialect" rather than "malformed body ''"; both reject the line, so the round-record accept/reject verdict is identical). No schema, spec, or contract change.