Releases: Imbad0202/academic-research-skills
Release list
v3.21.1 — Bounded workflow substrates, sealed bakeoffs, and transport hardening
What's Changed
- Add OrcaRouter to the community directory — #782
- Update cross-model recommendation surfaces for generation currency — #784
- Repair the Codex ChatGPT-subscription citation transport for codex-cli 0.147.0 — #786
- Record the first Promotion Bakeoff and validate
gpt-5.6-solfor that transport — #788 - Consolidate Markdown parsing helpers and harden link/heading semantics — #791
- Register write-scope guard launcher degradations — #792
- Re-derive
data_access_levelfor standalone paper and reviewer skills — #793 - Add sealed bakeoff and default-off workflow-profile contracts — #795
- Add the opt-in inquiry ledger, alternative-register design freeze, and bounded review-criteria proving set — #796
- Release preparation and version alignment — #797
Highlights
- Repairs the contained ChatGPT-subscription citation transport for Codex CLI 0.147.0, covering stderr authentication attestation, provider-incompatible schema keywords, and the code-mode-host/web-search interaction.
- Adds sealed preregistration for future Promotion Bakeoffs and records the first counterbalanced 180-call bakeoff validating
gpt-5.6-solfor the ChatGPT-subscription citation transport. - Ships a default-off research-workflow profile substrate and the opt-in
ARS_INQUIRY_LEDGER=1alpha, with deterministic contracts, append-only receipts, bounded budgets, fail-visible staleness, and CI-gated conformance. - Freezes the profile-relevant alternative-register design and adds one bounded, source-backed MSR 2027 review-criteria proving set with exact-axis resolution and three-consumer digest binding.
- Consolidates Markdown link/heading grammar, records write-scope guard degradation paths, re-derives data-access annotations, and refreshes cross-model recommendation and community-directory surfaces.
Important boundaries
- The research-workflow profile ships as a default-off, offline-only substrate: no manuscript inference, pipeline hook, or family-specific shipped profile is included, and behavioral usefulness remains
NOT_RUN. - The inquiry ledger is opt-in via
ARS_INQUIRY_LEDGER=1and defaults OFF; its usefulness and usability evidence remainNOT_RUN. - The alternative register is design-only in this release; no runtime implementation ships.
- The review-criteria fixture is one illustrative proving set, not general venue/discipline coverage or a real-author attestation. #575 remains open pending #684's two independent blinded human experts and separate blind adjudication.
- The
gpt-5.6-solresult is transport-qualified: validated for the ChatGPT-subscription citation route only; the first-party API route remains provisional. - No new general research-outcome, safety, reviewer-correctness, or effectiveness claim ships.
Full technical details: CHANGELOG.md
v3.21.0 — ISO/IEC 42001-spirit transparency, verifiability, and feasibility track
What's Changed
- ISO/IEC 42001-spirit gap assessment (dual-track, 2026-08-17) — #762
- Citation-surface version drift + version-consistency invariant 12 (#754) — #763
#743design-doc reset-boundary co-location hotfix — #764- Distribution-surface claims aligned with evidence ceilings (#753) — #766
- Per-channel control-availability matrix (#757) — #768
- DATA_FLOWS.md: single map of network touchpoints + local stores (#758) — #770
academic-pipelinedata_access_levelcorrected toraw+ per-skill pins (#756) — #772- CI workflow enforcement-class table + inventory lint (#755) — #774
- Lightweight risk register + RR-1..3 mirroring lint (#759) — #777
- Solo-maintainer governance statement + SECURITY triage procedure (#760) — #778
- Stage capability/evidence matrix with enforceable claim ceilings (#745) — #751 (+ design freezes #750/#752)
- Pipeline wiring for the #655 claim-standing probe (PR-C) — #733
- Release prep — #779
Highlights
The ISO/IEC 42001-spirit track (epic #761) ships complete. ARS does not pursue certification; it adopts three distilled operating principles — transparency, verifiability, feasibility — with informative anchors to ISO/IEC 42001, and this release closes every finding of the 2026-08-17 dual-track audit:
- Four standing transparency artifacts, each defended by a CI lint:
docs/CONTROL_AVAILABILITY.md(which controls operate in your install channel),docs/DATA_FLOWS.md(what leaves your machine, what is stored, for how long, and how to turn each path off),docs/ARCHITECTURE.md§7.1 (what each CI workflow actually enforces), anddocs/RISK_REGISTER.md(risk → control → evidence status → residual gap, statuses mirrored verbatim from the capability matrix). - Claims aligned with evidence ceilings: distribution manifests drop unlicensed language; the pipeline's no-bypass prose now names its recorded override routes (trust-based controls with audit trails — the five-language README guarantee line included); citation surfaces are lint-pinned to the suite version.
- Governance and security response stated honestly: root
GOVERNANCE.mdrecords decision authority, per-operator cross-model scope (an error-detection control, not organizational independence), release authority, and end-of-life posture, plus the principles↔42001 mapping with Annex C not-applicable assessments;SECURITY.mdbacks its 7-day acknowledgement with a severity-tiered, solo-runnable triage procedure. - Also in this tag: the #745 stage capability/evidence matrix with enforceable claim ceilings, and the #655 claim-standing probe's pipeline wiring (consent-gated, advisory-only, no live provider).
No new effectiveness numbers are claimed in this release; unmeasured surfaces remain explicitly NOT_RUN/DESIGNED in the capability matrix and risk register.
Full changelog: CHANGELOG.md § 3.21.0
v3.20.1 — Contract-honesty hardening and bounded evaluation substrates
This patch bundles all release-worthy changes since v3.20.0.
Highlights
- Hardens human-control and integrity contracts across research, revision, and review: an explicit exit from non-generating Socratic RQ mode, reason-bound E6 dispositions, replayable Claim Registry coverage, fail-visible read-scope resolution, criterion-bound categorical reviewer judgements, and replay-valid six-axis review-panel provenance.
- Adds the bounded claim-standing evaluation substrate: consent-bound discovery adapters, versioned stance and evidence contracts, exact-span evidence replay, inert presentation rendering, a bilingual synthetic seed set, and deterministic scoring.
- Adds a closed first-round assignment-ledger gate for the #659 blind bundle, including pair-level exposure blocking, sealed receipts, and isolated write-once delivery.
- Documents an opt-in post-v3.20 roadmap for bounded domain profiles, inquiry branches, alternative registers, and outcome evaluation while preserving the simple default path.
Important boundaries
- These contract changes do not establish improved scientific outcomes, reviewer correctness, complete semantic detection, authenticated human identity, or independent error processes.
- Claim-standing classification remains UNMEASURED: no live stance provider, pipeline hook, relevance assessor, model run, or baseline exists; the discovery adapters have not yet been exercised against live providers. #655 remains open.
- The assignment gate proves structural exposure constraints only; it cannot authenticate that pseudonymous handles represent distinct or independent people. No external sessions, human judgements, model calls, or network runs occurred. #659 remains open.
- The roadmap is planning work, not default-on product behavior; structural expansion remains opt-in and subject to usability and outcome gates.
Full technical details are in the v3.20.1 changelog.
Compare v3.20.0...v3.20.1
v3.20.0 — Evidence-bound review and revision, contained transports, hermetic evaluation substrates
v3.20.0 bundles every release-worthy commit merged since v3.19.0. It strengthens evidence, provenance, and authority boundaries across review, revision, citation, human-subjects, and submission workflows while preserving ARS's human-in-the-loop positioning.
Highlights
Evidence-bound review and revision — reviewer and re-review contracts gain role-scoped criteria, typed evidence anchors, evidence-before-persuasion gates, deterministic receipts, complete retry evidence, author-controlled non-ranking revision roadmaps, and source-bound evidence rows. Unified criteria can now follow one review target across formative, internal, and external review without turning rubric conformance into a substitute for human judgment.
Research-integrity and authority substrates — new or hardened layers cover bibliographic integrity, retraction observations, cross-document consistency, content coverage, human-subjects authority/pathway traces, deterministic submission-packet manifests, committee-correspondence concern accounting, and optional cross-run adjudication-activity observability. These remain advisory or procedural support; institutional determinations and author decisions are not delegated to the suite.
Contained transport and ingestion boundaries — the ChatGPT-subscription citation transport has a closed, bounded protocol with no API fallback; EOF-complete draining rejects late protocol activity and reaps the process group. An opt-in process-isolated PDF text/OCR classifier remains a structural advisory, and the offline claim-standing candidate-ledger substrate supplies contracts and a local finalizer for consent-bound retrieval-evidence records without adding a discovery adapter, live probe, stance judge, or measurement.
Hermetic evaluation substrates — frozen synthetic fixtures, closed schemas, dry-run materializers, durable first-stop evidence, and no-call envelopes now cover role topology, ideation diversity, indirect prompt injection, claim standing, review criteria, and tortured-phrase screening. They make study plans reproducible; they do not by themselves establish safety, efficacy, accuracy, or behavioral improvement.
Access and platform work — Chinese-literature resolution, explicit /ars-* command aliases, Pi integration, clinical-reporting guidance, and platform documentation are expanded without changing the suite's human-in-the-loop model.
Versions
- suite /
academic-pipeline: v3.20.0 deep-research: v2.12.0academic-paper: v3.3.0academic-paper-reviewer: v1.11.0
Important limits
Held-out suites that require independent human experts, judges, or adjudicators remain unmeasured until those people complete the frozen protocols. No API fallback, autonomous OCR-to-anchor gate, institutional authorization, or autonomous publication path is introduced.
Full detail for every change is in CHANGELOG.md.
v3.19.0 — Revision-round claim-drift guards, PDF read-integrity preflight, read-scope attestation
Three advisory-or-opt-in integrity layers plus a launcher fix, bundling every release-worthy commit merged since v3.18.0. All new mechanisms preserve the human-in-the-loop positioning: nothing gates by default.
Highlights
Revision-round claim-drift guards (#569 / #570) — closes the epistemic and token halves of the #390 honest-claim residual: the block-anchored patch confined silent-distortion exposure to touched blocks but never checked a touched block's interior (DELEGATE-52, arXiv:2604.15597). Two complementary layers:
- a claim-strength ladder (
is consistent with < is associated with < predicts < contributes to < affects/leads to < causes) whose invariant is "no silent move, either direction, without an authorizing roadmap item" — wired into revision drafting and a new advisory Phase E6; - a deterministic numeric/citation token-conservation checker (
scripts/check_revision_token_conservation.py) as its necessary-but-not-sufficient complement.
Rather than cite an earlier-generation-model study as motivation, the current frontier model's baseline was measured first (evals/heldout/revision_claim_drift/). Mechanism shape credited to Yila-AI/sci-ssci-skills by @MissOrangePeel.
PDF read-integrity preflight (#512) — a three-signal page-count cross-check (scripts/pdf_read_preflight.py) so a truncated or mispaginated PDF read cannot mint an apparently-valid page anchor. FAIL on positive truncation evidence; UNAVAILABLE (explicit-warning advisory) when verification simply cannot run.
read_scope honest-coverage attestation (#513) — an optional declaration on the human-read ledger (full_text / sections / abstract_only / toc_only) that makes the finalizer's citation promotion read-scope-aware; absent means unknown, never backfilled.
Write-scope guard launcher fix (#545) — removes a watchdog pipe-stall that blocked every healthy PreToolUse write-scope-guard call for the full wall-clock bound (~6 s → ~0.15 s), plus flaky-test margin fixes.
Docs (#564) — de-drifted the SETUP Method 4a description-length figure to a durable comparative form.
Notes
academic-pipeline tracks the suite at v3.19.0; the three underlying skill versions (deep-research, academic-paper, academic-paper-reviewer) are unchanged. Full detail per change in CHANGELOG.md.
v3.18.0 — Self-improvement survey integration
External motivation: Ren et al. (2026), Self-Improvements in Modern Agentic Systems: A Survey (arXiv:2607.13104). Eight issues (#539–#542, #547–#550) derived from a two-reader (Claude + codex) full read of the survey, shipped via PRs #551/#552/#554/#555/#556/#557/#558/#559 — every mechanism advisory-or-opt-in, preserving ARS's human-in-the-loop positioning. This release also carries one independent feature: the #544 plugin update reminder.
Highlights
- SessionStart update-available reminder (#543 → #544) — plugin installs now get a one-line reminder pointing at
/plugin update academic-research-skillswhen the installed version is behindmain(24 h cache, 3 s network ceiling,ARS_UPDATE_CHECK=0kill switch). Third-party marketplaces default auto-update OFF and surface no behind-signal — this closes that gap, and this v3.18.0 release is the first it will announce. - Scope-conformance advisory (#547) — per-sub-question scope bindings in the RQ Brief; Phase E4 flags claims that silently broaden beyond their inherited scope as
ADV-E4rows, displayed per-row at the MANDATORY integrity checkpoints. Advisory-only, never gates. - Search-bounded novelty claims (#548) — "first study to…" language defaults to the bounded form filled from the documented search strategy (+ new
last_searched_at); Phase E5 classifiesSUPPORTED_WITHIN_SEARCH/UNRESOLVED, never "globally verified"; the bounding qualifier is compression-protected end-to-end (draft → abstract → formatter). - Risk-stratified Stage 2.5 claim verification (#549) — 100% of HIGH-IMPACT claims + a random sentinel replaces the uniform 30% spot-check, extending the #518 reference tiers to claim level.
- Cache staleness + live re-validation (#541) — cache-through wired into the citation-verification gate by default (closing the v3.11 Delta-2 forward-decl), with an age-based staleness advisory (
ARS_CACHE_STALE_ADVISORY_DAYS) and opt-in per-row live re-validation (ARS_CACHE_REVALIDATE=1). - Cross-model reviewer track (#540) — consent-gated: one seat of the fixed five-seat panel runs on the second model family (explicitly NOT the retired 6th reviewer); single-family runs disclose the correlated-error caveat in a new Review Panel Provenance block.
- Re-review judge independence + Judge Record (#539) — Stage 3' verdicts get an independent judgment-specific cross-model pass with a transparent Judge Record (judge identity, Round-1 panel provenance, rubric surfaces, judging budget).
- Pipeline behavior robustness evals (#550) — metamorphic paired eval seed set for the routing/gate layer; also shipped the reviewer skill's missing zh-TW trigger aliases (審查論文/模擬審查/幫我審這篇…).
- Literature anchor (#542) — the survey joins The AI Scientist and Zhao et al. as the third human-in-the-loop anchor, cited as design rationale, not proof.
Full details: CHANGELOG.md
Quality record: 49 codex gpt-5.6-sol xhigh review rounds across the 8 survey feature PRs + the release PR, all converged to 0 P1/P2; 8 independent security scans, 0 findings; full local suite 3421 passed at tag time.
v3.17.0 — pipeline boundary semantics + canonical cross-model handoff envelope
Security
- Tools allowlist for the three top-level plugin agents (#514, implemented in PR #521 by @madtriceps). The three plugin-exposed agents (
synthesis_agent,research_architect_agent,report_compiler_agent; deep-research sources + byte-identicalagents/mirrors, six files) now declaretools: Read, Write, Edit, Grep, Globin frontmatter — no Bash, no WebFetch/WebSearch — so dispatch-time capability is least-privilege even in hook-less installs, complementing the runtime Bucket A Bash deny (scripts/ars_write_scope_guard.py), which keys on agent name and continues unchanged. Retrospective entry: the code merged just after the v3.16.0 tag; documented here per the changelog-covers-merges gate.
Fixed
-
Blind-checkpoint transport moved to the dispatching layer (#523). The #518 blind disagreement checkpoints told their Bucket A primary owners (
research_architect_agent,editorial_synthesizer_agent) to execute the cross-model curl transport themselves — unexecutable under the runtime Bash deny and, for the architect, the #514 dispatch-time allowlist, so on every hook-active run the check at an irreversible decision silently degraded to single-model, indistinguishable from a transient API outage. New Transport ownership (#523) contract inshared/cross_model_verification.md§ Blind Disagreement Checkpoints: the owner commits its structured decision and emits the sanitized cross-model input as a handoff artifact; the dispatching layer (the main session running the skill, orpipeline_orchestrator_agentin pipeline Mode A — neither is Bucket A) executes § API Call Patterns, applies the mechanical enum comparison, and re-invokes the owner only for the divergence rebuttal; the editorial checkpoint's dispatched shape is an explicit, justified exception to the before-the-roadmap ordering (safe because the sprint-contract boundary keeps cross-model drivers out of the roadmap). The rule generalizes to any Bucket A cross-model owner —devils_advocate_reviewer_agent's independent DA critique routes the same way, with every successful response returned to the owner (no mechanical comparison exists for the dispatcher to resolve); non-fenced owners (integrity_verification_agentat the Stage 2.5/4.5 gates, deep-researchdevils_advocate_agent, the main session) execute directly, unchanged. No Bash/WebFetch re-added to any fenced agent (resolution (a); (c) rejected). Converged 0 P1/P2 across first-party security review + two codexgpt-5.6-solxhigh rounds. -
Pipeline prompt-surface contradictions from the #528 Mode-A replay (#529). The two genuine contradictions of the four ambiguities the replay surfaced: (1) the Methodology Blueprint was listed as a Stage 1 deliverable and in the Material Dependency Matrix but omitted from all three Stage 1→2 handoff surfaces — added to
academic-pipeline/SKILL.md,references/pipeline_state_machine.md, andagents/pipeline_orchestrator_agent.md; (2) the post-review coaching trigger read "After Stage 3 or Stage 3' completion, Decision = Minor/Major", but routing sends a Stage 3' Minor directly to Stage 4.5 — the trigger is now split by stage (Stage 3 = Minor/Major; Stage 3' = Major only) and the Coaching Rules exclusion list extended to match. Text-only. Retrospective entry: the code merged before this entry was written; documented here per the changelog-covers-merges gate. -
Stage 5 / Stage 6 boundary semantics — the two under-specified boundaries from the #528 Mode-A replay (#528). New authority section
references/pipeline_state_machine.md§ Stage 5 and Stage 6 Boundary Semantics, mirrored inacademic-pipeline/SKILL.mdandagents/pipeline_orchestrator_agent.md+references/process_summary_protocol.md. Stage 5: "Before finalization: always MANDATORY" now names exactly one checkpoint — the entry gate between Stage 4.5 PASS and the Stage 5 dispatch, carrying the format decisions; the in-stage content confirmation before the final PDF is Stage 5 execution (not a pipeline checkpoint), and the Stage 5 completion checkpoint (Final Paper delivered, before Stage 6) is FULL — never SLIM — but not MANDATORY. Stage 6: the state machine previously ended atStage 5 → ENDwith no Stage 6 at all; it now defines the Stage 5→6 transition, the decline path (Stage 6 is non-mandatory: declining marks itskippedand the pipeline still terminatescompleted), the terminal checkpoint after the Process Record is delivered, and the canonical terminal-acknowledgement vocabulary (finish/end/done/confirm, or an unambiguous natural-language equivalent) whose acceptance sets the pipeline global state tocompleted. All derived from existing text — no architecture change, no checkpoint relaxed. Thirteen codexgpt-5.6-solxhigh review rounds drove the consequential closure across the wider surface set: thestate_tracker_agentcontract gains Stage 6 (stage_id enum, SSOT block, prerequisite rows, terminal/decline action pairs), the FULL checkpoint-type row stops claiming "before finalization" (it collided with the MANDATORY row), Stage 5 execution consumes the entry-gate citation-style decision instead of re-asking, Stage 6 joins the explicitly-skippable list (the skip validator would otherwise reject the pinned decline path), the engagement-tracking SLIM downgrade gains its FULL-checkpoint exception, and the whole-pipeline collaboration-observer pass is re-timed to Stage 6 record compilation (before delivery — "at pipeline completion" could not coexist with completion-after-acknowledgement). Newscripts/check_pipeline_boundary_semantics.pydefrift lock pins all four #528 resolutions across the five surfaces with 66 mutation tests (one adverse-value witness per closed gap), wired into spec-consistency.yml and the unified pytest manifest; because twelve rounds showed sentence-level pins alone cannot converge on prompt surfaces, all five files also carry bibliography_agent-style whole-file sha256 content locks — any byte change fails CI until the pinned hash is updated in the same commit. Closes #528.
Added
- Canonical cross-model handoff envelope + dispatcher consumer contract (#527). The #523 owner→dispatcher→owner transport path was internally coherent but enforced by prose only — no canonical delimiter, no machine-stable schema, no malformed-result mapping, no pinned consumer trigger, so every test could stay green while a dispatcher silently treated a checkpoint owner's handoff as an ordinary deliverable. #527 closes that: one canonical
[CROSS-MODEL-HANDOFF v1]envelope (checkpoint_kind / owner_agent / correlation_id / expected_result / owner_decision-outside-payload / payload) defined inshared/cross_model_verification.md§ Cross-model handoff envelope, withscripts/cross_model_handoff.pyas the NORMATIVE grammar (parse + outcome routing as pure functions) and a deterministic owner→dispatcher→owner fixture suite on a fake transport (scripts/test_cross_model_handoff.py— no external API, no manuscript upload; literal pins guard the module constants against self-referential testing). The three checkpoint owners (research_architect_agent,editorial_synthesizer_agent,devils_advocate_reviewer_agent) emit the envelope with their closed kind/result-shape pair; the Mode-A orchestrator pins the consumer contract (recognition as a transport request never a deliverable; malformed envelope/result →[CROSS-MODEL-ERROR: malformed_*]→ outcomeunavailable, never a fabricated judgment; agreement → mechanical fill with NO owner re-invocation; divergence → re-invoke the original owner with the minimum return context, the dispatcher never authors the rebuttal; DA full-return: every successful response goes back to the owner;ARS_CROSS_MODELunset stays byte-equivalent). Newscripts/check_cross_model_handoff_contract.pypins the contract across all five surfaces (including a prose-enums-follow-the-module invariant) with a per-branch mutation-witness suite; both suites wired into spec-consistency.yml + the unified pytest manifest. Closes #527. - Defrift lock for the #514 tools allowlist (#524). New
scripts/check_tools_allowlist.py+ a 74-test suite (a failing witness per invariant branch), wired into spec-consistency.yml and the unified pytest manifest. YAML is the authority, not a line scan: every semantic decision reads a duplicate-preserving node tree (yaml.compose, which keeps a shadowed duplicate key visible and resolves an alias into shared node identity), the frontmatter fence is a column-0---only (an indented---inside a block scalar can't truncate the block and hide keys below it), and any frontmatter that will not compose to a mapping or uses a merge key (<<) / alias is a fail-closed error. Invariant 1 pins thetools: Read, Write, Edit, Grep, Globvalue on all six #514 surfaces: the node tree must carry exactly onetoolskey whose value normalizes to exactly the canonical five, plus an additive byte-exact raw-line witness (CR-sensitive, so a symmetric LF→CRLF conversion is drift; fires when the verbatim pinned line is absent). This closes the drift scenario where a future PR edits a source+mirror pair symmetrically (re-adding Bash, dropping a tool, or typoing a name) and passes every CI gate green, becausecheck_agents_mirror_sync.pypins only pairwise byte-equality and the runtime guard keys on agent name, never frontmatter; changing the allowlist now requires touching the lint's pinned value in the same commit (standard lock semantics). Invariant 2 reconciles the frontmatter channel against the runtime channel: any agent whosenameis a Bucket A key inscripts/ars_phase_scope_manifest.jsonmust not declare Bash in atools:key in any YAML-legal form — comma string, quoted string, flow/block list, inline comment,Bash(...)permission specifier (BashOutputis a different tool and not flagged) — failing closed on a missing/non-map...
v3.16.0 — Model tiering, cross-model gate hardening, WP advisory sharpening
Added
-
Model tiering: judgment/execution split with two opt-in directions, default untouched (#517). New
ARS_MODEL_TIERINGenv switch and canonicalshared/model_tiering.md, motivated by Lance Martin's "Cost effective harnesses with Fable" (2026-07-10; advisor-checkpoint configs measured ~90% of frontier-solo quality at ~34% of token cost, with delegation paying only when workers absorb enough tokens to offset per-handoff coordination cost). Default (unset): byte-equivalent pre-#517 behavior — every agent staysmodel: inherit(same opt-in philosophy asterminal_policies).economy(frontier-tier session): the 13 execution-type agents dispatch exactly one tier below the session model, floor Opus-class, never Sonnet (academic-prose tolerance is untested; the article's numbers came from ML tuning) —draft_writerexplicitly flagged as the highest-savings / most quality-sensitive downgrade point.quality-boost(below-frontier session): the judgment-type agents dispatched at the Stage 2.5/4.5 integrity gates and the final-review surfaces step up to the frontier tier; nothing is ever downgraded. Both directions have explicit no-op conditions with a one-line announcement; unknown values warn once and behave as unset (fail-open to the safe default). Tiers are relative positions, never hard-pinned model ids (the v3.7.0opuscommand floor retired in the Fable 5 harness pass is the cited precedent). Because many ARS roles execute inline today (no per-role model choice), the mechanism is dispatch-shaped: when a direction applies to a role, the session dispatches it as a subagent pinned to the target tier — inline roles included — and falls open to inline-on-session-model with a one-line announcement where subagent dispatch is impossible;docs/PERFORMANCE.md's (en/zh-TW) "no separate model routing layer" sentence is reconciled with a pointer. The frozen 39-agent classification (26 judgment / 13 execution — the issue header's 25/12 arithmetic corrected, membership unchanged) lives twice on purpose: a machine-readablescripts/model_tiering_manifest.jsonand the canonical doc's table, pinned to each other AND to the*_agent.mdfiles on disk by newscripts/check_model_tiering.py— set equality with a repo-wide stray sweep (a new skill dir can't smuggle unclassified agents), tier-enum + duplicate checks, and EXACT per-(tier, skill) token-set comparison against the doc table (missing/extra/duplicate tokens, per-row counts, duplicate rows all fail; 15 mutation tests; wired into spec-consistency.yml + the local pytest manifest). Prompt-caching guidance (when a direction is active, route repeated same-stage calls to the SAME worker so its cache accumulates) documented in the canonical doc and each of the fourSKILL.mdfiles' compact## Model Tiering (#517, optional)dispatch block — scoped so the unset default stays byte-equivalent, dispatch shapes included. No agent-file edits (the sha256-lockedbibliography_agent.mduntouched), no schema change, no hook. Spec:docs/design/2026-07-12-517-model-tiering-spec.md. -
Cross-model gate hardening: risk-stratified sampling, blind disagreement checkpoints, id-status allowlist, promotion bakeoff (#518). Four upgrades to
shared/cross_model_verification.mdand its consumers, from a 2026-07-11 cross-model consult (gpt-5.6-sol, xhigh). (1) The integrity-gate cross-model sample moves from uniform random 30% (min 5, max 15) to risk stratification across four mutually-exclusive tiers (highest-precedence tier wins, one verification per reference): HIGH-IMPACT references (headline conclusions, numerical claims, causal claims, methods-critical, disputed) verified 100% uncapped at both gates; a 10% RANDOM sample of the remainder at Stage 2.5 (round-up, min 3, max 10); at Stage 4.5, NEW-CHANGED references (behind claims new or changed since 2.5) verified 100% uncapped plus a 10% CONTROL sample of the unchanged remainder replacing RANDOM — verification budget concentrates where the paper's weight rests, and the results table gains a Tier column (integrity_verification_agentupdated in lockstep). (2) The two irreversible checkpoints — research-design freeze (research_architect_agent) and final editorial decision (editorial_synthesizer_agent) — gain optional blind disagreement checks: the primary commits its own decision in the same structured form first (the architect in a new Design-Freeze Checkpoint Audit blueprint section; the synthesizer's is its emitted decision), the cross-model then produces an independent structured decision from the same inputs (never seeing the primary's decision — same anchoring-prevention rule as the integrity samples; the editorial input is the panel'spanel_sizeN usable reviewer cards, never a hardcoded five), differing enum values trigger a targeted rebuttal addressing each cross-model driver against the evidence on file, and divergence escalates to the user — a review trigger, never a vote, never averaged; under a sprint contract the check runs strictly post-Step-3 against the mechanical protocol'seditorial_decisionand its drivers never enter the scoring matrix. (3) The "6th reviewer — Planned" section is retired, not deferred: the consult's counterproductive-conditions list (score averaging, role duplication, findings treated as confirmed defects, majority-vote false confidence, synthesizer context burn) matches ARS's documented anti-patterns one-for-one; the blind checkpoints are the replacement design, and the live mirrors (.claude/CLAUDE.md,shared/raise_framework.md, SETUP feature tables en/zh-TW) drop the "remains planned" claim. (4) The model-detection snippet separates "which provider endpoint" from "is this id known-good": newCROSS_MODEL_ID_STATUS=validated|provisional|unlistedannouncement with an explicit warning for unlisted first-party-prefix ids (gpt-made-upno longer passes silently) — routing itself is byte-identical, an unlisted id still takes the grounded route and never falls through to the ungrounded compatible branch. Plus a § Promotion Bakeoff operationalizing thegpt-5.6-solprovisional→validated criteria: a 30-reference paired same-day run (20 real / 10 fabricated; committed as a versioned, labeled, sha256-recorded fixture before any run counts; 3 repeats with ≥2/3 majority verdict, a 1–1–1 split scored conservatively against the model that produced it) against five non-inferiority thresholds (grounded-search completion, mismatch recall, false-disagreement rate, jq-guard shape stability — a hard requirement, p95 latency), entry-gated byscripts/cross_model_smoke_test.sh, results recorded underaudits/either way — and a deliberate two-step outcome: a full pass makes the idvalidated, while the recommended default flips only with an additionally stated superiority or operational-benefit reason. Spec:docs/design/2026-07-12-518-cross-model-gate-hardening-spec.md. Related: #517 (model tiering) will reference the checkpoint surfaces added here. -
GPT-5.6 Sol listed as provisional cross-model verifier + explicit reasoning-effort control (#515). OpenAI's
gpt-5.6-sol(released 2026-07-08) joins the canonical model table inshared/cross_model_verification.mdas provisional pending ARS validation — endpoint support (Responses API + hostedweb_search), the reasoning-effort enum (none|low|medium|high|xhigh|max, defaultmedium), and pricing (same standard rates as GPT-5.5; premium isreasoning: {mode: "pro"}on the standard slug, NOT a-promodel id) were verified first-party against OpenAI's model page and GPT-5.6 guide, but ARS-specific behavior (grounded-search completion rate, citation-mismatch recall, false-disagreement rate, jq-guard response-shape stability, p95 latency) has no operating history, so GPT-5.5 stays the recommended default. The documented OpenAI Responses call pattern gains an explicit reasoning-effort control via newARS_CROSS_MODEL_REASONING_EFFORT— set, it is passed asreasoning.effortso the run's effort is visible and reproducible; unset, the field is omitted entirely and each model's own provider default applies (forcing one value would silently change behavior for existinggpt-5.5-pro/legacy setups, a codex-review P2) — and both SETUP quick-setup blocks (en/zh-TW, parity-linted) mirror the new example lines. Newscripts/cross_model_smoke_test.sh— a live, manual (not CI; needsOPENAI_API_KEY) promotion gate asserting HTTP 2xx, a completedweb_search_call, a single verdict token, VERIFIED-carries-source, model echo, and effort echo — is the prerequisite for ever flipping the default to Sol. The canonical doc's Chat-Completions-web-search claim was re-verified against OpenAI's current web-search guide and deliberately left unchanged (a cross-model review suggested it was stale; first-party docs confirm it is still accurate). -
WP advisory held-out miss-rate measurement + acceptance set, Part 2 (#501; direction from the PR #468 review thread, @brycewang-stanford). New
evals/heldout/rq_framing_offlist/: a 48-item held-out set (32 shells outside the WP01-WP20 surface forms and the four in-prompt examples — 23 family variants + 9 off-list — plus 16 domain-native hard negatives), generated cross-model (gpt-5.6-sol), shell items regex-filtered (four negatives intentionally carry listed surface substrings as hard-negative material), dual-annotated with documented drops, English-only per the #468 caveat. Scored against the runtime LLM judge (isolatedclaude-sonnet-5sub-agents, verbatim advisory section only, pre-#503 vs post-#503 variants, two post replicates): overall miss rate 0.34-0.38 (above the inherited FNR < 0.30 line), concentrated in decorated compound-title off-list shells (7/9 missed, stable across replicates; judges read generic topical nouns as the exemption's "specific mechanism"), family-variant generalization under the line post-#503 (0.17-0.22), false-fire 0/1...
v3.15.0 — Release-gate hardening, prompt-debt retirement round 2, defrift locks
A release-discipline-and-hygiene release; no skill-behavior changes.
Added
- Three CI release gates: CHANGELOG-covers-merges pre-tag gate — every release-worthy merge since the previous tag must be documented before tagging (#483); version-consistency invariants 9-11 (release-notes body ≥100 chars, Last-Updated within ±7 days, Key-Additions heading matches suite version) plus a
tag-version-match.ymlgate that re-runs the full lint at tag time (#487); command-invariants gate pinning the SessionStart announce list to the actual 16-command inventory (#486). - Two defrift locks (#491 → #492): the Phase Boundary enforcement sentence is pinned verbatim across all 23 Bucket A agent blocks (
CANONICAL_ENFORCEMENT, version-matched, per-file tails free — the drift class that sat factually wrong for a month now fails CI); newcheck_setup_cross_model_parity.pypins the SETUP en/zh-TWARS_CROSS_MODELexamples to each other and to the canonical model tables. Local pytest manifest gains the v3.9.4 temporal test (58 → 60 entries), closing a local-green/CI-red coverage split.
Changed
- Prompt-debt retirement round 2 (#489 → #490): deep-scanned the 17 agents the first pass deferred, via 4 parallel audit batches + an independent codex cross-model challenge. 13 findings (2 P1 + 11 P2; 2 user-rejected and recorded). Both
socratic_mentoragents carried live self-contradictions — stale "quit after 15 rounds" rules against a documented typical 20-30-round run — now merged to a single auto-end authority per file (threshold 30). The repo-wide stale "hook deferred to #134" enforcement sentence (false since PR #294) rewritten at 29 surfaces; few-shot and duplicated-process scaffolds trimmed across 7 agents. The 2026-06-10 F-007 negative-framing deferred item closes as verified-no-rewrite-needed. Audit report:audits/harness-retirement-2026-07-04.md.
Fixed
- DOI badge served from shields.io; link target stays the concept DOI (#482).
- SessionStart announce updated to the full 16-command set.
academic-pipeline tracks the suite at v3.15.0; the other three skill versions are unchanged. Full details in CHANGELOG.md.
v3.14.0 — Claude Science importability, eval-comment rendering, prompt-debt retirement
What's Changed
- docs: recommend auto permission mode over Skip Permissions by @Imbad0202 in #464
- docs: add GitHub Copilot repository instructions by @Imbad0202 in #465
- docs: add native-reviewed Korean README by @devCharlotte in #469
- docs: credit devCharlotte for Korean README translation by @Imbad0202 in #471
- ci: add platform-port reminder (remind, don't block) by @Imbad0202 in #473
- chore(audit): harness-retirement 2026-07 report (#476) by @Imbad0202 in #477
- chore(prompts): retire expired writing-harness scaffolds in 4 Bucket A agents (#476 P2) by @Imbad0202 in #478
- ci(eval-harness): render PR comment as verdict + table, fold raw JSON by @Imbad0202 in #479
- fix(plugin): declare explicit skill paths in marketplace.json by @Imbad0202 in #480
- docs(release): align all doc surfaces for v3.14.0 by @Imbad0202 in #481
New Contributors
- @devCharlotte made their first contribution in #469
Full Changelog: v3.13.0...v3.14.0
Highlights
- Claude Science importability (#480) — the marketplace manifest now declares explicit skill paths, so Claude Science's "Import from GitHub" finds all four skills (previously zero: GitHub-API importers cannot traverse the symlinked
skills/directory). Verified end-to-end on Claude Science; import guide in README +docs/SETUP.mdMethod 5. Claude Code installs are unaffected. - Readable eval-harness PR comments (#479) — a one-line verdict + per-task table with the raw JSON folded into
<details>, replacing the raw report dump. Display layer only;run_evals, the threshold gate, and the ack contract are byte-identical. - Prompt-debt retirement (#477/#478) — expired writing-harness scaffolds removed from four writer-surface agents (net −111 prompt lines) after the 2026-07 harness-retirement audit, with three-track verification (sub-agent audit + cross-model review + eval harness at 100%).
- Changelog backlog rollup — 16
[Unreleased]entries whose code shipped before the v3.13.0 tag (diff/patch revision mode #390, submission-package verifier #394, eval gold sets #215/#216, and more) are now versioned inCHANGELOG.mdunder a provenance note.