Skip to content

Releases: Imbad0202/academic-research-skills

v3.21.1 — Bounded workflow substrates, sealed bakeoffs, and transport hardening

Choose a tag to compare

@Imbad0202 Imbad0202 released this 24 Aug 11:21
Immutable release. Only release title and notes can be modified.
v3.21.1
127ff85

What's Changed

  • Add OrcaRouter to the community directory — #782
  • Update cross-model recommendation surfaces for generation currency — #784
  • Repair the Codex ChatGPT-subscription citation transport for codex-cli 0.147.0 — #786
  • Record the first Promotion Bakeoff and validate gpt-5.6-sol for that transport — #788
  • Consolidate Markdown parsing helpers and harden link/heading semantics — #791
  • Register write-scope guard launcher degradations — #792
  • Re-derive data_access_level for standalone paper and reviewer skills — #793
  • Add sealed bakeoff and default-off workflow-profile contracts — #795
  • Add the opt-in inquiry ledger, alternative-register design freeze, and bounded review-criteria proving set — #796
  • Release preparation and version alignment — #797

Highlights

  • Repairs the contained ChatGPT-subscription citation transport for Codex CLI 0.147.0, covering stderr authentication attestation, provider-incompatible schema keywords, and the code-mode-host/web-search interaction.
  • Adds sealed preregistration for future Promotion Bakeoffs and records the first counterbalanced 180-call bakeoff validating gpt-5.6-sol for the ChatGPT-subscription citation transport.
  • Ships a default-off research-workflow profile substrate and the opt-in ARS_INQUIRY_LEDGER=1 alpha, with deterministic contracts, append-only receipts, bounded budgets, fail-visible staleness, and CI-gated conformance.
  • Freezes the profile-relevant alternative-register design and adds one bounded, source-backed MSR 2027 review-criteria proving set with exact-axis resolution and three-consumer digest binding.
  • Consolidates Markdown link/heading grammar, records write-scope guard degradation paths, re-derives data-access annotations, and refreshes cross-model recommendation and community-directory surfaces.

Important boundaries

  • The research-workflow profile ships as a default-off, offline-only substrate: no manuscript inference, pipeline hook, or family-specific shipped profile is included, and behavioral usefulness remains NOT_RUN.
  • The inquiry ledger is opt-in via ARS_INQUIRY_LEDGER=1 and defaults OFF; its usefulness and usability evidence remain NOT_RUN.
  • The alternative register is design-only in this release; no runtime implementation ships.
  • The review-criteria fixture is one illustrative proving set, not general venue/discipline coverage or a real-author attestation. #575 remains open pending #684's two independent blinded human experts and separate blind adjudication.
  • The gpt-5.6-sol result is transport-qualified: validated for the ChatGPT-subscription citation route only; the first-party API route remains provisional.
  • No new general research-outcome, safety, reviewer-correctness, or effectiveness claim ships.

Full technical details: CHANGELOG.md

Compare v3.21.0...v3.21.1

v3.21.0 — ISO/IEC 42001-spirit transparency, verifiability, and feasibility track

Choose a tag to compare

@Imbad0202 Imbad0202 released this 18 Aug 05:46
Immutable release. Only release title and notes can be modified.
v3.21.0
2b639c1

What's Changed

  • ISO/IEC 42001-spirit gap assessment (dual-track, 2026-08-17) — #762
  • Citation-surface version drift + version-consistency invariant 12 (#754) — #763
  • #743 design-doc reset-boundary co-location hotfix — #764
  • Distribution-surface claims aligned with evidence ceilings (#753) — #766
  • Per-channel control-availability matrix (#757) — #768
  • DATA_FLOWS.md: single map of network touchpoints + local stores (#758) — #770
  • academic-pipeline data_access_level corrected to raw + per-skill pins (#756) — #772
  • CI workflow enforcement-class table + inventory lint (#755) — #774
  • Lightweight risk register + RR-1..3 mirroring lint (#759) — #777
  • Solo-maintainer governance statement + SECURITY triage procedure (#760) — #778
  • Stage capability/evidence matrix with enforceable claim ceilings (#745) — #751 (+ design freezes #750/#752)
  • Pipeline wiring for the #655 claim-standing probe (PR-C) — #733
  • Release prep — #779

Highlights

The ISO/IEC 42001-spirit track (epic #761) ships complete. ARS does not pursue certification; it adopts three distilled operating principles — transparency, verifiability, feasibility — with informative anchors to ISO/IEC 42001, and this release closes every finding of the 2026-08-17 dual-track audit:

  • Four standing transparency artifacts, each defended by a CI lint: docs/CONTROL_AVAILABILITY.md (which controls operate in your install channel), docs/DATA_FLOWS.md (what leaves your machine, what is stored, for how long, and how to turn each path off), docs/ARCHITECTURE.md §7.1 (what each CI workflow actually enforces), and docs/RISK_REGISTER.md (risk → control → evidence status → residual gap, statuses mirrored verbatim from the capability matrix).
  • Claims aligned with evidence ceilings: distribution manifests drop unlicensed language; the pipeline's no-bypass prose now names its recorded override routes (trust-based controls with audit trails — the five-language README guarantee line included); citation surfaces are lint-pinned to the suite version.
  • Governance and security response stated honestly: root GOVERNANCE.md records decision authority, per-operator cross-model scope (an error-detection control, not organizational independence), release authority, and end-of-life posture, plus the principles↔42001 mapping with Annex C not-applicable assessments; SECURITY.md backs its 7-day acknowledgement with a severity-tiered, solo-runnable triage procedure.
  • Also in this tag: the #745 stage capability/evidence matrix with enforceable claim ceilings, and the #655 claim-standing probe's pipeline wiring (consent-gated, advisory-only, no live provider).

No new effectiveness numbers are claimed in this release; unmeasured surfaces remain explicitly NOT_RUN/DESIGNED in the capability matrix and risk register.

Full changelog: CHANGELOG.md § 3.21.0

v3.20.1 — Contract-honesty hardening and bounded evaluation substrates

Choose a tag to compare

@Imbad0202 Imbad0202 released this 16 Aug 01:51
Immutable release. Only release title and notes can be modified.
v3.20.1
6837b4d

This patch bundles all release-worthy changes since v3.20.0.

Highlights

  • Hardens human-control and integrity contracts across research, revision, and review: an explicit exit from non-generating Socratic RQ mode, reason-bound E6 dispositions, replayable Claim Registry coverage, fail-visible read-scope resolution, criterion-bound categorical reviewer judgements, and replay-valid six-axis review-panel provenance.
  • Adds the bounded claim-standing evaluation substrate: consent-bound discovery adapters, versioned stance and evidence contracts, exact-span evidence replay, inert presentation rendering, a bilingual synthetic seed set, and deterministic scoring.
  • Adds a closed first-round assignment-ledger gate for the #659 blind bundle, including pair-level exposure blocking, sealed receipts, and isolated write-once delivery.
  • Documents an opt-in post-v3.20 roadmap for bounded domain profiles, inquiry branches, alternative registers, and outcome evaluation while preserving the simple default path.

Important boundaries

  • These contract changes do not establish improved scientific outcomes, reviewer correctness, complete semantic detection, authenticated human identity, or independent error processes.
  • Claim-standing classification remains UNMEASURED: no live stance provider, pipeline hook, relevance assessor, model run, or baseline exists; the discovery adapters have not yet been exercised against live providers. #655 remains open.
  • The assignment gate proves structural exposure constraints only; it cannot authenticate that pseudonymous handles represent distinct or independent people. No external sessions, human judgements, model calls, or network runs occurred. #659 remains open.
  • The roadmap is planning work, not default-on product behavior; structural expansion remains opt-in and subject to usability and outcome gates.

Full technical details are in the v3.20.1 changelog.
Compare v3.20.0...v3.20.1

v3.20.0 — Evidence-bound review and revision, contained transports, hermetic evaluation substrates

Choose a tag to compare

@Imbad0202 Imbad0202 released this 14 Aug 01:51
Immutable release. Only release title and notes can be modified.
v3.20.0
3af9f03

v3.20.0 bundles every release-worthy commit merged since v3.19.0. It strengthens evidence, provenance, and authority boundaries across review, revision, citation, human-subjects, and submission workflows while preserving ARS's human-in-the-loop positioning.

Highlights

Evidence-bound review and revision — reviewer and re-review contracts gain role-scoped criteria, typed evidence anchors, evidence-before-persuasion gates, deterministic receipts, complete retry evidence, author-controlled non-ranking revision roadmaps, and source-bound evidence rows. Unified criteria can now follow one review target across formative, internal, and external review without turning rubric conformance into a substitute for human judgment.

Research-integrity and authority substrates — new or hardened layers cover bibliographic integrity, retraction observations, cross-document consistency, content coverage, human-subjects authority/pathway traces, deterministic submission-packet manifests, committee-correspondence concern accounting, and optional cross-run adjudication-activity observability. These remain advisory or procedural support; institutional determinations and author decisions are not delegated to the suite.

Contained transport and ingestion boundaries — the ChatGPT-subscription citation transport has a closed, bounded protocol with no API fallback; EOF-complete draining rejects late protocol activity and reaps the process group. An opt-in process-isolated PDF text/OCR classifier remains a structural advisory, and the offline claim-standing candidate-ledger substrate supplies contracts and a local finalizer for consent-bound retrieval-evidence records without adding a discovery adapter, live probe, stance judge, or measurement.

Hermetic evaluation substrates — frozen synthetic fixtures, closed schemas, dry-run materializers, durable first-stop evidence, and no-call envelopes now cover role topology, ideation diversity, indirect prompt injection, claim standing, review criteria, and tortured-phrase screening. They make study plans reproducible; they do not by themselves establish safety, efficacy, accuracy, or behavioral improvement.

Access and platform work — Chinese-literature resolution, explicit /ars-* command aliases, Pi integration, clinical-reporting guidance, and platform documentation are expanded without changing the suite's human-in-the-loop model.

Versions

  • suite / academic-pipeline: v3.20.0
  • deep-research: v2.12.0
  • academic-paper: v3.3.0
  • academic-paper-reviewer: v1.11.0

Important limits

Held-out suites that require independent human experts, judges, or adjudicators remain unmeasured until those people complete the frozen protocols. No API fallback, autonomous OCR-to-anchor gate, institutional authorization, or autonomous publication path is introduced.

Full detail for every change is in CHANGELOG.md.

v3.19.0 — Revision-round claim-drift guards, PDF read-integrity preflight, read-scope attestation

Choose a tag to compare

@Imbad0202 Imbad0202 released this 22 Jul 04:00
Immutable release. Only release title and notes can be modified.
v3.19.0
828ef3b

Three advisory-or-opt-in integrity layers plus a launcher fix, bundling every release-worthy commit merged since v3.18.0. All new mechanisms preserve the human-in-the-loop positioning: nothing gates by default.

Highlights

Revision-round claim-drift guards (#569 / #570) — closes the epistemic and token halves of the #390 honest-claim residual: the block-anchored patch confined silent-distortion exposure to touched blocks but never checked a touched block's interior (DELEGATE-52, arXiv:2604.15597). Two complementary layers:

  • a claim-strength ladder (is consistent with < is associated with < predicts < contributes to < affects/leads to < causes) whose invariant is "no silent move, either direction, without an authorizing roadmap item" — wired into revision drafting and a new advisory Phase E6;
  • a deterministic numeric/citation token-conservation checker (scripts/check_revision_token_conservation.py) as its necessary-but-not-sufficient complement.

Rather than cite an earlier-generation-model study as motivation, the current frontier model's baseline was measured first (evals/heldout/revision_claim_drift/). Mechanism shape credited to Yila-AI/sci-ssci-skills by @MissOrangePeel.

PDF read-integrity preflight (#512) — a three-signal page-count cross-check (scripts/pdf_read_preflight.py) so a truncated or mispaginated PDF read cannot mint an apparently-valid page anchor. FAIL on positive truncation evidence; UNAVAILABLE (explicit-warning advisory) when verification simply cannot run.

read_scope honest-coverage attestation (#513) — an optional declaration on the human-read ledger (full_text / sections / abstract_only / toc_only) that makes the finalizer's citation promotion read-scope-aware; absent means unknown, never backfilled.

Write-scope guard launcher fix (#545) — removes a watchdog pipe-stall that blocked every healthy PreToolUse write-scope-guard call for the full wall-clock bound (~6 s → ~0.15 s), plus flaky-test margin fixes.

Docs (#564) — de-drifted the SETUP Method 4a description-length figure to a durable comparative form.

Notes

academic-pipeline tracks the suite at v3.19.0; the three underlying skill versions (deep-research, academic-paper, academic-paper-reviewer) are unchanged. Full detail per change in CHANGELOG.md.

v3.18.0 — Self-improvement survey integration

Choose a tag to compare

@Imbad0202 Imbad0202 released this 18 Jul 07:17
Immutable release. Only release title and notes can be modified.
v3.18.0
bbc0659

External motivation: Ren et al. (2026), Self-Improvements in Modern Agentic Systems: A Survey (arXiv:2607.13104). Eight issues (#539#542, #547#550) derived from a two-reader (Claude + codex) full read of the survey, shipped via PRs #551/#552/#554/#555/#556/#557/#558/#559 — every mechanism advisory-or-opt-in, preserving ARS's human-in-the-loop positioning. This release also carries one independent feature: the #544 plugin update reminder.

Highlights

  • SessionStart update-available reminder (#543#544) — plugin installs now get a one-line reminder pointing at /plugin update academic-research-skills when the installed version is behind main (24 h cache, 3 s network ceiling, ARS_UPDATE_CHECK=0 kill switch). Third-party marketplaces default auto-update OFF and surface no behind-signal — this closes that gap, and this v3.18.0 release is the first it will announce.
  • Scope-conformance advisory (#547) — per-sub-question scope bindings in the RQ Brief; Phase E4 flags claims that silently broaden beyond their inherited scope as ADV-E4 rows, displayed per-row at the MANDATORY integrity checkpoints. Advisory-only, never gates.
  • Search-bounded novelty claims (#548) — "first study to…" language defaults to the bounded form filled from the documented search strategy (+ new last_searched_at); Phase E5 classifies SUPPORTED_WITHIN_SEARCH/UNRESOLVED, never "globally verified"; the bounding qualifier is compression-protected end-to-end (draft → abstract → formatter).
  • Risk-stratified Stage 2.5 claim verification (#549) — 100% of HIGH-IMPACT claims + a random sentinel replaces the uniform 30% spot-check, extending the #518 reference tiers to claim level.
  • Cache staleness + live re-validation (#541) — cache-through wired into the citation-verification gate by default (closing the v3.11 Delta-2 forward-decl), with an age-based staleness advisory (ARS_CACHE_STALE_ADVISORY_DAYS) and opt-in per-row live re-validation (ARS_CACHE_REVALIDATE=1).
  • Cross-model reviewer track (#540) — consent-gated: one seat of the fixed five-seat panel runs on the second model family (explicitly NOT the retired 6th reviewer); single-family runs disclose the correlated-error caveat in a new Review Panel Provenance block.
  • Re-review judge independence + Judge Record (#539) — Stage 3' verdicts get an independent judgment-specific cross-model pass with a transparent Judge Record (judge identity, Round-1 panel provenance, rubric surfaces, judging budget).
  • Pipeline behavior robustness evals (#550) — metamorphic paired eval seed set for the routing/gate layer; also shipped the reviewer skill's missing zh-TW trigger aliases (審查論文/模擬審查/幫我審這篇…).
  • Literature anchor (#542) — the survey joins The AI Scientist and Zhao et al. as the third human-in-the-loop anchor, cited as design rationale, not proof.

Full details: CHANGELOG.md

Quality record: 49 codex gpt-5.6-sol xhigh review rounds across the 8 survey feature PRs + the release PR, all converged to 0 P1/P2; 8 independent security scans, 0 findings; full local suite 3421 passed at tag time.

v3.17.0 — pipeline boundary semantics + canonical cross-model handoff envelope

Choose a tag to compare

@Imbad0202 Imbad0202 released this 16 Jul 10:23
Immutable release. Only release title and notes can be modified.
v3.17.0
039d94f

Security

  • Tools allowlist for the three top-level plugin agents (#514, implemented in PR #521 by @madtriceps). The three plugin-exposed agents (synthesis_agent, research_architect_agent, report_compiler_agent; deep-research sources + byte-identical agents/ mirrors, six files) now declare tools: Read, Write, Edit, Grep, Glob in frontmatter — no Bash, no WebFetch/WebSearch — so dispatch-time capability is least-privilege even in hook-less installs, complementing the runtime Bucket A Bash deny (scripts/ars_write_scope_guard.py), which keys on agent name and continues unchanged. Retrospective entry: the code merged just after the v3.16.0 tag; documented here per the changelog-covers-merges gate.

Fixed

  • Blind-checkpoint transport moved to the dispatching layer (#523). The #518 blind disagreement checkpoints told their Bucket A primary owners (research_architect_agent, editorial_synthesizer_agent) to execute the cross-model curl transport themselves — unexecutable under the runtime Bash deny and, for the architect, the #514 dispatch-time allowlist, so on every hook-active run the check at an irreversible decision silently degraded to single-model, indistinguishable from a transient API outage. New Transport ownership (#523) contract in shared/cross_model_verification.md § Blind Disagreement Checkpoints: the owner commits its structured decision and emits the sanitized cross-model input as a handoff artifact; the dispatching layer (the main session running the skill, or pipeline_orchestrator_agent in pipeline Mode A — neither is Bucket A) executes § API Call Patterns, applies the mechanical enum comparison, and re-invokes the owner only for the divergence rebuttal; the editorial checkpoint's dispatched shape is an explicit, justified exception to the before-the-roadmap ordering (safe because the sprint-contract boundary keeps cross-model drivers out of the roadmap). The rule generalizes to any Bucket A cross-model owner — devils_advocate_reviewer_agent's independent DA critique routes the same way, with every successful response returned to the owner (no mechanical comparison exists for the dispatcher to resolve); non-fenced owners (integrity_verification_agent at the Stage 2.5/4.5 gates, deep-research devils_advocate_agent, the main session) execute directly, unchanged. No Bash/WebFetch re-added to any fenced agent (resolution (a); (c) rejected). Converged 0 P1/P2 across first-party security review + two codex gpt-5.6-sol xhigh rounds.

  • Pipeline prompt-surface contradictions from the #528 Mode-A replay (#529). The two genuine contradictions of the four ambiguities the replay surfaced: (1) the Methodology Blueprint was listed as a Stage 1 deliverable and in the Material Dependency Matrix but omitted from all three Stage 1→2 handoff surfaces — added to academic-pipeline/SKILL.md, references/pipeline_state_machine.md, and agents/pipeline_orchestrator_agent.md; (2) the post-review coaching trigger read "After Stage 3 or Stage 3' completion, Decision = Minor/Major", but routing sends a Stage 3' Minor directly to Stage 4.5 — the trigger is now split by stage (Stage 3 = Minor/Major; Stage 3' = Major only) and the Coaching Rules exclusion list extended to match. Text-only. Retrospective entry: the code merged before this entry was written; documented here per the changelog-covers-merges gate.

  • Stage 5 / Stage 6 boundary semantics — the two under-specified boundaries from the #528 Mode-A replay (#528). New authority section references/pipeline_state_machine.md § Stage 5 and Stage 6 Boundary Semantics, mirrored in academic-pipeline/SKILL.md and agents/pipeline_orchestrator_agent.md + references/process_summary_protocol.md. Stage 5: "Before finalization: always MANDATORY" now names exactly one checkpoint — the entry gate between Stage 4.5 PASS and the Stage 5 dispatch, carrying the format decisions; the in-stage content confirmation before the final PDF is Stage 5 execution (not a pipeline checkpoint), and the Stage 5 completion checkpoint (Final Paper delivered, before Stage 6) is FULL — never SLIM — but not MANDATORY. Stage 6: the state machine previously ended at Stage 5 → END with no Stage 6 at all; it now defines the Stage 5→6 transition, the decline path (Stage 6 is non-mandatory: declining marks it skipped and the pipeline still terminates completed), the terminal checkpoint after the Process Record is delivered, and the canonical terminal-acknowledgement vocabulary (finish / end / done / confirm, or an unambiguous natural-language equivalent) whose acceptance sets the pipeline global state to completed. All derived from existing text — no architecture change, no checkpoint relaxed. Thirteen codex gpt-5.6-sol xhigh review rounds drove the consequential closure across the wider surface set: the state_tracker_agent contract gains Stage 6 (stage_id enum, SSOT block, prerequisite rows, terminal/decline action pairs), the FULL checkpoint-type row stops claiming "before finalization" (it collided with the MANDATORY row), Stage 5 execution consumes the entry-gate citation-style decision instead of re-asking, Stage 6 joins the explicitly-skippable list (the skip validator would otherwise reject the pinned decline path), the engagement-tracking SLIM downgrade gains its FULL-checkpoint exception, and the whole-pipeline collaboration-observer pass is re-timed to Stage 6 record compilation (before delivery — "at pipeline completion" could not coexist with completion-after-acknowledgement). New scripts/check_pipeline_boundary_semantics.py defrift lock pins all four #528 resolutions across the five surfaces with 66 mutation tests (one adverse-value witness per closed gap), wired into spec-consistency.yml and the unified pytest manifest; because twelve rounds showed sentence-level pins alone cannot converge on prompt surfaces, all five files also carry bibliography_agent-style whole-file sha256 content locks — any byte change fails CI until the pinned hash is updated in the same commit. Closes #528.

Added

  • Canonical cross-model handoff envelope + dispatcher consumer contract (#527). The #523 owner→dispatcher→owner transport path was internally coherent but enforced by prose only — no canonical delimiter, no machine-stable schema, no malformed-result mapping, no pinned consumer trigger, so every test could stay green while a dispatcher silently treated a checkpoint owner's handoff as an ordinary deliverable. #527 closes that: one canonical [CROSS-MODEL-HANDOFF v1] envelope (checkpoint_kind / owner_agent / correlation_id / expected_result / owner_decision-outside-payload / payload) defined in shared/cross_model_verification.md § Cross-model handoff envelope, with scripts/cross_model_handoff.py as the NORMATIVE grammar (parse + outcome routing as pure functions) and a deterministic owner→dispatcher→owner fixture suite on a fake transport (scripts/test_cross_model_handoff.py — no external API, no manuscript upload; literal pins guard the module constants against self-referential testing). The three checkpoint owners (research_architect_agent, editorial_synthesizer_agent, devils_advocate_reviewer_agent) emit the envelope with their closed kind/result-shape pair; the Mode-A orchestrator pins the consumer contract (recognition as a transport request never a deliverable; malformed envelope/result → [CROSS-MODEL-ERROR: malformed_*] → outcome unavailable, never a fabricated judgment; agreement → mechanical fill with NO owner re-invocation; divergence → re-invoke the original owner with the minimum return context, the dispatcher never authors the rebuttal; DA full-return: every successful response goes back to the owner; ARS_CROSS_MODEL unset stays byte-equivalent). New scripts/check_cross_model_handoff_contract.py pins the contract across all five surfaces (including a prose-enums-follow-the-module invariant) with a per-branch mutation-witness suite; both suites wired into spec-consistency.yml + the unified pytest manifest. Closes #527.
  • Defrift lock for the #514 tools allowlist (#524). New scripts/check_tools_allowlist.py + a 74-test suite (a failing witness per invariant branch), wired into spec-consistency.yml and the unified pytest manifest. YAML is the authority, not a line scan: every semantic decision reads a duplicate-preserving node tree (yaml.compose, which keeps a shadowed duplicate key visible and resolves an alias into shared node identity), the frontmatter fence is a column-0 --- only (an indented --- inside a block scalar can't truncate the block and hide keys below it), and any frontmatter that will not compose to a mapping or uses a merge key (<<) / alias is a fail-closed error. Invariant 1 pins the tools: Read, Write, Edit, Grep, Glob value on all six #514 surfaces: the node tree must carry exactly one tools key whose value normalizes to exactly the canonical five, plus an additive byte-exact raw-line witness (CR-sensitive, so a symmetric LF→CRLF conversion is drift; fires when the verbatim pinned line is absent). This closes the drift scenario where a future PR edits a source+mirror pair symmetrically (re-adding Bash, dropping a tool, or typoing a name) and passes every CI gate green, because check_agents_mirror_sync.py pins only pairwise byte-equality and the runtime guard keys on agent name, never frontmatter; changing the allowlist now requires touching the lint's pinned value in the same commit (standard lock semantics). Invariant 2 reconciles the frontmatter channel against the runtime channel: any agent whose name is a Bucket A key in scripts/ars_phase_scope_manifest.json must not declare Bash in a tools: key in any YAML-legal form — comma string, quoted string, flow/block list, inline comment, Bash(...) permission specifier (BashOutput is a different tool and not flagged) — failing closed on a missing/non-map...
Read more

v3.16.0 — Model tiering, cross-model gate hardening, WP advisory sharpening

Choose a tag to compare

@Imbad0202 Imbad0202 released this 12 Jul 00:43
Immutable release. Only release title and notes can be modified.
v3.16.0
73c898c

Added

  • Model tiering: judgment/execution split with two opt-in directions, default untouched (#517). New ARS_MODEL_TIERING env switch and canonical shared/model_tiering.md, motivated by Lance Martin's "Cost effective harnesses with Fable" (2026-07-10; advisor-checkpoint configs measured ~90% of frontier-solo quality at ~34% of token cost, with delegation paying only when workers absorb enough tokens to offset per-handoff coordination cost). Default (unset): byte-equivalent pre-#517 behavior — every agent stays model: inherit (same opt-in philosophy as terminal_policies). economy (frontier-tier session): the 13 execution-type agents dispatch exactly one tier below the session model, floor Opus-class, never Sonnet (academic-prose tolerance is untested; the article's numbers came from ML tuning) — draft_writer explicitly flagged as the highest-savings / most quality-sensitive downgrade point. quality-boost (below-frontier session): the judgment-type agents dispatched at the Stage 2.5/4.5 integrity gates and the final-review surfaces step up to the frontier tier; nothing is ever downgraded. Both directions have explicit no-op conditions with a one-line announcement; unknown values warn once and behave as unset (fail-open to the safe default). Tiers are relative positions, never hard-pinned model ids (the v3.7.0 opus command floor retired in the Fable 5 harness pass is the cited precedent). Because many ARS roles execute inline today (no per-role model choice), the mechanism is dispatch-shaped: when a direction applies to a role, the session dispatches it as a subagent pinned to the target tier — inline roles included — and falls open to inline-on-session-model with a one-line announcement where subagent dispatch is impossible; docs/PERFORMANCE.md's (en/zh-TW) "no separate model routing layer" sentence is reconciled with a pointer. The frozen 39-agent classification (26 judgment / 13 execution — the issue header's 25/12 arithmetic corrected, membership unchanged) lives twice on purpose: a machine-readable scripts/model_tiering_manifest.json and the canonical doc's table, pinned to each other AND to the *_agent.md files on disk by new scripts/check_model_tiering.py — set equality with a repo-wide stray sweep (a new skill dir can't smuggle unclassified agents), tier-enum + duplicate checks, and EXACT per-(tier, skill) token-set comparison against the doc table (missing/extra/duplicate tokens, per-row counts, duplicate rows all fail; 15 mutation tests; wired into spec-consistency.yml + the local pytest manifest). Prompt-caching guidance (when a direction is active, route repeated same-stage calls to the SAME worker so its cache accumulates) documented in the canonical doc and each of the four SKILL.md files' compact ## Model Tiering (#517, optional) dispatch block — scoped so the unset default stays byte-equivalent, dispatch shapes included. No agent-file edits (the sha256-locked bibliography_agent.md untouched), no schema change, no hook. Spec: docs/design/2026-07-12-517-model-tiering-spec.md.

  • Cross-model gate hardening: risk-stratified sampling, blind disagreement checkpoints, id-status allowlist, promotion bakeoff (#518). Four upgrades to shared/cross_model_verification.md and its consumers, from a 2026-07-11 cross-model consult (gpt-5.6-sol, xhigh). (1) The integrity-gate cross-model sample moves from uniform random 30% (min 5, max 15) to risk stratification across four mutually-exclusive tiers (highest-precedence tier wins, one verification per reference): HIGH-IMPACT references (headline conclusions, numerical claims, causal claims, methods-critical, disputed) verified 100% uncapped at both gates; a 10% RANDOM sample of the remainder at Stage 2.5 (round-up, min 3, max 10); at Stage 4.5, NEW-CHANGED references (behind claims new or changed since 2.5) verified 100% uncapped plus a 10% CONTROL sample of the unchanged remainder replacing RANDOM — verification budget concentrates where the paper's weight rests, and the results table gains a Tier column (integrity_verification_agent updated in lockstep). (2) The two irreversible checkpoints — research-design freeze (research_architect_agent) and final editorial decision (editorial_synthesizer_agent) — gain optional blind disagreement checks: the primary commits its own decision in the same structured form first (the architect in a new Design-Freeze Checkpoint Audit blueprint section; the synthesizer's is its emitted decision), the cross-model then produces an independent structured decision from the same inputs (never seeing the primary's decision — same anchoring-prevention rule as the integrity samples; the editorial input is the panel's panel_size N usable reviewer cards, never a hardcoded five), differing enum values trigger a targeted rebuttal addressing each cross-model driver against the evidence on file, and divergence escalates to the user — a review trigger, never a vote, never averaged; under a sprint contract the check runs strictly post-Step-3 against the mechanical protocol's editorial_decision and its drivers never enter the scoring matrix. (3) The "6th reviewer — Planned" section is retired, not deferred: the consult's counterproductive-conditions list (score averaging, role duplication, findings treated as confirmed defects, majority-vote false confidence, synthesizer context burn) matches ARS's documented anti-patterns one-for-one; the blind checkpoints are the replacement design, and the live mirrors (.claude/CLAUDE.md, shared/raise_framework.md, SETUP feature tables en/zh-TW) drop the "remains planned" claim. (4) The model-detection snippet separates "which provider endpoint" from "is this id known-good": new CROSS_MODEL_ID_STATUS=validated|provisional|unlisted announcement with an explicit warning for unlisted first-party-prefix ids (gpt-made-up no longer passes silently) — routing itself is byte-identical, an unlisted id still takes the grounded route and never falls through to the ungrounded compatible branch. Plus a § Promotion Bakeoff operationalizing the gpt-5.6-sol provisional→validated criteria: a 30-reference paired same-day run (20 real / 10 fabricated; committed as a versioned, labeled, sha256-recorded fixture before any run counts; 3 repeats with ≥2/3 majority verdict, a 1–1–1 split scored conservatively against the model that produced it) against five non-inferiority thresholds (grounded-search completion, mismatch recall, false-disagreement rate, jq-guard shape stability — a hard requirement, p95 latency), entry-gated by scripts/cross_model_smoke_test.sh, results recorded under audits/ either way — and a deliberate two-step outcome: a full pass makes the id validated, while the recommended default flips only with an additionally stated superiority or operational-benefit reason. Spec: docs/design/2026-07-12-518-cross-model-gate-hardening-spec.md. Related: #517 (model tiering) will reference the checkpoint surfaces added here.

  • GPT-5.6 Sol listed as provisional cross-model verifier + explicit reasoning-effort control (#515). OpenAI's gpt-5.6-sol (released 2026-07-08) joins the canonical model table in shared/cross_model_verification.md as provisional pending ARS validation — endpoint support (Responses API + hosted web_search), the reasoning-effort enum (none|low|medium|high|xhigh|max, default medium), and pricing (same standard rates as GPT-5.5; premium is reasoning: {mode: "pro"} on the standard slug, NOT a -pro model id) were verified first-party against OpenAI's model page and GPT-5.6 guide, but ARS-specific behavior (grounded-search completion rate, citation-mismatch recall, false-disagreement rate, jq-guard response-shape stability, p95 latency) has no operating history, so GPT-5.5 stays the recommended default. The documented OpenAI Responses call pattern gains an explicit reasoning-effort control via new ARS_CROSS_MODEL_REASONING_EFFORT — set, it is passed as reasoning.effort so the run's effort is visible and reproducible; unset, the field is omitted entirely and each model's own provider default applies (forcing one value would silently change behavior for existing gpt-5.5-pro/legacy setups, a codex-review P2) — and both SETUP quick-setup blocks (en/zh-TW, parity-linted) mirror the new example lines. New scripts/cross_model_smoke_test.sh — a live, manual (not CI; needs OPENAI_API_KEY) promotion gate asserting HTTP 2xx, a completed web_search_call, a single verdict token, VERIFIED-carries-source, model echo, and effort echo — is the prerequisite for ever flipping the default to Sol. The canonical doc's Chat-Completions-web-search claim was re-verified against OpenAI's current web-search guide and deliberately left unchanged (a cross-model review suggested it was stale; first-party docs confirm it is still accurate).

  • WP advisory held-out miss-rate measurement + acceptance set, Part 2 (#501; direction from the PR #468 review thread, @brycewang-stanford). New evals/heldout/rq_framing_offlist/: a 48-item held-out set (32 shells outside the WP01-WP20 surface forms and the four in-prompt examples — 23 family variants + 9 off-list — plus 16 domain-native hard negatives), generated cross-model (gpt-5.6-sol), shell items regex-filtered (four negatives intentionally carry listed surface substrings as hard-negative material), dual-annotated with documented drops, English-only per the #468 caveat. Scored against the runtime LLM judge (isolated claude-sonnet-5 sub-agents, verbatim advisory section only, pre-#503 vs post-#503 variants, two post replicates): overall miss rate 0.34-0.38 (above the inherited FNR < 0.30 line), concentrated in decorated compound-title off-list shells (7/9 missed, stable across replicates; judges read generic topical nouns as the exemption's "specific mechanism"), family-variant generalization under the line post-#503 (0.17-0.22), false-fire 0/1...

Read more

v3.15.0 — Release-gate hardening, prompt-debt retirement round 2, defrift locks

Choose a tag to compare

@Imbad0202 Imbad0202 released this 04 Jul 06:32
Immutable release. Only release title and notes can be modified.
v3.15.0
f86d68a

A release-discipline-and-hygiene release; no skill-behavior changes.

Added

  • Three CI release gates: CHANGELOG-covers-merges pre-tag gate — every release-worthy merge since the previous tag must be documented before tagging (#483); version-consistency invariants 9-11 (release-notes body ≥100 chars, Last-Updated within ±7 days, Key-Additions heading matches suite version) plus a tag-version-match.yml gate that re-runs the full lint at tag time (#487); command-invariants gate pinning the SessionStart announce list to the actual 16-command inventory (#486).
  • Two defrift locks (#491#492): the Phase Boundary enforcement sentence is pinned verbatim across all 23 Bucket A agent blocks (CANONICAL_ENFORCEMENT, version-matched, per-file tails free — the drift class that sat factually wrong for a month now fails CI); new check_setup_cross_model_parity.py pins the SETUP en/zh-TW ARS_CROSS_MODEL examples to each other and to the canonical model tables. Local pytest manifest gains the v3.9.4 temporal test (58 → 60 entries), closing a local-green/CI-red coverage split.

Changed

  • Prompt-debt retirement round 2 (#489#490): deep-scanned the 17 agents the first pass deferred, via 4 parallel audit batches + an independent codex cross-model challenge. 13 findings (2 P1 + 11 P2; 2 user-rejected and recorded). Both socratic_mentor agents carried live self-contradictions — stale "quit after 15 rounds" rules against a documented typical 20-30-round run — now merged to a single auto-end authority per file (threshold 30). The repo-wide stale "hook deferred to #134" enforcement sentence (false since PR #294) rewritten at 29 surfaces; few-shot and duplicated-process scaffolds trimmed across 7 agents. The 2026-06-10 F-007 negative-framing deferred item closes as verified-no-rewrite-needed. Audit report: audits/harness-retirement-2026-07-04.md.

Fixed

  • DOI badge served from shields.io; link target stays the concept DOI (#482).
  • SessionStart announce updated to the full 16-command set.

academic-pipeline tracks the suite at v3.15.0; the other three skill versions are unchanged. Full details in CHANGELOG.md.

v3.14.0 — Claude Science importability, eval-comment rendering, prompt-debt retirement

Choose a tag to compare

@Imbad0202 Imbad0202 released this 02 Jul 02:33
Immutable release. Only release title and notes can be modified.
v3.14.0
8157a15

What's Changed

  • docs: recommend auto permission mode over Skip Permissions by @Imbad0202 in #464
  • docs: add GitHub Copilot repository instructions by @Imbad0202 in #465
  • docs: add native-reviewed Korean README by @devCharlotte in #469
  • docs: credit devCharlotte for Korean README translation by @Imbad0202 in #471
  • ci: add platform-port reminder (remind, don't block) by @Imbad0202 in #473
  • chore(audit): harness-retirement 2026-07 report (#476) by @Imbad0202 in #477
  • chore(prompts): retire expired writing-harness scaffolds in 4 Bucket A agents (#476 P2) by @Imbad0202 in #478
  • ci(eval-harness): render PR comment as verdict + table, fold raw JSON by @Imbad0202 in #479
  • fix(plugin): declare explicit skill paths in marketplace.json by @Imbad0202 in #480
  • docs(release): align all doc surfaces for v3.14.0 by @Imbad0202 in #481

New Contributors

Full Changelog: v3.13.0...v3.14.0

Highlights

  • Claude Science importability (#480) — the marketplace manifest now declares explicit skill paths, so Claude Science's "Import from GitHub" finds all four skills (previously zero: GitHub-API importers cannot traverse the symlinked skills/ directory). Verified end-to-end on Claude Science; import guide in README + docs/SETUP.md Method 5. Claude Code installs are unaffected.
  • Readable eval-harness PR comments (#479) — a one-line verdict + per-task table with the raw JSON folded into <details>, replacing the raw report dump. Display layer only; run_evals, the threshold gate, and the ack contract are byte-identical.
  • Prompt-debt retirement (#477/#478) — expired writing-harness scaffolds removed from four writer-surface agents (net −111 prompt lines) after the 2026-07 harness-retirement audit, with three-track verification (sub-agent audit + cross-model review + eval harness at 100%).
  • Changelog backlog rollup — 16 [Unreleased] entries whose code shipped before the v3.13.0 tag (diff/patch revision mode #390, submission-package verifier #394, eval gold sets #215/#216, and more) are now versioned in CHANGELOG.md under a provenance note.