Releases: caty-ai/caty-agent-harness
Release list
v0.23.8
v0.23.7: pin family-dev-handbook reusable workflows to ci-v1 peeled SHA (#245)
v0.23.7: pin family-dev-handbook reusable workflows to ci-v1 peeled SHA (#245)
v0.23.6 — docs: python3 3.9+ floor in AGENTS.md and plugin-convention (#248)
v0.23.6 — docs: python3 3.9+ floor in AGENTS.md and plugin-convention (#248)
Completes the floor declaration from v0.23.5 in the two files not claimed by
other lanes; the README table rows are folded into the #231 front redesign.
Evidence: PR #249 (completion record in body; unanimous 3-seat GO in comments)
#249
v0.23.5 — apply-promotions: Python 3.9-compatible stub write; python3 3.9+ floor declared in docs (#240)
v0.23.4 — docs: the v0.23.1–v0.23.3 sentinel surface in reference/engineering EN+JA (#238)
v0.23.4 — docs: the v0.23.1–v0.23.3 sentinel surface in reference/engineering EN+JA (#238)
Documents the operator-facing surface shipped by the sentinel serial
group: the OVF_MODEL_ALIASES / OVF_MODEL_THRESHOLDS contract rows
(load-time validation, placeholder defaults, ships-empty), corrected
OVF_T_ABS/OVF_W_PCT forwarding/provenance semantics, and the
regime_change / tap_drift event families with the new attempt_end and
task-end receipt fields. Docs-only; no behavior changes.
Evidence: PR #239 (L1-7 completion record inline; 3-seat claim-accuracy
review record: #238 (comment))
v0.23.3 — sentinel: tap_drift, the third instrument state (#218)
v0.23.3 — sentinel: tap_drift, the third instrument state (#218)
The instrument now distinguishes 'the tap lied' from 'didn't fire' and
'couldn't see': the claude-code adapter declares drift_reference:
derived and persists raw_usage per turn; core replays normalization
over ledger-held values and reconciles paired cumulative sums over R_e
with episode-triggered, direction-agnostic tap_drift events plus
schema-signature change detection, all regime-epoch scoped. v1 is
log-only. attempt_end and the task-end receipt carry tap_drift_count /
drift_reference_status (weakest across regimes); regime_change carries
real per-regime capability values.
Adopted from the external design review by @pm25coder (issue #159,
comment of 2026-08-27).
Evidence: PR #236 (L1-7 completion record inline; 3-seat review record:
#218 (comment))
v0.23.2 — sentinel: per-model threshold injection point (#219)
v0.23.2 — sentinel: per-model threshold injection point (#219)
The per-model threshold table (DESIGN v0.6.3 §5-1) now sits between the
explicit config override and the product default, per key, with frozen
deterministic matching (canonical exact > longest-literal-prefix glob,
star only) and load-time validation (duplicates/unknown/range errors —
never silent first-match). New OVF_MODEL_THRESHOLDS config knob carrying
the table plus N_drift/theta_drift placeholders for the #218 lane. The
table ships empty; P2 window-probe numbers enter later as pure config
writeback. threshold_sources can now report per-model; regime changes
re-consult the table.
Evidence: PR #234 (L1-7 completion record inline; 3-seat review record:
#219 (comment))
v0.23.1 — sentinel: reset MA/slope on model/runtime regime change (#217)
v0.23.1 — sentinel: reset MA/slope on model/runtime regime change (#217)
A mid-task model/runtime change now triggers exactly one regime reset:
identity-first comparison, series/hysteresis clear, full re-resolution of
model-keyed config, reset-before-fire. New regime_change event (DESIGN
v0.6.3 §6), regime_change_resets in attempt_end, runtime on turn events,
OVF_MODEL_ALIASES knob, threshold provenance via OVF_T_ABS/OVF_W_PCT
explicitness.
Adopted from the external design review by @pm25coder (issue #159,
comment of 2026-08-27).
Evidence: PR #232 (L1-7 completion record inline; 3-seat heterogeneous
review record: #217 (comment))
v0.23.0 — startup self-check: fail-closed ordering asserts for threshold pairs at all four entry points (#108)
v0.23.0 — startup self-check: fail-closed ordering asserts for threshold pairs at all four entry points (#108)
v0.22.6 — docs: publish P2-WIN (third sealed experiment) in README x4 + benchmark EN/JA
v0.22.6 — docs: publish P2-WIN (third sealed experiment) in README x4 + benchmark EN/JA