Skip to content

release prep: version 3.7.0, changelog for both adapters, clamp-branch test#127

Merged
b7n0de merged 2 commits into
mainfrom
release/v3.7.0-changelog-version
Jul 23, 2026
Merged

release prep: version 3.7.0, changelog for both adapters, clamp-branch test#127
b7n0de merged 2 commits into
mainfrom
release/v3.7.0-changelog-version

Conversation

@b7n0de

@b7n0de b7n0de commented Jul 23, 2026

Copy link
Copy Markdown
Owner

Release preparation per maintenance spec 2026-07-23 (version decision 3.7.0 confirmed):

Verified before this PR: full suite 1983 tests OK, mypy clean, conformance corpus green, 4/4 adapter negative probes per review spec. Tag/Release/PyPI are explicitly a separate owner-gated step.

kraxo added 2 commits July 23, 2026 10:16
…h test

- CHANGELOG: [3.7.0] section (lm-eval sample-count provenance #116 by @tuodijihua closing #115,
  conformance-crossimpl acceptance target folded from [Unreleased], CONFORMANCE/COMMERCIAL_BOUNDARY
  docs #107, Dependabot consolidation #119-#126); editorial note in [3.6.3] documenting that #112
  (inspect_ai scorer provenance, @tuodijihua) shipped in 3.6.3 but was not recorded at release time.
- Version 3.6.3 -> 3.7.0 single-sourced across pyproject.toml, __init__.py, CITATION.cff.
- One test for the clamp branch: effective > original yields skipped_samples 0 with raw counts visible.

Release, tag and PyPI publish are explicitly OUT of scope here (separate owner-gated step).
Six-lens falsification-first audit with refute-to-kill jury on commit 02509ca (version bump
PR head): zero confirmed findings; four standing regression targets plus the two learned
register-integrity classes attacked by name and confirmed fail-closed. Two honest residuals
recorded (register freshness binding P2, dependency-free wording) as pre-publication follow-ups.
Satisfies the version-coupled audit-record requirement (F7 / pre_tag_audit_gate).
@b7n0de
b7n0de merged commit defc3ee into main Jul 23, 2026
21 checks passed
@b7n0de
b7n0de deleted the release/v3.7.0-changelog-version branch July 23, 2026 11:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[feat] lm-eval adapter: record the scored/unscored sample breakdown in provenance (parity with #112)

1 participant