Releases: Quantile-Labs/touchstone
Release list
v0.6.0
What's Changed
- add uncertainty budget and assumption checks by @Adeniyikayodee in #57
- release 0.6.0 by @Adeniyikayodee in #58
Full Changelog: v0.5.0...v0.6.0
v0.5.0
What's Changed
- state conformance against nist ai 800-2 by @Adeniyikayodee in #51
- record run cost in environment.json by @Adeniyikayodee in #52
- add paired difference and plain grade output by @Adeniyikayodee in #53
- refuse writes to sealed bundles by @Adeniyikayodee in #54
- state the ceiling in inconclusive grade reasons by @Adeniyikayodee in #55
- release 0.5.0 by @Adeniyikayodee in #56
Full Changelog: v0.4.0...v0.5.0
v0.4.0
What's Changed
- name 0.3.0 as the current release by @Adeniyikayodee in #40
- add the open source governance files by @Adeniyikayodee in #41
- split replicate variance into its two components by @Adeniyikayodee in #42
- document how an agent drives the pipeline by @Adeniyikayodee in #43
- ignore the roadmap alongside its sources by @Adeniyikayodee in #44
- remove the agent section from the readme by @Adeniyikayodee in #45
- pull the base image the backend tests run by @Adeniyikayodee in #46
- name the estimand the clustered interval reports by @Adeniyikayodee in #47
- generate json schemas from the contracts by @Adeniyikayodee in #48
- add --json to the commands that check something by @Adeniyikayodee in #49
- release 0.4.0 by @Adeniyikayodee in #50
Full Changelog: v0.3.0...v0.4.0
v0.3.0
Intervals over repeated items are wider. A bundle sealed by 0.2.1 with replicates above one holds numbers this version does not reproduce from the same rows.
Asking an item twice does not give two independent observations of the system, but the rollup counted one row per item-trial and gave that denominator to Wilson. Over 200 items, a nominal 95% interval held:
| replicates | 0.2.1 | 0.3.0 |
|---|---|---|
| 1 | 95% | 95% |
| 5 | 79% | 96% |
| 20 | 54% | 95% |
The interval narrowed with every replicate added while the rate it covers did not move.
What changed
- The interval is computed over items. An outcome with repeated items reports
estimator: wilson_clusteredand carrieseffective_nanddesign_effect. - The BCa bootstrap resamples items rather than rows.
- The printed rate is the interval the bundle stores, rather than one recomputed from
kandn, which would have disagreed with the file once any interval was widened. - The confident-and-wrong rate counts an item once rather than once per replicate.
A run with one replicate per item is unchanged, and still reports estimator: wilson.
Upgrading
Re-read a bundle rather than comparing a 0.2.1 estimate to a 0.3.0 one. The rows are untouched and estimate recomputes from them, so an old bundle can be brought forward without rerunning the evaluation.
See Rates and Wilson and Status.
Still classified Development Status :: 2 - Pre-Alpha, which is accurate.
v0.2.1
What's Changed
- name a memory kill from the exit code by @Adeniyikayodee in #34
- release 0.2.1 by @Adeniyikayodee in #35
Full Changelog: v0.2.0...v0.2.1
v0.2.0
What's Changed
- describe the egress proxy in the pack docs by @Adeniyikayodee in #26
- roll up each stratum key on its own by @Adeniyikayodee in #24
- add a documentation site and deploy it from vercel by @Adeniyikayodee in #27
- publish the docs site from github pages by @Adeniyikayodee in #28
- name only inspect as the tool for exploring by @Adeniyikayodee in #29
- name only inspect in the readme too by @Adeniyikayodee in #30
- map dqi indicators to regulatory clauses by @Adeniyikayodee in #31
- cite the cbn cybersecurity framework by @Adeniyikayodee in #32
- release 0.2.0 by @Adeniyikayodee in #33
Full Changelog: v0.1.0...v0.2.0
v0.1.0
What's Changed
- work on branches and land through pull requests by @Adeniyikayodee in #1
- lead the readme with what it is for by @Adeniyikayodee in #2
- say plainly that this evaluates ai systems by @Adeniyikayodee in #3
- put trust at the front, not the sceptic by @Adeniyikayodee in #4
- write the readme in plain english by @Adeniyikayodee in #5
- add two commit message rules by @Adeniyikayodee in #6
- reject slogan subjects in the commit linter by @Adeniyikayodee in #7
- pin the readme validate and verify output by @Adeniyikayodee in #8
- add a golden bundle for estimate serialisation by @Adeniyikayodee in #9
- widen a worst stratum interval for selection by @Adeniyikayodee in #10
- state what the numbers do not prove by @Adeniyikayodee in #11
- reorder the readme around the person checking by @Adeniyikayodee in #12
- copy the plan only when it is somewhere else by @Adeniyikayodee in #13
- make the readme legible cold by @Adeniyikayodee in #14
- refuse to seal a run that never finished by @Adeniyikayodee in #15
- check for an issue reference in the subject by @Adeniyikayodee in #16
- print the value an expression graded by @Adeniyikayodee in #17
- grade the indicators a person assesses by @Adeniyikayodee in #18
- grade movement against a prior bundle by @Adeniyikayodee in #19
- test shutdown against a killed harness by @Adeniyikayodee in #20
- cut the readme to what it has to say by @Adeniyikayodee in #21
- drop the em dashes the readme cut introduced by @Adeniyikayodee in #22
- release 0.1.0 by @Adeniyikayodee in #23
New Contributors
- @Adeniyikayodee made their first contribution in #1
Full Changelog: v0.0.1...v0.1.0
v0.0.1
First release, and it claims the name more than it does the work.
Working: validate, bundle, verify, version. A bundle can be sealed, re-checked
offline, and re-checked without this tool using shasum and jq.
Not working: freeze, run, estimate, grade. They exit 2.
Requires Python 3.12 or later. See the README for what each command does.
Full Changelog: https://github.com/Quantile-Labs/touchstone/commits/v0.0.1