Repository navigation
v0.1.1 is the review release. A full review of the repository filed 43 issues against v0.1.0, and every one of them is fixed here. The local arm has also run against a real checkpoint for the first time, which found four more faults, also fixed. Nothing from the v0.2 milestone is in this release: score support, batching, and per-label scaling are still what v0.2 is for.
It is also the first release with distributions attached. They are built, checked, and attested by the release workflow, never on a laptop.
Installing
pip install "plumbline @ git+https://github.com/TMHSDigital/plumbline@v0.1.1"
Or install the wheel attached below, after checking where it came from:
gh attestation verify plumbline-0.1.1-py3-none-any.whl -R TMHSDigital/plumbline
Not pip install plumbline: that name on PyPI belongs to an unrelated project.
Fixed: what could have cost you money
- An unedited pricing template priced every call at $0, so
--max-cost-usdlet any run through (#28). The template's prices are now null, and a copy is refused until it is filled in. - Failures no retry can fix were retried anyway, each retry another billed call, with the SDKs retrying underneath (#26). Only transport failures are retried now, by one layer, so the artifact's attempt count is the call count.
- Two workers answering the same text raced on one cache file (#27). On Windows the loser crashed the run and lost its paid calls. A duplicate now waits for the first answer and is served from the cache.
- An endpoint set by environment variable changed which server answered but reached neither the cache key nor the artifact (#40).
- A row whose response named no model was never priced, so the guard and the report disagreed about the same run (#36).
- New:
plumbline run --dry-runloads, checks, and prices a run, then sends nothing (#58).
Fixed: what could have given you a wrong number
- The cascade chose its threshold and scored it on the same rows, so its coverage and savings were a best case (#30). It also never considered escalating every case (#34, #35).
- Equal-count binning split tied predictions by input order, so the same rows could give an ECE of 0.3 in one order and 0.2 in another (#38).
- A cached answer could be served for a differently ordered prompt on the arms that list options in order (#37).
- Option descriptions were parsed and then dropped (#39).
- An inverted score far below its null read as INCONCLUSIVE with advice to collect more rows (#33).
- A distribution holding NaN, a negative, or a value above 1 was accepted when its entries summed to about 1 (#31, #32).
plumbline reportcombined runs over different datasets as if their figures were comparable (#43). It now refuses unless asked with--allow-mixed.- The site printed some figures one step off from the report near a rounding tie (#48).
New
- The site, at https://tmhsdigital.github.io/plumbline/: a floor calculator that reproduces the Python's figures to 1e-9, a sample-size planner, the docs with search, and a box to paste your own predictions and see their ECE against its floor, entirely in the browser (#56).
- Run options:
--base-url,--timeout, and--devicereach an adapter that takes them and are recorded in the artifact, and--semanticsnow overrides any adapter that accepts it (#58, #3). - A report rebuilt from artifacts keeps its Dataset section (#58).
- Ties for the top option are recorded in the artifact and counted in the report (#10).
- The report states the grid probabilities arrived on, such as the hosted vendor's 0.01 (#9).
Measured
- The local arm ran for real, against Qwen2.5-1.5B-Instruct pinned by commit, on a GPU (#3). It scored 39 of the fixture's 105 rows and refused 66 whose options are multi-token. It landed at chance accuracy, which is what a null should do, and the run found four faults, now fixed: the checkpoint loaded once per worker, the command line had no device option, kernel setup was timed as latency, and two report lines were misworded for arms with failures.
- A two-decimal grid does not raise the ECE floor (#9). A calibrated model reported on the grid clears the floor's 95th percentile at the nominal 5 percent, from 40 to 10,000 rows.
- Reporting no distribution is a fixed bias, not a small-sample effect (#8). The top-line form's leftover miscalibration stays about 0.18 on an underconfident model from 500 to 20,000 rows, while the floor keeps falling.
Both measurements are reproducible from scripts in scripts/, and METHODOLOGY carries the tables.
Hardening
- The dependency floors are now the oldest releases the suite passes on, and CI tests them (#47). A broken SDK takes down only the adapter that needs it.
- CI runs on Ubuntu, Windows, and macOS, Python 3.12 through 3.14. Every action is pinned to a commit, and Dependabot keeps the pins current (#52).
- The release workflow builds with a read-only token and attests what it attaches, and a job holding the only write token runs no project code (#51).
- A credential inside a list no longer reaches the artifact, and a report no longer carries your absolute paths (#46).
scripts/build_site.py --outdeletes only a directory it built (#50).
Known limitations
- The generative transport has still never run outside the test suite (#3).
- Choice and yes/no questions only. Ordinal score rows load, are marked, and are excluded.
- One request per case, so cost and latency are conservative relative to batched use.
- Temperature scaling only, and it needs 200 held-out rows.
- Cost needs a pricing table you supply, for vendors whose terms keep pricing confidential.
- The local arm scores only options that are a single token for the checkpoint, and refuses the rest by name.
The full list is in CHANGELOG.md.