Skip to content

Releases: TMHSDigital/plumbline

plumbline v0.1.1

Choose a tag to compare

@TMHSDigital TMHSDigital released this 25 Sep 21:06
9542d44

v0.1.1 is the review release. A full review of the repository filed 43 issues against v0.1.0, and every one of them is fixed here. The local arm has also run against a real checkpoint for the first time, which found four more faults, also fixed. Nothing from the v0.2 milestone is in this release: score support, batching, and per-label scaling are still what v0.2 is for.

It is also the first release with distributions attached. They are built, checked, and attested by the release workflow, never on a laptop.

Installing

pip install "plumbline @ git+https://github.com/TMHSDigital/plumbline@v0.1.1"

Or install the wheel attached below, after checking where it came from:

gh attestation verify plumbline-0.1.1-py3-none-any.whl -R TMHSDigital/plumbline

Not pip install plumbline: that name on PyPI belongs to an unrelated project.

Fixed: what could have cost you money

  • An unedited pricing template priced every call at $0, so --max-cost-usd let any run through (#28). The template's prices are now null, and a copy is refused until it is filled in.
  • Failures no retry can fix were retried anyway, each retry another billed call, with the SDKs retrying underneath (#26). Only transport failures are retried now, by one layer, so the artifact's attempt count is the call count.
  • Two workers answering the same text raced on one cache file (#27). On Windows the loser crashed the run and lost its paid calls. A duplicate now waits for the first answer and is served from the cache.
  • An endpoint set by environment variable changed which server answered but reached neither the cache key nor the artifact (#40).
  • A row whose response named no model was never priced, so the guard and the report disagreed about the same run (#36).
  • New: plumbline run --dry-run loads, checks, and prices a run, then sends nothing (#58).

Fixed: what could have given you a wrong number

  • The cascade chose its threshold and scored it on the same rows, so its coverage and savings were a best case (#30). It also never considered escalating every case (#34, #35).
  • Equal-count binning split tied predictions by input order, so the same rows could give an ECE of 0.3 in one order and 0.2 in another (#38).
  • A cached answer could be served for a differently ordered prompt on the arms that list options in order (#37).
  • Option descriptions were parsed and then dropped (#39).
  • An inverted score far below its null read as INCONCLUSIVE with advice to collect more rows (#33).
  • A distribution holding NaN, a negative, or a value above 1 was accepted when its entries summed to about 1 (#31, #32).
  • plumbline report combined runs over different datasets as if their figures were comparable (#43). It now refuses unless asked with --allow-mixed.
  • The site printed some figures one step off from the report near a rounding tie (#48).

New

  • The site, at https://tmhsdigital.github.io/plumbline/: a floor calculator that reproduces the Python's figures to 1e-9, a sample-size planner, the docs with search, and a box to paste your own predictions and see their ECE against its floor, entirely in the browser (#56).
  • Run options: --base-url, --timeout, and --device reach an adapter that takes them and are recorded in the artifact, and --semantics now overrides any adapter that accepts it (#58, #3).
  • A report rebuilt from artifacts keeps its Dataset section (#58).
  • Ties for the top option are recorded in the artifact and counted in the report (#10).
  • The report states the grid probabilities arrived on, such as the hosted vendor's 0.01 (#9).

Measured

  • The local arm ran for real, against Qwen2.5-1.5B-Instruct pinned by commit, on a GPU (#3). It scored 39 of the fixture's 105 rows and refused 66 whose options are multi-token. It landed at chance accuracy, which is what a null should do, and the run found four faults, now fixed: the checkpoint loaded once per worker, the command line had no device option, kernel setup was timed as latency, and two report lines were misworded for arms with failures.
  • A two-decimal grid does not raise the ECE floor (#9). A calibrated model reported on the grid clears the floor's 95th percentile at the nominal 5 percent, from 40 to 10,000 rows.
  • Reporting no distribution is a fixed bias, not a small-sample effect (#8). The top-line form's leftover miscalibration stays about 0.18 on an underconfident model from 500 to 20,000 rows, while the floor keeps falling.

Both measurements are reproducible from scripts in scripts/, and METHODOLOGY carries the tables.

Hardening

  • The dependency floors are now the oldest releases the suite passes on, and CI tests them (#47). A broken SDK takes down only the adapter that needs it.
  • CI runs on Ubuntu, Windows, and macOS, Python 3.12 through 3.14. Every action is pinned to a commit, and Dependabot keeps the pins current (#52).
  • The release workflow builds with a read-only token and attests what it attaches, and a job holding the only write token runs no project code (#51).
  • A credential inside a list no longer reaches the artifact, and a report no longer carries your absolute paths (#46).
  • scripts/build_site.py --out deletes only a directory it built (#50).

Known limitations

  • The generative transport has still never run outside the test suite (#3).
  • Choice and yes/no questions only. Ordinal score rows load, are marked, and are excluded.
  • One request per case, so cost and latency are conservative relative to batched use.
  • Temperature scaling only, and it needs 200 held-out rows.
  • Cost needs a pricing table you supply, for vendors whose terms keep pricing confidential.
  • The local arm scores only options that are a single token for the checkpoint, and refuses the rest by name.

The full list is in CHANGELOG.md.

plumbline v0.1.0

Choose a tag to compare

@TMHSDigital TMHSDigital released this 22 Sep 00:02

This release ships no wheel or sdist, on purpose. Distributions are now built and attached only by the release workflow, which checks each one before it goes out: the tag and the package version agree, the full test gate passes at the tag, and the wheel installs clean with its py.typed marker. The v0.1.0 tag predates that marker, so its wheel would hide plumbline's type annotations from downstream type checkers, and the workflow refused it. Nothing was attached rather than something that fails the checks.

To use v0.1.0, install it from the tag:

pip install "plumbline @ git+https://github.com/TMHSDigital/plumbline@v0.1.0"

Not pip install plumbline: that name on PyPI belongs to an unrelated project. The next release will carry a wheel and an sdist with signed build provenance, which gh attestation verify can check.

plumbline measures whether a decision model's reported probabilities actually hold up on your own labeled data, and turns that into three decisions: whether to use the model, where to set your confidence threshold, and what a cascade at that threshold saves you. Every figure it prints is reported against its own null, and it says plainly when a value is indistinguishable from a perfectly calibrated model rather than handing you a number that looks like a finding.

This is not a leaderboard

plumbline does not rank vendors and publishes no combined score. For cross-vendor ranking of Jev-class decision models, go to JevBench and Benchmark Heaven. That is their job and they do it properly, on a shared dataset with a published methodology. Numbers from plumbline are never comparable with theirs: different harness, different prompts, different scoring.

The question plumbline answers is the one a ranking structurally cannot. A leaderboard tells you how a model did on someone else's rows. Only your rows can tell you whether its probabilities mean anything where you intend to use them.

Why the nulls matter

Expected Calibration Error has a floor that is not zero, and that floor depends on how many rows you have. A perfectly calibrated model measured on a few hundred rows does not score 0; it scores some positive number set by binning noise and sample size. If you do not know that number, you cannot read your own. This is not a rounding concern: on a few hundred rows a calibration claim is frequently not measurable at all. So plumbline computes the null for every figure and states when the observed value sits inside it. In the shipped example report, ECE is 0.0740 against a floor of 0.0707 over 105 rows, and the tool declines to draw a conclusion.

What v0.1.0 supports

  • Choice and Noul questions. A yes/no row is asked as a Noul where the transport has one, and every record carries both what the row asks and how it was asked, so the two are never averaged together.
  • Three adapter transports. A wire-format HTTP adapter for hosted and self-hosted endpoints, an option-token logits adapter for local checkpoints, and a generative control arm. Adding a vendor is a config entry, not a new module.
  • Three probability semantics classes. Calibrated claims, restricted softmax, and no probability at all. The report groups on this field and refuses to place figures from different classes side by side.
  • Temperature scaling with a refusal gate. A temperature is fitted on a held-out split, and when the residual says temperature is the wrong correction, the tool emits no temperature rather than one that does not fit. Below 200 held-out rows it refuses outright.
  • Cascade threshold selection. Given what one escalation and one wrong answer cost you, it states where to set the threshold and what that buys. Without both numbers it refuses, because no benchmark can know them.
  • Cost and latency with provenance. Every pricing entry carries its source and the date it was read, a blank cost column names which of four reasons made it blank, and latency percentiles use nearest rank.
  • A loader that refuses rather than repairs. A row whose gold label is not among its own options is refused with its line number, because scoring it would mark every system wrong and read as a model failure.

Known limitations

  • Choice and Noul only. Ordinal Score rows load, are marked, and are excluded from every figure.
  • One request per case, no batching, so cost and latency are conservative relative to batched use.
  • Temperature scaling only. Per-label and vector scaling are not fitted.
  • Recalibration needs 200 held-out rows and refuses below that. Most datasets people try first will not reach it.
  • Cost requires a pricing table you supply, for vendors whose terms treat pricing as confidential. Without one, cost reports as unpriced and the spend guard refuses the run rather than bounding it, because a guard cannot bound a run it cannot cost.
  • The METHODOLOGY numbers derived from the seeded mock are properties of the harness's resolution, not measurements of any vendor, and are labeled as such.
  • Adapters reporting no distribution are excluded from multiclass Brier and recalibrate materially worse.
  • Probabilities from a hosted API may arrive quantized, which bounds the resolution of any threshold or bin computed from them.
  • Verified on Windows and Ubuntu, Python 3.12 and 3.13.

Planned for v0.2

  • Score and ordinal support. Rank-aware metrics, because every metric here currently treats wrong-by-one and wrong-by-three identically.
  • Batching. Packing many questions against one shared state in a single call is a materially different cost and latency profile, and is the largest measurement gap in v0.1.
  • Per-label and vector scaling. So that a refused temperature leaves the user with something to apply instead of nothing.

Getting started

Requires Python 3.12 or later and uv. The quickstart runs against a vendored public fixture with a deterministic mock arm, needs no API key, and spends nothing. See the README.

Apache-2.0. The vendored JevBench fixture is MIT and attributed in datasets/public/README.md.