Skip to content

plumbline v0.1.0

Choose a tag to compare

@TMHSDigital TMHSDigital released this 22 Sep 00:02

This release ships no wheel or sdist, on purpose. Distributions are now built and attached only by the release workflow, which checks each one before it goes out: the tag and the package version agree, the full test gate passes at the tag, and the wheel installs clean with its py.typed marker. The v0.1.0 tag predates that marker, so its wheel would hide plumbline's type annotations from downstream type checkers, and the workflow refused it. Nothing was attached rather than something that fails the checks.

To use v0.1.0, install it from the tag:

pip install "plumbline @ git+https://github.com/TMHSDigital/plumbline@v0.1.0"

Not pip install plumbline: that name on PyPI belongs to an unrelated project. The next release will carry a wheel and an sdist with signed build provenance, which gh attestation verify can check.

plumbline measures whether a decision model's reported probabilities actually hold up on your own labeled data, and turns that into three decisions: whether to use the model, where to set your confidence threshold, and what a cascade at that threshold saves you. Every figure it prints is reported against its own null, and it says plainly when a value is indistinguishable from a perfectly calibrated model rather than handing you a number that looks like a finding.

This is not a leaderboard

plumbline does not rank vendors and publishes no combined score. For cross-vendor ranking of Jev-class decision models, go to JevBench and Benchmark Heaven. That is their job and they do it properly, on a shared dataset with a published methodology. Numbers from plumbline are never comparable with theirs: different harness, different prompts, different scoring.

The question plumbline answers is the one a ranking structurally cannot. A leaderboard tells you how a model did on someone else's rows. Only your rows can tell you whether its probabilities mean anything where you intend to use them.

Why the nulls matter

Expected Calibration Error has a floor that is not zero, and that floor depends on how many rows you have. A perfectly calibrated model measured on a few hundred rows does not score 0; it scores some positive number set by binning noise and sample size. If you do not know that number, you cannot read your own. This is not a rounding concern: on a few hundred rows a calibration claim is frequently not measurable at all. So plumbline computes the null for every figure and states when the observed value sits inside it. In the shipped example report, ECE is 0.0740 against a floor of 0.0707 over 105 rows, and the tool declines to draw a conclusion.

What v0.1.0 supports

  • Choice and Noul questions. A yes/no row is asked as a Noul where the transport has one, and every record carries both what the row asks and how it was asked, so the two are never averaged together.
  • Three adapter transports. A wire-format HTTP adapter for hosted and self-hosted endpoints, an option-token logits adapter for local checkpoints, and a generative control arm. Adding a vendor is a config entry, not a new module.
  • Three probability semantics classes. Calibrated claims, restricted softmax, and no probability at all. The report groups on this field and refuses to place figures from different classes side by side.
  • Temperature scaling with a refusal gate. A temperature is fitted on a held-out split, and when the residual says temperature is the wrong correction, the tool emits no temperature rather than one that does not fit. Below 200 held-out rows it refuses outright.
  • Cascade threshold selection. Given what one escalation and one wrong answer cost you, it states where to set the threshold and what that buys. Without both numbers it refuses, because no benchmark can know them.
  • Cost and latency with provenance. Every pricing entry carries its source and the date it was read, a blank cost column names which of four reasons made it blank, and latency percentiles use nearest rank.
  • A loader that refuses rather than repairs. A row whose gold label is not among its own options is refused with its line number, because scoring it would mark every system wrong and read as a model failure.

Known limitations

  • Choice and Noul only. Ordinal Score rows load, are marked, and are excluded from every figure.
  • One request per case, no batching, so cost and latency are conservative relative to batched use.
  • Temperature scaling only. Per-label and vector scaling are not fitted.
  • Recalibration needs 200 held-out rows and refuses below that. Most datasets people try first will not reach it.
  • Cost requires a pricing table you supply, for vendors whose terms treat pricing as confidential. Without one, cost reports as unpriced and the spend guard refuses the run rather than bounding it, because a guard cannot bound a run it cannot cost.
  • The METHODOLOGY numbers derived from the seeded mock are properties of the harness's resolution, not measurements of any vendor, and are labeled as such.
  • Adapters reporting no distribution are excluded from multiclass Brier and recalibrate materially worse.
  • Probabilities from a hosted API may arrive quantized, which bounds the resolution of any threshold or bin computed from them.
  • Verified on Windows and Ubuntu, Python 3.12 and 3.13.

Planned for v0.2

  • Score and ordinal support. Rank-aware metrics, because every metric here currently treats wrong-by-one and wrong-by-three identically.
  • Batching. Packing many questions against one shared state in a single call is a materially different cost and latency profile, and is the largest measurement gap in v0.1.
  • Per-label and vector scaling. So that a refused temperature leaves the user with something to apply instead of nothing.

Getting started

Requires Python 3.12 or later and uv. The quickstart runs against a vendored public fixture with a deterministic mock arm, needs no API key, and spends nothing. See the README.

Apache-2.0. The vendored JevBench fixture is MIT and attributed in datasets/public/README.md.