Skip to content

v2.0.0

Choose a tag to compare

@m-sanchez m-sanchez released this 01 Sep 19:45
· 5 commits to main since this release

ECE was not a function of the data. Equal-mass binning - the strategy the README recommends for exactly the top-heavy regime where it broke - split runs of tied confidences across bin boundaries. 1,000 predictions at confidence 0.9 with 60% correct (a true gap of 0.30) read 0.4146, and MCE swung 0.1111 to 0.1570 across 50 shuffles of the same rows. Cuts now snap to value boundaries, and a permutation-invariance property test locks it.

  • effectiveBins reports how many bins actually carried data, so a reliability diagram cannot be labelled with a bin count it does not have.
  • New eceInterval() (seeded bootstrap CI) and nullEce(), the noise floor. ECE is positively biased: a perfectly calibrated n=100 at 15 bins reads a median of 0.0874. The README publishes that table, because a bare ECE of 0.08 means nothing without it.
  • calibrationError([]) and brier([]) now return NaN rather than 0 - no data was quietly passing an ece <= 0.1 ship bar as perfect calibration.
  • fitTemperature reports atBound when the optimum is pinned to its own bracket, and validates labels and logits instead of returning NaN.
  • softmax no longer blows the stack on a real vocabulary (tested at 200,000 classes).

36 tests.