Skip to content

v0.26.0

Choose a tag to compare

@bbstats bbstats released this 30 Jul 15:31
· 330 commits to main since this release

Added

  • ChimeraBoostQuantileRegressor: a whole predictive distribution from one
    booster, with predictions that cannot cross.
    One tree structure per round
    serves every level in quantiles (default 0.05 ... 0.95), each leaf holding
    a K-vector whose entries are the exact per-level empirical quantiles of the
    leaf's residuals. predict returns the grid, or a central interval, or the
    mean by integrating the quantile function.

    The ordering guarantee is structural rather than a repair applied at predict
    time: the model starts from the sorted global quantiles and every leaf
    vector is projected onto increments that cannot reorder anything, so
    diff(Q, axis=1) >= 0 holds exactly, at every intermediate staged_predict
    stage too. Measured crossing rate 0.0000 against 0.18-0.21 for 19
    independently fitted LightGBM quantile boosters. Intervals can still be
    narrower than the pooled one where the data is quiet — a narrowing budget
    buys that back, which a plain monotone-increment construction cannot express
    at all.

    The split search runs once per round instead of once per level, so the
    saving grows with data width: 3.4x the fit speed of 19 independent
    boosters at 5 features, 4.8x at 32, 7.8x at 128
    , with pinball loss within
    3% throughout and better on wide data. conformalize=True calibrates the
    intervals on a fold held out before the early-stopping split (worst coverage
    error 0.7 points at n = 10 000). New chimeraboost.quantile_metrics scores
    a predicted grid: per-level pinball, CRPS, and coverage with width.
    Full record in benchmarks/QUANTILE_PLAN.md; defaults elsewhere are
    untouched.

Fixed

  • max_bins below 16 no longer crashes on a dominated column. The greedy
    border pass floored the light region's bin budget at max_bins // 16, which
    is zero under 16; a column whose mass sits overwhelmingly on one value then
    divided by zero. max_bins of 3, 4 and 6 all failed outright on a 90%-zeros
    column. Budgets of 16 and up keep their exact allocation.
  • A bagged classifier survives a one-row member draw. The rare-class guard
    added below overwrote a random drawn row with a donor of the missing class;
    when the draw held exactly one row that replaced the only row, so the member
    still saw a single class and raised "Need at least 2 classes" anyway. Tiny-n
    bags with a small max_samples now grow the draw to two rows instead.
  • A column dominated by its minimum value no longer loses its bins. When
    one value holds more than an even bin's share of the mass (sparse count
    features: mostly zeros), the quantile borders collapsed onto that value and
    the whole column silently binned as constant. Colliding quantile levels now
    trigger a greedy border pass that isolates heavy values in bins of their own
    and spreads the remaining budget over the rest by mass. Columns without
    collisions keep bit-identical borders.
  • eval_set targets are validated like training targets. A NaN/inf in the
    validation y (or a value outside the loss's domain, e.g. a zero with
    loss="Gamma", or a custom eval_metric returning NaN) made every
    validation score NaN, and early stopping silently kept a one-tree model.
    Both are errors now, raised with the cause named.
  • A reordered or renamed eval_set DataFrame raises at fit, matching the
    predict-time guard. It was consumed positionally and silently corrupted
    early stopping, temperature scaling, and the conformal offset. The
    shap_values background matrix gets the same check.
  • Zero-weight rows can no longer steer the post-fit calibrations. The
    classifier's temperature and the quantile regressor's conformal offset now
    honor validation-row weights, as the sample_weight contract promises.
  • conformalize=True on asymmetric quantile grids no longer breaks the
    non-crossing guarantee.
    Unpaired levels kept scale 1.0 while their
    neighbors shrank and could be jumped; they now interpolate their factor from
    the paired levels. A grid with no symmetric pair at all raises instead of
    silently skipping calibration.
  • A bagged binary fit survives a bootstrap member missing the rare class
    (one row of the missing class is injected) instead of crashing with "Need
    at least 2 classes".
  • ordered_boosting=True with l2_leaf_reg=0 no longer crashes on
    singleton leaves (ZeroDivisionError in the leave-one-out step).
  • Refitting on a plain array clears the previous fit's feature names, so
    the column-order guard no longer misfires against stale names.
  • The multi-quantile split search weights the hessian by sample_weight,
    matching the scalar path; weighted fits previously optimized a different
    objective in the structure than in the leaves.
  • groups is honored under bagging: a member's out-of-bag early-stopping
    rows now exclude every group present in its training sample (falling back
    to the member's group-aware auto-split when none remain). Previously the
    group boundary was silently ignored for n_ensembles > 1.
  • mkdocs build --strict failed on two pre-existing warnings. The
    ChimeraBoostQuantileRegressor docstring closed its Parameters section
    with a free prose paragraph, which griffe parsed as three malformed parameters
    (Other, defaults, grid) and rendered as garbage; it is a Notes
    section now. And docs/benchmarks.md pointed at ../images/public_pareto.png,
    outside the docs tree, so the chart did not render on the published page.

Changed

  • Small-batch predict is up to ~1.4x faster. The serial/parallel kernel
    dispatch threshold assumed the parallel forest walk overtakes serial at about
    5 rows; re-measuring both kernels on the same packed forest puts the crossover
    between 32 and 64, so every 5-to-32-row predict was paying thread fork/join
    for nothing. The threshold moves from 4 to 32. The two kernels are
    bit-identical, so predictions are unchanged. warmup() now derives its
    parallel-batch row count from the threshold rather than hardcoding it, so a
    future change cannot silently leave the parallel kernel uncompiled.

  • Binning is up to ~4x faster on zero-inflated columns and ~3x faster on
    dense ones.
    The greedy border pass walked every distinct value in a Python
    loop; it is now a numba kernel, its per-value mass is read off the sort's run
    lengths instead of a searchsorted + np.add.at pass, and the quantile
    probe partitions the already-sorted copy in place. Borders are bit-identical
    throughout (pinned against transcriptions of the old code). 200k x 30
    zero-inflated: 1.34 s -> 0.33 s; dense: 0.29 s -> 0.10 s.

  • Bagged members draw whole groups when groups is passed. The
    group-disjoint out-of-bag eval set introduced above was correct but dead in
    practice: a typical 80% row draw touches essentially every group, so it came
    back empty and every grouped member fell back to its auto-split. Drawing
    max_samples of the groups (a cluster bootstrap at 1.0) always holds at
    least one group out, so members early-stop on groups they never saw and no
    longer carve a validation slice out of their own sample. Held-out-group
    strength is unchanged on a 24-config synthetic panel (11W-13L, median
    +0.05%) with ~20% faster bagged fits. groups=None draws are byte-identical
    to before.

  • refit_full now defaults to "replay": the same full-data refit for
    about two thirds of the fit time.
    Refitting the early-stopping winner on
    100% of the rows has been on by default since 0.25.0, and fresh attribution
    puts that second, from-scratch fit at 37-49% of every default fit. But
    growing trees is 83-85% of a fit and is a SEARCH, and the refit already
    knows the structures it is rediscovering. "replay" replays the winner's
    splits round by round against gradients computed on all rows and refits only
    the leaf values (and the linear-leaf coefficients), so the held-out rows
    still shape every leaf value without the split search being paid for twice.

    Measured against refit_full=True at 3 seeds, accuracy is a wash on both
    decision suites while fit time falls sharply: Grinsztajn (59 datasets)
    27W-32L, mean +0.005%, median -0.005%, fit time -34.8% and faster on 58 of
    59
    ; high-cardinality (14 datasets) 3W-6L-5T, mean -0.017%, median +0.000%,
    fit time -15.2%.

    refit_full=True still selects the from-scratch refit. Scalar boosters
    only: multiclass grows one vector-leaf tree per round through a separate
    loop and keeps the from-scratch refit, so "replay" is an exact no-op
    there. quality=1/2 already disable refitting, and quality=4/5
    are unaffected because refit_full is a no-op inside bagged members — so
    the only rung this moves is 3, the default. Like refit_full=True it
    does nothing with an explicit eval_set, early_stopping=False, or
    loss="Quantile". Evidence: benchmarks/REPLAY_PLAN.md.

Docs

  • README and user docs rewritten for readability. README restructured into
    Install / Quickstart / What it is / Documentation / Why / Citations, with the
    TabArena chart captioned in words. Across the docs, benchmark-report content
    (win-loss records, suite names, seed counts) and version-history asides give
    way to direct guidance, since that history lives here. parameters.md cells
    are shortened and carry scikit-learn style "See the User Guide" links into
    recipes.md; the estimator docstrings gained the matching "Read more in the
    User Guide" line.
  • API reference split into one page per public name (an api/index.md
    overview plus a page each for the three estimators, CustomObjective,
    quantile_metrics and warmup), replacing a single page that had grown
    to 197 KB. The overview keeps the old /api/ URL. navigation.indexes is
    on so the section header links to it.
  • docs/benchmarks.md is now in the site nav, under Reference. It was
    reachable only through a GitHub link from the README.
  • New "Cross features" section in concepts.md explaining why an oblivious
    tree needs an x1 - x2 column to express a comparison between two features.
    That rationale previously existed only inside a parameters table cell.