Skip to content

Metrics regression

github-actions[bot] edited this page Aug 22, 2026 · 29 revisions

Regression metrics — Lodestar.Metrics

Your model predicted a number and the truth was a different number. How wrong is that? Every type on this page answers it, and they disagree about what "wrong" should cost. Squaring the miss makes one bad prediction dominate; taking the median makes it disappear entirely; dividing by the truth makes being 1 out on 10 as bad as being 100 out on 1000. None of these is the correct answer — the right one is the one that matches what a mistake actually costs you — and reporting the wrong one is the usual reason a model that looks good on a metric behaves badly in use.

Two families sit here, and it is worth knowing which one you are reading.

  • An error — everything with Error or Loss in its name — is 0 when the prediction is perfect and grows without bound. It is in the units of your target, or their square, so 0.5 means nothing until you know what the target was measured in.
  • A scoreR2 and ExplainedVariance — is 1 when the prediction is perfect, 0 for a model no better than always predicting the mean, and negative for one that is worse than that. It is unitless, so it can be compared across problems, which is precisely what an error cannot do.

How the parameters fit together

Everything here except MaxError shares one shape, and reading it once saves reading it sixteen times. The three deviances read it with one parameter more — a power that moves the domain of validity as well as the formula — and TweedieDeviance has that table.

flowchart TD
    A["<b>yTrue</b>, <b>yPred</b> — one flat span each.<br/>With more than one output they are row-major:<br/>sample 0's outputs, then sample 1's."] --> B["<b>outputCount</b> says where one row ends"]
    B --> C["per sample and per output, a residual,<br/>charged the way this metric charges it"]
    C --> D["<b>sampleWeight</b> weights the <i>rows</i><br/>— one weight per sample, not per value"]
    D --> E["one number per output"]
    E --> F["<b>PerOutput</b><br/>returns that array as it is"]
    E --> G["<b>Score</b><br/>reduces it: a plain mean, or a mean<br/>weighted by <b>outputWeights</b>"]
    E --> H["<b>VarianceWeighted</b><br/>reduces it by each output's own variance<br/>— <i>R2 and ExplainedVariance only</i>"]
Loading

outputCount defaults to 1, which is the ordinary case: one target per sample, one number out. There is no two-dimensional overload because a ReadOnlySpan<T> cannot carry one, and PerOutput is a method rather than an enum member because it changes the return type — decision 0021.

Two refusals every metric here shares, both reproducing the message their Python layer prints. A sampleWeight that is zero throughout is refused — the rule is every weight, not the sum, so [-1, -2, -3] still scores — and outputWeights whose sum is zero are refused, so [1, -1] is refused and [-1, -1] scores. Both arrive as ArgumentException. Non-finite values in yTrue or yPred are refused too.

Classification metrics — how often a label was right — are on the classification page, not here. ZeroDivision, which R2 takes, is documented there.

Which one do I report?

flowchart TD
    A["What are you reporting?"] --> B{"An error in the target's units,<br/>or a unitless score?"}
    B -->|a score, to compare across problems| C{"Should a constant bias<br/>count against the model?"}
    C -->|yes, it is a real error| D["R2"]
    C -->|no, only the spread matters| E["ExplainedVariance"]
    B -->|an error| F{"How should one very<br/>bad prediction count?"}
    F -->|more than its share| G{"In the target's units?"}
    G -->|yes| H["RootMeanSquaredError"]
    G -->|no, squared is fine| I["MeanSquaredError"]
    F -->|exactly its share| J["MeanAbsoluteError"]
    F -->|not at all| K["MedianAbsoluteError"]
    F -->|it is the only thing that matters| L["MaxError"]
    A --> M{"Is the target a count or a<br/>quantity spanning orders of magnitude?"}
    M -->|yes, and relative error is what hurts| N["MeanAbsolutePercentageError"]
    M -->|yes, and under-prediction hurts more| O["MeanSquaredLogError,<br/>RootMeanSquaredLogError"]
    A --> P{"Are you predicting a quantile<br/>rather than a mean?"}
    P -->|yes| Q["PinballLoss"]
Loading
Type What it measures
ExplainedVariance The share of the truth's spread the prediction accounts for, ignoring a constant bias.
MaxError The single worst prediction, and nothing else.
MeanAbsoluteError The average miss, in the target's own units.
MeanAbsolutePercentageError The average miss as a fraction of the truth.
MeanSquaredError The average squared miss — the one big errors dominate.
MeanSquaredLogError The same, on log(1 + y), so a ratio matters more than a difference.
MedianAbsoluteError The typical miss, immune to any number of outliers.
PinballLoss The loss for a quantile prediction, charging over- and under-shooting differently.
R2 How much better than always predicting the mean, as a unitless score.
RootMeanSquaredError MeanSquaredError back in the target's units.
RootMeanSquaredLogError MeanSquaredLogError back in log units.
TweedieDeviance The deviance a squared error becomes when the target is not normally distributed.
PoissonDeviance TweedieDeviance at power 1 — the deviance for a count.
GammaDeviance TweedieDeviance at power 2 — scale-invariant, for a positive quantity.
D2Tweedie What share of a deviance the model explains, against a constant baseline.
D2Pinball The same for a quantile prediction, against the best constant quantile.
D2AbsoluteError D2Pinball at the median — R2's question with an outlier-proof baseline.

Lodestar

Project

Clone this wiki locally