-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics 0.2.0 regression
Lodestar.Metrics 0.2.0. This page is frozen at that release. Read the current documentation for what
mainsays now. A link to a decision or a migration page followsmain, and leaves the archive.
Your model predicted a number and the truth was a different number. How wrong is that? Every type on this page answers it, and they disagree about what "wrong" should cost. Squaring the miss makes one bad prediction dominate; taking the median makes it disappear entirely; dividing by the truth makes being 1 out on 10 as bad as being 100 out on 1000. None of these is the correct answer — the right one is the one that matches what a mistake actually costs you — and reporting the wrong one is the usual reason a model that looks good on a metric behaves badly in use.
Two families sit here, and it is worth knowing which one you are reading.
- An error — everything with
ErrororLossin its name — is0when the prediction is perfect and grows without bound. It is in the units of your target, or their square, so0.5means nothing until you know what the target was measured in. - A score —
R2andExplainedVariance— is1when the prediction is perfect,0for a model no better than always predicting the mean, and negative for one that is worse than that. It is unitless, so it can be compared across problems, which is precisely what an error cannot do.
Everything here except MaxError shares one shape, and reading it once saves reading it eleven
times.
flowchart TD
A["<b>yTrue</b>, <b>yPred</b> — one flat span each.<br/>With more than one output they are row-major:<br/>sample 0's outputs, then sample 1's."] --> B["<b>outputCount</b> says where one row ends"]
B --> C["per sample and per output, a residual,<br/>charged the way this metric charges it"]
C --> D["<b>sampleWeight</b> weights the <i>rows</i><br/>— one weight per sample, not per value"]
D --> E["one number per output"]
E --> F["<b>PerOutput</b><br/>returns that array as it is"]
E --> G["<b>Score</b><br/>reduces it: a plain mean, or a mean<br/>weighted by <b>outputWeights</b>"]
E --> H["<b>VarianceWeighted</b><br/>reduces it by each output's own variance<br/>— <i>R2 and ExplainedVariance only</i>"]
outputCount defaults to 1, which is the ordinary case: one target per sample, one number out.
There is no two-dimensional overload because a ReadOnlySpan<T> cannot carry one, and PerOutput
is a method rather than an enum member because it changes the return type —
decision 0021.
Two refusals every metric here shares, both reproducing the message their Python layer prints. A
sampleWeight that is zero throughout is refused — the rule is every weight, not the sum, so
[-1, -2, -3] still scores — and outputWeights whose sum is zero are refused, so [1, -1] is
refused and [-1, -1] scores. Both arrive as ArgumentException. Non-finite values in yTrue or
yPred are refused too.
Classification metrics — how often a label was right — are on the
classification page, not here. ZeroDivision, which R2 takes, is documented
there.
flowchart TD
A["What are you reporting?"] --> B{"An error in the target's units,<br/>or a unitless score?"}
B -->|a score, to compare across problems| C{"Should a constant bias<br/>count against the model?"}
C -->|yes, it is a real error| D["R2"]
C -->|no, only the spread matters| E["ExplainedVariance"]
B -->|an error| F{"How should one very<br/>bad prediction count?"}
F -->|more than its share| G{"In the target's units?"}
G -->|yes| H["RootMeanSquaredError"]
G -->|no, squared is fine| I["MeanSquaredError"]
F -->|exactly its share| J["MeanAbsoluteError"]
F -->|not at all| K["MedianAbsoluteError"]
F -->|it is the only thing that matters| L["MaxError"]
A --> M{"Is the target a count or a<br/>quantity spanning orders of magnitude?"}
M -->|yes, and relative error is what hurts| N["MeanAbsolutePercentageError"]
M -->|yes, and under-prediction hurts more| O["MeanSquaredLogError,<br/>RootMeanSquaredLogError"]
A --> P{"Are you predicting a quantile<br/>rather than a mean?"}
P -->|yes| Q["PinballLoss"]
| Type | What it measures |
|---|---|
ExplainedVariance |
The share of the truth's spread the prediction accounts for, ignoring a constant bias. |
MaxError |
The single worst prediction, and nothing else. |
MeanAbsoluteError |
The average miss, in the target's own units. |
MeanAbsolutePercentageError |
The average miss as a fraction of the truth. |
MeanSquaredError |
The average squared miss — the one big errors dominate. |
MeanSquaredLogError |
The same, on log(1 + y), so a ratio matters more than a difference. |
MedianAbsoluteError |
The typical miss, immune to any number of outliers. |
PinballLoss |
The loss for a quantile prediction, charging over- and under-shooting differently. |
R2 |
How much better than always predicting the mean, as a unitless score. |
RootMeanSquaredError |
MeanSquaredError back in the target's units. |
RootMeanSquaredLogError |
MeanSquaredLogError back in log units. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels