-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics r2 score
Development build. This page describes
main, not a released package. The latest published Lodestar.Metrics is 0.3.0 — read its documentation.
One score for the whole prediction: 1 minus the residual variance over the truth's variance.
public static double Score(ReadOnlySpan<double> yTrue, ReadOnlySpan<double> yPred, int outputCount = 1, ReadOnlySpan<double> sampleWeight = default, ReadOnlySpan<double> outputWeights = default, bool forceFinite = true, ZeroDivision zeroDivision = ZeroDivision.NaN)Parameters — yTrue and yPred are the true and predicted values, row-major when there is
more
than one output. outputCount is how many outputs each row holds, sampleWeight weights the
rows,
and outputWeights weights the outputs in the reduction. forceFinite answers the case where the
truth has no variance over two or more samples, clamping to 1 or 0 instead of nan or -inf.
zeroDivision answers the different case of fewer than two samples, and defaults to
ZeroDivision.NaN, which is scikit-learn's value.
Returns — double at most 1: 1 for a perfect prediction, 0 for one exactly as good as
always predicting the mean, and negative — with no lower bound — for one that is worse.
Exceptions — ArgumentException when a length disagrees with the shape, the input is empty,
or
it holds a non-finite value; ArgumentOutOfRangeException when outputCount is below one;
UndefinedMetricException when there are fewer than two samples and zeroDivision is
ZeroDivision.Throw.
Example — four predictions, close but not exact.
using Lodestar.Metrics;
double[] yTrue = [3.0, -0.5, 2.0, 7.0];
double[] yPred = [2.5, 0.0, 2.0, 8.0];
double score = R2.Score(yTrue, yPred); // => 0.9486…Remarks — this is the number to report when a reader has to judge the model without knowing
the
target's units, and the only one on this page that can rank two different problems. Its zero point
is
the thing to hold on to: 0 is what you get by ignoring the inputs entirely and predicting the
mean
of the truth every time. A model that scores 0.1 is barely doing anything; a model that scores
below 0 is worse than that baseline, which is possible and much more common on held-out data
than
people expect. There is no floor — a bad enough model scores -14.
The trap that catches people moving from ExplainedVariance.Score is that R2 charges for a
constant bias. Predicting y + 1 for every sample tracks the truth perfectly and scores -0.5
here, where explained variance scores 1. That is the intended behaviour: an offset is a real
error
unless you are going to remove it.
The two undefined cases are deliberately kept apart and must not be merged.
-
Fewer than two samples is
zeroDivision's case: the truth has no variance to compare against because there is only one of it, and scikit-learn returnsnanhere whateverforce_finitesays.R2.Score([2.0], [1.0])isNaN. -
A constant truth over two or more samples is
forceFinite's case. WithforceFinite: true— the default — a perfect prediction of that constant scores1and any other scores0; withfalseyou getnanand-infinstead.
Decision 0026 has the argument for keeping them separate. Both passes are Neumaier-compensated, which is load-bearing on an ill-conditioned target: a sequential sum was measured 357 times outside the oracle's tolerance — decision 0033.
Applies to — net10.0, netstandard2.0.
See also — R2.PerOutput, R2.VarianceWeighted, ExplainedVariance.Score,
decision
0026,
the ZeroDivision entry,
the Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels