-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics balancedaccuracy score
The mean recall over the classes that have at least one true sample.
public static double Score(ConfusionMatrix cm, bool adjusted = false)
public static double Score(ReadOnlySpan<int> yTrue, ReadOnlySpan<int> yPred, bool adjusted = false, ReadOnlySpan<int> labels = default, ReadOnlySpan<double> sampleWeight = default)Parameters — cm is a matrix already counted, or pass yTrue and yPred and let it be
counted
here. adjusted rescales the result so that chance scores 0 instead of 1/k. labels fixes
the
label set and its order; omit it for the sorted union of both inputs. sampleWeight gives each
sample its own weight.
Returns — double in [0, 1] normally, 1 meaning every class was recalled perfectly. With
adjusted: true the range becomes [-1/(k-1) … 1], so a below-chance model returns a negative
number.
Exceptions — ArgumentException when the label spans disagree in length or are empty;
ArgumentNullException when cm is null.
Example — ten samples, two of them positive, and a model that predicts the majority class
every
time. Accuracy.Score on this data is 0.8.
using Lodestar.Metrics;
int[] yTrue = [0, 0, 0, 0, 0, 0, 0, 0, 1, 1];
int[] yPred = [0, 0, 0, 0, 0, 0, 0, 0, 0, 0];
double balanced = BalancedAccuracy.Score(yTrue, yPred); // => 0.5Remarks — reach for this the moment the classes are unbalanced and you still want one number.
It
is the honest version of accuracy for that case: a model that ignores the minority class cannot
get
above 0.5 on two classes however large the majority is, because each class contributes its own
recall and nothing else. On more than two classes it is exactly macro-averaged recall, so
BalancedAccuracy.Score and Recall.Score(…, Averaging.Macro) are two names for one number when
no
label subset is in play.
adjusted: true answers a different complaint: that 0.5 on two classes and 0.333… on three
both
mean "no better than guessing", and cannot be compared. Adjusting maps chance to 0 and perfect
to
1 whatever k is, at the price of a range that now goes negative.
Two traps, and the second is subtle. The average runs over the classes that appear in the
truth,
not the classes you asked for — a class named in labels with no true sample is dropped rather
than
scored 0, which is scikit-learn's behaviour and means the divisor is not always labels.Length.
And when only one class survives that filter, adjusted: true divides by 1 - 1/1, so the result
is NaN or -∞ rather than a number; that is left to IEEE 754 on purpose, and the reasoning is
in
decision
0029.
The ConfusionMatrix overload divides each recall by its own row sum in the Labels-sized view,
where Recall.Score divides by scikit-learn's true_sum over every observed label. The two agree
whenever nothing was dropped, and part company on a matrix built with an explicit label subset.
Applies to — net10.0, netstandard2.0.
See also — Accuracy.Score, Recall.Score, CohenKappa.Score,
decision
0029,
the Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels