-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics accuracy score
The share of samples whose predicted label equals the true one.
public static double Score(ReadOnlySpan<int> yTrue, ReadOnlySpan<int> yPred, bool normalize = true, ReadOnlySpan<double> sampleWeight = default)
public static double Score(ConfusionMatrix cm, bool normalize = true)Parameters — yTrue and yPred are the true and predicted labels, one per sample and the
same
length. cm is the alternative to both: a matrix already counted, which is what you pass when
several metrics are being read off one dataset. normalize chooses between the fraction (true,
the default) and the raw weight of the correct samples (false). sampleWeight gives each sample
its own weight; omit it and every sample counts 1.
Returns — double in [0, 1] when normalize is true, 1 meaning every sample was right.
With normalize: false it is a count instead — a weight, not a fraction, and unbounded.
Exceptions — ArgumentException when the two label spans disagree in length or are empty;
ArgumentNullException when cm is null.
Example — four spam messages and four legitimate ones; the filter caught two of the four and raised one false alarm.
using Lodestar.Metrics;
int[] yTrue = [1, 1, 1, 1, 0, 0, 0, 0];
int[] yPred = [1, 1, 0, 0, 1, 0, 0, 0];
double share = Accuracy.Score(yTrue, yPred); // => 0.625Remarks — this is the right metric when the classes are roughly balanced and every mistake costs about the same. Both conditions matter, and the first one is where people get hurt.
The trap has a number attached. Take ten samples of which two belong to the class you care about,
predict the majority class for all ten, and this returns 0.8 while
BalancedAccuracy.Score returns 0.5 — the score a coin gets. Accuracy is a weighted average of
the per-class recalls in which each class is weighted by how common it is, so a class that is 2%
of
the data moves it by at most 0.02. If your positive class is rare, this number is measuring the
negative class and telling you about it.
The ConfusionMatrix overload has one behaviour of its own worth knowing. It is accuracy over the
samples the matrix kept: a matrix built with an explicit labels subset drops every sample
whose true or predicted label falls outside that subset, so on a three-class problem restricted to
two labels this can read 0.75 where the same data scored over every sample reads 0.7142…. That
is not a bug in either number — they are answers to different questions — but only the span
overload matches accuracy_score.
Applies to — net10.0, netstandard2.0.
See also — BalancedAccuracy.Score, ConfusionMatrix.Compute,
ClassificationReport.Compute,
the Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels