-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics f1 score
Development build. This page describes
main, not a released package. The latest published Lodestar.Metrics is 0.3.0 — read its documentation.
The harmonic mean of precision and recall, reduced to one number by average.
public static double Score(ConfusionMatrix cm, Averaging average = Averaging.Binary, int posLabel = 1, ZeroDivision zeroDivision = ZeroDivision.Zero)
public static double Score(ReadOnlySpan<int> yTrue, ReadOnlySpan<int> yPred, Averaging average = Averaging.Binary, int posLabel = 1, ZeroDivision zeroDivision = ZeroDivision.Zero, ReadOnlySpan<int> labels = default, ReadOnlySpan<double> sampleWeight = default)Parameters — cm is a matrix already counted, or pass yTrue and yPred. average is how
the
per-class scores are reduced, Averaging.Binary by default. posLabel is the class reported
under
Averaging.Binary, 1 by default. zeroDivision decides what an undefined score becomes.
labels
fixes the label set and its order, and sampleWeight weights the samples.
Returns — double in [0, 1], larger meaning better.
Exceptions — ArgumentNullException when cm is null; ArgumentException when
Averaging.Binary is used on more than two classes, or posLabel does not occur;
UndefinedMetricException when the metric is undefined and zeroDivision is
ZeroDivision.Throw.
Example — the spam filter, whose precision is 0.6666… and whose recall is 0.5.
using Lodestar.Metrics;
int[] yTrue = [1, 1, 1, 1, 0, 0, 0, 0];
int[] yPred = [1, 1, 0, 0, 1, 0, 0, 0];
double f1 = F1.Score(yTrue, yPred); // => 0.5714…Remarks — this is the number to report when you have one class you care about, both kinds of
mistake matter, and you do not want to argue about which matters more. Being a harmonic mean is
the whole design: it sits much closer to the smaller of the two than an ordinary average would, so
a
model with precision 1.0 and recall 0.01 scores 0.0198, not 0.505. You cannot buy an F1 by
being perfect at one thing.
The trap is that F1 is not symmetric in the way people assume it is. It ignores TN completely,
so
it is not invariant under swapping which class you call positive: on the same predictions,
posLabel: 0 gives 0.6666… here where posLabel: 1 gives 0.5714…. Fix which class is
positive
before you compare two models, and say so in the report.
Averaging.Binary is the default and throws on more than two classes rather than guessing, so a
multiclass call has to name Macro, Weighted or Micro. For a beta other than 1 — recall worth
more than precision, or less — use FBeta.Score rather than post-processing this.
Applies to — net10.0, netstandard2.0.
See also — F1.PerClass, FBeta.Score, Precision.Score, Recall.Score,
the Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels