-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics normalizedmutualinformation score
Development build. This page describes
main, not a released package. The latest published Lodestar.Metrics is 0.3.0 — read its documentation.
How much knowing one labelling tells you about the other, scaled into [0, 1].
public static double Score(ReadOnlySpan<int> labelsTrue, ReadOnlySpan<int> labelsPred)Parameters — labelsTrue is the reference partition and labelsPred the one being
scored, one label per sample and the same length. The label values carry no meaning: only which
samples share one does.
Returns — double in [0, 1]. 1 when each labelling determines the other, 0 when neither says
anything about the other.
Exceptions — ArgumentException when the two labellings disagree in length. An empty
input is not an error: it scores 1.
Example — the same two clusterings, scored without the correction for chance.
using Lodestar.Metrics;
int[] truth = [0, 0, 1, 1];
int[] alone = [0, 1, 2, 3];
double score = NormalizedMutualInformation.Score(truth, alone); // => 0.6666…Remarks — the example is the reason this is not the default choice, and the reason is not
that the split clustering says nothing. It says a great deal: every cluster holds one sample, so
knowing the cluster tells you the class exactly, and the mutual information is genuinely high. What
it does not do is beat chance — a partition that fine agrees with any truth about that well by
accident, which is what AdjustedRand.Score subtracts and this does not. Read the two together and
the gap between them is the correction for chance.
The normalizer is the arithmetic mean of the two entropies, scikit-learn's default
average_method; min, geometric and max are absent rather than refused, because the frozen
corpus holds no row for them and an unproven normalizer is not parity. That same choice makes this
the identical number to VMeasure.Score on every input — the cancellation is written out in
that entry, and it is worth reading before reporting both.
Applies to — net10.0, netstandard2.0.
See also — AdjustedRand.Score, VMeasure.Score, the Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels