-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics mutualinformation score
Development build. This page describes
main, not a released package. The latest published Lodestar.Metrics is 0.3.0 — read its documentation.
The information two labellings share, in nats.
public static double Score(ReadOnlySpan<int> labelsTrue, ReadOnlySpan<int> labelsPred)Parameters — labelsTrue is the reference labelling and labelsPred the one being scored,
one label per sample and the same length. The label values carry no meaning: only which samples
share one does.
Returns — double in nats, never negative and not bounded above.
Exceptions — ArgumentException when the two labellings disagree in length. An empty input
is not an error: it returns 0, unlike scikit-learn — see the divergence below.
Example — every sample in its own cluster still shares one bit of information with a two-cluster truth.
using Lodestar.Metrics;
int[] truth = [0, 0, 1, 1];
int[] alone = [0, 1, 2, 3];
double shared = MutualInformation.Score(truth, alone); // => 0.6931471805599452
double scaled = NormalizedMutualInformation.Score(truth, alone); // => 0.6666666666666666Remarks — 0.693… is ln 2, exactly one bit — the unit is nats, natural logarithms, not
bits, because that is what scikit-learn uses. This is the raw form of
NormalizedMutualInformation.Score, which divides the
same quantity by the mean entropy to land in [0, 1]; that page's own remark has why 0.667 there
is not the same statement as 0.693 here.
Unbounded above is the practical difference from every other metric in this namespace. Two
scores are comparable only between labellings of the same data at the same sizes — splitting a
labelling into more clusters can only raise this number, never lower it, so it cannot rank
clusterings of different sizes fairly. Reach for
NormalizedMutualInformation.Score or
AdjustedMutualInformation.Score across datasets instead.
One divergence from scikit-learn, deliberate. mutual_info_score([], []) raises
ValueError in scikit-learn 1.9.0 — a log(0) inside an unguarded logarithm, not a documented
refusal. This method returns 0 there instead, matching every sibling clustering metric, which
all treat an empty input as a case rather than an error.
Decision 0039 has
the measurement and the reasoning. A single sample also returns 0 — unlike the six agreement
metrics this family started with, which score 1 there, because a single sample carries no
information to share rather than no disagreement to find.
Applies to — net10.0, netstandard2.0.
See also — NormalizedMutualInformation.Score,
AdjustedMutualInformation.Score,
the Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels