-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics clustering
You clustered some samples. Most of the types on this page score the result against a reference partition — labels from a human, from an earlier model, or from a dataset that came with them — and none of them cares what the clusters are called: swapping the names of two clusters changes nothing, which is exactly what separates these from the classification metrics.
Three of them need no reference at all, and score a clustering from the samples themselves. They take a feature matrix rather than two label vectors, and they are what you reach for when there is nothing to compare against — choosing how many clusters to ask for, for instance. They have their own section below.
The reference-partition metrics disagree on what "agree" means, and the disagreement is the reason there are several.
-
Corrected for chance or not. Split every sample into a cluster of its own and
Homogeneityscores a perfect1, because every cluster does hold one class.AdjustedRandscores0on the same input, because that is what random labelling achieves. When a clustering looks suspiciously good, this is the pair to read together. -
Symmetric or not.
HomogeneityandCompletenessare the same measurement with the two labellings exchanged, and they pull in opposite directions: splitting raises one and merging raises the other.VMeasureis their harmonic mean, for when you want one number.
The degenerate cases answer surprisingly, and it is deliberate. An empty input scores 1 on
every metric here — agreeing about nothing is agreeing — and so does a single sample. Two
independent partitions of four samples score -0.5 on AdjustedRand, not 0: the correction for
chance is a subtraction, and it can go below zero. Every one of those numbers is scikit-learn's,
measured against 1.9.0 and frozen in the oracle corpus rather than reasoned about.
flowchart TD
A["You clustered some samples"] --> B{"Is there a reference<br/>partition to score against?"}
B -->|"no, only the samples"| C{"Features, or a distance<br/>matrix you computed?"}
C -->|"a distance matrix"| C1["Silhouette<br/>the only one that takes one"]
C -->|"a feature matrix"| C2["Silhouette, CalinskiHarabasz<br/>higher is better<br/>DaviesBouldin — lower is better"]
B -->|yes| D{"What do you want to learn?"}
D -->|"one number, and the two<br/>cluster counts differ"| E["AdjustedMutualInformation"]
D -->|"one number, comparable<br/>cluster counts"| F{"Corrected for chance?"}
F -->|yes| F1["AdjustedRand"]
F -->|"no — and read it<br/>knowing that"| F2["RandIndex, FowlkesMallows,<br/>MutualInformation"]
D -->|"which way it fails:<br/>split, or merged"| G["Homogeneity — one class per cluster<br/>Completeness — one cluster per class<br/>VMeasure — their harmonic mean"]
D -->|"shared information,<br/>scaled into 0..1"| I["NormalizedMutualInformation"]
D -->|"the pair counts<br/>underneath the others"| H["PairConfusionMatrix"]
Corrected for chance is the branch to get right, and it is the one above that costs most when
missed: the uncorrected three are easy to reach for by name and rarely what is wanted. The
paragraph above has the worked case — a clustering that scores a perfect 1 on Homogeneity and
0 on AdjustedRand for the same input.
| Type | What it measures |
|---|---|
AdjustedRand |
How many pairs of samples the two partitions agree about, minus what chance would give. |
RandIndex |
The same pair count as AdjustedRand, uncorrected for chance. |
MutualInformation |
Shared information between the two labellings, unnormalised and in nats. |
PairConfusionMatrix |
The four pair counts AdjustedRand and RandIndex are both built from. |
AdjustedMutualInformation |
Shared information between the two labellings, minus what chance would give — the one to use across different cluster counts. |
FowlkesMallows |
The geometric mean of pair precision and pair recall, uncorrected for chance. |
Completeness |
Whether every sample of one class landed in the same cluster. |
Homogeneity |
Whether each cluster holds samples of a single class. |
NormalizedMutualInformation |
How much knowing one labelling tells you about the other, scaled into [0, 1]. |
VMeasure |
Homogeneity and completeness as one number, their harmonic mean. |
These three read the samples, not a second labelling. All of them refuse a label count outside
[2, n - 1] with the same sentence — one cluster leaves nothing to compare against, one cluster per
sample leaves nothing inside one — and none of them ever answers a non-finite value on an input
scikit-learn accepts.
They do not all read in the same direction. A clustering that improves moves
Silhouette and CalinskiHarabasz
up and DaviesBouldin down. Reading a table of the three
without knowing that gets one of them backwards.
Only Silhouette accepts a distance matrix you computed yourself. The
other two read cluster centroids, which a distance matrix does not carry, so a reader arriving from
silhouette_score(metric='precomputed') will look for the equivalent and find none — the reference
has none either.
| Type | What it measures | Direction |
|---|---|---|
Silhouette |
How well each sample sits in its own cluster rather than the nearest other one. | higher is better |
CalinskiHarabasz |
How far the clusters sit from each other against how spread they are inside. | higher is better |
DaviesBouldin |
The worst pairing each cluster is in, averaged. | lower is better |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels