-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics 0.2.0 clustering
Lodestar.Metrics 0.2.0. This page is frozen at that release. Read the current documentation for what
mainsays now. A link to a decision or a migration page followsmain, and leaves the archive.
You clustered some samples and you have a reference partition to compare against — labels from a human, from an earlier model, or from a dataset that came with them. Every type on this page scores how well the two partitions agree, and none of them cares what the clusters are called: swapping the names of two clusters changes nothing, which is exactly what separates these from the classification metrics.
They disagree on what "agree" means, and the disagreement is the reason there are five.
-
Corrected for chance or not. Split every sample into a cluster of its own and
Homogeneityscores a perfect1, because every cluster does hold one class.AdjustedRandscores0on the same input, because that is what random labelling achieves. When a clustering looks suspiciously good, this is the pair to read together. -
Symmetric or not.
HomogeneityandCompletenessare the same measurement with the two labellings exchanged, and they pull in opposite directions: splitting raises one and merging raises the other.VMeasureis their harmonic mean, for when you want one number.
The degenerate cases answer surprisingly, and it is deliberate. An empty input scores 1 on
every metric here — agreeing about nothing is agreeing — and so does a single sample. Two
independent partitions of four samples score -0.5 on AdjustedRand, not 0: the correction for
chance is a subtraction, and it can go below zero. Every one of those numbers is scikit-learn's,
measured against 1.9.0 and frozen in the oracle corpus rather than reasoned about.
| Type | What it measures |
|---|---|
AdjustedRand |
How many pairs of samples the two partitions agree about, minus what chance would give. |
Completeness |
Whether every sample of one class landed in the same cluster. |
Homogeneity |
Whether each cluster holds samples of a single class. |
Silhouette |
How well each sample sits in its own cluster rather than the nearest other one — no reference partition needed. |
NormalizedMutualInformation |
How much knowing one labelling tells you about the other, scaled into [0, 1]. |
VMeasure |
Homogeneity and completeness as one number, their harmonic mean. |
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels