-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics cohenkappa score
Development build. This page describes
main, not a released package. The latest published Lodestar.Metrics is 0.3.0 — read its documentation.
Observed agreement minus expected agreement, scaled so that perfect agreement is 1.
public static double Score(ConfusionMatrix cm, KappaWeighting weighting = KappaWeighting.None, ZeroDivision zeroDivision = ZeroDivision.NaN)
public static double Score(ReadOnlySpan<int> yTrue, ReadOnlySpan<int> yPred, KappaWeighting weighting = KappaWeighting.None, ZeroDivision zeroDivision = ZeroDivision.NaN, ReadOnlySpan<int> labels = default, ReadOnlySpan<double> sampleWeight = default)Parameters — cm is a matrix already counted, or pass yTrue and yPred. weighting says
how
far apart two different classes count as being — KappaWeighting.None by default, which charges
every disagreement the same. zeroDivision decides the answer when the expected agreement
collapses,
and defaults to ZeroDivision.NaN, which is scikit-learn's value here rather than the Zero the
precision family defaults to. labels fixes the label set and its order, and sampleWeight
weights
the samples.
Returns — double at most 1: 1 for total agreement, 0 for agreement no better than
chance, and negative for agreement worse than chance.
Exceptions — ArgumentNullException when cm is null; ArgumentOutOfRangeException when
weighting is not one of the three defined values; ArgumentException when the label spans
disagree in length or are empty; UndefinedMetricException when the expected agreement collapses
and zeroDivision is ZeroDivision.Throw.
Example — a model scored against a human rater on a three-point scale, with the disagreements charged by how far apart the two ratings were.
using Lodestar.Metrics;
int[] rater = [1, 1, 2, 2, 3, 3, 1, 3];
int[] model = [1, 3, 2, 1, 3, 2, 1, 3];
double flat = CohenKappa.Score(rater, model); // => 0.4285…
double linear = CohenKappa.Score(rater, model, KappaWeighting.Linear); // => 0.4666…
double quadratic = CohenKappa.Score(rater, model, KappaWeighting.Quadratic); // => 0.5Remarks — kappa is the metric for "two annotators, how much do they really agree" and, by
extension, for a model scored against a human. Its whole point is the subtraction: two raters who
both say "no" 95% of the time agree 90% of the time by accident, and accuracy will report that as
0.9 while kappa reports something near 0. Use it when the class distribution is skewed enough
that plain agreement flatters everyone.
weighting is what makes it usable on an ordinal scale — a five-point severity, a star rating
—
where confusing 1 with 2 is a smaller error than confusing 1 with 5. Linear charges the distance
in positions, Quadratic its square, so quadratic weighting forgives near misses much more than
it
forgives distant ones. Above, the same predictions score 0.4285… flat and 0.5 quadratic,
because
most of the disagreements are one step wide.
The trap is that distance is measured between positions in the label order, not between label
values, so any weighting other than None depends on the order of labels. Reorder the same
three
labels as [3, 1, 2] and the quadratic score above becomes 0.3846…; the unweighted score does
not
move at all. If your labels are ordinal, pass labels in the ordinal order every time, and never
let it default to the sorted union without checking that sorted is the ordinal order. The
reasoning, and the expected-matrix orientation this keeps from scikit-learn, are in
decision
0030.
The parameter is named weighting and not scikit-learn's weights because sampleWeight sits in
the same signature and the two are unrelated senses of the word.
Applies to — net10.0, netstandard2.0.
See also — KappaWeighting, MatthewsCorrelation.Score, BalancedAccuracy.Score,
decision
0030,
the Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels