-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics labelrankingaverageprecision score
Development build. This page describes
main, not a released package. The latest published Lodestar.Metrics is 0.3.0 — read its documentation.
The mean, over relevant labels, of how much of the ranking above them is relevant —
sklearn.metrics.label_ranking_average_precision_score.
public static double Score(ReadOnlySpan<bool> yTrue, ReadOnlySpan<double> yScore, int labelCount, ReadOnlySpan<double> sampleWeight = default)Parameters — yTrue says whether each label is relevant and yScore holds the scores the
ranking was made from, both row-major: one row per sample, labelCount values each, and the same
length. labelCount is how many labels a row holds; unlike the other two metrics of this family, a
labelCount of 1 is accepted. sampleWeight is one weight per sample, or empty — the default —
for an unweighted mean.
Returns — double in [0, 1] for non-negative weights. 1 when every relevant label outranks
every irrelevant one in every sample, and 1 as well for a sample where all labels or no labels are
relevant — such a ranking carries no information, and the reference scores it perfect rather than
dropping it from the average.
The answer is NaN when sampleWeight sums to zero, where
CoverageError.Score and
LabelRankingLoss.Score throw on the same input. The reference divides
by the weight sum directly on this path instead of going through numpy.average, which is the only
one of the three that refuses a zero sum.
Exceptions — ArgumentException in four shapes, each of them a refusal the reference also
makes: labelCount below 1; yTrue and yScore disagreeing in length; yTrue empty, or not a
whole number of rows of labelCount; and a non-empty sampleWeight whose length is not the row
count. A labelCount of exactly 1 is not refused here, and is refused by the other two.
Example — two samples over three labels. The first sample's relevant label ranks second of three, the second sample's ranks last.
using Lodestar.Metrics;
bool[] truth = [true, false, false, false, false, true];
double[] scores = [0.75, 0.5, 1.0, 1.0, 0.2, 0.1];
double lrap = LabelRankingAveragePrecision.Score(truth, scores, labelCount: 3); // => 0.4166…Half of the first row's ranking above its relevant label is relevant and a third of the second row's
is, so the mean is 0.4166…. And the single label column the other two metrics refuse:
using Lodestar.Metrics;
double single = LabelRankingAveragePrecision.Score([true], [0.7], labelCount: 1); // => 1Remarks — the rank of a label is rankdata(-y_score, "max"): 1 is the best score, and every
member of a tied group takes the group's worst rank. The order within a tie is never observed —
ranks are computed by counting how many scores are at least as high — so no permutation of equally
scored labels can change the answer, at any width.
A negative weight is accepted, as numpy.average accepts one, and takes the result outside [0, 1]
— measured, weights [-1, 2] on the worked example give -0.3333…. Report this beside
CoverageError.Score: a good average precision with a large coverage
means most rankings are clean and a few rows hide a relevant label at the bottom.
Applies to — net10.0, netstandard2.0.
See also — CoverageError.Score,
LabelRankingLoss.Score, the
Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels