-
Notifications
You must be signed in to change notification settings - Fork 0
Metrics ndcg score
Development build. This page describes
main, not a released package. The latest published Lodestar.Metrics is 0.3.0 — read its documentation.
The mean normalized discounted gain over the rows, in [0, 1] — sklearn.metrics.ndcg_score.
public static double Score(ReadOnlySpan<double> yTrue, ReadOnlySpan<double> yScore, int labelCount, int? k = null, bool ignoreTies = false, ReadOnlySpan<double> sampleWeight = default)Parameters — yTrue is the relevance of each document and yScore the scores the ranking was
made from, both row-major: one row per query, labelCount values each, and the same length.
k scores only the first k positions, or null for all of them; a k past labelCount scores
the whole row rather than raising. ignoreTies ranks equal scores in descending index order instead
of averaging over their permutations. sampleWeight carries one weight per query, or is empty for
an unweighted mean; over a single query it cancels, since it multiplies both halves of the mean.
There is no logBase, because ndcg_score has none: the discount cancels in the ratio only when
both halves share a base, and scikit-learn shares base 2. Pass one to
Dcg.Score instead, where it changes the answer.
Returns — double in [0, 1]. 1 when the ranking is as good as its relevance allows, and 0
when no document in the row is relevant — there is no ideal to divide by, and the answer is 0
rather than a division by zero.
The [0, 1] holds for every weight vector but one: a negative sampleWeight takes the mean
outside it, which the reference does too rather than refusing — frozen in ranking_weighted.json,
-0.7039180890341348 at k = 2 on weights [-1, 2]. Dcg.Score is unbounded
above and so has nothing to lose here.
Exceptions — ArgumentException when labelCount is below 2 (scikit-learn's own sentence,
"Computing NDCG is only meaningful when there is more than 1 document."), when sampleWeight is
neither empty nor one value per query, when it sums to zero — numpy.average's own refusal — when
yTrue and yScore
disagree in length, when the length is not a whole number of rows of labelCount, or when any
relevance is negative — "ndcg_score should not be used on negative y_true values.", which is
scikit-learn's refusal and the reason the [0, 1] above holds. ArgumentOutOfRangeException when
k is below 1; scikit-learn refuses the same value.
Example — the same relevance ranked perfectly, then backwards.
using Lodestar.Metrics;
double[] relevance = [3, 2, 1, 0];
double[] best = [0.9, 0.5, 0.4, 0.1];
double[] worst = [0.1, 0.4, 0.5, 0.9];
double perfect = Ndcg.Score(relevance, best, labelCount: 4); // => 1
double backwards = Ndcg.Score(relevance, worst, labelCount: 4); // => 0.6138…Remarks — the worst possible ordering scores 0.6138…, not 0. That is not a bug: the
logarithmic discount is shallow, so even a reversed list collects most of the ideal gain, and the
floor of this metric on a row with several relevant documents is well above zero. Read NDCG as a
comparison between rankings of the same rows, never as a fraction of "how much better than random".
The ideal is computed without tie averaging, as scikit-learn does — ranking a row by its own relevance leaves ties only between equal gains, which no ordering can separate.
Applies to — net10.0, netstandard2.0.
See also — Dcg.Score, ReciprocalRank.Score, the
Python equivalence table.
- 0001-target-framework
- 0002-unicode-comparison-unit
- 0003-provenance-and-licensing
- 0004-levenshtein-myers-backlog
- 0005-hamming-jellyfish-divergence
- 0006-ratcliff-autojunk
- 0007-metaphone-scope
- 0008-italian-enza-nltk-divergence
- 0009-sample-consumes-a-local-feed
- 0010-stop-word-list-provenance
- 0011-persistence-format
- 0012-per-package-versioning
- 0013-sentencepiece-parity-scope
- 0014-precompiled-normalizer
- 0015-sonar-rules-in-the-build
- 0016-metrics-package-placement
- 0017-bpe-parity-scope
- 0018-multiclass-roc-auc-parallelism-is-opt-in
- 0019-the-net-analysers-run-in-the-build-too
- 0020-normalize-is-a-projection-not-a-parameter
- 0021-multioutput-is-a-method-not-an-enum
- 0022-added-token-matching-flags
- 0023-byte-level-decode-substitutes
- 0024-weighted-median-averages-within-scikit-learns-epsilon
- 0025-quickselect-replaces-a-full-sort-for-the-median
- 0026-r2-and-explainedvariance-split-their-undefined-cases-differently
- 0027-r2-and-explainedvariance-vectorize-only-a-single-output
- 0028-log1p-is-kahans-identity-not-math-log-1-plus-x
- 0029-balanced-accuracy-adjusted-is-left-to-ieee-754-at-the-edge
- 0030-cohen-kappa-keeps-scikit-learns-expected-matrix-orientation
- 0031-nosamplecorrect-mirrors-numpys-float64-upcast
- 0032-fbeta-substitutes-tp-predicted-and-support-algebraically
- 0033-compensated-sum-is-neumaiers-variant
- 0034-dropout-is-refused-for-want-of-a-user
- 0035-a-null-pre-split-is-removed-with-invert-not-isolated
- 0036-a-member-may-ship-without-an-oracle-if-it-says-so
- 0037-the-guards-run-before-the-commit
- 0038-the-gate-confronts-an-exception-tag-with-the-page-that-documents-it
- 0039-mutual-information-returns-zero-on-an-empty-input
- 0040-a-curve-is-a-sealed-class-per-curve
- 0041-one-sample-file-per-public-class
- 0042-phonetic-encoders-refuse-a-null-word
- 0043-the-equality-table-is-sized-to-the-pattern
- 0044-compression-belongs-to-the-caller
- 0045-a-console-call-carries-its-reason-on-the-line
- 0046-check-adr-immutable-runs-in-ci-only
- 0047-one-gate-per-kernel-not-one-per-alphabet
- 0048-the-gate-depends-on-the-kernel-and-the-alphabet
- 0049-two-gates-per-kernel-tested-where-the-width-is-known
- 0050-the-sentencepiece-bpe-lineage-stays-a-bpe-model
- benchmark_latest
- decisions
- equivalence
- matplotlib
- migration
- nightly_run
- numpy
- pandas
- performance
- pytorch
- seaborn
- sklearn
- statsmodels