Skip to content

Metrics reciprocalrank score

github-actions[bot] edited this page Aug 26, 2026 · 26 revisions

Development build. This page describes main, not a released package. The latest published Lodestar.Metrics is 0.3.0 — read its documentation.

ReciprocalRank.Score

The mean of 1 / rank over the queries, where rank is the position of the first relevant document.

Not verified against a reference. There is no reciprocal function in sklearn.metrics to freeze a corpus from, so this member's definition is pinned by tests rather than by an oracle — decision 0036 is the rule that admits it, and says what would retire the exception.

public static double Score(ReadOnlySpan<double> relevance, ReadOnlySpan<double> yScore, int labelCount)

Parametersrelevance says whether each document is relevant and yScore holds the scores the ranking was made from, both row-major: one row per query, labelCount values each, and the same length. Relevance is read as a judgement, not a magnitude: any non-zero value is relevant, and 3 counts no more than 1. labelCount is how many documents each row holds.

Returnsdouble in [0, 1]. 1 when every query puts a relevant document first, 0 when no query retrieves one at all.

ExceptionsArgumentException when labelCount is below 2, when relevance and yScore disagree in length, or when the length is not a whole number of rows of labelCount.

Example — two queries, the first relevant document second and then first.

using Lodestar.Metrics;

double[] relevance = [0, 1, 0, 0, 1, 0, 0, 0];
double[] scores = [0.9, 0.5, 0.4, 0.1, 0.9, 0.5, 0.4, 0.1];

double mrr = ReciprocalRank.Score(relevance, scores, labelCount: 4);  // => 0.75

Remarks — the definition, in the three clauses the tests pin one by one: the reciprocal of the rank of the first relevant document, averaged over queries, with a query holding no relevant document contributing 0 rather than being dropped from the average. That last clause is the one implementations disagree about — dropping such queries raises the score and makes two runs over different query sets incomparable.

Everything after the first relevant document is invisible to this number, which is what makes it the wrong metric when the reader consumes the whole list. Report it beside Ndcg.Score, not instead of one.

Applies to — net10.0, netstandard2.0.

See alsoNdcg.Score, Dcg.Score, the Python equivalence table.

Lodestar

Project

Clone this wiki locally