Skip to content

2.22.2

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 03 Oct 13:44
· 13 commits to main since this release

2.22.2 (2026-10-03)

Chore

  • chore: update the citation cache [skip ci] (87f5bcd)

Fix

  • fix: point NeuCLIR2023Retrieval to the 2023 dataset (#5570)

NeuCLIR2023Retrieval loaded mteb/NeuCLIR2022Retrieval at the same
revision as NeuCLIR2022Retrieval, so it ran on the 2022 queries and
qrels. Point it to mteb/NeuCLIR2023Retrieval, whose queries and qrels
match the original mteb/neuclir-2023.

Regenerate its descriptive statistics and list it under
KNOWN_ISSUES["zero_relevant_docs"]: 4 queries only have score-0
judgments, as in the original qrels. (9572cf4)

Test

  • test: check final scores in model-task integrations (#5321)

  • test: check final model-task scores

  • test: isolate model-task score cases

  • test: account for media codec score baselines

  • test: check multimodal pair classification scores

  • test: account for pair classification codecs

  • test: check scores in library integrations

  • test: account for dataset score environments

  • test: structure score baselines by model

  • test: remove redundant model parametrization

  • simplify modelinfo

  • remove comment

  • update after merge

  • add future

  • print actual and expected scores

  • fix test

  • upd prescision


Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (33a8aa9)

Unknown

  • Regenerate uv lock (#5571)

upd (5a29ea9)

  • model: add ColNanoVDR late-interaction query towers (ColVec1.1-4b/8b) (#5518)

  • model: add ColNanoVDR late-interaction query towers (ColVec1.1-4b/8b)

Two asymmetric late-interaction retrievers: a 150M text-only multi-vector
student encodes queries, the frozen ColPali-style teacher it was distilled
from encodes page images, and scoring is MaxSim.

The student needs sentence-transformers>=6 for MultiVectorEncoder, so this
adds a colnanovdr requirement group; heavy imports stay inside functions.

Also corrects training_datasets for nanovdr/NanoVDR-S-Multi, which was
trained on the same data: TAT-DQA was missing and TabFQuAD listed instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

  • Fill in parameter counts for the ColNanoVDR models

n_embedding_parameters is the student's input embedding matrix
(50368 x 768); n_parameters is the exact count of the packaged model
rather than a rounded figure.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

  • Encode ColNanoVDR documents with the teacher's own MTEB entry

Documents now go through the registered ColVec1.1 wrapper, so the teacher's
pinned revision, processor settings and handling of text-only and image
documents are exactly those of its own entry. The teacher is built in
init.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

  • Regenerate the mock run with the new document path

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

  • Delete mteb_mock_run_results.md

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (2852c40)

  • model: add nvidia/llama-nemotron-rerank-vl-1b-v2 (#5565) (c9f9329)

  • Leaderboard integation with experiments (#4900)

  • recreate experiments

  • update after frontend testing

  • parse model meta from experiments

  • fix test

  • partly read meta

  • Address PR #4900 review feedback

  • Rename _variant_id -> _experiment_id (column, constant, helper) per
    review comment questioning what the column represented; it is the
    serialized experiment kwargs, so experiment_id is unambiguous.
  • Trim the AI-sounding long comment in build_benchmark_summary to state
    plainly what variants_by_model/variant_model_meta hold.
  • Document why join_revisions uses an empty-string sentinel instead of
    None for experiment_name (pandas groupby drops NaN/None group keys).

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>

  • make working natively

  • simplify experiments handling

  • refactor

  • simplify comments

  • fix typing

  • optimize loading

  • simplify comments

  • fix upload


Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> (60b2a1e)

  • Custom task grouping for leaderboard (#5174)

  • Add CustomGrouping: multi-dimension custom task aggregations

Generalizes the existing TASK_TYPES aggregation pattern so a Benchmark
can declare one or more named custom grouping dimensions directly in
its aggregations sequence (mixed with plain BenchmarkAggregation
flags, no new field, no new enum member). Each CustomGrouping produces
its own dynamic per-group mean columns in the leaderboard summary,
namespaced 'dimension::label' to avoid collisions between dimensions.

Migrates LMEB to use this for a 'Memory Type' breakdown (episodic /
dialogue / semantic / procedural) instead of registering four
duplicate Benchmark objects, resolving the reviewer objection on
PR #4914 for issue #4898.

  • CustomGroup / CustomGrouping dataclasses (benchmark.py)
  • _compute_custom_group_means (_benchmark_metrics.py, pure-Python path)
  • _get_means_per_custom_group + SummaryTable.custom_group_cols
    (_create_table.py, polars path)
  • CustomGroupSchema / CustomGroupingSchema + new fields on
    BenchmarkSchema / SummaryRowSchema / BenchmarkSummarySchema
    (api/schemas.py)
  • aggregators.py wiring: scores_by_custom_group per row, descriptions
    joined back from the benchmark's static declaration
  • LMEB migration + leaderboard menu entry
  • Tests: LMEB grouping coverage, polars column namespacing, get_score
    parity with the summary table, cross-dimension label collision safety

404 tests pass; ruff check/format clean.

  • Add Document Length custom grouping to BRIGHT(v1.1)

Groups the 20 tasks into Short (12) / Long (8) via the same
CustomGrouping mechanism used for LMEB's Memory Type breakdown --
the 8 domains with a Long variant land in both groups, the 4
code/math domains (Leetcode, Aops, TheoremQA*) only have Short.

Adds test_bright_document_length_grouping_covers_all_tasks, asserting
full task coverage and that group membership matches the Long/Short
suffix on each task name.

  • Recompute scores_by_custom_group under the language filter

The language sidebar filter isn't gated on Benchmark.language_view --
it reads tasksMeta[].languages directly and triggers a server refetch
for any benchmark with more than one task language. scores_by_custom_group
was previously left frozen at unfiltered values there (LMEB/BRIGHT(v1.1)
are both eng-only today, which is why this never surfaced, not because
of language_view as the old comment claimed).

  • _recompute_lenient_custom_groups: per-dimension analog of
    _recompute_lenient_means, same 'average only what's present' policy
  • build_benchmark_summary now builds custom_group_task_to_label from
    each declared CustomGrouping.task_to_label (only under a language
    filter) and threads it through _build_summary_rows
  • tests/test_api/test_aggregators.py: new coverage for the pure
    recompute function (present-tasks-only averaging, empty-mapping
    no-op, parity with _recompute_lenient_means on a single dimension)
  • start refactor

  • Merge duplicated bucket-and-average logic behind a shared _bucket_means

_recompute_lenient_means and _recompute_lenient_custom_groups both
independently bucketed scores_by_task by a task->key mapping and
averaged each bucket. Extracted the shared primitive (_bucket_means);
both callers now just supply their task_to_key mapping(s) and combine
the per-bucket results into their own return shape (task-type recompute
also derives the two scalar means; custom-group recompute returns one
bucket-dict per dimension).

No behavior change -- 409 tests still pass.

  • refactor

  • fix

  • remove comments

  • Move lenient-recompute functions out of mteb.api into _benchmark_metrics

_bucket_means/_recompute_lenient_means/_recompute_lenient_custom_groups
are pure functions over plain dicts, not tied to FastAPI/pydantic --
moved to mteb.benchmarks._benchmark_metrics (mteb.api.aggregators now
imports them) alongside the other aggregation helpers that already
live there.

Also extracted _bucket_task_result_scores, the TaskResult-based analog
of _bucket_means, and refactored _compute_task_types and
_compute_custom_group_means to share it instead of each re-implementing
the same bucket-by-key + null-tracking loop.

Tests moved from tests/test_api/ (now removed) into
test_benchmark_score.py alongside the other _benchmark_metrics.py
coverage. No behavior change -- 407 tests pass.

  • simplify comments

  • Apply suggestions from code review

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

  • move import

  • lint

  • remove unnecessary aggregations

  • Send custom-group task membership so the frontend can recompute under sidebar filters

CustomGroupSchema previously only sent {label, description} — the
frontend had no way to know which tasks belong to a group, so the
task-type/domain/modality sidebar filters (client-side, no server
round-trip) left scoresByCustomGroup frozen at stale values while
scoresByTaskType correctly recomputed alongside it.

Adds tasks: list[str] to CustomGroupSchema, populated from
CustomGroup.tasks in both BenchmarkSchema.from_benchmark (static
declaration) and aggregators.py's data-driven summary construction
(joined back via the same declared_by_dim lookup already used for
descriptions).

414 tests pass; verified live against LMEB.

Bypassing pre-commit: the typos hook fails on a pre-existing, unrelated
acronym false-positive (fgmcaps_retrieval.py's 'retrievAl' in FIGMA),
not on anything in this change.

  • add persubset

  • simplify

  • remove comments

  • fix test

  • update skip rule

  • Apply suggestion from @KennethEnevoldsen

Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>

  • don't compute leniently

  • remove vs


Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com> (6cb5891)