Repository navigation
2.19.4
2.19.4 (2026-08-17)
Fix
-
fix: Enable B905 and B028 Ruff rules (#5205)
-
Enable B905 and B028 Ruff rules
-
made strict=False for BUCC (
510fff6)
Unknown
-
model: add Amazon Nova Multimodal Embeddings (text, image, audio, video) (#5206)
-
model: add Amazon Nova Multimodal Embeddings (text, image)
Bedrock synchronous InvokeModel wrapper for
amazon.nova-2-multimodal-embeddings-v1:0.
Uses embeddingPurpose for asymmetric retrieval (GENERIC_RETRIEVAL for
queries, GENERIC_INDEX for documents) and maps Classification and
Clustering tasks to their dedicated purposes.
Audio and video are deferred to a follow-up: mteb decodes both before
they reach the encoder, so sending them would require re-encoding
decoded media back into a container.
Closes #5204
-
style: satisfy ruff format and no-self-use
-
refactor: move Nova wrapper into bedrock_models and share helpers
Per review, NovaMultimodalEmbeddingsModel now lives alongside BedrockModel
instead of in its own file.
Extracts three pieces both classes use:
- get_bedrock_runtime_client() for client construction
- CHARS_PER_TOKEN for the pre-truncation heuristic
- read_response_body() for decoding InvokeModel responses
Nova stays a separate class rather than a third provider branch on
BedrockModel: BedrockModel.encode is text-only by construction, and Nova's
request body and response shape differ from Titan's.
- model: add audio support to Nova wrapper
Addresses review:
- show_progress_bar is now an explicit keyword arg rather than pulled
from **kwargs - adds the audio path and declares audio in modalities
Audio arrives from mteb as a float array plus sampling rate, so it can be
encoded to WAV losslessly via the stdlib wave module. Verified against
Bedrock: text, image and audio all return embeddings at the requested dim.
Video remains unsupported: mteb decodes video to a frame tensor with no
frame rate and no audio track before the encoder sees it, so reconstructing
a container would require choosing an arbitrary fps.
- model: add video support to Nova wrapper
Video arrives as a torchcodec VideoDecoder, so frames are re-encoded to mp4
at the source frame rate (metadata.average_fps) and sent inline as base64.
Pass-through of the source container is not an option: Nova validates the
container against the declared format and MSVD ships AVI, which is not in
the accepted enum.
Uses AUDIO_VIDEO_COMBINED, which returns a single vector per item.
AUDIO_VIDEO_SEPARATE would return one vector per stream and break the
one-embedding-per-item contract.
Two guards on the decode: num_frames from the container header can overshoot
what actually decodes, so trailing frames are dropped until the read
succeeds (same approach as FramesCollator), and frames are cropped to even
dimensions since h264 requires them. Segments are capped at 30s, Nova's
limit.
MSVDT2VRetrieval nDCG@10 = 0.846 over 660 videos.
Co-authored-by: Hubert Lu <hubielu@email.com> (cd6cd8f)
-
model: add VideoMAE video encoders (#5094)
-
model: add VideoMAE and TimeSformer video encoders
-
model: add VideoMAE video encoders
Split TimeSformer into a separate file and PR per review.
Load the checkpoint state dict directly instead of via a lazy closure.
-
model: pin VideoMAE to transformers v4
-
model: return the encoder output for VideoMAE
Drop fc_norm per review; it belongs to the classification head. That also
removes _load_checkpoint_tensors, unused now that the transformers v4 pin
handles the attention biases.
-
fix: pass frames as HWC lists for the transformers v4 image processor
-
Update mteb/models/model_implementations/videomae_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
-
fix: restore logger import
-
model: use AutoVideoProcessor and drop the transformers v4 pin
Restore q_bias/v_bias in the wrapper instead, which v5 drops on load.
Same HMDB51Clustering score as the pinned path and roughly twice as fast.
- chore: declare the transformers-v5 requirement for VideoMAE
AutoVideoProcessor only resolves videomae from transformers 5.0.0 onward.
Co-authored-by: Hubert Lu <hubielu@email.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (32072a9)
-
[MOEB research]fix: strip empty caption segments in Clotho retrieval tasks (#5062)
-
fix: strip empty caption segments in Clotho retrieval tasks
ClothoT2ARetrieval and ClothoA2TRetrieval build queries/corpus by
splitting captions on ".", but never stripped or filtered the
resulting segments. Captions ending in a period produced a trailing
empty-string segment that became a real query/document with its own
qrel entry.
Add .strip() + skip-if-empty to both load_data() implementations and
regenerate the committed descriptive_stats JSON to reflect the cleaned
data (num_queries/num_documents: 5585 -> 4680 on the affected side,
905 empty segments removed, zero non-empty duplicates).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- fix: create v2 versions of Clotho retrieval tasks with empty-query fix
Per review feedback: existing model results were already scored under
ClothoT2ARetrieval/ClothoA2TRetrieval as originally defined, so we
should not silently mutate those tasks in place. Instead, following
MTEB's established versioning convention (see STS22 -> STS22.v2,
PoemSentimentClassification.v2), this reverts the in-place fix and
adds ClothoT2ARetrievalV2/ClothoA2TRetrievalV2 (name suffix ".v2")
with the corrected load_data().
- ClothoT2ARetrieval / ClothoA2TRetrieval: reverted to original
(buggy) load_data(); added superseded_by pointing to the .v2 task. - ClothoT2ARetrievalV2 / ClothoA2TRetrievalV2 (new): load_data()
strips whitespace and skips empty caption segments produced by
splitting on "."; adapted_from points back to the original task. - Regenerated descriptive_stats JSON: old tasks restored to their
original (pre-fix) committed values; new .v2 tasks get fresh stats
(num_queries/num_documents: 5585 -> 4680 on the affected side, 905
empty segments removed, zero non-empty duplicates, min_text_length
0 -> 33). - Registered both new classes in
mteb/tasks/retrieval/eng/init.py.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Revert "fix: create v2 versions of Clotho retrieval tasks with empty-query fix"
This reverts commit 940ed4b.
- fix: create v2 versions of Clotho retrieval tasks with empty-query fix
Per review feedback: existing model results were already scored under
ClothoT2ARetrieval/ClothoA2TRetrieval as originally defined, so we
should not silently mutate those tasks in place. Instead, following
MTEB's established versioning convention (see STS22 -> STS22.v2,
PoemSentimentClassification.v2, DBPediaHardNegatives.v2), this adds
ClothoT2ARetrievalV2/ClothoA2TRetrievalV2 (name suffix ".v2") with
the corrected load_data(), leaving the original tasks untouched.
- ClothoT2ARetrieval / ClothoA2TRetrieval: unchanged (original,
buggy) load_data(); added superseded_by pointing to the .v2 task. - ClothoT2ARetrievalV2 / ClothoA2TRetrievalV2 (new): load_data()
strips whitespace and skips empty caption segments produced by
splitting on "."; adapted_from points back to the original task. - Regenerated descriptive_stats JSON: old tasks restored to their
original (pre-fix) committed values; new .v2 tasks get fresh stats
via the real load_data() pipeline (num_queries/num_documents:
5585 -> 4680 on the affected side, 905 empty segments removed,
zero non-empty duplicates, min_text_length 0 -> 33). - Registered both new classes in
mteb/tasks/retrieval/eng/init.py.
Not included in this commit: updating mteb/benchmarks/benchmarks.py
to point MAEB(beta) at the .v2 task (left as a separate decision).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- Shorten ClothoT2ARetrieval.v2 description
docs: shorten ClothoT2ARetrieval.v2 description per review suggestion
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me>
- docs: shorten ClothoA2TRetrieval.v2 description
Mirrors the ClothoT2ARetrieval.v2 description shortened per
KennethEnevoldsen's review suggestion (applied directly on GitHub in
4f7e166). Also fixes the pre-existing "datasetst" typo in both v2
descriptions.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
- fix: point Clotho v2 retrieval tasks at pre-built HF datasets
ClothoT2ARetrieval.v2 and ClothoA2TRetrieval.v2 now load from
lxercode/clotho_t2a_v2 and lxercode/clotho_a2t_v2 (pinned revisions)
instead of running a custom load_data() against mteb/Clotho. These
repos contain the same empty-query-fixed data the custom loaders
already produced, materialized and uploaded ahead of time, so the
tasks no longer need a load_data() override.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Xu Liu <lxer@Xus-MacBook-Pro.local>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Kenneth Enevoldsen <kenevoldsen@pm.me> (dc1ac0a)
-
model: add UNITE (friedrichor/Unite-Base-Qwen2-VL-2B) (#5196)
-
model: add UNITE (friedrichor/Unite-Base-Qwen2-VL-2B)
-
model: fix UNITE metadata (memory, revision, training data)
-
model: lint fixes for UNITE wrapper
-
model: load UNITE via base Qwen2VL class, subclassing breaks key remapping on transformers 5.x
-
model: cap UNITE video frames at 360*420 to match reference inference
-
model: fix UNITE video path, tensor-aware frame resize and fps sampling
-
style: ruff format unite_models.py
-
refactor: drop unused num_frames param, fps mode only
-
Update mteb/models/model_implementations/unite_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
-
model: move video cap into processor, batch encode
-
Update mteb/models/model_implementations/unite_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
-
docs: note why UNITE loads via the base class
-
model: use UNITE subclass with transformers v4
-
model: use 32 frames for UNITE video evaluation
-
Update mteb/models/model_implementations/unite_models.py
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
- style: ruff format after removing _embed_one
Co-authored-by: Hubert Lu <hubielu@email.com>
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (a7c7086)
-
dataset: add IncompeBench (#5194)
-
dataset: add IncompeBench
-
add IncompeBenchLenientRetrieval and rename the strict task (
fa7ce6b) -
[MOEB] Add VELOCITI video pair classification task (#4980) (#4982)
-
[MOEB] Add VELOCITI video pair classification task (#4980)
-
fix: annotate mutable class attributes with ClassVar for lint
-
fix: rebuild VELOCITI-PC to avoid video duplication, rewrite description
Addresses review on PR #4982:
- Rebuilt yaswanth169/VELOCITI-PC to store video_id (string) + a single
videos.zip of the 864 unique files, instead of materializing full video
bytes into every one of the 17,584 rows. Repo size drops from ~32GB to
~1.6GB. load_data() now downloads the zip once and builds the local
Video()-typed dataset by path reference. - Rewrote the task description to lead with what VELOCITI measures rather
than assuming the reader already knows what it is.
- fix: store VELOCITI videos as a parquet config instead of a zip asset
Addresses review on PR #4982: instead of downloading a raw videos.zip
and manually extracting it, videos now live in a second HF dataset
config (video_id, video) with 864 unique rows using a proper Video()
feature, joined against the (video_id, text, label) table at load time
via concatenate_datasets(axis=1) -- an Arrow-level column join, no
decode/re-encode round trip. Removes zipfile/tempfile handling
entirely; simplifies both the wrapper code and standalone usage of the
dataset outside mteb, per Samoed's review comment.
- fix: deduplicate VELOCITI-PC (video, text) pairs, drop 1 label conflict
CI's test_dataset_quality caught 3,469 duplicated (video_id, text) pairs
in the reshaped dataset -- VELOCITI's source data reuses captions across
different test categories/events for the same video. Deduplicated
17,584 -> 11,669 rows, and dropped the single pair that had conflicting
labels across source rows (unsafe to keep either). Labels are no longer
perfectly balanced (7,463 label=0 / 4,206 label=1) as a result -- a real
property of the deduplicated data, not a bug.
- fix: force fresh download of VELOCITI-PC metadata table
CI's persistent HF cache (workflow key 'Linux-hf', never invalidated
across commits) was silently serving a pre-dedup Arrow snapshot of the
small (video_id, text, label) table with zero network calls, producing
stale duplicate-pair failures on the correct, verified-clean data.
Confirmed via job log: 'Cache restored from key: Linux-hf' followed by
no download activity for VELOCITI-PC at all. force_redownload makes
this cheap 401KB table immune to that class of staleness going forward.
- fix: load VELOCITI-PC metadata table via direct parquet download
The previous force_redownload fix addressed stale-cache reuse but not
the real bug: CI runs the suite with pytest-xdist (-n auto), and
datasets' Arrow-cache build for this table isn't safe when a concurrent
test worker touches the same shared HF cache dir at the same time --
that's a race, not staleness, so force_redownload could never have
fixed it (confirmed: it didn't -- CI still failed identically on that
commit).
Switched to hf_hub_download + pandas for this small (401KB) table
instead of load_dataset, since a single atomic revision-pinned file
fetch has no cache-build step for another worker to race. Verified
correct under 5 concurrent processes hitting the same cache
directory simultaneously, in addition to a normal single-process load.
- fix: regenerate stale VELOCITI-PC descriptive stats
test_dataset_quality reads task.metadata.descriptive_stats, a committed
JSON file under mteb/descriptive_stats/ -- it never calls load_data().
That file was generated once, before any of the zip->parquet rebuild,
dedup, or cache fixes, and was never regenerated afterward, so it kept
reporting the original 17,584-row/11,670-unique-pair pre-dedup numbers
regardless of what the dataset actually contained. Every prior fix
(force_redownload, then hf_hub_download) was correct for load_data()
but irrelevant to this specific failure.
Regenerated via task.calculate_descriptive_statistics(overwrite_results=True)
against the live, correct dataset: num_samples/unique_pairs now 11669/11669
(zero duplicates), unique_videos still 864 as expected.
- fix: simplify VELOCITI-PC loading back to plain load_dataset
The hf_hub_download switch was based on an unconfirmed race-condition
theory (CI runs pytest -n auto, multiple workers sharing one HF cache
dir). It was never actually proven necessary: the test that was failing
(test_dataset_quality) doesn't call load_data() at all, it reads a
committed descriptive-stats JSON, which was the real fix. Reverting to
plain load_dataset() per review -- simpler, and there's no evidence the
extra complexity was ever solving anything.
Verified end-to-end on a fully cleared cache: 11,669 rows, zero
duplicate pairs, correct columns. (b7f1504)