Changelog
Features 🏎️
- #592 - add
google/embeddinggemma-300mby @joein - #651 - add
nomic-ai/nomic-embed-vision-v1.5andnomic-ai/nomic-embed-vision-v1.5-Q, which share an embedding space withnomic-ai/nomic-embed-text-v1.5for text-to-image search by @Dylancouzon - #652 - add inference-free SPLADE
opensearch-project/opensearch-neural-sparse-encoding-doc-v3-gte, which runs the model on documents only and encodes queries without inference by @joein - #678, #680 - add
Qwen/Qwen3-Embedding-0.6BandQwen/Qwen3-Embedding-0.6B-Q(the quantized model requiresonnxruntime>=1.23), plusPoolingType.LAST_TOKENfor custom models by @joein - #683 - add
google/siglip2-base-patch16-224toTextEmbeddingandImageEmbeddingby @joein - #692 - add Model2Vec static models
minishlab/potion-base-8M,minishlab/potion-retrieval-32M, andminishlab/potion-multilingual-128Mby @stephantul @Dylancouzon
Fixes 🔧
- #593, #707 - use canonical Hugging Face repo IDs for
BAAI/bge-base-en-v1.5andBAAI/bge-small-en-v1.5to avoid a redirect that broke downloads behind some proxies (see Upgrade Notes) by @Harnas @rastagan-git @joein - #623 - run the fp32 version of
jinaai/jina-embeddings-v2-base-deinstead of fp16, which fails ononnxruntime>=1.23(the download grows from 0.32 GB to 0.64 GB) by @joein - #624 - check that model files exist before using a cached model, so several variants of one repo (for example fp32 and quantized) can share a cache dir by @joein
- #629 - add a timeout to downloads from Google Cloud Storage (GCS), so a stalled connection fails instead of hanging indefinitely by @joein
- #645 - fix a
KeyErrorwhen loading a custom text model with different casing than it was registered with by @CODING-DARSH @joein - #647 - fix a path traversal vulnerability in model archive extraction that could write files outside the cache dir by @he-yufeng @joein
- #682 - fix normalization of batched
(N, C, H, W)image arrays, which ran along the wrong axis by @serhiizghama @joein - #693 - make
config.jsonandspecial_tokens_map.jsonoptional when loading a tokenizer by @libaojiang @joein - #697 - fix the
Resizetransform swapping height and width for non-square sizes by @Ramnath0521 @joein - #714 - fix custom text embedding and cross-encoder models (
add_custom_model) failing whenparallelis set by @joein @S0rryHorizon - #716, #717 - always pad a batch to its longest sequence, fixing a
ValueErroron mixed-length batches for models whose tokenizer ships a fixed padding length, such asthenlper/gte-base(regression in 0.8.0) by @joein @mohmedmm - #718 - stage each GCS download in its own temporary dir, so a failed or concurrent download can't delete files from other downloads by @joein
Upgrade Notes
BAAI/bge-small-en-v1.5(the default model) andBAAI/bge-base-en-v1.5now resolve toQdrant/bge-small-en-v1.5-onnx-QandQdrant/bge-base-en-v1.5-onnx-Q. The cache dir name follows the repo ID's casing, so on case-sensitive filesystems (typically Linux) the existing cache isn't reused and both models download again once. Offline setups (HF_HUB_OFFLINE=1orlocal_files_only=True) need to refresh their cache before upgrading. You can delete the oldmodels--qdrant--bge-*-onnx-qdirs afterwards.jinaai/jina-embeddings-v2-base-denow loadsonnx/model.onnxinstead ofonnx/model_fp16.onnx, so offline setups need to download it first.
Thanks to everyone who contributed to this release @CODING-DARSH @Dylancouzon @Harnas @he-yufeng @libaojiang @mohmedmm @Ramnath0521 @rastagan-git @S0rryHorizon @serhiizghama @stephantul @joein