Skip to content

v0.8.1

Latest

Choose a tag to compare

@joein joein released this 22 Sep 20:52
· 27 commits to main since this release
8de28b8

Changelog

Features 🏎️

  • #592 - add google/embeddinggemma-300m by @joein
  • #651 - add nomic-ai/nomic-embed-vision-v1.5 and nomic-ai/nomic-embed-vision-v1.5-Q, which share an embedding space with nomic-ai/nomic-embed-text-v1.5 for text-to-image search by @Dylancouzon
  • #652 - add inference-free SPLADE opensearch-project/opensearch-neural-sparse-encoding-doc-v3-gte, which runs the model on documents only and encodes queries without inference by @joein
  • #678, #680 - add Qwen/Qwen3-Embedding-0.6B and Qwen/Qwen3-Embedding-0.6B-Q (the quantized model requires onnxruntime>=1.23), plus PoolingType.LAST_TOKEN for custom models by @joein
  • #683 - add google/siglip2-base-patch16-224 to TextEmbedding and ImageEmbedding by @joein
  • #692 - add Model2Vec static models minishlab/potion-base-8M, minishlab/potion-retrieval-32M, and minishlab/potion-multilingual-128M by @stephantul @Dylancouzon

Fixes 🔧

  • #593, #707 - use canonical Hugging Face repo IDs for BAAI/bge-base-en-v1.5 and BAAI/bge-small-en-v1.5 to avoid a redirect that broke downloads behind some proxies (see Upgrade Notes) by @Harnas @rastagan-git @joein
  • #623 - run the fp32 version of jinaai/jina-embeddings-v2-base-de instead of fp16, which fails on onnxruntime>=1.23 (the download grows from 0.32 GB to 0.64 GB) by @joein
  • #624 - check that model files exist before using a cached model, so several variants of one repo (for example fp32 and quantized) can share a cache dir by @joein
  • #629 - add a timeout to downloads from Google Cloud Storage (GCS), so a stalled connection fails instead of hanging indefinitely by @joein
  • #645 - fix a KeyError when loading a custom text model with different casing than it was registered with by @CODING-DARSH @joein
  • #647 - fix a path traversal vulnerability in model archive extraction that could write files outside the cache dir by @he-yufeng @joein
  • #682 - fix normalization of batched (N, C, H, W) image arrays, which ran along the wrong axis by @serhiizghama @joein
  • #693 - make config.json and special_tokens_map.json optional when loading a tokenizer by @libaojiang @joein
  • #697 - fix the Resize transform swapping height and width for non-square sizes by @Ramnath0521 @joein
  • #714 - fix custom text embedding and cross-encoder models (add_custom_model) failing when parallel is set by @joein @S0rryHorizon
  • #716, #717 - always pad a batch to its longest sequence, fixing a ValueError on mixed-length batches for models whose tokenizer ships a fixed padding length, such as thenlper/gte-base (regression in 0.8.0) by @joein @mohmedmm
  • #718 - stage each GCS download in its own temporary dir, so a failed or concurrent download can't delete files from other downloads by @joein

Upgrade Notes

  • BAAI/bge-small-en-v1.5 (the default model) and BAAI/bge-base-en-v1.5 now resolve to Qdrant/bge-small-en-v1.5-onnx-Q and Qdrant/bge-base-en-v1.5-onnx-Q. The cache dir name follows the repo ID's casing, so on case-sensitive filesystems (typically Linux) the existing cache isn't reused and both models download again once. Offline setups (HF_HUB_OFFLINE=1 or local_files_only=True) need to refresh their cache before upgrading. You can delete the old models--qdrant--bge-*-onnx-q dirs afterwards.
  • jinaai/jina-embeddings-v2-base-de now loads onnx/model.onnx instead of onnx/model_fp16.onnx, so offline setups need to download it first.

Thanks to everyone who contributed to this release @CODING-DARSH @Dylancouzon @Harnas @he-yufeng @libaojiang @mohmedmm @Ramnath0521 @rastagan-git @S0rryHorizon @serhiizghama @stephantul @joein