Skip to content

2.19.0

Choose a tag to compare

@KennethEnevoldsen KennethEnevoldsen released this 14 Aug 12:06
· 257 commits to main since this release

2.19.0 (2026-08-14)

Feature

  • feat: add OpenAIEndpointWrapper for HTTP-based vLLM servers (#4834)

  • feat: add VllmEndpointWrapper for HTTP-based vLLM servers

Add a new wrapper to enable MTEB benchmarking against remote vLLM
servers via HTTP API, complementing the existing VllmEncoderWrapper
which requires local in-process instantiation.

The existing VllmEncoderWrapper only supports local vLLM instantiation,
which has several limitations:

  • Requires loading model weights on every benchmark run
  • Assumes GPU availability (gpu_memory_utilization, tensor_parallel_size)
  • Cannot test remote or production vLLM deployments
  • Cannot reuse running vLLM servers across multiple test runs

New VllmEndpointWrapper class that:

  • Communicates with vLLM servers via OpenAI-compatible /v1/embeddings API

  • Works with any vLLM backend (CPU or GPU)

  • Supports remote endpoints, authentication, and SSL configuration

  • Auto-detects max_length from model metadata

  • Includes retry logic and response validation

  • Enables server reuse across multiple benchmark runs

  • Benchmarking production vLLM deployments

  • Testing remote vLLM CPU servers

  • Evaluating vLLM endpoints without local instantiation overhead

  • CI/CD integration with existing vLLM infrastructure

  • mteb/models/vllm_endpoint_wrapper.py: New HTTP-based wrapper

  • mteb/models/init.py: Export VllmEndpointWrapper

  • tests/test_models/test_vllm_endpoint_wrapper.py: Basic tests

  • docs/get_started/advanced_usage/vllm_wrapper.md: Documentation

Validated against production vLLM CPU servers running:

  • RedHatAI/granite-embedding-english-r2
  • RedHatAI/all-MiniLM-L6-v2
  • Multiple MTEB tasks (classification, retrieval, STS)

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • fix: address CI failures for VllmEndpointWrapper
  • Fix ruff formatting issues
  • Fix mypy type errors (missing return statement, payload type annotation)
  • Fix test_ssl_verification by mocking requests to avoid DNS resolution
  • Refactor _verify_server to reduce nesting complexity
  • Extract _detect_max_length_from_models to separate method
  • Add PLR0913 noqa for init (many parameters needed for flexibility)
  • Add PLR6301 to test file ignores (pytest test methods pattern)

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • refactor: address PR review feedback

Changes per maintainer review:

  1. Renamed VllmEndpointWrapper -> OpenAIAPIWrapper to reflect OpenAI-compatible API support
  2. Added self.mteb_model_meta using ModelMeta.create_empty() for MTEB compatibility
  3. Removed unused revision parameter from init
  4. Updated encode() signature to properly use Unpack[EncodeKwargs] for batch_size, show_progress_bar, precision
  5. Added tqdm progress bar to batch processing with show_progress_bar control
  6. Removed environment-dependent fixtures from tests - all tests now fully mocked
  7. Added test_encode_with_mock_task test with actual data returns for query/corpus encoding
  8. Added test_batch_size_override to verify batch_size can be overridden per encode() call
  9. Fixed missing self.model_name assignment before _verify_server() call
  10. Refactored _verify_server() to reduce nesting (extracted _detect_max_length_from_models)
  11. Updated documentation to reflect OpenAI-compatible API support (vLLM, OpenAI, etc.)
  12. Added PLR6301 to test file ignores for pytest test methods pattern

All 8 tests pass. Tested successfully with live vLLM server running ibm-granite/granite-embedding-english-r2.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • feat: add OpenAIRerankWrapper and refactor to shared base class

Refactored OpenAI-compatible API wrappers to eliminate code duplication and
add reranking support:

New Features

  • Added OpenAIRerankWrapper for reranking models via /v1/rerank endpoint
  • Supports both vLLM and OpenAI reranking endpoints
  • Implements CrossEncoderProtocol for MTEB reranking tasks

Refactoring

  • Extracted OpenAIBaseWrapper with shared HTTP logic (connection, retry, error handling)
  • Refactored OpenAIAPIWrapper to inherit from OpenAIBaseWrapper
  • Eliminated ~100 lines of duplicate code
  • Single source of truth for HTTP communication

Architecture

OpenAIBaseWrapper (shared HTTP logic)
├── OpenAIAPIWrapper (embeddings via /v1/embeddings)
└── OpenAIRerankWrapper (reranking via /v1/rerank)

Testing

  • All 15 tests pass (8 embedding + 7 reranking)
  • Fully mocked tests - no live server required
  • Validated with live vLLM endpoints:
    • Embeddings: ibm-granite/granite-embedding-english-r2 @ :8000 ✅
    • Reranking: BAAI/bge-reranker-v2-m3 @ :8001 ✅

Files Changed

  • mteb/models/openai_wrappers.py: New unified module (540 lines)
    • OpenAIBaseWrapper: Base class with HTTP logic
    • OpenAIAPIWrapper: Embeddings wrapper (refactored)
    • OpenAIRerankWrapper: Reranking wrapper (new)
  • tests/test_models/test_openai_rerank_wrapper.py: New test suite (7 tests)
  • tests/test_models/test_openai_api_wrapper.py: Updated imports
  • mteb/models/init.py: Export both wrappers
  • Removed: mteb/models/openai_api_wrapper.py (replaced)

Addresses reviewer feedback for reranking API support while maintaining clean,
maintainable architecture.

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • docs: update vLLM documentation for OpenAI-compatible API wrappers
  • Update vllm_wrapper.md to document OpenAIAPIWrapper and OpenAIRerankWrapper
  • Add examples for both embedding and reranking usage
  • Fix mypy type annotation issue in openai_wrappers.py

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • style: format test file

Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • refactor: address PR review feedback on API wrapper design

Changes based on review from @Samoed:

  • Move mteb_model_meta to OpenAIBaseWrapper base class (shared)
  • Remove batch_size from init - now passed to encode()/predict()
  • Add batch_size and show_progress_bar as direct parameters
  • Remove top_k from init - now passed to predict() per task
  • Fix predict() to only accept equal-length queries and documents
  • Update all tests to match new signatures

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • fix: handle empty input in encode() and predict() methods
  • Return empty array for empty inputs instead of crashing
  • Prevents numpy vstack error on empty list
  • Matches expected behavior for edge cases

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • fix: resolve CI failures - mypy and test assertions
  • Declare mteb_model_meta as class attribute with ModelMeta | None type
    to match AbsEncoder protocol and fix mypy incompatibility error
  • Use pytest.approx() for float comparisons in reranking tests to handle
    numpy float32 vs Python float precision differences

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • style: remove unused noqa directives for PLR6301

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • refactor: rename OpenAI wrapper classes for clarity
  • OpenAIAPIWrapper → OpenAIAPIEncodeWrapper
  • OpenAIRerankWrapper → OpenAIAPIRerankWrapper

This makes it clear that OpenAIAPIEncodeWrapper is for encoding (embeddings)
and adds "API" to OpenAIAPIRerankWrapper for consistency. Both classes wrap
OpenAI-compatible HTTP APIs.

Updated all references in:

  • Source code (mteb/models/openai_wrappers.py)
  • Public exports (mteb/models/init.py)
  • Documentation (docs/get_started/advanced_usage/vllm_wrapper.md)
  • Tests (tests/test_models/test_openai_*.py)

Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • docs: move OpenAI API wrappers to separate documentation page

Address PR review feedback:

  • Move OpenAI-Compatible API content from vllm_wrapper.md to its own
    page (openai_api_wrappers.md) with lucide/plug icon
  • Remove confusing "local in-process" line from vLLM page
  • Remove redundant "Additional Options" section (API reference suffices)
  • Add note that CLI does not yet support OpenAI API wrappers
  • Add cross-link from vLLM page to new OpenAI API page
  • Add new page to mkdocs.yml nav

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • chore: sync uv.lock with main after rebase

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>

  • add whats new

  • add multimodality support

  • fix typing

  • run tasks


Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (dc7176b)

Unknown

  • data: add FGMCaps (#5191)

  • data: add FGMCaps

  • lint format (2270401)

  • Remove dead dataloader code (#5192)

remove dead dataloader code (3347cd1)

  • Enables B (flake8-bugbear) ruff rule (#5189) (f1037c6)

  • Add SSW60 dataset (#5154)

  • Add SSW60 first commit

  • Modify the license

  • Add descriptive stats

  • Remove not required Exception handling

  • ran lint

  • pin revision

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>

  • run uv check --fix

Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (6e38137)

  • add concurrency in dataset loading (#5182) (54494de)

  • Enables ISC (flake8-implicit-str-concat) ruff rule (#5181) (8c68e18)