2.19.0
2.19.0 (2026-08-14)
Feature
-
feat: add OpenAIEndpointWrapper for HTTP-based vLLM servers (#4834)
-
feat: add VllmEndpointWrapper for HTTP-based vLLM servers
Add a new wrapper to enable MTEB benchmarking against remote vLLM
servers via HTTP API, complementing the existing VllmEncoderWrapper
which requires local in-process instantiation.
The existing VllmEncoderWrapper only supports local vLLM instantiation,
which has several limitations:
- Requires loading model weights on every benchmark run
- Assumes GPU availability (gpu_memory_utilization, tensor_parallel_size)
- Cannot test remote or production vLLM deployments
- Cannot reuse running vLLM servers across multiple test runs
New VllmEndpointWrapper class that:
-
Communicates with vLLM servers via OpenAI-compatible /v1/embeddings API
-
Works with any vLLM backend (CPU or GPU)
-
Supports remote endpoints, authentication, and SSL configuration
-
Auto-detects max_length from model metadata
-
Includes retry logic and response validation
-
Enables server reuse across multiple benchmark runs
-
Benchmarking production vLLM deployments
-
Testing remote vLLM CPU servers
-
Evaluating vLLM endpoints without local instantiation overhead
-
CI/CD integration with existing vLLM infrastructure
-
mteb/models/vllm_endpoint_wrapper.py: New HTTP-based wrapper
-
mteb/models/init.py: Export VllmEndpointWrapper
-
tests/test_models/test_vllm_endpoint_wrapper.py: Basic tests
-
docs/get_started/advanced_usage/vllm_wrapper.md: Documentation
Validated against production vLLM CPU servers running:
- RedHatAI/granite-embedding-english-r2
- RedHatAI/all-MiniLM-L6-v2
- Multiple MTEB tasks (classification, retrieval, STS)
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- fix: address CI failures for VllmEndpointWrapper
- Fix ruff formatting issues
- Fix mypy type errors (missing return statement, payload type annotation)
- Fix test_ssl_verification by mocking requests to avoid DNS resolution
- Refactor _verify_server to reduce nesting complexity
- Extract _detect_max_length_from_models to separate method
- Add PLR0913 noqa for init (many parameters needed for flexibility)
- Add PLR6301 to test file ignores (pytest test methods pattern)
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- refactor: address PR review feedback
Changes per maintainer review:
- Renamed VllmEndpointWrapper -> OpenAIAPIWrapper to reflect OpenAI-compatible API support
- Added self.mteb_model_meta using ModelMeta.create_empty() for MTEB compatibility
- Removed unused revision parameter from init
- Updated encode() signature to properly use Unpack[EncodeKwargs] for batch_size, show_progress_bar, precision
- Added tqdm progress bar to batch processing with show_progress_bar control
- Removed environment-dependent fixtures from tests - all tests now fully mocked
- Added test_encode_with_mock_task test with actual data returns for query/corpus encoding
- Added test_batch_size_override to verify batch_size can be overridden per encode() call
- Fixed missing self.model_name assignment before _verify_server() call
- Refactored _verify_server() to reduce nesting (extracted _detect_max_length_from_models)
- Updated documentation to reflect OpenAI-compatible API support (vLLM, OpenAI, etc.)
- Added PLR6301 to test file ignores for pytest test methods pattern
All 8 tests pass. Tested successfully with live vLLM server running ibm-granite/granite-embedding-english-r2.
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- feat: add OpenAIRerankWrapper and refactor to shared base class
Refactored OpenAI-compatible API wrappers to eliminate code duplication and
add reranking support:
New Features
- Added OpenAIRerankWrapper for reranking models via /v1/rerank endpoint
- Supports both vLLM and OpenAI reranking endpoints
- Implements CrossEncoderProtocol for MTEB reranking tasks
Refactoring
- Extracted OpenAIBaseWrapper with shared HTTP logic (connection, retry, error handling)
- Refactored OpenAIAPIWrapper to inherit from OpenAIBaseWrapper
- Eliminated ~100 lines of duplicate code
- Single source of truth for HTTP communication
Architecture
OpenAIBaseWrapper (shared HTTP logic)
├── OpenAIAPIWrapper (embeddings via /v1/embeddings)
└── OpenAIRerankWrapper (reranking via /v1/rerank)
Testing
- All 15 tests pass (8 embedding + 7 reranking)
- Fully mocked tests - no live server required
- Validated with live vLLM endpoints:
- Embeddings: ibm-granite/granite-embedding-english-r2 @ :8000 ✅
- Reranking: BAAI/bge-reranker-v2-m3 @ :8001 ✅
Files Changed
- mteb/models/openai_wrappers.py: New unified module (540 lines)
- OpenAIBaseWrapper: Base class with HTTP logic
- OpenAIAPIWrapper: Embeddings wrapper (refactored)
- OpenAIRerankWrapper: Reranking wrapper (new)
- tests/test_models/test_openai_rerank_wrapper.py: New test suite (7 tests)
- tests/test_models/test_openai_api_wrapper.py: Updated imports
- mteb/models/init.py: Export both wrappers
- Removed: mteb/models/openai_api_wrapper.py (replaced)
Addresses reviewer feedback for reranking API support while maintaining clean,
maintainable architecture.
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- docs: update vLLM documentation for OpenAI-compatible API wrappers
- Update vllm_wrapper.md to document OpenAIAPIWrapper and OpenAIRerankWrapper
- Add examples for both embedding and reranking usage
- Fix mypy type annotation issue in openai_wrappers.py
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- style: format test file
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- refactor: address PR review feedback on API wrapper design
Changes based on review from @Samoed:
- Move mteb_model_meta to OpenAIBaseWrapper base class (shared)
- Remove batch_size from init - now passed to encode()/predict()
- Add batch_size and show_progress_bar as direct parameters
- Remove top_k from init - now passed to predict() per task
- Fix predict() to only accept equal-length queries and documents
- Update all tests to match new signatures
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- fix: handle empty input in encode() and predict() methods
- Return empty array for empty inputs instead of crashing
- Prevents numpy vstack error on empty list
- Matches expected behavior for edge cases
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- fix: resolve CI failures - mypy and test assertions
- Declare mteb_model_meta as class attribute with ModelMeta | None type
to match AbsEncoder protocol and fix mypy incompatibility error - Use pytest.approx() for float comparisons in reranking tests to handle
numpy float32 vs Python float precision differences
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- style: remove unused noqa directives for PLR6301
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- refactor: rename OpenAI wrapper classes for clarity
- OpenAIAPIWrapper → OpenAIAPIEncodeWrapper
- OpenAIRerankWrapper → OpenAIAPIRerankWrapper
This makes it clear that OpenAIAPIEncodeWrapper is for encoding (embeddings)
and adds "API" to OpenAIAPIRerankWrapper for consistency. Both classes wrap
OpenAI-compatible HTTP APIs.
Updated all references in:
- Source code (mteb/models/openai_wrappers.py)
- Public exports (mteb/models/init.py)
- Documentation (docs/get_started/advanced_usage/vllm_wrapper.md)
- Tests (tests/test_models/test_openai_*.py)
Co-Authored-By: Claude Sonnet 4.5 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- docs: move OpenAI API wrappers to separate documentation page
Address PR review feedback:
- Move OpenAI-Compatible API content from vllm_wrapper.md to its own
page (openai_api_wrappers.md) with lucide/plug icon - Remove confusing "local in-process" line from vLLM page
- Remove redundant "Additional Options" section (API reference suffices)
- Add note that CLI does not yet support OpenAI API wrappers
- Add cross-link from vLLM page to new OpenAI API page
- Add new page to mkdocs.yml nav
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
- chore: sync uv.lock with main after rebase
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
-
add whats new
-
add multimodality support
-
fix typing
-
run tasks
Signed-off-by: Maryam Tahhan <mtahhan@redhat.com>
Co-authored-by: Claude Sonnet 4.5 <noreply@anthropic.com>
Co-authored-by: Kenneth Enevoldsen <kennethcenevoldsen@gmail.com>
Co-authored-by: Roman Solomatin <36135455+Samoed@users.noreply.github.com> (dc7176b)
Unknown
-
data: add FGMCaps (#5191)
-
data: add FGMCaps
-
lint format (
2270401) -
Remove dead dataloader code (#5192)
remove dead dataloader code (3347cd1)
-
Add SSW60 dataset (#5154)
-
Add SSW60 first commit
-
Modify the license
-
Add descriptive stats
-
Remove not required Exception handling
-
ran lint
-
pin revision
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com>
- run uv check --fix
Co-authored-by: Roman Solomatin <samoed.roman@gmail.com> (6e38137)