OpenMultimodalLab v1.0.0
OpenMultimodalLab v1.0.0
OpenMultimodalLab v1.0.0 is the first public source release of a local-first, reproducible benchmark toolkit for vision-language models.
Highlights
- One versioned task contract and CLI across image, document, and short-video inputs.
- 102 project-generated, Apache-2.0-provenance tasks covering images, documents, short videos, and controlled visual robustness; the complete corpus was human-checked.
- Formal runs for two real open-source model families on one 8 GB consumer GPU: Qwen3-VL-2B and SmolVLM2-500M.
- One warm-up followed by three complete measured repetitions for each model: 612 measured attempts with no runtime failures.
- Preserved quality, TTFT, throughput, latency, peak VRAM, environment, failure, and integrity evidence.
- Deterministic Markdown, CSV, failure-data, SVG, and manifest reports that rebuild byte for byte from committed results.
- English and Simplified Chinese documentation, a copyable first-run tutorial, and a reproducible short-video demo.
- Windows Python 3.11/3.13 validation, GitHub Linux Python 3.11/3.12 CI, fresh-wheel checks, security review, and dependency-license audit.
Formal comparison
| Model | Mean task score | Median TTFT | Median task latency | Peak allocated VRAM | Runtime failures |
|---|---|---|---|---|---|
| Qwen3-VL-2B | 0.784 | 120.5 ms | 212.9 ms | 4,180.5 MiB | 0/306 |
| SmolVLM2-500M | 0.690 | 260.0 ms | 471.5 ms | 1,265.3 MiB | 0/306 |
These results describe the pinned tasks, revisions, hardware, and decoding configuration. They are not a universal model ranking.
Start here
- English README
- 简体中文说明
- First reproducible benchmark tutorial
- Byte-rebuildable 102-task report
- Release-readiness evidence
Reproducibility identity
- Release commit:
ad443bc73bbfd1a2bbb81aa1e83324dc8a98afff - Formal 102-task input SHA-256:
d18e6dce941cfac1fee0d637449229d786d7d6b601c063c0af2266b7e2d7a5a8 - Final Windows wheel evidence SHA-256:
463d679593239b14a8b6a7c5684c8b6b83a44bb7c1b53ee7b8690931d6c80d64 - Executable release gate:
19/19checks passed - Final main CI: Python 3.11, Python 3.12, and repository-quality jobs passed
Distribution and licensing boundary
This Release intentionally contains only GitHub-generated source archives. It does not attach model weights, Python environments, local run directories, or PyAV/FFmpeg runtime binaries. Model weights remain separately downloaded upstream artifacts governed by their own recorded licenses and immutable revisions.
Known limitations
- Formal measurements come from one RTX 4060 Laptop GPU and should not be generalized to other hardware without rerunning the protocol.
- The benchmark is intentionally small and controlled; most media are deterministic synthetic assets.
- Token throughput is tokenizer-dependent and is not directly comparable across model families.
- Real-model dependencies are optimized for Python 3.11/3.12; the core remains tested on Python 3.13 with optional PyAV coverage skipped where unavailable.
No adoption, user-feedback, or GitHub Star claims are made in this release.