Skip to content

OpenMultimodalLab v1.0.0

Choose a tag to compare

@AlbertXXuu AlbertXXuu released this 10 Aug 07:38
· 62 commits to main since this release
ad443bc

OpenMultimodalLab v1.0.0

OpenMultimodalLab v1.0.0 is the first public source release of a local-first, reproducible benchmark toolkit for vision-language models.

Highlights

  • One versioned task contract and CLI across image, document, and short-video inputs.
  • 102 project-generated, Apache-2.0-provenance tasks covering images, documents, short videos, and controlled visual robustness; the complete corpus was human-checked.
  • Formal runs for two real open-source model families on one 8 GB consumer GPU: Qwen3-VL-2B and SmolVLM2-500M.
  • One warm-up followed by three complete measured repetitions for each model: 612 measured attempts with no runtime failures.
  • Preserved quality, TTFT, throughput, latency, peak VRAM, environment, failure, and integrity evidence.
  • Deterministic Markdown, CSV, failure-data, SVG, and manifest reports that rebuild byte for byte from committed results.
  • English and Simplified Chinese documentation, a copyable first-run tutorial, and a reproducible short-video demo.
  • Windows Python 3.11/3.13 validation, GitHub Linux Python 3.11/3.12 CI, fresh-wheel checks, security review, and dependency-license audit.

Formal comparison

Model Mean task score Median TTFT Median task latency Peak allocated VRAM Runtime failures
Qwen3-VL-2B 0.784 120.5 ms 212.9 ms 4,180.5 MiB 0/306
SmolVLM2-500M 0.690 260.0 ms 471.5 ms 1,265.3 MiB 0/306

These results describe the pinned tasks, revisions, hardware, and decoding configuration. They are not a universal model ranking.

Start here

Reproducibility identity

  • Release commit: ad443bc73bbfd1a2bbb81aa1e83324dc8a98afff
  • Formal 102-task input SHA-256: d18e6dce941cfac1fee0d637449229d786d7d6b601c063c0af2266b7e2d7a5a8
  • Final Windows wheel evidence SHA-256: 463d679593239b14a8b6a7c5684c8b6b83a44bb7c1b53ee7b8690931d6c80d64
  • Executable release gate: 19/19 checks passed
  • Final main CI: Python 3.11, Python 3.12, and repository-quality jobs passed

Distribution and licensing boundary

This Release intentionally contains only GitHub-generated source archives. It does not attach model weights, Python environments, local run directories, or PyAV/FFmpeg runtime binaries. Model weights remain separately downloaded upstream artifacts governed by their own recorded licenses and immutable revisions.

Known limitations

  • Formal measurements come from one RTX 4060 Laptop GPU and should not be generalized to other hardware without rerunning the protocol.
  • The benchmark is intentionally small and controlled; most media are deterministic synthetic assets.
  • Token throughput is tokenizer-dependent and is not directly comparable across model families.
  • Real-model dependencies are optimized for Python 3.11/3.12; the core remains tested on Python 3.13 with optional PyAV coverage skipped where unavailable.

No adoption, user-feedback, or GitHub Star claims are made in this release.