Releases: AlbertXXuu/OpenMultimodalLab
Release list
OpenMultimodalLab v1.1.2
OpenMultimodalLab v1.1.2 is a bounded software maintenance patch. The immutable research and evidence baseline remains v1.0.0.
Fixed and clarified
- Removed framework-added navigation padding and button margins so the packaged Studio retains the canonical AlvenX Header geometry.
- Clarified that the current CUDA memory field measures PyTorch allocator memory.
- Added an observed-problem and matched-comparison gate for any future alternative-runtime study.
- Updated current software identity and quick-start refs to v1.1.2.
No model, task, dependency, runtime behavior, protocol version, evidence file, result or research conclusion changed.
Provenance
- Merge commit:
0357d411b8c623331cb2e3e778af40cbdf680f03 - Pull request: #57
- Exact-main CI: 33376654428
Distribution SHA-256
openmultimodal_lab-1.1.2-py3-none-any.whl:4101eed7125cbc501640610a4d1826f328fbcc5fb53cfd5279057a678e136fb5openmultimodal_lab-1.1.2.tar.gz:f78559f9b84b6d578065cc7154df3a917646b5c36484a92dd3522fdb4fb58dc6
Both artifacts were installed in separate fresh Python 3.11 environments; package identity, dependency checks, CLI doctor, packaged Studio assets and the offline smoke workflow passed.
OpenMultimodalLab v1.1.1
v1.1.1 is a documentation correction over the v1.1.0 presentation and maintenance closure; the immutable v1.0.0 research/evidence baseline, benchmark results, protocols, behavior, dependencies, and claims are unchanged.
This patch corrects the current-software Quick Start to clone the release that contains Studio and aligns package, CLI, Studio, changelog, and maintenance identity with v1.1.1. The uploaded wheel and source distribution were built from main commit 94ad8006842d93193d461f636030de9a69472fcf and installed in fresh Python 3.11 environments.
v1.1.0 — Presentation and Maintenance Closure
v1.1.0 is the presentation and maintenance closure release built on the unchanged v1.0.0 research/evidence baseline.
This release finalizes the Studio presentation, shared AlvenX interface, contributor and maintenance documentation, portability checks, and installable Python packaging. The frozen v1.0.0 benchmark results and research claims are unchanged.
SHA-256:
openmultimodal_lab-1.1.0-py3-none-any.whl:4198d837c0f5caf9972f36eaf9866510aa2082eb4824332b53a018d6406872d9openmultimodal_lab-1.1.0.tar.gz:d80e74ab2174f92c17cba5148055a94945a3fc03aeb2785632032d77fc755094
OpenMultimodalLab v1.0.0
OpenMultimodalLab v1.0.0
OpenMultimodalLab v1.0.0 is the first public source release of a local-first, reproducible benchmark toolkit for vision-language models.
Highlights
- One versioned task contract and CLI across image, document, and short-video inputs.
- 102 project-generated, Apache-2.0-provenance tasks covering images, documents, short videos, and controlled visual robustness; the complete corpus was human-checked.
- Formal runs for two real open-source model families on one 8 GB consumer GPU: Qwen3-VL-2B and SmolVLM2-500M.
- One warm-up followed by three complete measured repetitions for each model: 612 measured attempts with no runtime failures.
- Preserved quality, TTFT, throughput, latency, peak VRAM, environment, failure, and integrity evidence.
- Deterministic Markdown, CSV, failure-data, SVG, and manifest reports that rebuild byte for byte from committed results.
- English and Simplified Chinese documentation, a copyable first-run tutorial, and a reproducible short-video demo.
- Windows Python 3.11/3.13 validation, GitHub Linux Python 3.11/3.12 CI, fresh-wheel checks, security review, and dependency-license audit.
Formal comparison
| Model | Mean task score | Median TTFT | Median task latency | Peak allocated VRAM | Runtime failures |
|---|---|---|---|---|---|
| Qwen3-VL-2B | 0.784 | 120.5 ms | 212.9 ms | 4,180.5 MiB | 0/306 |
| SmolVLM2-500M | 0.690 | 260.0 ms | 471.5 ms | 1,265.3 MiB | 0/306 |
These results describe the pinned tasks, revisions, hardware, and decoding configuration. They are not a universal model ranking.
Start here
- English README
- 简体中文说明
- First reproducible benchmark tutorial
- Byte-rebuildable 102-task report
- Release-readiness evidence
Reproducibility identity
- Release commit:
ad443bc73bbfd1a2bbb81aa1e83324dc8a98afff - Formal 102-task input SHA-256:
d18e6dce941cfac1fee0d637449229d786d7d6b601c063c0af2266b7e2d7a5a8 - Final Windows wheel evidence SHA-256:
463d679593239b14a8b6a7c5684c8b6b83a44bb7c1b53ee7b8690931d6c80d64 - Executable release gate:
19/19checks passed - Final main CI: Python 3.11, Python 3.12, and repository-quality jobs passed
Distribution and licensing boundary
This Release intentionally contains only GitHub-generated source archives. It does not attach model weights, Python environments, local run directories, or PyAV/FFmpeg runtime binaries. Model weights remain separately downloaded upstream artifacts governed by their own recorded licenses and immutable revisions.
Known limitations
- Formal measurements come from one RTX 4060 Laptop GPU and should not be generalized to other hardware without rerunning the protocol.
- The benchmark is intentionally small and controlled; most media are deterministic synthetic assets.
- Token throughput is tokenizer-dependent and is not directly comparable across model families.
- Real-model dependencies are optimized for Python 3.11/3.12; the core remains tested on Python 3.13 with optional PyAV coverage skipped where unavailable.
No adoption, user-feedback, or GitHub Star claims are made in this release.