Skip to content

Releases: AlbertXXuu/OpenMultimodalLab

OpenMultimodalLab v1.1.2

Choose a tag to compare

@AlbertXXuu AlbertXXuu released this 31 Aug 09:22
0357d41

OpenMultimodalLab v1.1.2 is a bounded software maintenance patch. The immutable research and evidence baseline remains v1.0.0.

Fixed and clarified

  • Removed framework-added navigation padding and button margins so the packaged Studio retains the canonical AlvenX Header geometry.
  • Clarified that the current CUDA memory field measures PyTorch allocator memory.
  • Added an observed-problem and matched-comparison gate for any future alternative-runtime study.
  • Updated current software identity and quick-start refs to v1.1.2.

No model, task, dependency, runtime behavior, protocol version, evidence file, result or research conclusion changed.

Provenance

  • Merge commit: 0357d411b8c623331cb2e3e778af40cbdf680f03
  • Pull request: #57
  • Exact-main CI: 33376654428

Distribution SHA-256

  • openmultimodal_lab-1.1.2-py3-none-any.whl: 4101eed7125cbc501640610a4d1826f328fbcc5fb53cfd5279057a678e136fb5
  • openmultimodal_lab-1.1.2.tar.gz: f78559f9b84b6d578065cc7154df3a917646b5c36484a92dd3522fdb4fb58dc6

Both artifacts were installed in separate fresh Python 3.11 environments; package identity, dependency checks, CLI doctor, packaged Studio assets and the offline smoke workflow passed.

OpenMultimodalLab v1.1.1

Choose a tag to compare

@AlbertXXuu AlbertXXuu released this 30 Aug 17:24
94ad800

v1.1.1 is a documentation correction over the v1.1.0 presentation and maintenance closure; the immutable v1.0.0 research/evidence baseline, benchmark results, protocols, behavior, dependencies, and claims are unchanged.

This patch corrects the current-software Quick Start to clone the release that contains Studio and aligns package, CLI, Studio, changelog, and maintenance identity with v1.1.1. The uploaded wheel and source distribution were built from main commit 94ad8006842d93193d461f636030de9a69472fcf and installed in fresh Python 3.11 environments.

v1.1.0 — Presentation and Maintenance Closure

Choose a tag to compare

@AlbertXXuu AlbertXXuu released this 30 Aug 17:02

v1.1.0 is the presentation and maintenance closure release built on the unchanged v1.0.0 research/evidence baseline.

This release finalizes the Studio presentation, shared AlvenX interface, contributor and maintenance documentation, portability checks, and installable Python packaging. The frozen v1.0.0 benchmark results and research claims are unchanged.

SHA-256:

  • openmultimodal_lab-1.1.0-py3-none-any.whl: 4198d837c0f5caf9972f36eaf9866510aa2082eb4824332b53a018d6406872d9
  • openmultimodal_lab-1.1.0.tar.gz: d80e74ab2174f92c17cba5148055a94945a3fc03aeb2785632032d77fc755094

OpenMultimodalLab v1.0.0

Choose a tag to compare

@AlbertXXuu AlbertXXuu released this 10 Aug 07:38
ad443bc

OpenMultimodalLab v1.0.0

OpenMultimodalLab v1.0.0 is the first public source release of a local-first, reproducible benchmark toolkit for vision-language models.

Highlights

  • One versioned task contract and CLI across image, document, and short-video inputs.
  • 102 project-generated, Apache-2.0-provenance tasks covering images, documents, short videos, and controlled visual robustness; the complete corpus was human-checked.
  • Formal runs for two real open-source model families on one 8 GB consumer GPU: Qwen3-VL-2B and SmolVLM2-500M.
  • One warm-up followed by three complete measured repetitions for each model: 612 measured attempts with no runtime failures.
  • Preserved quality, TTFT, throughput, latency, peak VRAM, environment, failure, and integrity evidence.
  • Deterministic Markdown, CSV, failure-data, SVG, and manifest reports that rebuild byte for byte from committed results.
  • English and Simplified Chinese documentation, a copyable first-run tutorial, and a reproducible short-video demo.
  • Windows Python 3.11/3.13 validation, GitHub Linux Python 3.11/3.12 CI, fresh-wheel checks, security review, and dependency-license audit.

Formal comparison

Model Mean task score Median TTFT Median task latency Peak allocated VRAM Runtime failures
Qwen3-VL-2B 0.784 120.5 ms 212.9 ms 4,180.5 MiB 0/306
SmolVLM2-500M 0.690 260.0 ms 471.5 ms 1,265.3 MiB 0/306

These results describe the pinned tasks, revisions, hardware, and decoding configuration. They are not a universal model ranking.

Start here

Reproducibility identity

  • Release commit: ad443bc73bbfd1a2bbb81aa1e83324dc8a98afff
  • Formal 102-task input SHA-256: d18e6dce941cfac1fee0d637449229d786d7d6b601c063c0af2266b7e2d7a5a8
  • Final Windows wheel evidence SHA-256: 463d679593239b14a8b6a7c5684c8b6b83a44bb7c1b53ee7b8690931d6c80d64
  • Executable release gate: 19/19 checks passed
  • Final main CI: Python 3.11, Python 3.12, and repository-quality jobs passed

Distribution and licensing boundary

This Release intentionally contains only GitHub-generated source archives. It does not attach model weights, Python environments, local run directories, or PyAV/FFmpeg runtime binaries. Model weights remain separately downloaded upstream artifacts governed by their own recorded licenses and immutable revisions.

Known limitations

  • Formal measurements come from one RTX 4060 Laptop GPU and should not be generalized to other hardware without rerunning the protocol.
  • The benchmark is intentionally small and controlled; most media are deterministic synthetic assets.
  • Token throughput is tokenizer-dependent and is not directly comparable across model families.
  • Real-model dependencies are optimized for Python 3.11/3.12; the core remains tested on Python 3.13 with optional PyAV coverage skipped where unavailable.

No adoption, user-feedback, or GitHub Star claims are made in this release.