Skip to content

InferBridge 0.9.3-beta.1

Pre-release
Pre-release

Choose a tag to compare

@Quazmoz Quazmoz released this 05 Aug 17:25
· 423 commits to main since this release

InferBridge 0.9.3-beta.1

This is a beta pre-release and its artifacts are unsigned. Windows will report an
unknown publisher. Install it only if you accept that. Nothing in this release has been
Authenticode signed, and no signing claim is made anywhere in its metadata.

It is a model lifecycle release. A model conversion that was cancelled, ran out of disk, or
was interrupted by an upgrade could previously leave partially written OpenVINO output in
the live model directory, and that output was then treated as a converted model. This
release makes conversion output transactional, validates it before it is published, and
refuses several unsafe filesystem operations that could have followed a link out of the
model directory.

Converted models are now published transactionally

Optimum writes many files over a long export. Writing them directly into the live model
directory means any interruption mixes partial output with a previously runnable model.

  • Exports are staged, then published. Conversion writes into a sibling staging
    directory. The result is validated, and only then does it replace the live directory,
    through a bounded backup and rollback window. A previously working model survives a
    failed or cancelled reconversion.
  • Publication is hardened for Windows. Replacement retries transient sharing
    violations and access denials rather than failing on the first antivirus or Explorer
    lock, and clears read-only attributes it encounters.
  • Packaged conversions run inside the same transaction. The packaged converter path
    previously bypassed staging; conversions launched from the installed or portable build
    now get the same guarantee as source runs.

Incomplete model directories are no longer treated as models

A partially written export can contain an IR XML file before its weights or configuration
are on disk. That directory previously counted as "downloaded", so retries skipped
conversion and handed an incomplete artifact to the native OpenVINO loader.

  • One shared readiness check. runtime/model_artifacts.py validates that the IR XML is
    a complete document, that its paired weights file exists and is non-empty, and that the
    model configuration is present and parseable. Catalog, recovery, and load paths all use
    it, so they can no longer disagree about whether a model is ready.
  • The check stays cheap. It probes bounded regions of the IR XML rather than parsing
    multi-gigabyte files, so catalog listing is not slowed down.
  • Recovery classifies staged output correctly. Interrupted staging is now reported as
    recoverable output instead of being missed, and recovery cleans it up.

Preparing the same model twice is now blocked across processes

Two InferBridge processes — an installed instance and a portable one, for example — could
previously convert the same model into the same directory at the same time. Model
preparation now takes a cross-process lock on the output directory. The second process
fails fast with a message telling you to wait or close the other instance.

Symbolic links and Windows junctions are refused

Deletion and recovery cleanup previously followed reparse points, so a model directory
that was really a junction could have caused deletion outside model storage. Junctions are
also invisible to is_symlink(), so link detection now reads file attributes directly.

  • Model deletion refuses to act through a symbolic link or junction and asks you to remove
    the link manually after confirming its target.
  • Recovery cleanup refuses the same, refuses paths that are not directories, and refuses
    staged paths that fall outside the configured model directory.

Device switches no longer interrupt the loaded model

Switching a loaded model to another device previously took the model's lock before
compiling the replacement, so the model was unavailable for the entire compilation.

  • The loaded engine stays usable while the replacement compiles. The lock is now taken
    only for the short handoff at the end, and it is re-validated in case the engine was
    replaced while waiting.
  • The catalog reports the switch honestly. A loaded model that is switching devices is
    shown as loading and cannot be unloaded, instead of appearing idle and unloadable.
  • Unload is rejected during a switch. Unloading a model that is still loading or
    switching devices previously raced the load task; it now returns a clear error. Shutdown
    can still force cleanup after its generation drain timeout.
  • Progress copy is accurate on first load. "The currently loaded model remains
    available…" was shown even when nothing was loaded yet. A first load now says "First load
    can take several minutes…".

Clearer conversion failures

  • Disk exhaustion is named as such, with a note that cached downloads and any
    previously working model are preserved where possible.
  • Windows file locking produces a message about closing other instances, Explorer
    windows, and antivirus scans, rather than a raw WinError 32.
  • A concurrent preparation in another process is reported as exactly that.
  • Repeated wrappers are collapsed. Errors no longer read
    Conversion failed: RuntimeError: Conversion failed: ….
  • Converter diagnostics are bounded, so a converter that emits a large volume of output
    no longer grows memory without limit.

Continuous integration

A focused Windows Model Lifecycle workflow runs the download, convert, load, and recovery
regressions on windows-latest, because most of the behaviour above is Windows-specific
and was previously only covered on Linux.

Validation performed for this pre-release

This build was produced with -SkipTests. Lint, the Python test suite, the source mock
API contract validator, and the packaged mock smoke tests were not run as part of the
release build. The changes in this release were verified manually on Windows 11 build 26200
before the build was cut.

Release provenance, the model library manifest, and SHA-256 checksums are verified
independently of the build and did pass, so the artifact set is confirmed to be internally
consistent and built from this exact commit.

Continuous integration on this commit is red, with failures that are all in test code
or pre-existing:

  • tests/test_lifecycle_delete_safety.py has an unsorted import block that fails
    ruff check.
  • tests/test_model_recovery_cleanup.py::test_incomplete_output_rejects_a_dangling_link_before_exists_check
    fails because the test fakes is_symlink() on a path it never creates, so the
    lstat() that precedes the check raises and the guard correctly declines to fire. A real
    dangling link is still rejected; the fixture, not the shipped guard, is wrong.
  • Four Chromium browser tests fail. They also failed at the v0.9.2-beta.1 tag, and
    nothing under web/ or browser_tests/ changed in this release.

Not verified

Authenticode signing (these artifacts are unsigned), real Intel CPU, GPU, or NPU execution,
installer upgrade and downgrade on a real Windows installation, and conversion behaviour
against real disk-full and antivirus-lock conditions rather than simulated ones. Those
checks require the documented Windows release and hardware certification procedures. A
signed stable build must advance the version; see 0.9.0 for the planned stable
release record.