Skip to content

v0.7.0

Choose a tag to compare

@github-actions github-actions released this 15 Sep 17:33
· 65 commits to main since this release
v0.7.0
d703adf

[0.7.0] - 2026-09-16

Added

  • ltx2: added the ltx-2.5-distilled model (Lightricks/LTX-2.5) and made it
    the default. distilled (LTX-2) remains available. LTX-2.5 adds the request
    options auto_duration, enhance_prompt, pipeline="dfr", and
    video_decoder="diffusion". Run kiapi activate --family ltx2 to install the
    new mlx-video and the gated LTX-2.5 weights (about 83 GB including the prompt
    enhancer and DFR adapter).
  • setup: HfSnapshotResource accepts allow_patterns so a model can download
    only the files it loads from a large repo.
  • chat: added the qwen3.8-27b model (mlx-community/Qwen3.8-27B-4bit, aliases
    qwen3.8, qwen3_5, qwen3-vl, vlm). The last three moved from
    qwen3.6-27b, which keeps only qwen3.6. It runs on the existing qwen3_5
    handler (text + image, Hermes/XML tool calls, reasoning off by default).

Fixed

  • chat: image and video input failed on every model with
    TypeError: repeat(): incompatible function arguments since mlx 0.32. Fixed by
    updating mlx-vlm (below).
  • chat: Qwen3-Omni now receives its deepstack visual features. mlx-vlm 0.6.3
    computed them but dropped them before the decoder, so image/video input used
    less of the vision tower than intended.
  • chat: Qwen3-Omni video input no longer decodes garbage or crashes the server
    with a Metal GPU address fault on mlx-vlm 0.7.1. The deepstack mask and
    features are now windowed to each prefill chunk (patch H, upstream #2099).
  • chat: Qwen3-Omni placed the video deepstack rows at the wrong positions when a
    prompt had both an image and a video. Patch C now rewrites that join as in
    upstream PR #2257 instead of grafting mx.where / mx.scatter onto mlx.

Changed

  • chat: updated mlx-vlm from 0.6.3 to 0.7.1 (still pinned exactly). The
    streaming UTF-8 and mx.repeat patches are removed because upstream fixed
    both; the stereo-audio, mx.where/mx.scatter, and stream-text patches stay.
    mlx-lm is now declared directly because mlx-embeddings' Qwen3-VL model
    imports it and mlx-vlm no longer pulls it in.
  • Restructured the repository from a uv workspace (packages/kiapi) into a
    single package: the source now lives in src/kiapi/, the tests in tests/,
    and the release workflow builds and publishes kiapi directly. The package
    README and CHANGELOG are merged into the root ones, and the root VERSION
    file is replaced by the version in pyproject.toml.
  • Updated dependencies. FastAPI moves to >=0.141 (build_openapi now walks routing.iter_route_contexts to handle the lazy included routers of FastAPI 0.137+; the generated OpenAPI documents are unchanged), and the numpy<2.5 cap is lifted (numba >= 0.67 supports numpy 2.5; the out-of-band LTX-2 install needs numba>=0.67).
  • Refreshed the locked dependencies: torch 2.14, torchvision 0.29,
    huggingface-hub 1.30, tokenizers 0.23.2, anyio 4.15, and ruff 0.16.6.
  • Updated the GitHub Actions used by CI, the PyPI release, and the Pages deploy
    to their current majors. upload-pages-artifact now sets
    include-hidden-files: true so public/.nojekyll keeps being published.
  • Raised the mise pinned in CI and the release workflow from 2026.5.0 to 2026.9.1,
    the version mise run ci is verified against locally.
  • Replaced the Dependabot configuration: the npm ecosystem entry pointed at the
    package.json removed in 0.6.0, so it is dropped in favour of the uv and
    github-actions ecosystems. mlx-vlm is ignored there for the same reason it
    is pinned.

Removed

  • Removed the Node tooling (package.json / pnpm / firebase-tools) that existed only for the retired GCP relay setup task. This clears all open Dependabot alerts, which were transitive dependencies of firebase-tools.