You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
ltx2: added the ltx-2.5-distilled model (Lightricks/LTX-2.5) and made it
the default. distilled (LTX-2) remains available. LTX-2.5 adds the request
options auto_duration, enhance_prompt, pipeline="dfr", and video_decoder="diffusion". Run kiapi activate --family ltx2 to install the
new mlx-video and the gated LTX-2.5 weights (about 83 GB including the prompt
enhancer and DFR adapter).
setup: HfSnapshotResource accepts allow_patterns so a model can download
only the files it loads from a large repo.
chat: added the qwen3.8-27b model (mlx-community/Qwen3.8-27B-4bit, aliases qwen3.8, qwen3_5, qwen3-vl, vlm). The last three moved from qwen3.6-27b, which keeps only qwen3.6. It runs on the existing qwen3_5
handler (text + image, Hermes/XML tool calls, reasoning off by default).
Fixed
chat: image and video input failed on every model with TypeError: repeat(): incompatible function arguments since mlx 0.32. Fixed by
updating mlx-vlm (below).
chat: Qwen3-Omni now receives its deepstack visual features. mlx-vlm 0.6.3
computed them but dropped them before the decoder, so image/video input used
less of the vision tower than intended.
chat: Qwen3-Omni video input no longer decodes garbage or crashes the server
with a Metal GPU address fault on mlx-vlm 0.7.1. The deepstack mask and
features are now windowed to each prefill chunk (patch H, upstream #2099).
chat: Qwen3-Omni placed the video deepstack rows at the wrong positions when a
prompt had both an image and a video. Patch C now rewrites that join as in
upstream PR #2257 instead of grafting mx.where / mx.scatter onto mlx.
Changed
chat: updated mlx-vlm from 0.6.3 to 0.7.1 (still pinned exactly). The
streaming UTF-8 and mx.repeat patches are removed because upstream fixed
both; the stereo-audio, mx.where/mx.scatter, and stream-text patches stay. mlx-lm is now declared directly because mlx-embeddings' Qwen3-VL model
imports it and mlx-vlm no longer pulls it in.
Restructured the repository from a uv workspace (packages/kiapi) into a
single package: the source now lives in src/kiapi/, the tests in tests/,
and the release workflow builds and publishes kiapi directly. The package
README and CHANGELOG are merged into the root ones, and the root VERSION
file is replaced by the version in pyproject.toml.
Updated dependencies. FastAPI moves to >=0.141 (build_openapi now walks routing.iter_route_contexts to handle the lazy included routers of FastAPI 0.137+; the generated OpenAPI documents are unchanged), and the numpy<2.5 cap is lifted (numba >= 0.67 supports numpy 2.5; the out-of-band LTX-2 install needs numba>=0.67).
Refreshed the locked dependencies: torch 2.14, torchvision 0.29,
huggingface-hub 1.30, tokenizers 0.23.2, anyio 4.15, and ruff 0.16.6.
Updated the GitHub Actions used by CI, the PyPI release, and the Pages deploy
to their current majors. upload-pages-artifact now sets include-hidden-files: true so public/.nojekyll keeps being published.
Raised the mise pinned in CI and the release workflow from 2026.5.0 to 2026.9.1,
the version mise run ci is verified against locally.
Replaced the Dependabot configuration: the npm ecosystem entry pointed at the package.json removed in 0.6.0, so it is dropped in favour of the uv and github-actions ecosystems. mlx-vlm is ignored there for the same reason it
is pinned.
Removed
Removed the Node tooling (package.json / pnpm / firebase-tools) that existed only for the retired GCP relay setup task. This clears all open Dependabot alerts, which were transitive dependencies of firebase-tools.