Repository navigation
Releases: kiarina/kiapi
Release list
v0.8.0
[0.8.0] - 2026-09-26
Dependencies
- mflux 0.19.1 -> 0.20.0, pinned to a kiarina fork commit (
144a6bec) that adds
mflux-community/mflux#741 on top of 0.20.0. 0.20.0 also fixes the seedvr2
mx.repeatfailure.
Removed
- chat: removed the
qwen3.6-27bmodel (aliasqwen3.6).qwen3.8-27buses the
same handler, modalities and memory footprint;vlmselectsqwen3.8-flash-next.
Remove the downloaded weights withkiapi deactivate --repo mlx-community/Qwen3.6-27B-4bit.
Changed
- chat:
modelis now required on/v1/chat/completions; there is no default
chat model. Requests without it return HTTP 422. The chat models differ in
accepted input modalities (onlyqwen3-omnitakes audio/video) and memory
footprint (qwen3.8-flash-nextkeeps about 74 GiB resident), so an implicit
default could silently evict other models or reject media. - chat: the aliases
vlmandqwen3.8now selectqwen3.8-flash-next
(previouslyqwen3.8-27b).qwen3.8-flash-nextalso answers to
qwen3.8-flashandflash-next; itsqwen4_expalias was removed.
qwen3.8-27bkeepsqwen3_5andqwen3-vl. - BREAKING: chat: removed the
max_tokens_capsetting
(KIAPI_CHAT_MAX_TOKENS_CAP, 4096).max_completion_tokensis no longer
capped by the server; generation stops atmax_completion_tokensor when the
prompt plus the output fills the model's context window, whichever comes first. - chat: the default
max_completion_tokensis now 1024 (was 512). - chat: non-streaming requests now run through
stream_generatelike streaming
ones, so the context window bound applies to both.
Added
-
Web UI at
/: a dashboard (server, queue, memory, setup), model setup
status with copyablekiapi activatecommands, jobs, files with previews, and a
guide and API reference for every family, in light and dark themes. Every
family can be run from its Playground: forms are built from the family's
OpenAPI schema, generation runs as a job with live progress, inputs come from
uploads or stored files, and "Write with chat" drafts prompts with a chat model
that reads the family guide. Chat has its own conversation view with every
request parameter (tools, tool choice, parallel tool calls, token limits,
sampling, template kwargs, streaming on or off), tool-call results, and token
usage. A "?" beside each field shows its full description, and a floating
assistant answers questions about the current family from its OpenAPI document
and, when asked, fills in the form through tool calls (it never submits). -
GET /v1/setup: every model with its setup state, likekiapi status. -
qwen: added the
image-2.1model (Qwen/Qwen-Image-2.1, aliases
qwen-image-2.1,qwen-2.1). One resident model serves both/generate
(text-to-image) and/edit(up to 10 reference images), outputs RGBA, and
renders up to 2752 px. Defaults are 40 steps, guidance 1.0 and q8. It does not
takeinit_imageorloras. Its weights are under the non-commercial Qwen
Research License. Editing comes from the pinned mflux fork (144a6bec,
mflux-community/mflux#741); installs with official mflux do not register it. -
qwen:
jpegoutput flattens transparent areas onto white. -
chat: Qwen3.8-Flash-Next reuses unchanged history when images are appended,
like Qwen3.8-27B, through the pinned mlx-vlm fork (6581ba8c). -
chat: added the
qwen3.8-flash-nextmodel (mlx-community/Qwen3.8-Flash-Next-4bit,
aliasesqwen3.8-flash,flash-next,qwen4_exp). It loads through a view that
memory-maps its n-gram (PLE) table, which keeps about 80 GB resident instead of
111.5 GB, and raises the open-file limit the mapped table needs. -
chat: Omni reuses unchanged prefixes when images, audio clips or videos are
appended, including demuxed audiovisual inputs, through a pinned mlx-vlm fork. -
chat: the pinned engine supports multiple audio clips with independent feature
extraction/encoding and corrected CNN lengths and chunk masks. -
chat: Qwen3.8 reuses unchanged image history when new images are appended,
using a pinned mlx-vlm fork in uv-managed checkouts. Official-engine installs
retain the conservative whole-request fallback. -
chat: client disconnects now cancel queued work or stop running generation at
the next token boundary. Cancellation uses the existingcanceledjob state
and safely clears APC state when a generator closes early. -
chat: responses report APC reuse through the OpenAI-compatible
usage.prompt_tokens_details.cached_tokensfield. Streaming requests support
stream_options.include_usageand emit usage before[DONE]when requested. -
chat: bounded, memory-only automatic prefix caching for Qwen3.6 / Qwen3.8
text and images, and Qwen3-Omni text, image, audio, video, and image + video.
Media content hashes and video options protect cache identity, and Omni
restores complete-prompt positions when reusing media prefixes. -
chat:
GET /v1/modelsreturns each model'scontext_window, read from the
model'sconfig.json(nulluntil the model is set up).
Fixed
-
chat: suppress extra streamed tool names when
parallel_tool_calls=false. -
memory: include idle chat caches in cross-model eviction and transient reservations,
and derive active cache headroom from the configured APC capacity. -
chat:
finish_reasonis now"length"when generation stops at
max_completion_tokensor the context window. It was always"stop".
v0.7.0
[0.7.0] - 2026-09-16
Added
- ltx2: added the
ltx-2.5-distilledmodel (Lightricks/LTX-2.5) and made it
the default.distilled(LTX-2) remains available. LTX-2.5 adds the request
optionsauto_duration,enhance_prompt,pipeline="dfr", and
video_decoder="diffusion". Runkiapi activate --family ltx2to install the
new mlx-video and the gated LTX-2.5 weights (about 83 GB including the prompt
enhancer and DFR adapter). - setup:
HfSnapshotResourceacceptsallow_patternsso a model can download
only the files it loads from a large repo. - chat: added the
qwen3.8-27bmodel (mlx-community/Qwen3.8-27B-4bit, aliases
qwen3.8,qwen3_5,qwen3-vl,vlm). The last three moved from
qwen3.6-27b, which keeps onlyqwen3.6. It runs on the existingqwen3_5
handler (text + image, Hermes/XML tool calls, reasoning off by default).
Fixed
- chat: image and video input failed on every model with
TypeError: repeat(): incompatible function argumentssince mlx 0.32. Fixed by
updating mlx-vlm (below). - chat: Qwen3-Omni now receives its deepstack visual features. mlx-vlm 0.6.3
computed them but dropped them before the decoder, so image/video input used
less of the vision tower than intended. - chat: Qwen3-Omni video input no longer decodes garbage or crashes the server
with a Metal GPU address fault on mlx-vlm 0.7.1. The deepstack mask and
features are now windowed to each prefill chunk (patch H, upstream #2099). - chat: Qwen3-Omni placed the video deepstack rows at the wrong positions when a
prompt had both an image and a video. Patch C now rewrites that join as in
upstream PR #2257 instead of graftingmx.where/mx.scatteronto mlx.
Changed
- chat: updated
mlx-vlmfrom 0.6.3 to 0.7.1 (still pinned exactly). The
streaming UTF-8 andmx.repeatpatches are removed because upstream fixed
both; the stereo-audio,mx.where/mx.scatter, and stream-text patches stay.
mlx-lmis now declared directly because mlx-embeddings' Qwen3-VL model
imports it and mlx-vlm no longer pulls it in. - Restructured the repository from a uv workspace (
packages/kiapi) into a
single package: the source now lives insrc/kiapi/, the tests intests/,
and the release workflow builds and publisheskiapidirectly. The package
README and CHANGELOG are merged into the root ones, and the rootVERSION
file is replaced by theversioninpyproject.toml. - Updated dependencies. FastAPI moves to
>=0.141(build_openapinow walksrouting.iter_route_contextsto handle the lazy included routers of FastAPI 0.137+; the generated OpenAPI documents are unchanged), and thenumpy<2.5cap is lifted (numba >= 0.67 supports numpy 2.5; the out-of-band LTX-2 install needsnumba>=0.67). - Refreshed the locked dependencies: torch 2.14, torchvision 0.29,
huggingface-hub 1.30, tokenizers 0.23.2, anyio 4.15, and ruff 0.16.6. - Updated the GitHub Actions used by CI, the PyPI release, and the Pages deploy
to their current majors.upload-pages-artifactnow sets
include-hidden-files: truesopublic/.nojekyllkeeps being published. - Raised the mise pinned in CI and the release workflow from 2026.5.0 to 2026.9.1,
the versionmise run ciis verified against locally. - Replaced the Dependabot configuration: the npm ecosystem entry pointed at the
package.jsonremoved in 0.6.0, so it is dropped in favour of theuvand
github-actionsecosystems.mlx-vlmis ignored there for the same reason it
is pinned.
Removed
- Removed the Node tooling (
package.json/ pnpm /firebase-tools) that existed only for the retired GCP relay setup task. This clears all open Dependabot alerts, which were transitive dependencies offirebase-tools.
v0.6.0
[0.6.0] - 2026-09-01
Removed
- BREAKING: Removed the relay transport and retired the
kiapi-relayandkiapi-proxypackages.kiapi run --relay, therelay-gcpextra, therelayfield of/health, and theKIAPI_RELAY_*settings are gone. To reach kiapi from other machines, expose it over your own private network layer instead — for example, runtailscale serve --bg --https=8500 8500on the kiapi machine and point clients athttps://<machine>.<tailnet>.ts.net:8500. The publishedkiapi-relay/kiapi-proxydistributions remain on PyPI as-is but will receive no further updates. (The unreleased GCP RTDB watch read-timeout fix was removed together with the relay.)
v0.5.3
[0.5.3] - 2026-07-30
Fixed
- kiapi:
kiapi service installnow preservesPATHin the launchd property list so background capabilities can find external tools such as Homebrew FFmpeg.
v0.5.2
[0.5.2] - 2026-07-29
Fixed
- Publish each workspace package with its own PyPI Trusted Publishing token so a release can upload multiple projects.
v0.5.1
[0.5.1] - 2026-07-29
Added
- kiapi:
kiapi run --relay noneexplicitly disables the relay, overriding a relay enabled in user settings. - kiapi-relay:
KIAPI_RELAY_DEFAULT=none(and otherKIAPI_RELAY_*env vars set tonone) is now parsed as unset, so the relay can be disabled via the environment.
Fixed
- kiapi-relay: Reconnect the GCP RTDB request stream when Firebase reports an expired credential or canceled subscription, instead of leaving the relay unable to receive new requests.
v0.5.0
[0.5.0] - 2026-07-11
Fixed
- kiapi / kiapi-proxy:
service installnow also pinsXDG_CACHE_HOMEin the launchd property list so the background service uses the same cache directory and single-instance lock as the interactive CLI.
v0.4.0
[0.4.0] - 2026-07-11
Added
- kiapi: Added
kiapi service showto print the installed launchd property list.
Changed
- kiapi-relay: The
gcp:setuptask no longer passes--scopesto
gcloud auth application-default login. The default ADC scopes already
includecloud-platform, which covers the GCS and RTDB access the relay
needs, so the extra scopes were unnecessary. - kiapi-relay: The
gcp:setupImpersonation method now runs
gcloud auth application-default loginitself instead of only reminding the
user to, since ADC is the base credential the impersonation chain mints SA
tokens from, and grantsroles/iam.serviceAccountTokenCreatorto the actual
ADC principal rather than the active gcloud CLI account.
Fixed
- kiapi:
kiapi service installnow pins the currentXDG_CONFIG_HOMEandXDG_DATA_HOMEvalues in the launchd property list so the background service resolves the same user settings and data directories as the interactive CLI. - kiapi: The hot-reload worker subprocess (
kiapi run --debug) now loads the user settings file in the ASGI factory, so relays configured only in user settings (for example the GCP relay'sdatabase_url/bucket) start correctly. Previouslykiapi run --relay gcp --debugfailed at startup with "required field is not set" because the reload subprocess never ranload_user_settings().
v0.3.0
[0.3.0] - 2026-07-02
Added
- Manage Node dev tooling with pnpm through a root
package.json;firebase-toolsis now a project-local dev dependency installed bymise run setup(pnpm install) rather than a globalnpm install -g, and mise puts it onPATHvianode_modules/.bin. - kiapi-relay: Added a
gcp:setupmise task (run frompackages/kiapi-relay/) that interactively provisions the GCS bucket, Firebase Realtime Database instance, and authentication forGCPRelay, then prints the kiapi YAML to paste withkiapi config edit. It uses the project-localfirebase-tools. The task verifiesfirebase-toolshas its own login (viafirebase login:list) instead of relying onfirebase projects:list, which also succeeds through the Application Default Credentials fallback and then fails the Realtime Database calls with a quota-project 403; the RTDB creation failure message now points atfirebase-debug.logand lists both the missing-login and Blaze-plan causes. The GCP relay README was rewritten around this task. - kiapi:
GET /healthnow reports the status of the relay started with the server in arelayfield (name,running,failed), ornullwhen no relay is configured. - kiapi-relay: Added a
nameattribute to theRelayprotocol, populated by the relay registry through afactory_wrapperand shared via a newBaseRelaybase class.RelayRunner.status()returns a newRelayHealthview (name,running,failed). - kiapi: Added a
requestmethod to theRelayprotocol and implemented it onLocalRelayandGCPRelay, promoting the relay request client from the verification scripts into the relay packages. Responses are returned asRelayResponse, with binary bodies materialized to a temporary file the caller owns. - kiapi-relay: Relay participants derive a stable
node_idfrom a data directory (get_or_create_node_id), and discover a target node through liveness heartbeats published underliveness/{node_id}as part of thewatchlifecycle (heartbeat_interval_s/liveness_ttl_ssettings);Relay.requestfails withno_relay_nodewhen none is fresh. - kiapi-proxy: Expanded the CLI to mirror the
kiapicommand layout so the proxy is managed independently:config(init/show/edit/template) manages a user settings file separate from kiapi's (holdingkiapi_proxy.apiandkiapi_relaysettings, loaded on every command);check --relay local|gcpsends a single request (default/health, overridable with--path) through the relay to a live kiapi node and prints the response without starting the server, so relay connectivity can be verified as a health check (it reuses the persistent relaynode_idand holds the single-instance lock likerun, failing fast if the proxy server is already running);service(install/start/status/stop/uninstall) manages a launchd user agent (io.github.kiarina.kiapi-proxy) that runskiapi-proxy run(installpins theXDG_CONFIG_HOME/XDG_DATA_HOMEvalues present at install time into the plist, since launchd does not inherit them, so the service resolves the same config/data directories as the interactive shell and can find the user settings written byconfig edit). - kiapi / kiapi-proxy: Each server resolves a persistent relay
node_idfrom its user data directory, injects it into the relay, and acquires a single-instance lock (viakiarina-utils-app, scoped to the user cache directory) so a second process cannot share the same node identity. - kiapi-relay: Initial release of
kiapi-relay, extracted fromkiapi. Provides the relay protocol, request/response schemas, the in-processRelayRunner, and the relay request client, plus the local filesystem (kiapi_relay.impl.local) and GCP (kiapi_relay.impl.gcp, available via thegcpextra) relay backends. - kiapi-proxy: Initial release of
kiapi-proxy: a proxy server that forwards incoming HTTP requests to a kiapi instance over a relay (kiapi-relay) and returns the result. Supports JSON responses, file/binary responses, and chattext/event-streamresponses re-emitted as Server-Sent Events. Ships akiapi-proxyCLI and does not depend onkiapior MLX, so it runs on Linux, Windows, and resource-constrained machines.
Changed
- Reworked
mise run verifyinto a Python driver (scripts/verify.py) that selects a target (--kiapi/--kiapi-relay/--kiapi-proxy, or fzf-interactive), starts and stops the kiapi / kiapi-proxy servers it needs (stopping and restarting the launchd services if they are running), and runs the matching verification scripts. It replaces the old per-capability andrelay*task arguments; scope with--familyand--relayinstead, or run an individualscripts/capabilities/verify_*.pydirectly. Capability verify output now honoursKIAPI_VERIFY_DIR(default.verify), and the driver routes artifacts to.verify/kiapivs.verify/kiapi-proxyso direct and proxied runs no longer collide. The relay verify scripts were consolidated intoscripts/relay/verify_local.pyandscripts/relay/verify_gcp.py(which gain the core files/jobs CRUD checks); the capability-heavyscripts/relay/verify.pywas removed, since capabilities are now covered through the proxy path. TheMakefileverify targets are nowverify/verify-fast/verify-kiapi/verify-kiapi-relay/verify-kiapi-proxy. - Moved development and CI operations from Make recipes into package-aware mise tasks, including the new setup task and namespaced test-assets download task.
- Manage the repository version in a single root
VERSIONfile and release the whole workspace under one shared version. The release pipeline now detects packages with unreleased changelog entries, bumps and publishes only those, and is triggered by a singlev<version>tag instead of per-package<package>-v<version>tags. - kiapi-relay: The relay
node_idis now generated automatically and persisted per data directory instead of being configured. The manualnode_id/source_node_idsettings were removed from the local and GCP backends; clients discover a target node through liveness heartbeats and address responses with their own generatednode_id. - kiapi: Simplified the
Relayprotocol to a singlewatchmethod and moved listener tasks and the HTTP client into thewatchlifecycle, removing the explicitclosemethod. - kiapi: Reworked the relay verification scripts to issue requests through
Relay.requestvia the relay registry factories, removing the duplicated transport client inscripts/relay/_client.py. - kiapi: Converted the repository into a uv workspace and moved the
kiapipackage topackages/kiapi/with asrc/layout. Packaging and lint/test paths are now per-package. - kiapi: Extracted the relay subsystem into a separate
kiapi-relaypackage.kiapi.core.relayis nowkiapi_relay, andkiapi.relay.{local,gcp}are nowkiapi_relay.{local,gcp}. Therelay-gcpextra now pullskiapi-relay[gcp]. - kiapi / kiapi-proxy: User-directory resolution and single-instance locking are delegated to the shared
kiarina-utils-apppackage and used directly (kiarina.utils.app) rather than through acore.appre-export layer, removing the privateAppSettings/user-directory copy and the directplatformdirsdependency. kiapi'score.appnow provides only theAppContextschema; kiapi-proxy'score.appmodule was removed. Each server sets the app identity by callingkiarina.utils.app.configure(...)at its CLI entry point (and, for kiapi, the ASGI factory used by hot reload). The directory getters returnpathlib.Path. For kiapi, the usersettings.yamlsection for these settings moves fromkiapi.core.apptokiarina.utils.app; the override environment variables areKIARINA_UTILS_APP_.
v0.2.0
[0.2.0] - 2026-06-26
Added
- Added an optional plugin-based remote job relay with Firebase Realtime Database notifications, GCS request/response payloads, in-process ASGI dispatch, committed-response recovery, and atomic terminal delivery.
- Added a filesystem-backed LocalRelay for local relay verification without GCP services.
- Added relay verification scripts for LocalRelay, GCPRelay, and end-to-end LocalRelay capability checks.
Changed
- Moved GCP relay dependencies into the optional
relay-gcpextra. - Updated the minimum
pydantic-settings-managerdependency to 3.7.0.
Fixed
- Added the OAuth scopes required by the Firebase Realtime Database REST API to GCP relay credentials.
- Added relay support for multipart file uploads such as
POST /v1/files.