Repository navigation
[0.8.0] - 2026-09-26
Dependencies
- mflux 0.19.1 -> 0.20.0, pinned to a kiarina fork commit (
144a6bec) that adds
mflux-community/mflux#741 on top of 0.20.0. 0.20.0 also fixes the seedvr2
mx.repeatfailure.
Removed
- chat: removed the
qwen3.6-27bmodel (aliasqwen3.6).qwen3.8-27buses the
same handler, modalities and memory footprint;vlmselectsqwen3.8-flash-next.
Remove the downloaded weights withkiapi deactivate --repo mlx-community/Qwen3.6-27B-4bit.
Changed
- chat:
modelis now required on/v1/chat/completions; there is no default
chat model. Requests without it return HTTP 422. The chat models differ in
accepted input modalities (onlyqwen3-omnitakes audio/video) and memory
footprint (qwen3.8-flash-nextkeeps about 74 GiB resident), so an implicit
default could silently evict other models or reject media. - chat: the aliases
vlmandqwen3.8now selectqwen3.8-flash-next
(previouslyqwen3.8-27b).qwen3.8-flash-nextalso answers to
qwen3.8-flashandflash-next; itsqwen4_expalias was removed.
qwen3.8-27bkeepsqwen3_5andqwen3-vl. - BREAKING: chat: removed the
max_tokens_capsetting
(KIAPI_CHAT_MAX_TOKENS_CAP, 4096).max_completion_tokensis no longer
capped by the server; generation stops atmax_completion_tokensor when the
prompt plus the output fills the model's context window, whichever comes first. - chat: the default
max_completion_tokensis now 1024 (was 512). - chat: non-streaming requests now run through
stream_generatelike streaming
ones, so the context window bound applies to both.
Added
-
Web UI at
/: a dashboard (server, queue, memory, setup), model setup
status with copyablekiapi activatecommands, jobs, files with previews, and a
guide and API reference for every family, in light and dark themes. Every
family can be run from its Playground: forms are built from the family's
OpenAPI schema, generation runs as a job with live progress, inputs come from
uploads or stored files, and "Write with chat" drafts prompts with a chat model
that reads the family guide. Chat has its own conversation view with every
request parameter (tools, tool choice, parallel tool calls, token limits,
sampling, template kwargs, streaming on or off), tool-call results, and token
usage. A "?" beside each field shows its full description, and a floating
assistant answers questions about the current family from its OpenAPI document
and, when asked, fills in the form through tool calls (it never submits). -
GET /v1/setup: every model with its setup state, likekiapi status. -
qwen: added the
image-2.1model (Qwen/Qwen-Image-2.1, aliases
qwen-image-2.1,qwen-2.1). One resident model serves both/generate
(text-to-image) and/edit(up to 10 reference images), outputs RGBA, and
renders up to 2752 px. Defaults are 40 steps, guidance 1.0 and q8. It does not
takeinit_imageorloras. Its weights are under the non-commercial Qwen
Research License. Editing comes from the pinned mflux fork (144a6bec,
mflux-community/mflux#741); installs with official mflux do not register it. -
qwen:
jpegoutput flattens transparent areas onto white. -
chat: Qwen3.8-Flash-Next reuses unchanged history when images are appended,
like Qwen3.8-27B, through the pinned mlx-vlm fork (6581ba8c). -
chat: added the
qwen3.8-flash-nextmodel (mlx-community/Qwen3.8-Flash-Next-4bit,
aliasesqwen3.8-flash,flash-next,qwen4_exp). It loads through a view that
memory-maps its n-gram (PLE) table, which keeps about 80 GB resident instead of
111.5 GB, and raises the open-file limit the mapped table needs. -
chat: Omni reuses unchanged prefixes when images, audio clips or videos are
appended, including demuxed audiovisual inputs, through a pinned mlx-vlm fork. -
chat: the pinned engine supports multiple audio clips with independent feature
extraction/encoding and corrected CNN lengths and chunk masks. -
chat: Qwen3.8 reuses unchanged image history when new images are appended,
using a pinned mlx-vlm fork in uv-managed checkouts. Official-engine installs
retain the conservative whole-request fallback. -
chat: client disconnects now cancel queued work or stop running generation at
the next token boundary. Cancellation uses the existingcanceledjob state
and safely clears APC state when a generator closes early. -
chat: responses report APC reuse through the OpenAI-compatible
usage.prompt_tokens_details.cached_tokensfield. Streaming requests support
stream_options.include_usageand emit usage before[DONE]when requested. -
chat: bounded, memory-only automatic prefix caching for Qwen3.6 / Qwen3.8
text and images, and Qwen3-Omni text, image, audio, video, and image + video.
Media content hashes and video options protect cache identity, and Omni
restores complete-prompt positions when reusing media prefixes. -
chat:
GET /v1/modelsreturns each model'scontext_window, read from the
model'sconfig.json(nulluntil the model is set up).
Fixed
-
chat: suppress extra streamed tool names when
parallel_tool_calls=false. -
memory: include idle chat caches in cross-model eviction and transient reservations,
and derive active cache headroom from the configured APC capacity. -
chat:
finish_reasonis now"length"when generation stops at
max_completion_tokensor the context window. It was always"stop".