Skip to content

v0.8.0

Latest

Choose a tag to compare

@github-actions github-actions released this 25 Sep 21:14
· 20 commits to main since this release
v0.8.0
f32d8fa

[0.8.0] - 2026-09-26

Dependencies

  • mflux 0.19.1 -> 0.20.0, pinned to a kiarina fork commit (144a6bec) that adds
    mflux-community/mflux#741 on top of 0.20.0. 0.20.0 also fixes the seedvr2
    mx.repeat failure.

Removed

  • chat: removed the qwen3.6-27b model (alias qwen3.6). qwen3.8-27b uses the
    same handler, modalities and memory footprint; vlm selects qwen3.8-flash-next.
    Remove the downloaded weights with kiapi deactivate --repo mlx-community/Qwen3.6-27B-4bit.

Changed

  • chat: model is now required on /v1/chat/completions; there is no default
    chat model. Requests without it return HTTP 422. The chat models differ in
    accepted input modalities (only qwen3-omni takes audio/video) and memory
    footprint (qwen3.8-flash-next keeps about 74 GiB resident), so an implicit
    default could silently evict other models or reject media.
  • chat: the aliases vlm and qwen3.8 now select qwen3.8-flash-next
    (previously qwen3.8-27b). qwen3.8-flash-next also answers to
    qwen3.8-flash and flash-next; its qwen4_exp alias was removed.
    qwen3.8-27b keeps qwen3_5 and qwen3-vl.
  • BREAKING: chat: removed the max_tokens_cap setting
    (KIAPI_CHAT_MAX_TOKENS_CAP, 4096). max_completion_tokens is no longer
    capped by the server; generation stops at max_completion_tokens or when the
    prompt plus the output fills the model's context window, whichever comes first.
  • chat: the default max_completion_tokens is now 1024 (was 512).
  • chat: non-streaming requests now run through stream_generate like streaming
    ones, so the context window bound applies to both.

Added

  • Web UI at /: a dashboard (server, queue, memory, setup), model setup
    status with copyable kiapi activate commands, jobs, files with previews, and a
    guide and API reference for every family, in light and dark themes. Every
    family can be run from its Playground: forms are built from the family's
    OpenAPI schema, generation runs as a job with live progress, inputs come from
    uploads or stored files, and "Write with chat" drafts prompts with a chat model
    that reads the family guide. Chat has its own conversation view with every
    request parameter (tools, tool choice, parallel tool calls, token limits,
    sampling, template kwargs, streaming on or off), tool-call results, and token
    usage. A "?" beside each field shows its full description, and a floating
    assistant answers questions about the current family from its OpenAPI document
    and, when asked, fills in the form through tool calls (it never submits).

  • GET /v1/setup: every model with its setup state, like kiapi status.

  • qwen: added the image-2.1 model (Qwen/Qwen-Image-2.1, aliases
    qwen-image-2.1, qwen-2.1). One resident model serves both /generate
    (text-to-image) and /edit (up to 10 reference images), outputs RGBA, and
    renders up to 2752 px. Defaults are 40 steps, guidance 1.0 and q8. It does not
    take init_image or loras. Its weights are under the non-commercial Qwen
    Research License. Editing comes from the pinned mflux fork (144a6bec,
    mflux-community/mflux#741); installs with official mflux do not register it.

  • qwen: jpeg output flattens transparent areas onto white.

  • chat: Qwen3.8-Flash-Next reuses unchanged history when images are appended,
    like Qwen3.8-27B, through the pinned mlx-vlm fork (6581ba8c).

  • chat: added the qwen3.8-flash-next model (mlx-community/Qwen3.8-Flash-Next-4bit,
    aliases qwen3.8-flash, flash-next, qwen4_exp). It loads through a view that
    memory-maps its n-gram (PLE) table, which keeps about 80 GB resident instead of
    111.5 GB, and raises the open-file limit the mapped table needs.

  • chat: Omni reuses unchanged prefixes when images, audio clips or videos are
    appended, including demuxed audiovisual inputs, through a pinned mlx-vlm fork.

  • chat: the pinned engine supports multiple audio clips with independent feature
    extraction/encoding and corrected CNN lengths and chunk masks.

  • chat: Qwen3.8 reuses unchanged image history when new images are appended,
    using a pinned mlx-vlm fork in uv-managed checkouts. Official-engine installs
    retain the conservative whole-request fallback.

  • chat: client disconnects now cancel queued work or stop running generation at
    the next token boundary. Cancellation uses the existing canceled job state
    and safely clears APC state when a generator closes early.

  • chat: responses report APC reuse through the OpenAI-compatible
    usage.prompt_tokens_details.cached_tokens field. Streaming requests support
    stream_options.include_usage and emit usage before [DONE] when requested.

  • chat: bounded, memory-only automatic prefix caching for Qwen3.6 / Qwen3.8
    text and images, and Qwen3-Omni text, image, audio, video, and image + video.
    Media content hashes and video options protect cache identity, and Omni
    restores complete-prompt positions when reusing media prefixes.

  • chat: GET /v1/models returns each model's context_window, read from the
    model's config.json (null until the model is set up).

Fixed

  • chat: suppress extra streamed tool names when parallel_tool_calls=false.

  • memory: include idle chat caches in cross-model eviction and transient reservations,
    and derive active cache headroom from the configured APC capacity.

  • chat: finish_reason is now "length" when generation stops at
    max_completion_tokens or the context window. It was always "stop".