Releases: bisand/kvad
Release list
kvad 0.11.0
Video with sound, from a prompt: LTX-2.5
kvad now makes video with sound, with Lightricks' LTX-2.5 (22 B parameters): a prompt in, an MP4 out, through /v1/videos, kvad videos make or the new Videos page. It is LTX-2.5's distilled model in two stages, generated at half size and refined at full size, as the reference pipeline does, and every part was checked against Lightricks' own code on the same weights. (#53, #54, #57, #68, #70, #79, #90)
kvad videos make "a red fox trots through fresh snow" --size 768x512
Measured on an M5 Pro (48 GB), 121 frames (5 s), with the default decoder:
| Clip | Total | Peak memory |
|---|---|---|
| 768×512 | 77.6 s | 26.9 GB |
| 1536×1024 | 391 s | 31.4 GB |
/v1/videosfollows OpenAI's job shape. A request returns a queued video at once, and its status, progress, MP4 and thumbnail are fetched as OpenAI's SDK fetches them. Progress also streams as server-sent events. (#90, #93)- Start from a picture: OpenAI's
input_referencebecomes the first frame. (#94) - The model chooses the length from the prompt when a request gives none, with LTX-2.5's own duration head. (#95)
- Guidance:
steps,guidance_scaleor a negative prompt run LTX-2.5's dev model, with the reference's guidance, instead of the distilled one. Its files come fromkvad pull Lightricks/LTX-2.5 --dev(51 GB). (#97) - The diffusion decoder is the default, as in the reference. The convolutional one is still there with
--decoder conv. (#98) - 48 and 96 fps with DFR, the reference's production pipeline: generated keyframes, a detailing pass, and temporal rounds that double the frame rate once or twice. It runs whenever more than 30 fps is asked for, or with
--pipeline dfr. 768×512 × 241 frames at 48 fps took 260 s, and × 481 at 96 fps 621 s. Its detailing LoRA is in a separate gated repo, fetched by the first DFR request. (#101) - While it works: the Videos page shows a rough preview and the time left. Finished videos are compressed with
ffmpegwhen it is installed. (#91, #92) - Admission charges a video's peak memory, and sizes are capped at what has been measured. On a Mac, a Metal allocation past memory restarts the machine.
Made fast on the M5's neural accelerators
LTX-2.5's attention, projections and 3D convolutions run on the M5's matrix units, as FLUX's projections did in 0.10.0. (#84–#89)
- Attention is level with MLX's. (#86, #87)
- The q8 projections run at 21–25 TFLOP/s, MLX's rate. (#89)
- The convolutional decoder decodes a 1536×1024 clip in 31 s, where it took 278 s. (#88)
- Small element-wise ops are fused into single Metal kernels. (#85)
On other Macs these kernels give way to candle's. LTX-2.5 has only been run on an M5 Pro.
Before the first video
Lightricks/LTX-2.5is a gated repo: accept its terms on Hugging Face, with a token saved on the machine.- A first load fetches the files a generation reads, about 73 GB, and builds its q8 caches once.
The community's GGUFs of the image and video models
GGUF files of Qwen-Image's, FLUX.1-schnell's and LTX-2.5's denoisers load by name: the repo and the quantisation together. (#102)
kvad pull city96/Qwen-Image-gguf:Q4_K_S
kvad images make "a lighthouse at dusk" --model city96/Qwen-Image-gguf:Q4_K_S
- The base supplies the rest. The model card's
base_modelsupplies the text encoders, VAE and configs, and a pull fetches those, never the base's own transformer. - Unsupported bases cost almost nothing. The base is checked before anything large downloads, so a GGUF of a model kvad does not run costs a model card.
- Deleting a GGUF deletes that one file.
- FLUX's files keep Black Forest Labs' layout, and are read through a map of rows, so no block is quantised twice.
- The Q4_K, Q5_K and Q6_K blocks run on the M5's matrix units, in the same kernel as kvad's own q8.
| Model | kvad's q8 | GGUF | Memory saved by the GGUF |
|---|---|---|---|
| Qwen-Image Q4_K_S, 1024², a step | 7.10–7.61 s | 8.31–8.41 s | 9.1 GB |
| FLUX.1-schnell Q4_K_S, 1024², a step | 7.01–7.93 s | 7.56–7.77 s | 5.5 GB |
| LTX-2.5 Q4_K_M, 768×512 × 121, the denoise | 51.1–61.6 s | 59.0–62.1 s | about 4 GB |
A GGUF of LTX-2.5 runs the fast pipeline only: DFR and guidance need kvad's own DiT.
Fixes
- After a pull,
kvad pullsuggestskvad images makeorkvad videos makefor an image or video model, where it always saidkvad run. (#102)
For developers
- The plans:
docs/video-plan.mdanddocs/gguf-plan.mdrecord each step, and what it measured against the reference. - Reference checks:
scripts/ltx-fixtures.pywrites the reference outputs. Thecrates/gpu/examples/ltx_*examples check each part against them, andggufchecks a GGUF against its base. - Profiling:
ltx_costandconv3d_costmeasure where a generation's time goes. (#57, #84) - Tests: the attention kernel's tests are seeded, and no longer fail now and then on Metal. (#96)
- Build: the non-macOS stubs match their kernels again (#99), and the batched matmul no longer warns off aarch64 (#100).
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.10.1
Models load without doubling their memory
Loading a model used to hold its weights twice for a while, and on macOS that made the first image or token slow. (#80)
- FLUX and Qwen-Image checkpoints are read in the compute dtype. They used to be read as f32 and then cast on the GPU. A bf16 FLUX load of the text encoders alone went from 32.5 GB to 10.4 GB, and a whole bf16 FLUX pipeline now fits in 48 GB.
- Image checkpoints are read past the page cache. Memory-mapping them made macOS compress the weights it had just loaded. Dense FLUX now loads in 44–48 s instead of 94 s. The first image's encode takes 0.33 s instead of about 30 s, and its first step 1.96 s instead of 31–39 s.
- The quantised-weight cache is read the same way. A q8 load of Qwen-Image went from 71 s to 5.1 s, FLUX from 22 s to 3.1 s, and Qwen3-14B from 21 s to 3.2 s.
- Large load buffers stay out of macOS's allocator cache. Qwen3-14B's q8 load peaked 7.35 GB above what it holds once loaded. It now peaks 0.84 GB above.
Output is unchanged: greedy LLM output and a FLUX render are identical before and after.
An icon, a wordmark, and the name in runes
Kvad has an icon: an open book seen from its end, with its covers making a V. The web UI uses it as its favicon and home-screen icon, and the sidebar shows the new wordmark. Hover over the name and it turns into ᚴᚢᛅᛏ, kvad in the younger futhark, a letter at a time. The sources are in docs/brand. (#82)
For developers
Pull requests are now built and tested on macOS by .github/workflows/check.yml. It uses a ci build profile, which is release optimisation without thin LTO, so the tests link in seconds rather than minutes. (#81)
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.10.0
Image models and prefill use the M5's neural accelerators
candle's Metal kernels never reach the matrix units in each M5 GPU core. kvad now reaches them through Metal 4's tensor API.
- FLUX.1-schnell at q8, 1024²: 12.9–13.7 → 8.7–9.2 s per step, about 1.5× faster. The same kernel runs Qwen-Image's q8 weights. (#58)
- SDXL in f16, 1024²: 3.94 → 3.12 s per step, 1.26× faster. The convolutions are unchanged. (#59)
- bf16 prefill: a 2,079-token prompt on Qwen2.5-1.5B-Instruct went from 0.99 to 0.39 s, 2.5× faster, with the same output. Decode is unchanged, since a single row is still faster on candle's kernel. (#59)
This needs an M5 (Apple GPU family 10) and a macOS with Metal 4. Other Macs run as before. KVAD_GPU_MPP=0 turns it off.
Faster text generation
Measured on Qwen2.5-1.5B-Instruct on an M5 Pro:
- Sampling at a temperature above 0 used to sort all 151,936 logits for every token. It now picks the top k: q8 decode went from 96.6 to 120.7 tok/s. The same seed can pick a different token only where two logits tie exactly. (#61)
- Streaming text used to decode the whole conversation twice per token to find the new text. It now decodes a short window: 1 µs instead of 3 ms per token at 32k context, and 10% faster decode after an 8.4k-token prompt. (#62)
- The GPU attention cache grows in place instead of being copied whole every token: q8 decode is 11% faster at 2k context and 32% faster at 8.4k. (#63)
- Quantised models keep their attention cache in f16 on Metal, which halves what attention reads: another 12% at 8.4k context. Perplexity moved by less than 0.02%, and greedy output is unchanged.
KVAD_GPU_KV_F32=1keeps the f32 cache. (#64) - A decode step's small ops are fused into single Metal kernels, and the query, key and value projections run as one matmul, as do gate and up. A layer drops from 25 kernels to 10. q8 decode went from 114–117 to 126–129 tok/s and bf16 from 62–64 to 68–69. bf16 prefill of 1,812 tokens went from 0.36 to 0.31 s. Greedy output at q8 is unchanged. In bf16 it can drift after a few dozen tokens from rounding, which is no less accurate against an f32 run.
KVAD_GPU_FUSED=0uses candle's ops. (#66)
Large models gain less: a token of Qwen3-14B is dominated by reading 15 GB of weights.
The server charges GPU models for the memory they use
- The attention cache is charged at its real width. GPU models keep 2 bytes a number (bf16, or f16 for quantised models on Metal), but were charged 4. Qwen3-14B at 32k context is now charged 5.4 GB instead of 10.7 GB, so more models fit side by side. (#65)
- A recurrent state is charged at its real width too: bf16 for a bf16 Qwen3.5 or Qwen3-Next model. (#67)
For developers
New examples in crates/gpu/examples measure where time goes on Metal:
neural_accelcompares the accelerated matmuls against candle's.decode_costanddecode_breakdownsplit a decode step into its parts.kv_precisionmeasures what the f16 cache changes.
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.9.0
Change the server's settings from the web UI
Settings has a new Server section for administrators. It edits [server] in kvad.toml:
- load the active model at start
- load a model a request names
- memory for models
- context charged per model
- listen address
A save rewrites only the lines that changed and keeps comments. It's refused, with nothing written, when the result is a file the server couldn't start from: an open address with no auth, an address that's taken, or a layout the page can't edit safely. The page lists the saved settings that aren't in force yet and offers Restart now. The server restarts in place and keeps its process id. (#49)
Move the data from the web UI
A new Data section shows where the database, images, datasets, trained models and Hugging Face models are, and how much space each takes. Move the data… moves all of it somewhere else and restarts the server into the new place:
- Checked as you type: the page shows what will be renamed (same disk, instant) and what will be copied (another disk), and how much space is free. It refuses a target that isn't empty, overlaps the data, or doesn't have room.
- On another disk, the copy runs while the server keeps serving, with progress and a Cancel button. At the end, with the database locked for a moment, it copies the database, catches up anything that changed during the copy, writes
[data] dirtokvad.tomland restarts. - The old copy is kept until you delete it from the page. The confirmation lists exactly what goes, and it never touches anything the server is using.
- Hugging Face models come along unless
HF_HUB_CACHEorHF_HOMEis set. Your Hugging Face token stays where it is. (#50)
install.sh --data-dir still works for setting the data dir from the command line.
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.8.0
Loading a downloaded model no longer needs the internet
Loading a model already in the Hugging Face cache used to ask the Hub about every file first, even though the files were cached. When huggingface.co was unreachable, each file waited about 3 minutes for retries to run out, and a restart could load nothing for ten minutes.
- A cached model now loads from the cache without contacting the Hub. This covers the server's loads,
kvad --localandkvad-gpu, and applies to both single-file and split weights. kvad pullstill asks the Hub, so it still brings a cached model up to date. Loading no longer checks for a newer version.- The image-model check reads the cache too. (#47)
Changed default: [server] autoload is now false. A restart loads nothing until a request asks for a model. To keep loading the default model at startup, put this in kvad.toml:
[server]
autoload = trueA chosen data directory holds the Hugging Face models too
Setting [data] dir (new in 0.7.0) or KVAD_DATA_DIR now also moves models pulled from Hugging Face, into huggingface/hub inside it. Previously it covered the database, images, datasets and trained models, but not the largest thing on disk.
- Without a data dir set, models stay in
~/.cache/huggingface/hubas before, so upgrading doesn't strand an existing cache. HF_HUB_CACHEorHF_HOMEstill win when set.- The layout inside the data dir matches the default, so an existing cache moves with one
mv.install.sh --data-dirprints the command. - Your Hugging Face token stays where it is, so gated models keep working.
- The server's startup banner shows a
hub cacheline. (#48)
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.7.0
Generated pictures show the right image
Every new picture could show the first one ever made, and deleting a picture seemed to leave it on disk. Both had one cause. Deleting the newest image freed its id, the next image got that id back, and the browser kept showing its year-long cached copy of the old picture at that URL.
- Ids are never reused. A database migration (schema 9) makes the images table's ids only ever count up. Existing pictures keep their ids and files.
- Image links include when the picture was made,
/api/images/<id>.png?v=<time>. This gets past copies browsers already have cached, and past ids that start again at 1 after a data directory is wiped. (#45)
Choose where kvad keeps its data
The database, generated images, datasets and trained models can now live somewhere other than ~/.local/share/kvad:
# ~/.config/kvad/kvad.toml
[data]
dir = "/Volumes/Models/kvad"- The CLI and the server both read it, so they agree.
kvad-servetakes it from the file it was started with, including a--configone. ~/and paths relative to the file work.KVAD_DATA_DIRoverrides it for one process.install.sh --data-dir DIRwrites the setting and leaves the rest of the file alone.- Changing it moves nothing. The installer tells you where the old data is.
- The server's startup banner has a new
dataline. - Models pulled from Hugging Face aren't part of this. They stay in the Hub's cache, which
HF_HOMEmoves. (#46)
A second install leaves the service alone
install.sh --prefix DIR used to treat the background service as its own. It stopped the running server, and with --yes it repointed the service at DIR/kvad-serve. --uninstall --prefix DIR removed the service entirely.
Now an install or uninstall into a directory the service doesn't run from leaves it alone. That makes it safe to put a second copy beside the real one. --service moves the service to the new copy on purpose. (#46)
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.6.1
macOS builds are Apple Silicon only
There's no Intel Mac build from this release on. Across nine releases the Intel tarball got no downloads. It was cross-compiled and never run before release, and its Metal path was never tested on an Intel Mac's GPU.
- On an Intel Mac,
install.shnow stops with a clear error instead of failing on a missing download. v0.6.0 is the last release with an Intel build:sh -s -- --version v0.6.0still installs it. - Under Rosetta on Apple Silicon: a Terminal running under Rosetta reports
x86_64, soinstall.shused to install the Intel build there without saying so. It now installs the Apple Silicon build. - Release checks: all four tarballs are now run on their build machine before release. The Intel one was the only target that skipped this. (#32)
Nothing else has changed since 0.6.0.
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.6.0
FLUX.1-schnell
A third image model, alongside SDXL and Qwen-Image from 0.5.0: FLUX.1-schnell from Black Forest Labs. It makes a 1024² image in four steps and can spell. (#31)
kvad pull black-forest-labs/FLUX.1-schnell
kvad images make 'a wooden sign in the snow that says "FLUX"' --model black-forest-labs/FLUX.1-schnellThe repository is gated. Accept the licence on its Hugging Face page, then save a token for that account to ~/.cache/huggingface/token (mode 600) so that the CLI and the service can download it. Saving the token from the web UI is planned in #30.
What to expect (M5 Pro, 8-bit, single runs)
| per step | decode | |
|---|---|---|
| 1024², 4 steps | 13.1 s | 9.6 s |
| 768², 4 steps | 7.3 s | 5.3 s |
It needs about 18.3 GB of memory: 12.6 GB for the transformer and 5.1 GB for the T5 text encoder, quantised on the first load and cached after that. The download is about 34 GB.
How it is built
FLUX's first 19 transformer blocks are the same as Qwen-Image's, so that block now lives once, in mmdit.rs, and both models load it. What FLUX adds is a T5-XXL text encoder and 38 single-stream blocks. docs/image-plan.md has every shape, read off the checkpoint.
Also
- Guidance on a model that doesn't take it is refused up front. Schnell is distilled to run without guidance. A request that sets
guidance_scaleornegative_promptfor such a model now gets an immediate 400, instead of failing after the model has run. The Images page greys both fields out for it. FLUX.1-dev, which needs guidance, is refused by name until it is implemented. - SDXL and Qwen-Image produce the same images as in 0.5.0.
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.5.0
Images from text
kvad can now make pictures as well as text. Two pipelines run on the GPU: Stable Diffusion XL and Qwen-Image. Each is written out by hand in kvad-gpu: both CLIP encoders, the UNet, both VAEs, the MMDiT with its three-axis RoPE, both schedulers, and a PNG encoder. Nothing comes from candle-transformers. docs/image-plan.md records every shape, read off the checkpoints, and explains why there is no CPU path. (#8, #29)
kvad pull stabilityai/stable-diffusion-xl-base-1.0
kvad images make "a lighthouse on a cliff at dusk, oil painting" --out lighthouse.png
kvad images # everything made so far
kvad images rm 3On the server
POST /v1/images/generationstakes OpenAI's request shape, plussteps,guidance_scale,seedandnegative_prompt.stream: truesends each denoising step, with a preview whenpartial_imagesorpreviewasks for one.- Every image is kept, with the settings that made it (seed included), in
images/under the data directory./api/imagesis the gallery. Database migration 008 runs on its own. - A new Images page in the web UI has a prompt, the settings, a live preview and progress bar, Stop, and a gallery with Reuse and Delete.
/v1/modelssays whether each model ischatorimage. A language model asked for an image is refused before anything is loaded, and so is an image model asked to chat.
What to expect (M5 Pro, single runs)
| per step | decode | |
|---|---|---|
| SDXL, f16, 1024², 30 steps | 3.9 s | 9.4 s |
| Qwen-Image, q8, 1024², 20 steps | 28.3 s | 5.8 s |
Qwen-Image runs quantised. It needs about 30 GB (21.7 GB for the transformer at q8, 7.5 GB for the text tower), so on a 48 GB machine it is the only thing loaded. The first load quantises and caches it; loads after that take about a minute. Its download is 58 GB.
macOS only. Images are made on the GPU, and the Linux builds have no GPU backend.
Also
- A fix for every quantised GPU model: candle's Metal quantised matrix multiply ignored where a sliced input started. Nothing shipped before this release was affected, but Qwen-Image was, and
Proj::forwardnow guards against it for every model.
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.
kvad 0.4.1
Prompt suites work with instruct models again
Every prompt-suite case run against an instruct model came back with an empty answer and failed. This happened on every backend, while chat with the same model worked. Suites sent each prompt as raw text, and an instruct model reads a bare "Capital of France?" as a finished user message: the first token it chooses is end-of-turn, so it writes nothing.
A case is now asked through the model's own chat template, the same way chat asks it. Base models have no template and still get the prompt as written. Each case still starts from an empty cache, so a suite gives the same answers every time it runs. (#26)
A model that writes nothing says so
- Evals: a case the model answered with nothing shows
<the model wrote nothing: it ended its turn before the first token>instead of an empty answer. The note never counts as an answer, so it can't make a case pass. - Benchmarks: a round that generates nothing now fails the benchmark instead of being kept as a 0 tok/s sample in the median.
Worth knowing: benchmarks still send their prompt as raw text. If a benchmark prompt is a question, an instruct model will now fail the run with a message saying so. Use text to continue instead, such as the start of a sentence. (#27)
Also
- The chat endpoint is now tested against an instruct model trained inside the test. The tests cover the template, stopping at the model's end token, and reuse of the cache between requests. (#28)
Install
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh
It asks before putting anything on your PATH or starting anything at
login, and if the server should start at login it asks what address to
listen on, offering 127.0.0.1:5823. Name one it cannot serve from — an
address something already holds, or a non-loopback one, which
kvad-serve refuses while no auth mode is configured — and it says why
and asks again, rather than installing a service that cannot start.
For a machine with nobody watching, --yes never opens a terminal and
answers every question the quiet way: binaries only, nothing added to
your shell and nothing started. Ask for the rest explicitly, --bind
included, since a run that cannot stop to ask would otherwise take 5823:
curl -fsSL https://raw.githubusercontent.com/bisand/kvad/master/install.sh | sh -s -- --yes --service --bind 127.0.0.1:8123
--prefix DIR, --version vX.Y.Z and --uninstall do what they look
like; --uninstall takes the binaries and the service and leaves your
models and conversations alone. sh install.sh --help lists the rest.
macOS ships kvad, kvad-serve, kvad-tui and kvad-gpu. Linux ships
kvad and a CPU-only kvad-serve; the other two link Metal.
Downloads can be checked against SHA256SUMS below.