Releases: gpillon/ignis
Release list
v0.5.0
ignis v0.5.0
Changes
- Thinking budget: a forced close that answers, and the Qwen3.8 model card's sampling by default — ends the agent tool loops where every turn was forced and the model deferred with one more tool call (#297)
- hq banded prefill: a conversation past 262K keys no longer dies of a 64-byte workspace under-reservation; a failed chunk partway through a prompt ends its request at once, and a streaming engine error is an OpenAI error chunk (#296)
- Responses API: streaming, function tools and WebSocket mode (#282)
- Playground: talks to ignis over the Responses WebSocket, so parallel agents pass the browser's six-connection cap (#283)
- Chat surface: every OpenAI field honoured, refused or declared inert — plus
stop,max_completion_tokensandcached_tokens(#284) - Server:
POST /v1/tokenizeand/v1/detokenize— the prompt counted without being served (#285) - Kernel: an A16 NVFP4 call of 64+ columns runs on the tensor cores (the dflash2 drafter's context append; TTFT −12.8%)
- Core: a request cancelled while evicted takes its snapshot out of the host tier
- Playground: the Decide tab takes a file as evidence; streaming text lands once per frame for every session
- Tests: temp files a test writes carry the pid (#198); a contributed entailment fixture for
noul
Upgrade notes
- Sampling defaults changed. A chat completion or
/v1/responsesrequest that leaves a sampling field unset now takes the Qwen3.8 model card's value for its mode, field by field: thinkingtemperature1.0,top_p0.95,top_k20;enable_thinking: false0.7 / 0.80 / 20 withpresence_penalty1.5. It was greedy (temperature: 0). - Greedy is now
temperature: 0, sent explicitly; its unset fields are neutral, and a non-neutraltop_p/top_k/penalty sent with it is still a 400. A request that sendstop_portop_kwithouttemperatureis now served instead of refused. - A request with no
seeddraws a fresh one, so two identical seedless requests are two independent samples. It was seed 0. - The forced close a thinking budget emits is now "My thinking time is over. I must now write the complete final answer from what I already have, without calling any more tools." (it was the model card's "Considering the limited time by the user, …").
/v1/responsesechoestemperatureandtop_pas the decimals they were written as (0.95, not0.949999988079071).
Work in progress (issues still open)
- NVFP4 34816x5120 W4A4 at narrow prefill widths: three-stage schedule, weight-code L2 promotion, and a 256-token TMA floor (#223)
- Tools:
tool_choice: "required"and a named function — forcing the call with a permitted-set opener (#286) - Structured output:
response_format/ a grammar — the permitted set needs a second form (#287) - Research: teacher-forced scoring as an API — how likely is this text, with nothing generated (#295)
Follow-ups to issues an earlier release already carried
- Server: three specs for the OpenAI surface — the fields, the count, the forced call (#101)
- Server: spec corrections — the budget already forces and releases, and the real clients decide the table (#121)
- Playground: the conversation goes over the Responses WebSocket (#283) (#220)
Decisions
- ADR 0017 — Opt-in Prometheus metrics with zero inference-path work (amended)
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.5.0
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.4.0...v0.5.0
v0.4.0
ignis v0.4.0
Changes
- Memory: retained slots on the host —
--retained-host(default 16) and--retained-device(default 0); the KV pool gains ~42% (475K -> 677K tokens at the default load) at level throughput (#281) - Make: the default context is 524288 tokens, with
ROPE_SCALING=yarn:2to match - Decisions:
locateover very long texts — heads narrow, a labelledchoicedecides;kind,method,compression,found(spec 22) (#278) - Playground Decide tab: locate + examples (#277)
- Decisions:
locatephase A — calibrate the heads that point into text, and decide go/no-go (#274) - /v1/decide: a parts
statenever reuses its prefix — static text before an image is re-prefilled on every request and every question (#270) - Playground: the Monitor says where the retained slots live (#281)
Upgrade notes
--retained-slots/IGNIS_RETAINED_SLOTSare removed and refuse the start: size the retained slots with--retained-device(VRAM) and--retained-host(pinned host memory). The Makefile'sRETAINED_SLOTSknob becameRETAINED_DEVICEandRETAINED_HOST.- The default 16 host slots pin ~3.5 GiB of host RAM at load. Windows counts it as the process's shared GPU memory, beside the KV-RAM arena. A card with VRAM to spare can keep today's device copies with
--retained-device 16 --retained-host 0. makenow starts atMAX_CONTEXT=524288 ROPE_SCALING=yarn:2, which rescales the rotary table for every request;MAX_CONTEXT=262144 ROPE_SCALING=nonekeeps the trained table.- New metrics:
ignis_retained_host_slots,ignis_retained_host_bytes; theignis.runtime.vram_planevent carries the host block.
Work in progress (issues still open)
- Decisions: share one state across question kinds — measure the state before the instruction (#271)
- Capture a reuse boundary inside one forward instead of splitting the prefill (#272)
- Decisions:
locatephase B — a place in the text evidence, read from attention (#275) - Decisions:
locateresearch — a span read from attention (spec 19) (#276) - Research:
locateby a constrained copy —method: "copy"(future study) (#279)
Follow-ups to issues an earlier release already carried
Decisions
- ADR 0017 — Opt-in Prometheus metrics with zero inference-path work (amended)
- ADR 0029 — cross-request state reuse: content-matched retained state across three residency tiers (amended)
- ADR 0030 — device memory reserved at load, within a VRAM budget (amended: retained slots on the host)
- ADR 0038 — the
Computeseam carries one attention head's scores (amended) - ADR 0039 — the attention readout carries a head set (amended)
- ADR 0041 — the attention readout reads a text span, in whole rows
- ADR 0042 — a locate narrows by attention heads and decides by a labelled choice, by kind of text
Commits in the range that name no issue
- Make: ROPE_SCALING defaults to yarn:2, the envelope the 524288 context needs (6c1541b)
- Make: the default context is 524288 tokens (c0d6823)
- Decisions: spec 22 and the round-3 finding say what ran — spaced JSON, found-counted records numbers, the original-line step's cost (6ee2469)
- Decisions: locate at length and in real logs, the research (specs 20, 21) (cb6ea1e)
- Playground: a Decide example that asks the Three Laws (8551cc3)
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.4.0
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.3.3...v0.4.0
v0.3.3
ignis v0.3.3
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.3.3
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.3.2...v0.3.3
v0.3.2
ignis v0.3.2
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.3.2
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.3.1...v0.3.2
v0.3.1
ignis v0.3.1
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.3.1
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.3.0...v0.3.1
v0.3.0
ignis v0.3.0
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.3.0
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.2.1...v0.3.0
v0.2.1
ignis v0.2.1
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.2.1
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.2.0...v0.2.1
v0.1.2
ignis v0.1.2
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.1.2
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is published by its own
workflow, so check the image run if the tag above is not there
yet.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.1.1...v0.1.2
v0.1.1
ignis v0.1.1
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.1.1
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is published by its own
workflow, so check the image run if the tag above is not there
yet.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.1.0...v0.1.1
ignis v0.1.0
ignis v0.1.0
The first release that can be installed rather than built.
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.1.0
(docker: --gpus all in place of --device.) The image carries the CUDA
runtime but no driver — the host's NVIDIA driver is injected by the container
runtime. Without IGNIS_ARTIFACT it starts on the deterministic CPU mock, so
running it with no model and no GPU is a smoke test of the image itself.
Every server flag has an IGNIS_* environment variable; anything after the
image name is passed to the server, and --ui — a bare switch with no
environment variable — is the image's default command.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell consumer) only,
and the engine needs a .ninfer artifact, which is not part of this release.
The Linux tarball needs the CUDA 13 runtime on the host; the Windows zip
carries cudart64_13.dll. Both need an NVIDIA driver new enough for CUDA 13.
Each archive ships with a .sha256 sidecar.
Licensing
Apache-2.0 (LICENSE, NOTICE). The vendored reference ops keep their own
terms — NOTICE-kernel, and kernel/vendor/manifest.json in the repo pins
the reference commit and every vendored file's content hash.