ignis v0.5.0
Changes
- Thinking budget: a forced close that answers, and the Qwen3.8 model card's sampling by default — ends the agent tool loops where every turn was forced and the model deferred with one more tool call (#297)
- hq banded prefill: a conversation past 262K keys no longer dies of a 64-byte workspace under-reservation; a failed chunk partway through a prompt ends its request at once, and a streaming engine error is an OpenAI error chunk (#296)
- Responses API: streaming, function tools and WebSocket mode (#282)
- Playground: talks to ignis over the Responses WebSocket, so parallel agents pass the browser's six-connection cap (#283)
- Chat surface: every OpenAI field honoured, refused or declared inert — plus
stop,max_completion_tokensandcached_tokens(#284) - Server:
POST /v1/tokenizeand/v1/detokenize— the prompt counted without being served (#285) - Kernel: an A16 NVFP4 call of 64+ columns runs on the tensor cores (the dflash2 drafter's context append; TTFT −12.8%)
- Core: a request cancelled while evicted takes its snapshot out of the host tier
- Playground: the Decide tab takes a file as evidence; streaming text lands once per frame for every session
- Tests: temp files a test writes carry the pid (#198); a contributed entailment fixture for
noul
Upgrade notes
- Sampling defaults changed. A chat completion or
/v1/responsesrequest that leaves a sampling field unset now takes the Qwen3.8 model card's value for its mode, field by field: thinkingtemperature1.0,top_p0.95,top_k20;enable_thinking: false0.7 / 0.80 / 20 withpresence_penalty1.5. It was greedy (temperature: 0). - Greedy is now
temperature: 0, sent explicitly; its unset fields are neutral, and a non-neutraltop_p/top_k/penalty sent with it is still a 400. A request that sendstop_portop_kwithouttemperatureis now served instead of refused. - A request with no
seeddraws a fresh one, so two identical seedless requests are two independent samples. It was seed 0. - The forced close a thinking budget emits is now "My thinking time is over. I must now write the complete final answer from what I already have, without calling any more tools." (it was the model card's "Considering the limited time by the user, …").
/v1/responsesechoestemperatureandtop_pas the decimals they were written as (0.95, not0.949999988079071).
Work in progress (issues still open)
- NVFP4 34816x5120 W4A4 at narrow prefill widths: three-stage schedule, weight-code L2 promotion, and a 256-token TMA floor (#223)
- Tools:
tool_choice: "required"and a named function — forcing the call with a permitted-set opener (#286) - Structured output:
response_format/ a grammar — the permitted set needs a second form (#287) - Research: teacher-forced scoring as an API — how likely is this text, with nothing generated (#295)
Follow-ups to issues an earlier release already carried
- Server: three specs for the OpenAI surface — the fields, the count, the forced call (#101)
- Server: spec corrections — the budget already forces and releases, and the real clients decide the table (#121)
- Playground: the conversation goes over the Responses WebSocket (#283) (#220)
Decisions
- ADR 0017 — Opt-in Prometheus metrics with zero inference-path work (amended)
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.5.0
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.4.0...v0.5.0