Skip to content

Releases: gpillon/ignis

v0.5.0

Choose a tag to compare

@github-actions github-actions released this 01 Oct 22:49

ignis v0.5.0

Changes

  • Thinking budget: a forced close that answers, and the Qwen3.8 model card's sampling by default — ends the agent tool loops where every turn was forced and the model deferred with one more tool call (#297)
  • hq banded prefill: a conversation past 262K keys no longer dies of a 64-byte workspace under-reservation; a failed chunk partway through a prompt ends its request at once, and a streaming engine error is an OpenAI error chunk (#296)
  • Responses API: streaming, function tools and WebSocket mode (#282)
  • Playground: talks to ignis over the Responses WebSocket, so parallel agents pass the browser's six-connection cap (#283)
  • Chat surface: every OpenAI field honoured, refused or declared inert — plus stop, max_completion_tokens and cached_tokens (#284)
  • Server: POST /v1/tokenize and /v1/detokenize — the prompt counted without being served (#285)
  • Kernel: an A16 NVFP4 call of 64+ columns runs on the tensor cores (the dflash2 drafter's context append; TTFT −12.8%)
  • Core: a request cancelled while evicted takes its snapshot out of the host tier
  • Playground: the Decide tab takes a file as evidence; streaming text lands once per frame for every session
  • Tests: temp files a test writes carry the pid (#198); a contributed entailment fixture for noul

Upgrade notes

  • Sampling defaults changed. A chat completion or /v1/responses request that leaves a sampling field unset now takes the Qwen3.8 model card's value for its mode, field by field: thinking temperature 1.0, top_p 0.95, top_k 20; enable_thinking: false 0.7 / 0.80 / 20 with presence_penalty 1.5. It was greedy (temperature: 0).
  • Greedy is now temperature: 0, sent explicitly; its unset fields are neutral, and a non-neutral top_p/top_k/penalty sent with it is still a 400. A request that sends top_p or top_k without temperature is now served instead of refused.
  • A request with no seed draws a fresh one, so two identical seedless requests are two independent samples. It was seed 0.
  • The forced close a thinking budget emits is now "My thinking time is over. I must now write the complete final answer from what I already have, without calling any more tools." (it was the model card's "Considering the limited time by the user, …").
  • /v1/responses echoes temperature and top_p as the decimals they were written as (0.95, not 0.949999988079071).

Work in progress (issues still open)

  • NVFP4 34816x5120 W4A4 at narrow prefill widths: three-stage schedule, weight-code L2 promotion, and a 256-token TMA floor (#223)
  • Tools: tool_choice: "required" and a named function — forcing the call with a permitted-set opener (#286)
  • Structured output: response_format / a grammar — the permitted set needs a second form (#287)
  • Research: teacher-forced scoring as an API — how likely is this text, with nothing generated (#295)

Follow-ups to issues an earlier release already carried

  • Server: three specs for the OpenAI surface — the fields, the count, the forced call (#101)
  • Server: spec corrections — the budget already forces and releases, and the real clients decide the table (#121)
  • Playground: the conversation goes over the Responses WebSocket (#283) (#220)

Decisions

  • ADR 0017 — Opt-in Prometheus metrics with zero inference-path work (amended)

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.5.0

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.4.0...v0.5.0

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 01:34

ignis v0.4.0

Changes

  • Memory: retained slots on the host — --retained-host (default 16) and --retained-device (default 0); the KV pool gains ~42% (475K -> 677K tokens at the default load) at level throughput (#281)
  • Make: the default context is 524288 tokens, with ROPE_SCALING=yarn:2 to match
  • Decisions: locate over very long texts — heads narrow, a labelled choice decides; kind, method, compression, found (spec 22) (#278)
  • Playground Decide tab: locate + examples (#277)
  • Decisions: locate phase A — calibrate the heads that point into text, and decide go/no-go (#274)
  • /v1/decide: a parts state never reuses its prefix — static text before an image is re-prefilled on every request and every question (#270)
  • Playground: the Monitor says where the retained slots live (#281)

Upgrade notes

  • --retained-slots / IGNIS_RETAINED_SLOTS are removed and refuse the start: size the retained slots with --retained-device (VRAM) and --retained-host (pinned host memory). The Makefile's RETAINED_SLOTS knob became RETAINED_DEVICE and RETAINED_HOST.
  • The default 16 host slots pin ~3.5 GiB of host RAM at load. Windows counts it as the process's shared GPU memory, beside the KV-RAM arena. A card with VRAM to spare can keep today's device copies with --retained-device 16 --retained-host 0.
  • make now starts at MAX_CONTEXT=524288 ROPE_SCALING=yarn:2, which rescales the rotary table for every request; MAX_CONTEXT=262144 ROPE_SCALING=none keeps the trained table.
  • New metrics: ignis_retained_host_slots, ignis_retained_host_bytes; the ignis.runtime.vram_plan event carries the host block.

Work in progress (issues still open)

  • Decisions: share one state across question kinds — measure the state before the instruction (#271)
  • Capture a reuse boundary inside one forward instead of splitting the prefill (#272)
  • Decisions: locate phase B — a place in the text evidence, read from attention (#275)
  • Decisions: locate research — a span read from attention (spec 19) (#276)
  • Research: locate by a constrained copy — method: "copy" (future study) (#279)

Follow-ups to issues an earlier release already carried

  • Decisions: spec 22, locate by copy over a folded state; ADR 0042 proposed (#278) (#242)

Decisions

  • ADR 0017 — Opt-in Prometheus metrics with zero inference-path work (amended)
  • ADR 0029 — cross-request state reuse: content-matched retained state across three residency tiers (amended)
  • ADR 0030 — device memory reserved at load, within a VRAM budget (amended: retained slots on the host)
  • ADR 0038 — the Compute seam carries one attention head's scores (amended)
  • ADR 0039 — the attention readout carries a head set (amended)
  • ADR 0041 — the attention readout reads a text span, in whole rows
  • ADR 0042 — a locate narrows by attention heads and decides by a labelled choice, by kind of text

Commits in the range that name no issue

  • Make: ROPE_SCALING defaults to yarn:2, the envelope the 524288 context needs (6c1541b)
  • Make: the default context is 524288 tokens (c0d6823)
  • Decisions: spec 22 and the round-3 finding say what ran — spaced JSON, found-counted records numbers, the original-line step's cost (6ee2469)
  • Decisions: locate at length and in real logs, the research (specs 20, 21) (cb6ea1e)
  • Playground: a Decide example that asks the Three Laws (8551cc3)

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.4.0

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.3.3...v0.4.0

v0.3.3

Choose a tag to compare

@github-actions github-actions released this 25 Sep 13:07

ignis v0.3.3

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.3.3

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.3.2...v0.3.3

v0.3.2

Choose a tag to compare

@github-actions github-actions released this 25 Sep 00:44

ignis v0.3.2

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.3.2

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.3.1...v0.3.2

v0.3.1

Choose a tag to compare

@github-actions github-actions released this 24 Sep 22:40

ignis v0.3.1

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.3.1

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.3.0...v0.3.1

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 23 Sep 22:35

ignis v0.3.0

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.3.0

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.2.1...v0.3.0

v0.2.1

Choose a tag to compare

@github-actions github-actions released this 21 Sep 15:48

ignis v0.2.1

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.2.1

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.2.0...v0.2.1

v0.1.2

Choose a tag to compare

@github-actions github-actions released this 19 Sep 17:41

ignis v0.1.2

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.1.2

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is published by its own
workflow, so check the image run if the tag above is not there
yet.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.1.1...v0.1.2

v0.1.1

Choose a tag to compare

@github-actions github-actions released this 19 Sep 15:44

ignis v0.1.1

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.1.1

(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is published by its own
workflow, so check the image run if the tag above is not there
yet.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.

Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).

Full Changelog: v0.1.0...v0.1.1

ignis v0.1.0

Choose a tag to compare

@gpillon gpillon released this 19 Sep 11:51

ignis v0.1.0

The first release that can be installed rather than built.

Container image (linux/amd64)

podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
  -v /path/to/models:/models:ro \
  -e IGNIS_ARTIFACT=/models/<name>.ninfer \
  ghcr.io/gpillon/ignis:0.1.0

(docker: --gpus all in place of --device.) The image carries the CUDA
runtime but no driver — the host's NVIDIA driver is injected by the container
runtime. Without IGNIS_ARTIFACT it starts on the deterministic CPU mock, so
running it with no model and no GPU is a smoke test of the image itself.

Every server flag has an IGNIS_* environment variable; anything after the
image name is passed to the server, and --ui — a bare switch with no
environment variable — is the image's default command.

Binaries

The kernel is compiled for SM120a (RTX 5090 / Blackwell consumer) only,
and the engine needs a .ninfer artifact, which is not part of this release.
The Linux tarball needs the CUDA 13 runtime on the host; the Windows zip
carries cudart64_13.dll. Both need an NVIDIA driver new enough for CUDA 13.

Each archive ships with a .sha256 sidecar.

Licensing

Apache-2.0 (LICENSE, NOTICE). The vendored reference ops keep their own
terms — NOTICE-kernel, and kernel/vendor/manifest.json in the repo pins
the reference commit and every vendored file's content hash.