v0.4.0
ignis v0.4.0
Changes
- Memory: retained slots on the host —
--retained-host(default 16) and--retained-device(default 0); the KV pool gains ~42% (475K -> 677K tokens at the default load) at level throughput (#281) - Make: the default context is 524288 tokens, with
ROPE_SCALING=yarn:2to match - Decisions:
locateover very long texts — heads narrow, a labelledchoicedecides;kind,method,compression,found(spec 22) (#278) - Playground Decide tab: locate + examples (#277)
- Decisions:
locatephase A — calibrate the heads that point into text, and decide go/no-go (#274) - /v1/decide: a parts
statenever reuses its prefix — static text before an image is re-prefilled on every request and every question (#270) - Playground: the Monitor says where the retained slots live (#281)
Upgrade notes
--retained-slots/IGNIS_RETAINED_SLOTSare removed and refuse the start: size the retained slots with--retained-device(VRAM) and--retained-host(pinned host memory). The Makefile'sRETAINED_SLOTSknob becameRETAINED_DEVICEandRETAINED_HOST.- The default 16 host slots pin ~3.5 GiB of host RAM at load. Windows counts it as the process's shared GPU memory, beside the KV-RAM arena. A card with VRAM to spare can keep today's device copies with
--retained-device 16 --retained-host 0. makenow starts atMAX_CONTEXT=524288 ROPE_SCALING=yarn:2, which rescales the rotary table for every request;MAX_CONTEXT=262144 ROPE_SCALING=nonekeeps the trained table.- New metrics:
ignis_retained_host_slots,ignis_retained_host_bytes; theignis.runtime.vram_planevent carries the host block.
Work in progress (issues still open)
- Decisions: share one state across question kinds — measure the state before the instruction (#271)
- Capture a reuse boundary inside one forward instead of splitting the prefill (#272)
- Decisions:
locatephase B — a place in the text evidence, read from attention (#275) - Decisions:
locateresearch — a span read from attention (spec 19) (#276) - Research:
locateby a constrained copy —method: "copy"(future study) (#279)
Follow-ups to issues an earlier release already carried
Decisions
- ADR 0017 — Opt-in Prometheus metrics with zero inference-path work (amended)
- ADR 0029 — cross-request state reuse: content-matched retained state across three residency tiers (amended)
- ADR 0030 — device memory reserved at load, within a VRAM budget (amended: retained slots on the host)
- ADR 0038 — the
Computeseam carries one attention head's scores (amended) - ADR 0039 — the attention readout carries a head set (amended)
- ADR 0041 — the attention readout reads a text span, in whole rows
- ADR 0042 — a locate narrows by attention heads and decides by a labelled choice, by kind of text
Commits in the range that name no issue
- Make: ROPE_SCALING defaults to yarn:2, the envelope the 524288 context needs (6c1541b)
- Make: the default context is 524288 tokens (c0d6823)
- Decisions: spec 22 and the round-3 finding say what ran — spaced JSON, found-counted records numbers, the original-line step's cost (6ee2469)
- Decisions: locate at length and in real logs, the research (specs 20, 21) (cb6ea1e)
- Playground: a Decide example that asks the Three Laws (8551cc3)
Container image (linux/amd64)
podman run --rm --device nvidia.com/gpu=all -p 8000:8000 \
-v /path/to/models:/models:ro \
-e IGNIS_ARTIFACT=/models/<name>.ninfer \
ghcr.io/gpillon/ignis:0.4.0
(docker: --gpus all in place of --device.) The image carries
the CUDA runtime but no driver — the host's NVIDIA driver is
injected by the container runtime. It is built from the same
compile as the Linux tarball below, so the two always carry the
same binaries.
Binaries
The kernel is compiled for SM120a (RTX 5090 / Blackwell
consumer) only, and the engine needs a .ninfer artifact, which
is not part of this release. The Linux tarball needs the CUDA 13
runtime on the host; the Windows zip carries cudart64_*.dll.
Both need an NVIDIA driver new enough for CUDA 13.
Apache-2.0 (LICENSE, NOTICE); the vendored reference ops keep
their own terms (NOTICE-kernel).
Full Changelog: v0.3.3...v0.4.0