Releases: fmatsos/gekko
Release list
v0.8.0
Added
npu config schema backend|model|command|testprints the JSON Schema of each configuration
format, derived from the parser itself; point an editor at it (Taplo's#:schema) for
completion and unknown-key warnings (3b4b900)[partials]and{{ partials.<id> }}: a shared text fragment (style guide, glossary) lives
in<scope root>/partials/and is inserted verbatim in the body,systemor an example; a
missing partial exits2naming both files, before the input is read
(9a7fcf7)[output].extract = "/pointer"on a JSON command writes one value instead of the document — a
string bare, anything else as compact JSON — after the whole document passed its schema; a
pointer the answer lacks exits4. CLI stdout only: MCP andconfig testkeep the document
(1bd5cf9)NPU_STATS_FILE: every run of a configured command (CLI,config testcase, MCP tool call)
appends one JSON line with the requested and answering model, fallback, backend, duration,
token usage, finish reason and exit code — never the prompt, the answer or a header value. A
failed write is a warning and changes nothing else
(a225360,9ca7dda)max_concurrent = 1on a backend serializes its requests across everynpuprocess on the
machine (an advisory lock in the state directory): a second invocation waits instead of
failing with a5xxon a single NPU.--no-waitmakes a busy backend a backend failure
instead, which the model'sfallbackabsorbs or which exits3
(6831f83,b412aa0)protocol = "embeddings"on an[operations.<name>]table: the rendered prompt is sent as
inputand the answer is the vector, as the command's JSON output.chatstays the default,
so existing backends are unchanged (aedef68)[input] mode = "binary"andprotocol = "transcriptions": a command uploads a file (or stdin)
as bytes in a multipart request, the prompt as an optional hint, and prints the transcript
(npu transcribe memo.wav). Binary commands are not offered as MCP tools
(96b3caf,b412aa0)
Changed
- Breaking:
no-waitis now a reserved argument name, taken by the new--no-waitflag;
rename any[args."no-wait"]a command declares
(6831f83) - A command running an embeddings or transcriptions model is checked against that protocol when
it runs and bynpu doctor(chat-only keys such assystemare rejected, exit2); an
embeddings model rejectsfallbackand[generation], and a fallback must speak its model's
protocol, at load time (aedef68,96b3caf)
Fixed
npu mcp serveno longer sends a non-object JSON answer asstructuredContent, nor advertises
an output schema whose root is not an object (f466b41)
Full changelog: v0.7.1...v0.8.0
v0.7.1
Fixed
- The MCP pipeline error-envelope integration test now reaches the configured-command pipeline
instead of passing an undeclared CLI argument and testing clap's usage envelope
(d3856e6)
Full changelog: v0.7.0...v0.7.1
v0.6.1
Changed
- Linux x86-64 ships as one binary,
x86_64-unknown-linux-gnu, for every glibc-based
distribution; the separate Fedora and Arch builds are gone (they were the same program).
npu updateno longer reads/etc/os-release. A Fedora or Arch install on 0.6.0 or earlier
updates to this release as usual; one that skips it must reinstall from the releases page
once, as later manifests no longer listx86_64-fedoraorx86_64-arch
(8590376)
Full changelog: v0.6.0...v0.6.1
v0.6.0
Added
--dry-runon every configured command prints the requestnpuwould send — URL, header
names (values redacted) and body — as JSON on stdout, without sending it or starting a runtime.
The body is built by the same code as a real call, including"stream": truewhen the real
call would stream (3a21b65,1933062)--model <ID>on every configured command uses another configured model for that one call; an
unknown id exits2naming the available ones, before the input is read (5c0f366)--jsonondoctor/config check,backend statusandconfig models: the same report as
a JSON array (kindis"config"or"reachability") for a calling program (5a266c2)npu describewith no argument lists every describable path, built-in groups included
(5a266c2,566e600)--error-format jsonputs a command-line usage error on stderr as one JSON line
({"kind":"usage","message":...}), including a missing subcommand andnpu help <unknown>;
--helpand--versionare never affected (5a266c2,04cc703)- The project scope is found by walking up from the current directory to the nearest
.npu,
stopping at the repository root (.git) or at$HOME;--config-dir <DIR>or
NPU_CONFIG_DIRnames it directly, andnpu doctorreports which one was used
(6f09ed9,dea2e5f,d17661f) - Backends accept a
[headers]table sent with every request, for hosted or authenticated
OpenAI-compatible servers. Values may only use{{ env.NAME }}, are resolved before the input
is read, and are never logged (4226aa3,92498e4) - Commands accept a
systemprompt and[[examples]](few-shot user/assistant turns). A command
declaring neither sends exactly the same request as before (2c9c27e) - Model
[generation]gainsseed,top_p,stopand a free-form[generation.extra]
forwarded as-is (e.g.chat_template_kwargs); a command may override any of them for itself,
key by key (073e541,92498e4) [output] strip_reasoning = trueremoves a leading<think>…</think>block before the output
contract is applied (faab3f1)- The token usage reported by the server is logged at
--verbose info(7d4e719) - Shell completions:
COMPLETE=bash npu(orzsh,fish…) prints the registration script;
completion covers built-ins, configured commands and their arguments
(042f9f1,ef5230f,155ca00) backend tune --dry-runalso prints the NPU calibration constants its estimate uses
(a178c83,e22dbac)- A Cargo feature,
hardware-tooling(on by default), holdsmodel discoverandbackend tune;
building with--no-default-featuresleaves them out (06ed679,55f413d,d84d13b)
Changed
- Breaking: an answer cut at
max_tokens(finish_reason = "length") now exits4, naming
the model that answered and the limit in effect, instead of exiting0with a truncated
answer. Setallow_truncated = trueunder the command's[output]to keep the old behaviour.
A truncation never triggers the fallback (7d4e719,142600a) - Breaking:
dry-run,model,error-formatandconfig-dirare now reserved argument
names. A command file declaring one of them under[args]is rejected at load time, naming
the file; rename the argument (3a21b65,5c0f366,5a266c2,6f09ed9) - Breaking: a
.npuin a parent directory is now loaded whennpuruns from a
subdirectory. A parent.npuyou did not mean to use must be moved, or pass--config-dir
(6f09ed9) - Breaking: on Windows, the system and user scopes are
%ProgramData%\npuand
%APPDATA%\npu(else%USERPROFILE%\.config\npu). A configuration placed under
HOMEorXDG_CONFIG_HOMEas a workaround must move there (5cf5371) - A stream that reports an error, or ends with no content, now fails with exit
3instead of
returning a partial or empty answer (7d4e719) - The OpenVINO architecture registry used by
model discoveris pinned to optimum-intel
v2.2.0instead of its movingmainbranch (5a4f1c2)
Fixed
npu updateno longer warns that a valid configuration is invalid after every successful
update (64934dc)- An unreadable, missing, non-UTF-8 or oversized input (over 64 MiB) is reported naming the
file orstdin; still exit1(8b3f187,cfa24e4) model discovernow recognises architectures registered only through optimum-intel's shared
text-generation task list (5a4f1c2)- Usage lines name the binary
npuon Windows too, instead ofnpu.exe(b617677) - A backend error without a fallback keeps its URL and HTTP status, and a failing fallback is
reported under its own backend (142600a)
Full changelog: v0.5.1...v0.6.0
v0.5.1
Added
- A free-text answer (
format = "text", nomax_lines) is streamed to a terminal: each token
is printed as the model produces it, so the answer starts showing within a second instead of
at the end of the generation. The fallback still takes over while nothing is on screen (a
stopped container, a prompt the NPU refuses). Once a token is shown, a failure exits3
after the partial answer. A JSON ormax_linesanswer, and anything written to a pipe or a
file, still arrives in one piece, byte for byte as before
(ecce23c)
Full changelog: v0.5.0...v0.5.1
v0.5.0
Added
npu model discover [words]searches Hugging Face for the models this host can run, on CPU,
GPU or NPU. When thellmfitCLI is onPATH, a model is kept on its verdict (Perfector
Good, score at least--min-score). Otherwise, its INT4 weights must fit--max-memory
percent of the RAM.--backend openvino|llamacpp|mlxnarrows the list to one engine's packaging. It also
accepts a configured backend's identifier, whose engine is read from its runtime.openvino
keeps the architecturesoptimum-intelexports, read from its registry at run time, with
original, open weights.--npuadds the Intel NPU check, and exits3on a host without one.- A third-party GGUF inherits llmfit's verdict for the model it packages.
--sort column[:asc|desc],...orders the report by one or more columns.- On a terminal the report is coloured; a pipe gets a plain table.
- It needs no configuration, except to resolve a backend identifier.
(496d534) (fa8ad71) (86bf561) (aebbbcb) (00974f2) (634054c)
npu backend tunesizes the context of every exported model from the model'sconfig.json
and the host's RAM, and writes it tograph.pbtxt. It also sets each model file's
[generation].max_tokensto the answer length. Two limits apply:--max-memory(percent of
the RAM, default 50) and--max-models(how many run at once, default all).- On NPU, it writes
MAX_PROMPT_LENandMIN_RESPONSE_LEN, and enables
NPUW_LLM_ENABLE_PREFIX_CACHING. - On GPU, it bounds
cache_sizeand setsmax_num_seqsto 4.--kv-u8stores the KV cache
as u8, which holds about twice the context. --npuor--gputunes one device, so each can get its own limits.--dry-runprints the
plan without writing.- It replaces the
npu-context.pyscript of thenpu-exportskill.
(25504d3) (c3208b0) (974653c)
- On NPU, it writes
- A backend declaring
structured_output = truereceives the command's output schema as an
OpenAIresponse_format(json_schema). The model's decoding is constrained by it, and the
answer is still validated afterwards (9d4eae4) - Commands can declare a
[schemas]table and paste a schema into the prompt with
{{ schemas.<id> }}. An undeclared id is rejected at load time, naming the file, and
npu doctorchecks each entry (9d4eae4) - On a terminal, a command's answer is set apart from the command line: a blank line, and a
header naming the model that actually answered (the fallback, when it took over). A pipe or a
file still receives the answer alone, byte for byte (1886c98) (9b5fc96)
Changed
- Breaking:
modelis now a reserved name, for the newnpu modelgroup. Rename a
configured command whose file ismodel.mdor sits undermodel/
(496d534) - Breaking: a
schemavalue that is a bare name, with no/and no.json, now resolves
to<scope root>/schemas/<name>.json. A path (relative to the scope root) or an absolute path
resolves as before. A bare file name at the scope root must be written with its.json
extension (9d4eae4) - A fallback taking over is logged at
infoinstead ofwarn: it is the designed path, and the
command still succeeds. Use-v infoto see it (45418b6)
Fixed
- A probe on a stopped container no longer prints Docker's
No such containeron the terminal:
Docker's stderr is captured and added to the error, shown only when the error is reported
(45418b6)
Full changelog: v0.4.0...v0.5.0
v0.4.0
Added
npu describedescribes built-ins as well as configured commands, and takes the path as words
(npu describe backend serve,npu describe git review;git/reviewstill works). Every
description carrieskind(builtinorcommand); a built-in lists its arguments,
subcommands and whether it runs with a broken configuration — whichdescribeitself does for
built-ins; a configured command adds its resolvedbackend, itsfallbackand itssource
(the file that won and its scope). No existing field changes
(770867a)- On a terminal, a spinner while a model is waited on (relabelled when the fallback takes over)
and while a process backend starts, and a progress bar whilenpu updatedownloads. Drawn on
stderr only, only when stderr is a terminal and--verboseis aboveerror
(fd18b97) - Colours on a terminal: help and usage errors, the
warn/error/infolabels, the final error
line anddoctor's check marks. Through a pipe, or withNO_COLOR, output is byte for byte what
it was without them
(c2c04e9)
Changed
-
Breaking: the built-ins are grouped. Update scripts as follows:
Before Now npu serve,stop,status,logsnpu backend serve,stop,status,logsnpu modelsnpu config modelsnpu versionnpu --version, which printsnpu X.Y.Znpu doctoris unchanged and also available asnpu config check. The old names are no longer
recognised (exit2, nothing on stdout). The reserved command names shrink tobackend,
config,doctor,describe,updateandhelp: a command file may now be namedserve,
stop,status,logs,modelsorversion
(83130e3) -
npu --helplists the configured commands and the built-ins in two separate sections,
Commands:andBuilt-ins:;npu help <command>is listed among the built-ins
(c556e6e) (c2c04e9)
Full changelog: v0.3.1...v0.4.0
v0.3.1
Fixed
npu updateandnpu versionno longer warn about an invalid configuration: neither reads it.
After an update, the newly installed binary checks the configuration instead of the old one, so
a key introduced by the new release no longer looks invalid; if the new version does reject it,
a warning on stderr points to this changelog and the documentation, and the update still exits
0
(3e2fb18)
Full changelog: v0.3.0...v0.3.1
v0.3.0
Added
- A backend can be started as a local process instead of a container:
[runtime]with
type = "process", acommand, itsarguments, an optional[runtime.env]overlay and a
startup_timeout_secsreadiness budget.npu servespawns it, prints its pid only once its port
answers, andstop,statusandlogsfind it again through a small state record. This is
what lets a Mac runllama-serveron Metal, which a Linux container cannot reach. Unix only:
on Windows such a backend is rejected at load time, naming the file
(552881f,
92d3cc5) npu stopnever signals a process it cannot prove it started: the record keeps the pid and the
moment the system says it was born, and a recycled pid is forgotten, not killed. Two projects
that both declare a backendllamacppget separate records and logs, and neither can stop the
other's server
(552881f,
92d3cc5)fallbackon a model retries once on another model when the first one fails with a backend
error, which moves a prompt too long for an NPU-compiled graph onto a GPU-served twin. The
primary failure is always logged atwarn, so a backend down all day does not pass for a
working fallback. A fallback naming an unknown model, or itself, is rejected at load time
(bdd564c)porton a backend is declared once and read as{{ backend.port }}inbase_urland the
[runtime]lists, so the two can no longer drift apart.port = "auto"lets Docker pick a free
port and reads it back.npu servenow reports a backend that is already served, or a fixed
port held by something else, before starting anything (exit3)
(bdd564c)- Two Claude Code skills for the model side of an Intel NPU deployment:
npu-discoverfinds
Hugging Face models the NPU can actually run,npu-exportexports one withoptimum-cli,
checks it on CPU and writes the model files, GPU twin included
(8db677f,
bdd564c) - Two deployment guides: Intel NPU (export, quantization, OVMS, the GPU
twin) and Apple Silicon (llama-serveron Metal, started by
npu serve)
(b508085,
bdd564c,
f61c33d)
Changed
- Breaking:
npu statusprintsBACKEND RUNTIME INSTANCE URL STATEinstead of
BACKEND CONTAINER STATE. A program reading its columns by position must be updated: the
instance (container name or pid) is now column 3
(bdd564c,
ad5803e) - Breaking:
npu modelsgains a trailingFALLBACKcolumn (-when the model declares
none)
(bdd564c) - A backend's runtime is declared in a
[runtime]table tagged bytype("docker"or
"process"), so an unsupported family is rejected by name. The[docker]table of earlier
versions is still accepted and meanstype = "docker"; declaring both is rejected, naming the
file
(ad5803e) - Release binaries are built with thin LTO: the build takes half the time, and the binary is
about 1.5 MB larger
(2d3279e)
Full changelog: v0.2.0...v0.3.0
v0.2.0
Added
npu versionandnpu updatecheck, download, verify and install the latest GitHub release
binary. Both run in degraded mode likedoctor, since neither depends on the AI configuration
(8003b25)
Changed
- Breaking: a backend's
[timeouts]table (request_secs, in seconds) is now honoured
instead of being rejected as an unknown key, and the default request timeout rises from 30s to
120s to cover a fullmax_tokensgeneration on a slow accelerator.request_secs = 0is
rejected at load time, naming the file
(e40cd5e)
Fixed
- The
ovmsbackend fixture pins the OpenVINO Model Server image to2026.4.0(was the floating
:latesttag) and mounts a persistent--cache_dir, so NPU/GPU graph compilation happens once
per model instead of on every container restart
(829e8da,
c5b4a31)
Full changelog: v0.1.0...v0.2.0