Skip to content

Releases: fmatsos/gekko

v0.8.0

Choose a tag to compare

@github-actions github-actions released this 25 Sep 14:15

Added

  • npu config schema backend|model|command|test prints the JSON Schema of each configuration
    format, derived from the parser itself; point an editor at it (Taplo's #:schema) for
    completion and unknown-key warnings (3b4b900)
  • [partials] and {{ partials.<id> }}: a shared text fragment (style guide, glossary) lives
    in <scope root>/partials/ and is inserted verbatim in the body, system or an example; a
    missing partial exits 2 naming both files, before the input is read
    (9a7fcf7)
  • [output].extract = "/pointer" on a JSON command writes one value instead of the document — a
    string bare, anything else as compact JSON — after the whole document passed its schema; a
    pointer the answer lacks exits 4. CLI stdout only: MCP and config test keep the document
    (1bd5cf9)
  • NPU_STATS_FILE: every run of a configured command (CLI, config test case, MCP tool call)
    appends one JSON line with the requested and answering model, fallback, backend, duration,
    token usage, finish reason and exit code — never the prompt, the answer or a header value. A
    failed write is a warning and changes nothing else
    (a225360, 9ca7dda)
  • max_concurrent = 1 on a backend serializes its requests across every npu process on the
    machine (an advisory lock in the state directory): a second invocation waits instead of
    failing with a 5xx on a single NPU. --no-wait makes a busy backend a backend failure
    instead, which the model's fallback absorbs or which exits 3
    (6831f83, b412aa0)
  • protocol = "embeddings" on an [operations.<name>] table: the rendered prompt is sent as
    input and the answer is the vector, as the command's JSON output. chat stays the default,
    so existing backends are unchanged (aedef68)
  • [input] mode = "binary" and protocol = "transcriptions": a command uploads a file (or stdin)
    as bytes in a multipart request, the prompt as an optional hint, and prints the transcript
    (npu transcribe memo.wav). Binary commands are not offered as MCP tools
    (96b3caf, b412aa0)

Changed

  • Breaking: no-wait is now a reserved argument name, taken by the new --no-wait flag;
    rename any [args."no-wait"] a command declares
    (6831f83)
  • A command running an embeddings or transcriptions model is checked against that protocol when
    it runs and by npu doctor (chat-only keys such as system are rejected, exit 2); an
    embeddings model rejects fallback and [generation], and a fallback must speak its model's
    protocol, at load time (aedef68, 96b3caf)

Fixed

  • npu mcp serve no longer sends a non-object JSON answer as structuredContent, nor advertises
    an output schema whose root is not an object (f466b41)

Full changelog: v0.7.1...v0.8.0

v0.7.1

Choose a tag to compare

@github-actions github-actions released this 25 Sep 11:15

Fixed

  • The MCP pipeline error-envelope integration test now reaches the configured-command pipeline
    instead of passing an undeclared CLI argument and testing clap's usage envelope
    (d3856e6)

Full changelog: v0.7.0...v0.7.1

v0.6.1

Choose a tag to compare

@github-actions github-actions released this 24 Sep 12:00

Changed

  • Linux x86-64 ships as one binary, x86_64-unknown-linux-gnu, for every glibc-based
    distribution; the separate Fedora and Arch builds are gone (they were the same program).
    npu update no longer reads /etc/os-release. A Fedora or Arch install on 0.6.0 or earlier
    updates to this release as usual; one that skips it must reinstall from the releases page
    once, as later manifests no longer list x86_64-fedora or x86_64-arch
    (8590376)

Full changelog: v0.6.0...v0.6.1

v0.6.0

Choose a tag to compare

@github-actions github-actions released this 24 Sep 11:20

Added

  • --dry-run on every configured command prints the request npu would send — URL, header
    names (values redacted) and body — as JSON on stdout, without sending it or starting a runtime.
    The body is built by the same code as a real call, including "stream": true when the real
    call would stream (3a21b65, 1933062)
  • --model <ID> on every configured command uses another configured model for that one call; an
    unknown id exits 2 naming the available ones, before the input is read (5c0f366)
  • --json on doctor / config check, backend status and config models: the same report as
    a JSON array (kind is "config" or "reachability") for a calling program (5a266c2)
  • npu describe with no argument lists every describable path, built-in groups included
    (5a266c2, 566e600)
  • --error-format json puts a command-line usage error on stderr as one JSON line
    ({"kind":"usage","message":...}), including a missing subcommand and npu help <unknown>;
    --help and --version are never affected (5a266c2, 04cc703)
  • The project scope is found by walking up from the current directory to the nearest .npu,
    stopping at the repository root (.git) or at $HOME; --config-dir <DIR> or
    NPU_CONFIG_DIR names it directly, and npu doctor reports which one was used
    (6f09ed9, dea2e5f, d17661f)
  • Backends accept a [headers] table sent with every request, for hosted or authenticated
    OpenAI-compatible servers. Values may only use {{ env.NAME }}, are resolved before the input
    is read, and are never logged (4226aa3, 92498e4)
  • Commands accept a system prompt and [[examples]] (few-shot user/assistant turns). A command
    declaring neither sends exactly the same request as before (2c9c27e)
  • Model [generation] gains seed, top_p, stop and a free-form [generation.extra]
    forwarded as-is (e.g. chat_template_kwargs); a command may override any of them for itself,
    key by key (073e541, 92498e4)
  • [output] strip_reasoning = true removes a leading <think>…</think> block before the output
    contract is applied (faab3f1)
  • The token usage reported by the server is logged at --verbose info (7d4e719)
  • Shell completions: COMPLETE=bash npu (or zsh, fish…) prints the registration script;
    completion covers built-ins, configured commands and their arguments
    (042f9f1, ef5230f, 155ca00)
  • backend tune --dry-run also prints the NPU calibration constants its estimate uses
    (a178c83, e22dbac)
  • A Cargo feature, hardware-tooling (on by default), holds model discover and backend tune;
    building with --no-default-features leaves them out (06ed679, 55f413d, d84d13b)

Changed

  • Breaking: an answer cut at max_tokens (finish_reason = "length") now exits 4, naming
    the model that answered and the limit in effect, instead of exiting 0 with a truncated
    answer. Set allow_truncated = true under the command's [output] to keep the old behaviour.
    A truncation never triggers the fallback (7d4e719, 142600a)
  • Breaking: dry-run, model, error-format and config-dir are now reserved argument
    names. A command file declaring one of them under [args] is rejected at load time, naming
    the file; rename the argument (3a21b65, 5c0f366, 5a266c2, 6f09ed9)
  • Breaking: a .npu in a parent directory is now loaded when npu runs from a
    subdirectory. A parent .npu you did not mean to use must be moved, or pass --config-dir
    (6f09ed9)
  • Breaking: on Windows, the system and user scopes are %ProgramData%\npu and
    %APPDATA%\npu (else %USERPROFILE%\.config\npu). A configuration placed under
    HOME or XDG_CONFIG_HOME as a workaround must move there (5cf5371)
  • A stream that reports an error, or ends with no content, now fails with exit 3 instead of
    returning a partial or empty answer (7d4e719)
  • The OpenVINO architecture registry used by model discover is pinned to optimum-intel
    v2.2.0 instead of its moving main branch (5a4f1c2)

Fixed

  • npu update no longer warns that a valid configuration is invalid after every successful
    update (64934dc)
  • An unreadable, missing, non-UTF-8 or oversized input (over 64 MiB) is reported naming the
    file or stdin; still exit 1 (8b3f187, cfa24e4)
  • model discover now recognises architectures registered only through optimum-intel's shared
    text-generation task list (5a4f1c2)
  • Usage lines name the binary npu on Windows too, instead of npu.exe (b617677)
  • A backend error without a fallback keeps its URL and HTTP status, and a failing fallback is
    reported under its own backend (142600a)

Full changelog: v0.5.1...v0.6.0

v0.5.1

Choose a tag to compare

@github-actions github-actions released this 23 Sep 16:04

Added

  • A free-text answer (format = "text", no max_lines) is streamed to a terminal: each token
    is printed as the model produces it, so the answer starts showing within a second instead of
    at the end of the generation. The fallback still takes over while nothing is on screen (a
    stopped container, a prompt the NPU refuses). Once a token is shown, a failure exits 3
    after the partial answer. A JSON or max_lines answer, and anything written to a pipe or a
    file, still arrives in one piece, byte for byte as before
    (ecce23c)

Full changelog: v0.5.0...v0.5.1

v0.5.0

Choose a tag to compare

@github-actions github-actions released this 23 Sep 15:50

Added

  • npu model discover [words] searches Hugging Face for the models this host can run, on CPU,
    GPU or NPU. When the llmfit CLI is on PATH, a model is kept on its verdict (Perfect or
    Good, score at least --min-score). Otherwise, its INT4 weights must fit --max-memory
    percent of the RAM.
    • --backend openvino|llamacpp|mlx narrows the list to one engine's packaging. It also
      accepts a configured backend's identifier, whose engine is read from its runtime. openvino
      keeps the architectures optimum-intel exports, read from its registry at run time, with
      original, open weights.
    • --npu adds the Intel NPU check, and exits 3 on a host without one.
    • A third-party GGUF inherits llmfit's verdict for the model it packages.
    • --sort column[:asc|desc],... orders the report by one or more columns.
    • On a terminal the report is coloured; a pipe gets a plain table.
    • It needs no configuration, except to resolve a backend identifier.
      (496d534) (fa8ad71) (86bf561) (aebbbcb) (00974f2) (634054c)
  • npu backend tune sizes the context of every exported model from the model's config.json
    and the host's RAM, and writes it to graph.pbtxt. It also sets each model file's
    [generation].max_tokens to the answer length. Two limits apply: --max-memory (percent of
    the RAM, default 50) and --max-models (how many run at once, default all).
    • On NPU, it writes MAX_PROMPT_LEN and MIN_RESPONSE_LEN, and enables
      NPUW_LLM_ENABLE_PREFIX_CACHING.
    • On GPU, it bounds cache_size and sets max_num_seqs to 4. --kv-u8 stores the KV cache
      as u8, which holds about twice the context.
    • --npu or --gpu tunes one device, so each can get its own limits. --dry-run prints the
      plan without writing.
    • It replaces the npu-context.py script of the npu-export skill.
      (25504d3) (c3208b0) (974653c)
  • A backend declaring structured_output = true receives the command's output schema as an
    OpenAI response_format (json_schema). The model's decoding is constrained by it, and the
    answer is still validated afterwards (9d4eae4)
  • Commands can declare a [schemas] table and paste a schema into the prompt with
    {{ schemas.<id> }}. An undeclared id is rejected at load time, naming the file, and
    npu doctor checks each entry (9d4eae4)
  • On a terminal, a command's answer is set apart from the command line: a blank line, and a
    header naming the model that actually answered (the fallback, when it took over). A pipe or a
    file still receives the answer alone, byte for byte (1886c98) (9b5fc96)

Changed

  • Breaking: model is now a reserved name, for the new npu model group. Rename a
    configured command whose file is model.md or sits under model/
    (496d534)
  • Breaking: a schema value that is a bare name, with no / and no .json, now resolves
    to <scope root>/schemas/<name>.json. A path (relative to the scope root) or an absolute path
    resolves as before. A bare file name at the scope root must be written with its .json
    extension (9d4eae4)
  • A fallback taking over is logged at info instead of warn: it is the designed path, and the
    command still succeeds. Use -v info to see it (45418b6)

Fixed

  • A probe on a stopped container no longer prints Docker's No such container on the terminal:
    Docker's stderr is captured and added to the error, shown only when the error is reported
    (45418b6)

Full changelog: v0.4.0...v0.5.0

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 23 Sep 09:55

Added

  • npu describe describes built-ins as well as configured commands, and takes the path as words
    (npu describe backend serve, npu describe git review; git/review still works). Every
    description carries kind (builtin or command); a built-in lists its arguments,
    subcommands and whether it runs with a broken configuration — which describe itself does for
    built-ins; a configured command adds its resolved backend, its fallback and its source
    (the file that won and its scope). No existing field changes
    (770867a)
  • On a terminal, a spinner while a model is waited on (relabelled when the fallback takes over)
    and while a process backend starts, and a progress bar while npu update downloads. Drawn on
    stderr only, only when stderr is a terminal and --verbose is above error
    (fd18b97)
  • Colours on a terminal: help and usage errors, the warn/error/info labels, the final error
    line and doctor's check marks. Through a pipe, or with NO_COLOR, output is byte for byte what
    it was without them
    (c2c04e9)

Changed

  • Breaking: the built-ins are grouped. Update scripts as follows:

    Before Now
    npu serve, stop, status, logs npu backend serve, stop, status, logs
    npu models npu config models
    npu version npu --version, which prints npu X.Y.Z

    npu doctor is unchanged and also available as npu config check. The old names are no longer
    recognised (exit 2, nothing on stdout). The reserved command names shrink to backend,
    config, doctor, describe, update and help: a command file may now be named serve,
    stop, status, logs, models or version
    (83130e3)

  • npu --help lists the configured commands and the built-ins in two separate sections,
    Commands: and Built-ins:; npu help <command> is listed among the built-ins
    (c556e6e) (c2c04e9)

Full changelog: v0.3.1...v0.4.0

v0.3.1

Choose a tag to compare

@github-actions github-actions released this 23 Sep 08:02

Fixed

  • npu update and npu version no longer warn about an invalid configuration: neither reads it.
    After an update, the newly installed binary checks the configuration instead of the old one, so
    a key introduced by the new release no longer looks invalid; if the new version does reject it,
    a warning on stderr points to this changelog and the documentation, and the update still exits
    0
    (3e2fb18)

Full changelog: v0.3.0...v0.3.1

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 22 Sep 18:20

Added

  • A backend can be started as a local process instead of a container: [runtime] with
    type = "process", a command, its arguments, an optional [runtime.env] overlay and a
    startup_timeout_secs readiness budget. npu serve spawns it, prints its pid only once its port
    answers, and stop, status and logs find it again through a small state record. This is
    what lets a Mac run llama-server on Metal, which a Linux container cannot reach. Unix only:
    on Windows such a backend is rejected at load time, naming the file
    (552881f,
    92d3cc5)
  • npu stop never signals a process it cannot prove it started: the record keeps the pid and the
    moment the system says it was born, and a recycled pid is forgotten, not killed. Two projects
    that both declare a backend llamacpp get separate records and logs, and neither can stop the
    other's server
    (552881f,
    92d3cc5)
  • fallback on a model retries once on another model when the first one fails with a backend
    error, which moves a prompt too long for an NPU-compiled graph onto a GPU-served twin. The
    primary failure is always logged at warn, so a backend down all day does not pass for a
    working fallback. A fallback naming an unknown model, or itself, is rejected at load time
    (bdd564c)
  • port on a backend is declared once and read as {{ backend.port }} in base_url and the
    [runtime] lists, so the two can no longer drift apart. port = "auto" lets Docker pick a free
    port and reads it back. npu serve now reports a backend that is already served, or a fixed
    port held by something else, before starting anything (exit 3)
    (bdd564c)
  • Two Claude Code skills for the model side of an Intel NPU deployment: npu-discover finds
    Hugging Face models the NPU can actually run, npu-export exports one with optimum-cli,
    checks it on CPU and writes the model files, GPU twin included
    (8db677f,
    bdd564c)
  • Two deployment guides: Intel NPU (export, quantization, OVMS, the GPU
    twin) and Apple Silicon (llama-server on Metal, started by
    npu serve)
    (b508085,
    bdd564c,
    f61c33d)

Changed

  • Breaking: npu status prints BACKEND RUNTIME INSTANCE URL STATE instead of
    BACKEND CONTAINER STATE. A program reading its columns by position must be updated: the
    instance (container name or pid) is now column 3
    (bdd564c,
    ad5803e)
  • Breaking: npu models gains a trailing FALLBACK column (- when the model declares
    none)
    (bdd564c)
  • A backend's runtime is declared in a [runtime] table tagged by type ("docker" or
    "process"), so an unsupported family is rejected by name. The [docker] table of earlier
    versions is still accepted and means type = "docker"; declaring both is rejected, naming the
    file
    (ad5803e)
  • Release binaries are built with thin LTO: the build takes half the time, and the binary is
    about 1.5 MB larger
    (2d3279e)

Full changelog: v0.2.0...v0.3.0

v0.2.0

Choose a tag to compare

@github-actions github-actions released this 22 Sep 08:02

Added

  • npu version and npu update check, download, verify and install the latest GitHub release
    binary. Both run in degraded mode like doctor, since neither depends on the AI configuration
    (8003b25)

Changed

  • Breaking: a backend's [timeouts] table (request_secs, in seconds) is now honoured
    instead of being rejected as an unknown key, and the default request timeout rises from 30s to
    120s to cover a full max_tokens generation on a slow accelerator. request_secs = 0 is
    rejected at load time, naming the file
    (e40cd5e)

Fixed

  • The ovms backend fixture pins the OpenVINO Model Server image to 2026.4.0 (was the floating
    :latest tag) and mounts a persistent --cache_dir, so NPU/GPU graph compilation happens once
    per model instead of on every container restart
    (829e8da,
    c5b4a31)

Full changelog: v0.1.0...v0.2.0