Skip to content

Releases: ooguz/yonga

v0.5.0

Choose a tag to compare

@github-actions github-actions released this 11 Oct 02:48

The first public release, and the first on PyPI. It fixes everything an end-to-end campaign of
fresh-user runs on real hardware found (AppImage, pip, every command, GPU, NPU and CPU), keeps
OpenVINO's usage reporting from running, and installs the AppImage on PATH.

Added

  • yonga self install and yonga self uninstall. Run from the AppImage,
    ./yonga-X.Y.Z-x86_64.AppImage self install copies it to ~/.local/bin/yonga (or --bin-dir DIR): atomically (temporary file in the directory, fsync, 0755, rename), creating the
    directory when missing. The same file again is "already installed"; another yonga AppImage
    there is replaced after a question showing both versions where known (--yes answers it).
    Anything else at that name (a pip or pipx command, a symbolic link, a directory, a script) is
    refused and left alone; an AppImage is recognised by its ELF and AI\x02 magic bytes, never by
    running it. Afterwards it checks that yonga on PATH runs the copy and, if not, says what to do
    (log out and in, or export PATH="$HOME/.local/bin:$PATH") and exits 2; it never edits a shell
    startup file and never uses sudo (an unwritable directory exits 10 with the advice to pick one you
    own). Outside an AppImage it exits 70 and points pip users at their environment's yonga or
    pipx install 'yonga[openvino,export]'. self uninstall removes only an AppImage, after a
    question, and keeps models, logs, caches and settings, saying where they are.
    docs/CLI_SPEC.md section 12c.
  • yonga registry profiles [--device DEVICE] [--json] lists the conversion profiles
    --profile accepts, with device, weights, size band, priority and status. An unknown
    --profile now names the profiles for the requested device and points to this command.
  • The preparation plan always shows an Export environment row: a managed environment, or this
    environment with the interpreter path and the reason.
  • Issue rule gguf-reader-open-failed (export.gguf.open_failed, exit 10). GGUF import failures
    are now classified by the registry like optimum export failures instead of always being
    unknown_runtime_error.
  • Issue rule torchvision-torch-build-mismatch (export.torchvision.build_mismatch, exit 10) for
    operator torchvision::nms does not exist. prepare also imports the exporter in its own
    environment at preflight, so a broken torch/torchvision pair is reported before the download.
  • The probe job log records every probe result with its raw runtime text (probe.result), as the
    diagnostics already said it did.
  • yonga eval prints Saved: PATH, and --json includes result_path, as bench does.
  • doctor reports whether the exporter imports (exporter import), reusing the check prepare
    recorded, so a CPU torch next to PyPI's torchvision shows as a warning with the registry's
    remedy; --full checks a recorded failure again.
  • chat, bench, probe, eval and publish say which artifact a model id chose when several
    ready ones match ("Using ID (PROFILE, DEVICE), the newest of N ready artifacts for MODEL; pass the
    artifact id to choose another").
  • Registry schema 4 (needs yonga 0.5.0): an issue rule's blocks_export lists model types whose
    export is certain to fail when the rule's constraints hold for the interpreter that will run the
    export, and prepare refuses such an export before the download and before any question
    (overridable with --ignore-exporter-requirements, recorded in the manifest).

Changed

  • doctor reports how yonga is installed. A new Install line in the Host section names the
    AppImage file or the Python environment, and what yonga on PATH resolves to. From an AppImage
    it is a warning with self install as the remedy when yonga is not on PATH, or when it is
    another installation (named with its path, for example an older AppImage). For a pip install it
    is information only. doctor --json carries it as installation. docs/CLI_SPEC.md section 3.
  • A command line yonga cannot parse exits 70 instead of Click's 2: an unknown command or
    option, a missing argument, a value of the wrong type (--timeout abc, --older-than abc) and
    a bare yonga. Exit 2 is left meaning only warnings or an inconclusive probe; --help still
    exits 0.
  • prepare downloads only the files the conversion reads. ONNX copies, original/ checkpoints,
    training runs and a second weight format are skipped (SmolLM2-360M: 693 MiB instead of
    4.7 GiB), and the plan's expected download counts the same files.
  • Automatic profile selection skips a group-wise profile whose group size cannot divide the
    model's hidden_size (or intermediate_size at ratio 1.0) from config.json, and says why
    under "Chosen because". With no fitting profile for an explicit device, prepare stops before
    the download (exit 70) and names the devices that fit; --device auto moves on to the next
    device.
  • yonga logs lists each job's id prefix, local start time with its zone, command, model, device,
    outcome (ok/failed/cancelled/running/unfinished), duration and log size; --json adds these
    fields. One job's view shows the same facts and prints record times in local time instead of
    unmarked UTC.
  • yonga list shows the source revision as 7 characters under a Rev header.
  • While a job runs, warnings logged by huggingface_hub (such as the Hub's "You are sending
    unauthenticated requests" advice) go to the job log as library.log records instead of breaking
    into the progress line.
  • An explicit --profile whose group size cannot divide the model's config.json channel widths
    is refused at the plan, before any download (exit 70, export.quantization.group_size_mismatch).
    The message names the profiles that fit, as your command with --profile changed, and never
    swaps the profile itself. When no profile for that device fits (SmolLM2-360M with
    --device gpu --profile gpu-int4-asym-g128-r08), it offers the fitting profiles of the other
    devices, as your command with --device and --profile changed.
  • The context_length_exceeded message no longer repeats "The prompt". It names the offered tools
    and images when they count toward the limit, and says what to do: shorten the conversation, offer
    fewer tools or send fewer images, or use a GPU or CPU artifact, whose prompt length is not fixed.
  • The README quick-start doctor sample and its exit-2 explanation match the current output.

Fixed

  • --offline (and YONGA_OFFLINE / HF_HUB_OFFLINE) no longer contacts the Hugging Face Hub for
    inspect, prepare or revision resolution. A cached model is inspected and prepared from the
    local cache alone, including one downloaded by commit SHA with no refs/main. An uncached
    revision, or a snapshot with only metadata files, fails with hub.offline.missing without
    making any request.
  • yonga inspect no longer says a model needs --trust-remote-code when the registry declares a
    native Transformers implementation (Phi-4-mini, phi3). inspect and prepare now share one
    remote-code decision.
  • yonga --offline search reports hub.offline.search_unavailable ("Search needs the network")
    instead of a misleading cache-miss message.
  • Re-running prepare for a model that is already prepared reuses the artifact without asking
    anything, so it works in scripts without --yes. The reuse plan no longer lists a download,
    temp space, repairs or an export environment.
  • An explicit --device that the registry records as unsupported for the model (official or
    tested evidence) is refused before the download with exit 50 and the recommended fallback,
    instead of after a full download and conversion.
  • Without the OpenVINO runtime installed, prepare stops before the download with exit 10 and
    names pip install 'yonga[openvino]', instead of blaming the host for having no devices. It no
    longer offers to build a managed export environment then.
  • A gated or private repository with no Hugging Face token is flagged in the plan and refused
    before the download with exit 21.
  • In a session without a terminal and without --yes, prepare says the confirmation could not
    be asked and suggests --yes, instead of saying the user declined it.
  • A prepare that adopts a retained export, or copies published OpenVINO IR, reports this as a
    note and exits 0 instead of 2.
  • The plan's "Estimated artifact" is worked out from the profile's quantization and the model's
    parameter count (Qwen3-0.6B INT4: about 516 MiB, previously 1.4 GiB).
  • The plan's device decision for --device auto states the actual reason, such as the registry's
    recommendation.
  • NNCF's group-size refusal is classified (export.quantization.group_size_mismatch, exit 70) even
    when its list of refused nodes pushes the exception line out of the exporter's output tail; it
    used to end as a generic export.failed.
  • The group-size remediation no longer offers the NPU's channel-wise profile on the CPU or GPU,
    where its MAX_PROMPT_LEN fails the probe; it suggests cpu-int8-asym-ref on the CPU instead.
    The Python 3.14 optimum diagnostic now explains what AppImage users can do.
  • Remediation steps print the real repository and job id instead of MODEL and JOB_ID wherever
    they are known (prepare, probe, chat, bench, --json), and a step that repeats an
    earlier command is listed once.
  • "Chosen because" for a profile named by a model or architecture rule gives the parameter count
    and the profile the rule overrode, such as the size-banded default (automatic_profile_id in the
    manifest).
  • GGUF import no longer fails with "Failed to open ... with gguf_open" on a huggingface-hub 1.x
    cache (read-only blobs). OpenVINO GenAI's reader opens the file read-write, so yonga imports
    from a private writable copy in staging (a reflink where the filesystem supports one) instead of
    a link into the cache, and the disk plan...
Read more

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 09 Oct 17:00

Added

  • serve honours OpenAI stop. A string or up to 4 non-empty strings of at most 256
    characters; anything else is a 400 invalid_request naming stop. yonga matches them on the
    visible answer after the reasoning sanitiser, never inside a hidden reasoning block, across chunk
    boundaries, and ends the answer before the first stop sequence to complete with finish_reason
    stop. A stream holds back only the tail that could still become a match and the whitespace in
    front of it. A request with stop sequences is
    streamed internally even when the client did not ask for a stream, so generation ends at the
    match. docs/CLI_SPEC.md section 14.
  • serve takes images on a VLM artifact. OpenAI image_url content parts holding a
    data:image/(png|jpeg|webp|gif);base64, URL. Every other URL (http, https, file, ...) is
    refused with a 400 naming the content part, permanently: yonga never fetches an image. At most 4
    images per request, each at most 4096 x 4096 pixels, checked from the image header before
    decoding and refused rather than scaled down. Images go with the last user message only, because
    OpenVINO GenAI 2026.4 binds images to the newest turn of a stateless conversation; an image in an
    earlier message is a 400 serve.request.images_earlier_turn saying so. An LLM artifact refuses
    images with serve.request.images_unsupported. Every other check, stop and the generation
    settings included, comes before any image is read. Only the headers are read, off the event
    loop, before a request waits for the pipeline; the pixels are decoded once it holds the
    pipeline, so at most one request holds decoded images however many are queued. Any failure
    inside Pillow is a refusal naming the image, never a server error. EXIF rotation is applied as a browser
    shows it (an unreadable EXIF block leaves the image as stored), and 16-bit greyscale is scaled to
    8 bits rather than clipped. docs/CLI_SPEC.md section 14.
  • serve --carry-images-forward moves images of earlier messages onto the last user message,
    with a line telling the model they came earlier, for clients that resend every image of a
    conversation. The model then no longer knows which words each image went with.
  • serve --max-body-mb N sets the request body limit, 1 to 64 MiB (8 by default), since
    embedded images count towards it.
  • serve takes OpenAI tools on models the registry has seen calling them. A chat profile's
    new tool_calling entry (registry schema 3, below) says how a family writes calls; only a model
    whose profile declares one, and whose own chat template contains the entry's
    template_contains literal, accepts tools. Built in: Qwen2.5, Qwen3 and SmolLM3
    (<tool_call>{json}</tool_call>) and Granite (a JSON list), each seen producing a parseable call
    on the reference host on 2026-10-09 (CPU, one prompt each). Every other model refuses tools
    with a 400 serve.request.tools_unsupported naming tools and saying why. The tools, an
    assistant's tool_calls and tool results reach the artifact's own chat template through
    GenAI's ChatHistory.set_tools; a conversation may end on a tool result. The answer is read for
    calls outside reasoning, also with --show-reasoning, and returned as message.tool_calls with
    finish_reason tool_calls; streamed, text is released until a call could begin and each call
    then arrives as one complete delta, and a client that disconnects meanwhile still stops
    generation. An answer that cannot be read as calls to the offered functions is returned
    as text, never as a guessed call, and logged as such without its content. tool_choice auto
    and none are honoured; required and a named function are refused (they need constrained
    decoding). parallel_tool_calls: false keeps the first call. stop sequences match only the
    text outside the calls. docs/CLI_SPEC.md section 14.
  • Registry schema version 3: tool_calling on a chat profile. It selects one of the tool-call
    parsers compiled into yonga (tagged_json, json_list), fills in its markers and field names,
    and must state evidence tested or official with at least one source. Older files load
    unchanged. A rule using it needs schema_version: 3, inherited from its document as size bands
    are; tools/registry_bundle.py gives a schema-3 bundle min_tool_version 0.4.0, so older
    releases refuse it at registry update. registry validate and remote bundles refuse a format
    this yonga has no parser for. yonga registry explain shows the chosen profile's tool-call
    format. docs/REGISTRY_SPEC.md section 8b.
  • yonga probe --quick and --timeout SECONDS. Settling a deferred NPU compile means building
    the GenAI pipeline, which for a 9B model ran past 15 minutes. --quick never builds it;
    --timeout builds it in a child process (python -P -m yonga.adapters.smoke_child) and stops it
    at the limit, since a native compile cannot be interrupted in-process. Either way an unsettled
    device gets the new probe status inconclusive: never counted as working, never as a failure, and
    shown as such in the tables, the role matrix and --json ("inconclusive": true); a model card
    never counts it as working. A probe where nothing passed or failed outright exits 2. A child that
    dies without a result is a failure classified by the registry from its stderr. Without either
    option nothing changes, and prepare still always waits. python -m yonga.application.probe
    takes both options too. docs/CLI_SPEC.md section 11.
  • yonga cache list and yonga cache prune: the OpenVINO compile cache is no longer
    write-only.
    Every compile cache entry yonga uses now gets an ownership record,
    .yonga-entry.json (artifact id and path, repository, profile, device, OpenVINO version, runtime
    properties, created and last used); writing it never fails a load. cache list (or plain
    yonga cache) shows each entry's size, last use, device, OpenVINO version, owner and status:
    in-use, orphaned (the artifact is gone), stale-runtime (built by an OpenVINO release other
    than the installed one, so never reused) or unattributed. Entries made before the records are
    attributed by recomputing the keys of every artifact in the store and every artifact a saved
    benchmark names. Stale-runtime is relative to the OpenVINO running the command; an entry of
    another release used after this one first built an entry here belongs to another install
    sharing the cache and stays in-use. cache prune removes orphaned and stale-runtime entries,
    skipping any entry used after the plan was made (unless --all); --older-than DAYS,
    --unattributed and --all go further, --dry-run only shows the plan, and it asks first
    unless --yes. It deletes only key-named, real directories directly under the compile cache
    root, never follows a symbolic link, and reports the bytes freed. --json on both.
    docs/CLI_SPEC.md section 12b.
  • doctor reports the compile cache under Storage: entries, size and the reclaimable part,
    a warning naming yonga cache prune from 10 GiB reclaimable or 25 % of the free space on the
    cache filesystem (never below 1 GiB), and the numbers under compile_cache in --json.

Changed

  • Qwen2, SmolLM3 and Granite have their own chat profiles (chat-qwen2, chat-smollm3,
    chat-granite) instead of chat-default. They copy its policy exactly (thinking detection,
    <think> stripping, 4096 new tokens, no sampling defaults) and add only tool_calling, so chat
    behaves as before; registry explain names the new rule. chat-qwen2 matches the whole qwen2
    model type, but only Qwen2.5 was measured, and only an artifact whose template carries
    <tool_call> (the Qwen2.5 Instruct models, not the original Qwen2 Instruct) accepts tools.
  • Qwen3 MoE has its own chat profile, chat-qwen3_moe, split from chat-qwen3 with the same
    policy and no tool_calling: the MoE models share the template but were not measured writing
    calls, so they refuse tools. chat behaves as before.
  • serve accepts tool messages and assistant tool_calls in a conversation, alongside
    tools. Assistant tool_calls used to be dropped without an error, and tool messages were
    refused as an unknown role; now both are accepted alongside tools, and without tools they are
    a 400 naming the message. The legacy functions and function_call are refused by their own name
    instead of as tools.
  • The registry chat profile's stop_strings is applied. The field was parsed and never used; it
    now reaches OpenVINO GenAI's GenerationConfig.stop_strings as model-level end markers, left out
    of the answer. Empty and duplicate entries are dropped when a rule loads. A benchmark's settings
    fingerprint includes stop strings only when there are some, so every recorded fingerprint still
    matches. No built-in rule declares any, so nothing shipped behaves differently.
    docs/REGISTRY_SPEC.md section 8b.
  • The serve extra installs Pillow, which decodes images. It is imported only when an image
    arrives: without it text is served as before, serve warns at startup for a VLM artifact, and an
    image request is a 500 serve.images.dependency_missing naming pip install 'yonga[serve]'.
    The AppImage, which bundles the extra, grows by Pillow's size.
  • <ov_genai_image_N> and <ov_genai_video_N> are removed from text sent to a VLM artifact,
    in chat as in serve, repeatedly until none is left (removing one must not join a new one).
    GenAI reads these universal tags anywhere as "image N goes here", so typed text could claim an
    image that was not attached and fail generation. A model's own vision tokens typed as text are
    not removed; which they are is per-family knowledge the registry does not hold yet.
  • The VLM families follow the size bands. Qwen2.5-VL and Qwen3.5 ...
Read more

v0.3.0

Choose a tag to compare

@github-actions github-actions released this 09 Oct 12:47

Added

  • Registry size bands (schema version 2). A conversion profile can apply to a range of model
    sizes: applies_to: parameter_count: {below: 5B}, in the Hugging Face Hub's parameter count. An
    unknown size (no safetensors metadata, packed GPTQ/AWQ weights, a quantization_config) is
    outside every band, and a GGUF import or an OpenVINO copy never chooses a profile by size.
    Validation requires an unbanded fallback per device and reports a band that can never win.
    docs/REGISTRY_SPEC.md section 5.1.
  • "Chosen because". The plan, registry explain and the artifact manifest
    (export.profile_selection) say why the profile was chosen: the size band, the device default,
    a rule naming it, or --profile. inspect shows what the parameter count includes, or why it
    is not used.
  • tools/measure_profiles.py, which prepares and evaluates profiles across many models,
    resumably.
  • tools/registry_bundle.py fills min_tool_version from the registry schema a bundle uses
    (--min-tool-version to raise it), so an older yonga refuses a bundle it cannot parse.
  • SECURITY.md, for reporting a vulnerability privately through a GitHub security advisory, and
    CONTRIBUTING.md.
  • Every AppImage release carries yonga-X.Y.Z-x86_64.AppImage.packages.txt, the exact Python
    packages the image bundles and the sha256 of the yonga wheel; the image holds the same list as
    usr/share/yonga/packages.txt.

Changed

  • Small models get a more accurate default quantization. Below 5B parameters, prepare
    now chooses cpu-int4-asym-g128-r08 on the CPU, the new gpu-int4-asym-g128-r08 on the GPU
    (both INT4 with a fifth of the layers at INT8) and npu-int4-sym-g128 on the NPU. Measured
    against INT8 on twelve models from nine families, 350M to 8B: ratio 0.8 lowers the divergence on
    every one by 16-42 % for 9-18 % more disk and up to about a fifth of the GPU's tokens per second
    (Qwen3-1.7B's perplexity increase goes from +23.7 % to +8.5 %), and group-wise INT4 on the NPU
    roughly halves channel-wise's divergence or better (Qwen2.5-0.5B: +107 % perplexity to +38 %)
    for a tenth to a third of the NPU's tokens per second (--profile npu-int4-sym-cw is the fast
    choice). Models of 5B and above and models whose size is unknown get the unbanded defaults.
    docs/PROFILE_COMPARISON.md section 10 has the tables, and tools/measure_profiles.py
    reproduces them.
  • Exceptions, kept on purpose. The VLM families (Qwen3.5, Qwen2.5-VL) keep their CPU and GPU
    pins to ratio 1.0: no VLM was measured. The generic rule (Mistral and other unrecognised members
    of common families) and Phi-3 keep the NPU on channel-wise INT4: group-wise was never probed on
    the NPU for them.
  • Qwen3 on the NPU at 5B and above now uses channel-wise INT4. It was group-wise for every
    size, from a test of the 0.6B model only. At 8B channel-wise is a third faster (13.9 against
    10.6 tok/s) and much less accurate (+42.7 % against +9.4 % perplexity); the default follows the
    NPU documentation, and --profile npu-int4-sym-g128 keeps the accurate one. A Qwen3 OpenVINO
    copy for the NPU is now stored under npu-int4-sym-cw too, so it is copied again once.
  • A changed default is a new artifact, never a replacement. Preparing a model already prepared
    under the old default exports again under the new profile id, keeps the old artifact, and the
    plan says so beforehand with the --profile that would reuse it. chat, bench and probe
    given a model id use the newest ready artifact.
  • A profile named by an architecture or model rule is now a preference, not an override: it must
    be usable, satisfy its constraints and match its own applies_to, or resolution falls through to
    scoring. Before, a deprecated profile named by a rule was used anyway.
  • The text-model architecture rules no longer name the CPU and GPU profiles (they named the
    defaults), and the measured families no longer pin the NPU, so the size bands decide.
  • Upgrading with registry overrides. The banded profiles sit at priority 110. An override
    profile on the same device at exactly 110 now ties with them for models below 5B, and prepare
    stops with registry.ambiguous_rules (yonga registry validate names the pair): raise it above
    110. One between 101 and 110 still beats the unbanded default but now loses to the band below
    5B. One above 110 wins for every family that does not name a profile.
  • The all extra no longer installs the development tools (pytest, ruff, mypy, build, twine);
    they stay in dev.
  • Releases publish to PyPI only after the AppImage built from the same wheel passes its smoke
    test and the pypi environment is approved. A manual publish runs on the tag itself
    (gh workflow run release.yml --ref vX.Y.Z; docs/RELEASING.md).
  • prepare exits 10 when the host keeps it from using a device. A device failure whose
    diagnostic marks the environment as broken, such as npu.device.permission_denied or
    device.driver_missing, now exits 10 (environment unusable) instead of 50: the device would work
    once the host is fixed. Other device failures in prepare still exit 50, and probe, chat and
    bench keep the category's code.

Fixed

  • A registry file declaring a schema version newer than this build reports "too new" instead of a
    schema violation.
  • A signed registry bundle whose created_at or expires_at carries no UTC offset is refused as
    registry.remote.invalid. Before, a naive expires_at crashed verification.
  • serve gave every HTTP error the error code not_found, a wrong method included. A 404 now
    carries not_found, a 405 method_not_allowed, a 413 request_too_large and any other
    invalid_request.

Security

  • The AppImage imported Python modules from the current directory. Run inside a directory
    holding, say, a typer.py, it imported that file instead of its own. The image now runs
    python -P -m yonga, and every Python child process yonga starts (the optimum-cli export run as
    a module, venv and pip for managed export environments, the exporter-range lookup, the GGUF import)
    runs with -P too, so none of them puts the current directory on sys.path.
  • Registry data can no longer pass optimum-cli the options yonga controls. A profile's
    export_extra_args and an architecture's or model's export.extra_args may not name
    --model/-m, --task, --trust-remote-code, --model-kwargs, --cache_dir, --token,
    --quantization-statistics-path or a bare --, nor the quantization options yonga derives from
    the profile's quantization section (--weight-format, --sym, --ratio, --group-size,
    --backup-precision, --dataset, --all-layers, --awq, --scale-estimation). Abbreviations
    argparse accepts, the --name=value form and -m with its value attached count as well. A
    registry file that does fails to load with registry.schema_violation; a profile built in code
    is refused at export time with export.reserved_extra_args. Before, a profile could enable
    remote code without --trust-remote-code.
  • yonga logs JOB_ID --bundle rewrites the Hugging Face account name to <hf-user>, as it already
    rewrote the home directory to ~, the hostname to <host> and the login name to <user>; the
    host report included the account name.
  • serve read a request body of any size. A body over 8 MiB is now refused with 413 and the
    error code request_too_large, counted on the bytes received, so a chunked body is held to it
    too; with an API key configured, authentication runs first. A request with more than 4096
    messages is refused with 400 and param messages.
  • The serve extra requires FastAPI 0.132 or newer, the first that does not read a body sent
    without a Content-Type as JSON. With an older one and no YONGA_API_KEY, any web page could send
    a chat request to serve on 127.0.0.1, since a browser sends such a body without a CORS
    preflight.

v0.2.0

Choose a tag to compare

@github-actions github-actions released this 27 Sep 21:58

Added

  • Hardware tests. pytest -m intel_gpu and pytest -m intel_npu run on a real Intel GPU and
    NPU; they were empty markers before. They cover generation, compile-cache reuse, probing,
    streaming, serve over the real runtime, GPU logits against CPU, a real GGUF import (GPU
    generation, and the NPU refusal the registry names), and, opt-in with
    YONGA_TEST_VLM_ARTIFACT, multi-turn VLM chat. tools/make_test_artifact.py --npu builds the
    small real LLM the NPU generation tests need.

  • NPU profiles carry MAX_PROMPT_LEN=1024 / MIN_RESPONSE_LEN=128, so the manifest records the
    static-shape budget an NPU artifact was probed with.

  • ProbeStatus.DEFERRED, for a raw compile the pipeline was asked to settle.

  • Managed export environments. When the installed packages cannot meet an architecture's exporter
    requirements -- Qwen3.5 needs transformers==5.2.* -- prepare builds an isolated environment
    under the cache directory and runs the export in it, instead of refusing and leaving the user to
    reshape their own environment. It mirrors the packages yonga runs with and changes only the failing
    pin; torch and torchvision come from PyTorch's CPU index (1.44 GiB against a 6.2 GB CUDA install).
    It is confirmed in the plan, built and verified at preflight before the model download, never used
    half-built, reused by later jobs, and recorded in the manifest as export.environment_key.
    yonga env list|remove manage them and doctor reports their disk use. --no-managed-env restores
    the refusal; --ignore-exporter-requirements exports in the current environment and never builds
    one. Registry data can choose only a version for a package yonga lists, never a package of its own.

  • Ten verified architecture families. Registry rules with tested CPU, GPU and NPU claims for
    qwen3, qwen2, llama, gemma3-text, gemma2, olmo2, granite, smollm3, lfm2 and qwen2.5-vl (NPU
    refused), each measured on the reference host with a small model. phi3 is recorded as not yet
    exportable with optimum-intel 2.2.0.

  • optimum-intel's own Transformers ranges are honoured before the download. Where the registry
    pins nothing, the range optimum-intel's export config declares for the model type (for example
    <=5.0 for Gemma and Qwen2.5-VL) is read in a child process, cached, and satisfied with a
    managed export environment instead of failing after the download.

  • export.native_implementation lets the registry declare that Transformers implements an
    architecture itself. A repository's own modelling code (Phi-4-mini's modeling_phi3.py) is then
    reported as not run, and remote code stays off.

  • Direct GGUF import. A GGUF-only repository with no usable declared derivative is imported
    with OpenVINO GenAI's own GGUF reader when a new registry gguf rule verifies the file's
    architecture (llama, qwen2, qwen3) and every tensor type in it (F32, F16, Q4_0, Q4_K, Q6_K, Q8_0).
    Both are read from the file header over HTTP range requests before anything is downloaded, so a
    Q4_K_M file carrying Q5_0 tensors is passed over with the reason in the plan. Only the chosen
    file is fetched. The import runs in a child process from a symlink in staging, never writing to
    the Hub cache, and the result is validated, repaired and probed like any artifact. It is named by
    the file's quantization (gguf-q4_0-gpu). --gguf-file picks the file;
    eval --allow-different-sources compares a GGUF import with the original weights. A new issue
    rule classifies the NPU's refusal of GGUF-built graphs (npu.pipeline.static_reshape_failed).

  • yonga eval CANDIDATE --reference REF measures what a quantization cost. It compares the
    two artifacts' next-token distributions, teacher-forced, over a shipped corpus of original text
    in eight domains (including Python and Turkish), and reports mean KL divergence, top-1 agreement
    and perplexity, in total and per passage. It runs the IR directly through openvino.Core, on the
    CPU at f32 by default, with one model loaded at a time. It refuses artifacts of different weights
    or vocabularies. New cpu-int8-asym-ref and cpu-fp16-ref profiles prepare references and are
    never a device's default. On Qwen3.5-9B, npu-int4-sym-cw diverges from INT8 2.5 times as much as
    gpu-int4-asym-g128 (KL 0.197 vs 0.080, perplexity +22 % vs +7 %): see PROFILE_COMPARISON.md
    section 8.

  • cpu-int4-asym-g128-r08 profile, opt-in: INT4 with 20 % of layers kept at INT8. On
    Qwen3-0.6B it cuts the divergence from FP16 by 31 % (KL 0.245 vs 0.356) for 10.5 % more disk.
    It is never selected automatically.

  • Signed remote registry bundles: yonga registry update [--url | --file], status, reset.
    A bundle is verified against Ed25519 keys the user trusts (yonga ships none) and refused when
    expired, older than the installed one, or naming a path outside it. Every repair rule must name a
    built-in implementation. The installed bundle is re-verified on every load and ignored if it
    stops verifying. cryptography is now a dependency. tools/registry_bundle.py creates keys and
    signed bundles.

  • yonga publish ARTIFACT --repo NS/NAME uploads a prepared artifact to the Hugging Face Hub
    in one commit, with a model card generated from the manifest. It refuses before uploading unless
    the destination is writable, both source licenses permit redistribution (read fresh from the Hub;
    unrecognised ones need --license-reviewed), the artifact is ready and every file still matches
    its digest. The card says the artifact is a quantized conversion and not a fine-tune, discloses
    repairs, and claims only devices whose probe passed. Private by default; --dry-run writes the
    card locally; local paths, login name and hostname are anonymised.

  • yonga serve MODEL_OR_ARTIFACT [--host] [--port] [--device], an OpenAI-compatible HTTP API:
    GET /health, GET /v1/models, POST /v1/chat/completions with SSE streaming. It serves the same
    session chat opens, loads the model at startup, binds 127.0.0.1 by default, answers one
    request at a time, stops a generation whose streaming client disconnects, hides reasoning while
    streaming too, and refuses parameters it cannot honour with a 400. YONGA_API_KEY enables bearer
    authentication. Needs the new serve extra (FastAPI, uvicorn).

  • VLM generation is stateless where OpenVINO GenAI allows it. GenAI 2026.4's VLMPipeline
    accepts a ChatHistory, so the whole conversation now travels with every turn, as on the LLM
    path. Builds without that overload still use the hidden start_chat() session.

  • Reuse of a completed export after a failed job. A successful export now writes a checkpoint
    (fingerprint of every input, plus each file's size, mtime and small-file digest) into staging. When
    a later stage fails, an identical prepare adopts the retained export, restoring repaired files
    from their backups and rejecting anything that changed, and skips the download and conversion.
    The plan says so and the manifest records export.reused_from_job; --fresh-export declines. An
    export interrupted part-way still restarts: optimum-intel cannot resume one.

  • yonga search QUERY [--limit N] [--device D] searches the Hub in one request and runs each
    result through the registry's identification, marking GGUF-header identities. Compatibility is
    shown per architecture next to a format hint, because an MLX or AWQ repository can report a
    supported architecture it cannot be converted from. Works without OpenVINO installed.

  • yonga logs JOB_ID --bundle writes a bug-report archive: job log and state, artifact manifest
    and validation record, and a host report, redacted again and with the home directory, hostname and
    login name anonymised. yonga logs and the bundle now share one job resolver.

  • probe says before the wait when a pipeline build will compile more than 2 GiB of weights.

  • docs/PROFILE_COMPARISON.md — the same Qwen3.5-9B weights quantized twice
    (gpu-int4-asym-g128 and npu-int4-sym-cw), both converted and benchmarked on this
    host's GPU, with artifact sizes, conversion duration, cold-compile and warm-load
    times and per-prompt TTFT/tok/s/TPOT. It measures size and speed only and says so:
    output quality was not evaluated, and the two profiles differ in pipeline properties
    as well as in quantization, so the timings are not a clean comparison of the two
    quantizations.

  • AppImage. Every release carries yonga-X.Y.Z-x86_64.AppImage: yonga, a relocatable CPython
    3.14 and the openvino and serve extras in one 135 MB file, for x86_64 Linux with glibc 2.28 or
    newer and nothing else installed. tools/build_appimage.sh builds it with every download pinned
    and checksummed, chooses wheels for glibc 2.28 rather than the build host, and fails if a bundled
    library needs anything newer. CI builds and smoke-tests it on every push. Verified on the
    reference host from the image: doctor, bench on the GPU and the NPU, a GGUF import, a full
    conversion through managed environments, and serve.

  • Conversion without an export toolchain. When no optimum-intel sits beside yonga -- a plain
    pip install yonga, or the AppImage -- prepare builds a managed export environment with the
    toolchain yonga is tested with instead of failing at the export. Once it exists, the Transformers
    range optimum-intel declares for the model type is read inside it, and an environment within the
    range replaces one outside it, still before the download (Gemma-3: <=5.0). Inside an AppImage,
    environments are made from a copy of the bundled interpreter kept under the environments
    directory, because a virtual environment made from the image's own mount would break when it
    exits.

  • PyPI publishing through trusted publishing, from release.yml, running only while the
    repository is public, with a manual run to publish an earlier tag (...

Read more