Repository navigation
Releases: ooguz/yonga
Release list
v0.5.0
The first public release, and the first on PyPI. It fixes everything an end-to-end campaign of
fresh-user runs on real hardware found (AppImage, pip, every command, GPU, NPU and CPU), keeps
OpenVINO's usage reporting from running, and installs the AppImage on PATH.
Added
yonga self installandyonga self uninstall. Run from the AppImage,
./yonga-X.Y.Z-x86_64.AppImage self installcopies it to~/.local/bin/yonga(or--bin-dir DIR): atomically (temporary file in the directory,fsync,0755, rename), creating the
directory when missing. The same file again is "already installed"; another yonga AppImage
there is replaced after a question showing both versions where known (--yesanswers it).
Anything else at that name (a pip or pipx command, a symbolic link, a directory, a script) is
refused and left alone; an AppImage is recognised by its ELF andAI\x02magic bytes, never by
running it. Afterwards it checks thatyongaonPATHruns the copy and, if not, says what to do
(log out and in, orexport PATH="$HOME/.local/bin:$PATH") and exits 2; it never edits a shell
startup file and never uses sudo (an unwritable directory exits 10 with the advice to pick one you
own). Outside an AppImage it exits 70 and points pip users at their environment'syongaor
pipx install 'yonga[openvino,export]'.self uninstallremoves only an AppImage, after a
question, and keeps models, logs, caches and settings, saying where they are.
docs/CLI_SPEC.mdsection 12c.yonga registry profiles [--device DEVICE] [--json]lists the conversion profiles
--profileaccepts, with device, weights, size band, priority and status. An unknown
--profilenow names the profiles for the requested device and points to this command.- The preparation plan always shows an
Export environmentrow: a managed environment, or this
environment with the interpreter path and the reason. - Issue rule
gguf-reader-open-failed(export.gguf.open_failed, exit 10). GGUF import failures
are now classified by the registry like optimum export failures instead of always being
unknown_runtime_error. - Issue rule
torchvision-torch-build-mismatch(export.torchvision.build_mismatch, exit 10) for
operator torchvision::nms does not exist.preparealso imports the exporter in its own
environment at preflight, so a broken torch/torchvision pair is reported before the download. - The probe job log records every probe result with its raw runtime text (
probe.result), as the
diagnostics already said it did. yonga evalprintsSaved: PATH, and--jsonincludesresult_path, asbenchdoes.doctorreports whether the exporter imports (exporter import), reusing the checkprepare
recorded, so a CPU torch next to PyPI's torchvision shows as a warning with the registry's
remedy;--fullchecks a recorded failure again.chat,bench,probe,evalandpublishsay which artifact a model id chose when several
ready ones match ("Using ID (PROFILE, DEVICE), the newest of N ready artifacts for MODEL; pass the
artifact id to choose another").- Registry schema 4 (needs yonga 0.5.0): an issue rule's
blocks_exportlists model types whose
export is certain to fail when the rule's constraints hold for the interpreter that will run the
export, andpreparerefuses such an export before the download and before any question
(overridable with--ignore-exporter-requirements, recorded in the manifest).
Changed
doctorreports how yonga is installed. A newInstallline in the Host section names the
AppImage file or the Python environment, and whatyongaonPATHresolves to. From an AppImage
it is a warning withself installas the remedy whenyongais not onPATH, or when it is
another installation (named with its path, for example an older AppImage). For a pip install it
is information only.doctor --jsoncarries it asinstallation.docs/CLI_SPEC.mdsection 3.- A command line yonga cannot parse exits 70 instead of Click's 2: an unknown command or
option, a missing argument, a value of the wrong type (--timeout abc,--older-than abc) and
a bareyonga. Exit 2 is left meaning only warnings or an inconclusive probe;--helpstill
exits 0. preparedownloads only the files the conversion reads. ONNX copies,original/checkpoints,
training runs and a second weight format are skipped (SmolLM2-360M: 693 MiB instead of
4.7 GiB), and the plan's expected download counts the same files.- Automatic profile selection skips a group-wise profile whose group size cannot divide the
model'shidden_size(orintermediate_sizeat ratio 1.0) fromconfig.json, and says why
under "Chosen because". With no fitting profile for an explicit device,preparestops before
the download (exit 70) and names the devices that fit;--device automoves on to the next
device. yonga logslists each job's id prefix, local start time with its zone, command, model, device,
outcome (ok/failed/cancelled/running/unfinished), duration and log size;--jsonadds these
fields. One job's view shows the same facts and prints record times in local time instead of
unmarked UTC.yonga listshows the source revision as 7 characters under aRevheader.- While a job runs, warnings logged by huggingface_hub (such as the Hub's "You are sending
unauthenticated requests" advice) go to the job log aslibrary.logrecords instead of breaking
into the progress line. - An explicit
--profilewhose group size cannot divide the model'sconfig.jsonchannel widths
is refused at the plan, before any download (exit 70,export.quantization.group_size_mismatch).
The message names the profiles that fit, as your command with--profilechanged, and never
swaps the profile itself. When no profile for that device fits (SmolLM2-360M with
--device gpu --profile gpu-int4-asym-g128-r08), it offers the fitting profiles of the other
devices, as your command with--deviceand--profilechanged. - The
context_length_exceededmessage no longer repeats "The prompt". It names the offered tools
and images when they count toward the limit, and says what to do: shorten the conversation, offer
fewer tools or send fewer images, or use a GPU or CPU artifact, whose prompt length is not fixed. - The README quick-start
doctorsample and its exit-2 explanation match the current output.
Fixed
--offline(andYONGA_OFFLINE/HF_HUB_OFFLINE) no longer contacts the Hugging Face Hub for
inspect,prepareor revision resolution. A cached model is inspected and prepared from the
local cache alone, including one downloaded by commit SHA with norefs/main. An uncached
revision, or a snapshot with only metadata files, fails withhub.offline.missingwithout
making any request.yonga inspectno longer says a model needs--trust-remote-codewhen the registry declares a
native Transformers implementation (Phi-4-mini, phi3).inspectandpreparenow share one
remote-code decision.yonga --offline searchreportshub.offline.search_unavailable("Search needs the network")
instead of a misleading cache-miss message.- Re-running
preparefor a model that is already prepared reuses the artifact without asking
anything, so it works in scripts without--yes. The reuse plan no longer lists a download,
temp space, repairs or an export environment. - An explicit
--devicethat the registry records as unsupported for the model (official or
tested evidence) is refused before the download with exit 50 and the recommended fallback,
instead of after a full download and conversion. - Without the OpenVINO runtime installed,
preparestops before the download with exit 10 and
namespip install 'yonga[openvino]', instead of blaming the host for having no devices. It no
longer offers to build a managed export environment then. - A gated or private repository with no Hugging Face token is flagged in the plan and refused
before the download with exit 21. - In a session without a terminal and without
--yes,preparesays the confirmation could not
be asked and suggests--yes, instead of saying the user declined it. - A
preparethat adopts a retained export, or copies published OpenVINO IR, reports this as a
note and exits 0 instead of 2. - The plan's "Estimated artifact" is worked out from the profile's quantization and the model's
parameter count (Qwen3-0.6B INT4: about 516 MiB, previously 1.4 GiB). - The plan's device decision for
--device autostates the actual reason, such as the registry's
recommendation. - NNCF's group-size refusal is classified (
export.quantization.group_size_mismatch, exit 70) even
when its list of refused nodes pushes the exception line out of the exporter's output tail; it
used to end as a genericexport.failed. - The group-size remediation no longer offers the NPU's channel-wise profile on the CPU or GPU,
where itsMAX_PROMPT_LENfails the probe; it suggestscpu-int8-asym-refon the CPU instead.
The Python 3.14 optimum diagnostic now explains what AppImage users can do. - Remediation steps print the real repository and job id instead of
MODELandJOB_IDwherever
they are known (prepare,probe,chat,bench,--json), and a step that repeats an
earlier command is listed once. - "Chosen because" for a profile named by a model or architecture rule gives the parameter count
and the profile the rule overrode, such as the size-banded default (automatic_profile_idin the
manifest). - GGUF import no longer fails with "Failed to open ... with gguf_open" on a huggingface-hub 1.x
cache (read-only blobs). OpenVINO GenAI's reader opens the file read-write, so yonga imports
from a private writable copy in staging (a reflink where the filesystem supports one) instead of
a link into the cache, and the disk plan...
v0.4.0
Added
servehonours OpenAIstop. A string or up to 4 non-empty strings of at most 256
characters; anything else is a 400invalid_requestnamingstop. yonga matches them on the
visible answer after the reasoning sanitiser, never inside a hidden reasoning block, across chunk
boundaries, and ends the answer before the first stop sequence to complete withfinish_reason
stop. A stream holds back only the tail that could still become a match and the whitespace in
front of it. A request with stop sequences is
streamed internally even when the client did not ask for a stream, so generation ends at the
match.docs/CLI_SPEC.mdsection 14.servetakes images on a VLM artifact. OpenAIimage_urlcontent parts holding a
data:image/(png|jpeg|webp|gif);base64,URL. Every other URL (http,https,file, ...) is
refused with a 400 naming the content part, permanently: yonga never fetches an image. At most 4
images per request, each at most 4096 x 4096 pixels, checked from the image header before
decoding and refused rather than scaled down. Images go with the last user message only, because
OpenVINO GenAI 2026.4 binds images to the newest turn of a stateless conversation; an image in an
earlier message is a 400serve.request.images_earlier_turnsaying so. An LLM artifact refuses
images withserve.request.images_unsupported. Every other check,stopand the generation
settings included, comes before any image is read. Only the headers are read, off the event
loop, before a request waits for the pipeline; the pixels are decoded once it holds the
pipeline, so at most one request holds decoded images however many are queued. Any failure
inside Pillow is a refusal naming the image, never a server error. EXIF rotation is applied as a browser
shows it (an unreadable EXIF block leaves the image as stored), and 16-bit greyscale is scaled to
8 bits rather than clipped.docs/CLI_SPEC.mdsection 14.serve --carry-images-forwardmoves images of earlier messages onto the last user message,
with a line telling the model they came earlier, for clients that resend every image of a
conversation. The model then no longer knows which words each image went with.serve --max-body-mb Nsets the request body limit, 1 to 64 MiB (8 by default), since
embedded images count towards it.servetakes OpenAItoolson models the registry has seen calling them. A chat profile's
newtool_callingentry (registry schema 3, below) says how a family writes calls; only a model
whose profile declares one, and whose own chat template contains the entry's
template_containsliteral, acceptstools. Built in: Qwen2.5, Qwen3 and SmolLM3
(<tool_call>{json}</tool_call>) and Granite (a JSON list), each seen producing a parseable call
on the reference host on 2026-10-09 (CPU, one prompt each). Every other model refusestools
with a 400serve.request.tools_unsupportednamingtoolsand saying why. The tools, an
assistant'stool_callsandtoolresults reach the artifact's own chat template through
GenAI'sChatHistory.set_tools; a conversation may end on a tool result. The answer is read for
calls outside reasoning, also with--show-reasoning, and returned asmessage.tool_callswith
finish_reasontool_calls; streamed, text is released until a call could begin and each call
then arrives as one complete delta, and a client that disconnects meanwhile still stops
generation. An answer that cannot be read as calls to the offered functions is returned
as text, never as a guessed call, and logged as such without its content.tool_choiceauto
andnoneare honoured;requiredand a named function are refused (they need constrained
decoding).parallel_tool_calls: falsekeeps the first call.stopsequences match only the
text outside the calls.docs/CLI_SPEC.mdsection 14.- Registry schema version 3:
tool_callingon a chat profile. It selects one of the tool-call
parsers compiled into yonga (tagged_json,json_list), fills in its markers and field names,
and must stateevidencetestedorofficialwith at least one source. Older files load
unchanged. A rule using it needsschema_version: 3, inherited from its document as size bands
are;tools/registry_bundle.pygives a schema-3 bundlemin_tool_version0.4.0, so older
releases refuse it atregistry update.registry validateand remote bundles refuse a format
this yonga has no parser for.yonga registry explainshows the chosen profile's tool-call
format.docs/REGISTRY_SPEC.mdsection 8b. yonga probe --quickand--timeout SECONDS. Settling a deferred NPU compile means building
the GenAI pipeline, which for a 9B model ran past 15 minutes.--quicknever builds it;
--timeoutbuilds it in a child process (python -P -m yonga.adapters.smoke_child) and stops it
at the limit, since a native compile cannot be interrupted in-process. Either way an unsettled
device gets the new probe statusinconclusive: never counted as working, never as a failure, and
shown as such in the tables, the role matrix and--json("inconclusive": true); a model card
never counts it as working. A probe where nothing passed or failed outright exits 2. A child that
dies without a result is a failure classified by the registry from its stderr. Without either
option nothing changes, andpreparestill always waits.python -m yonga.application.probe
takes both options too.docs/CLI_SPEC.mdsection 11.yonga cache listandyonga cache prune: the OpenVINO compile cache is no longer
write-only. Every compile cache entry yonga uses now gets an ownership record,
.yonga-entry.json(artifact id and path, repository, profile, device, OpenVINO version, runtime
properties, created and last used); writing it never fails a load.cache list(or plain
yonga cache) shows each entry's size, last use, device, OpenVINO version, owner and status:
in-use,orphaned(the artifact is gone),stale-runtime(built by an OpenVINO release other
than the installed one, so never reused) orunattributed. Entries made before the records are
attributed by recomputing the keys of every artifact in the store and every artifact a saved
benchmark names. Stale-runtime is relative to the OpenVINO running the command; an entry of
another release used after this one first built an entry here belongs to another install
sharing the cache and staysin-use.cache pruneremoves orphaned and stale-runtime entries,
skipping any entry used after the plan was made (unless--all);--older-than DAYS,
--unattributedand--allgo further,--dry-runonly shows the plan, and it asks first
unless--yes. It deletes only key-named, real directories directly under the compile cache
root, never follows a symbolic link, and reports the bytes freed.--jsonon both.
docs/CLI_SPEC.mdsection 12b.doctorreports the compile cache under Storage: entries, size and the reclaimable part,
a warning namingyonga cache prunefrom 10 GiB reclaimable or 25 % of the free space on the
cache filesystem (never below 1 GiB), and the numbers undercompile_cachein--json.
Changed
- Qwen2, SmolLM3 and Granite have their own chat profiles (
chat-qwen2,chat-smollm3,
chat-granite) instead ofchat-default. They copy its policy exactly (thinking detection,
<think>stripping, 4096 new tokens, no sampling defaults) and add onlytool_calling, sochat
behaves as before;registry explainnames the new rule.chat-qwen2matches the whole qwen2
model type, but only Qwen2.5 was measured, and only an artifact whose template carries
<tool_call>(the Qwen2.5 Instruct models, not the original Qwen2 Instruct) accepts tools. - Qwen3 MoE has its own chat profile,
chat-qwen3_moe, split fromchat-qwen3with the same
policy and notool_calling: the MoE models share the template but were not measured writing
calls, so they refusetools.chatbehaves as before. serveacceptstoolmessages and assistanttool_callsin a conversation, alongside
tools. Assistanttool_callsused to be dropped without an error, andtoolmessages were
refused as an unknown role; now both are accepted alongsidetools, and withouttoolsthey are
a 400 naming the message. The legacyfunctionsandfunction_callare refused by their own name
instead of astools.- The registry chat profile's
stop_stringsis applied. The field was parsed and never used; it
now reaches OpenVINO GenAI'sGenerationConfig.stop_stringsas model-level end markers, left out
of the answer. Empty and duplicate entries are dropped when a rule loads. A benchmark's settings
fingerprint includes stop strings only when there are some, so every recorded fingerprint still
matches. No built-in rule declares any, so nothing shipped behaves differently.
docs/REGISTRY_SPEC.mdsection 8b. - The
serveextra installs Pillow, which decodes images. It is imported only when an image
arrives: without it text is served as before,servewarns at startup for a VLM artifact, and an
image request is a 500serve.images.dependency_missingnamingpip install 'yonga[serve]'.
The AppImage, which bundles the extra, grows by Pillow's size. <ov_genai_image_N>and<ov_genai_video_N>are removed from text sent to a VLM artifact,
inchatas inserve, repeatedly until none is left (removing one must not join a new one).
GenAI reads these universal tags anywhere as "image N goes here", so typed text could claim an
image that was not attached and fail generation. A model's own vision tokens typed as text are
not removed; which they are is per-family knowledge the registry does not hold yet.- The VLM families follow the size bands. Qwen2.5-VL and Qwen3.5 ...
v0.3.0
Added
- Registry size bands (schema version 2). A conversion profile can apply to a range of model
sizes:applies_to: parameter_count: {below: 5B}, in the Hugging Face Hub's parameter count. An
unknown size (no safetensors metadata, packed GPTQ/AWQ weights, aquantization_config) is
outside every band, and a GGUF import or an OpenVINO copy never chooses a profile by size.
Validation requires an unbanded fallback per device and reports a band that can never win.
docs/REGISTRY_SPEC.mdsection 5.1. - "Chosen because". The plan,
registry explainand the artifact manifest
(export.profile_selection) say why the profile was chosen: the size band, the device default,
a rule naming it, or--profile.inspectshows what the parameter count includes, or why it
is not used. tools/measure_profiles.py, which prepares and evaluates profiles across many models,
resumably.tools/registry_bundle.pyfillsmin_tool_versionfrom the registry schema a bundle uses
(--min-tool-versionto raise it), so an older yonga refuses a bundle it cannot parse.SECURITY.md, for reporting a vulnerability privately through a GitHub security advisory, and
CONTRIBUTING.md.- Every AppImage release carries
yonga-X.Y.Z-x86_64.AppImage.packages.txt, the exact Python
packages the image bundles and the sha256 of the yonga wheel; the image holds the same list as
usr/share/yonga/packages.txt.
Changed
- Small models get a more accurate default quantization. Below 5B parameters,
prepare
now choosescpu-int4-asym-g128-r08on the CPU, the newgpu-int4-asym-g128-r08on the GPU
(both INT4 with a fifth of the layers at INT8) andnpu-int4-sym-g128on the NPU. Measured
against INT8 on twelve models from nine families, 350M to 8B: ratio 0.8 lowers the divergence on
every one by 16-42 % for 9-18 % more disk and up to about a fifth of the GPU's tokens per second
(Qwen3-1.7B's perplexity increase goes from +23.7 % to +8.5 %), and group-wise INT4 on the NPU
roughly halves channel-wise's divergence or better (Qwen2.5-0.5B: +107 % perplexity to +38 %)
for a tenth to a third of the NPU's tokens per second (--profile npu-int4-sym-cwis the fast
choice). Models of 5B and above and models whose size is unknown get the unbanded defaults.
docs/PROFILE_COMPARISON.mdsection 10 has the tables, andtools/measure_profiles.py
reproduces them. - Exceptions, kept on purpose. The VLM families (Qwen3.5, Qwen2.5-VL) keep their CPU and GPU
pins to ratio 1.0: no VLM was measured. The generic rule (Mistral and other unrecognised members
of common families) and Phi-3 keep the NPU on channel-wise INT4: group-wise was never probed on
the NPU for them. - Qwen3 on the NPU at 5B and above now uses channel-wise INT4. It was group-wise for every
size, from a test of the 0.6B model only. At 8B channel-wise is a third faster (13.9 against
10.6 tok/s) and much less accurate (+42.7 % against +9.4 % perplexity); the default follows the
NPU documentation, and--profile npu-int4-sym-g128keeps the accurate one. A Qwen3 OpenVINO
copy for the NPU is now stored undernpu-int4-sym-cwtoo, so it is copied again once. - A changed default is a new artifact, never a replacement. Preparing a model already prepared
under the old default exports again under the new profile id, keeps the old artifact, and the
plan says so beforehand with the--profilethat would reuse it.chat,benchandprobe
given a model id use the newest ready artifact. - A profile named by an architecture or model rule is now a preference, not an override: it must
be usable, satisfy its constraints and match its ownapplies_to, or resolution falls through to
scoring. Before, a deprecated profile named by a rule was used anyway. - The text-model architecture rules no longer name the CPU and GPU profiles (they named the
defaults), and the measured families no longer pin the NPU, so the size bands decide. - Upgrading with registry overrides. The banded profiles sit at priority 110. An override
profile on the same device at exactly 110 now ties with them for models below 5B, andprepare
stops withregistry.ambiguous_rules(yonga registry validatenames the pair): raise it above
110. One between 101 and 110 still beats the unbanded default but now loses to the band below
5B. One above 110 wins for every family that does not name a profile. - The
allextra no longer installs the development tools (pytest, ruff, mypy, build, twine);
they stay indev. - Releases publish to PyPI only after the AppImage built from the same wheel passes its smoke
test and thepypienvironment is approved. A manual publish runs on the tag itself
(gh workflow run release.yml --ref vX.Y.Z;docs/RELEASING.md). prepareexits 10 when the host keeps it from using a device. A device failure whose
diagnostic marks the environment as broken, such asnpu.device.permission_deniedor
device.driver_missing, now exits 10 (environment unusable) instead of 50: the device would work
once the host is fixed. Other device failures inpreparestill exit 50, andprobe,chatand
benchkeep the category's code.
Fixed
- A registry file declaring a schema version newer than this build reports "too new" instead of a
schema violation. - A signed registry bundle whose
created_atorexpires_atcarries no UTC offset is refused as
registry.remote.invalid. Before, a naiveexpires_atcrashed verification. servegave every HTTP error the error codenot_found, a wrong method included. A 404 now
carriesnot_found, a 405method_not_allowed, a 413request_too_largeand any other
invalid_request.
Security
- The AppImage imported Python modules from the current directory. Run inside a directory
holding, say, atyper.py, it imported that file instead of its own. The image now runs
python -P -m yonga, and every Python child process yonga starts (theoptimum-cliexport run as
a module,venvandpipfor managed export environments, the exporter-range lookup, the GGUF import)
runs with-Ptoo, so none of them puts the current directory onsys.path. - Registry data can no longer pass optimum-cli the options yonga controls. A profile's
export_extra_argsand an architecture's or model'sexport.extra_argsmay not name
--model/-m,--task,--trust-remote-code,--model-kwargs,--cache_dir,--token,
--quantization-statistics-pathor a bare--, nor the quantization options yonga derives from
the profile'squantizationsection (--weight-format,--sym,--ratio,--group-size,
--backup-precision,--dataset,--all-layers,--awq,--scale-estimation). Abbreviations
argparse accepts, the--name=valueform and-mwith its value attached count as well. A
registry file that does fails to load withregistry.schema_violation; a profile built in code
is refused at export time withexport.reserved_extra_args. Before, a profile could enable
remote code without--trust-remote-code. yonga logs JOB_ID --bundlerewrites the Hugging Face account name to<hf-user>, as it already
rewrote the home directory to~, the hostname to<host>and the login name to<user>; the
host report included the account name.serveread a request body of any size. A body over 8 MiB is now refused with 413 and the
error coderequest_too_large, counted on the bytes received, so a chunked body is held to it
too; with an API key configured, authentication runs first. A request with more than 4096
messages is refused with 400 andparammessages.- The
serveextra requires FastAPI 0.132 or newer, the first that does not read a body sent
without a Content-Type as JSON. With an older one and noYONGA_API_KEY, any web page could send
a chat request toserveon 127.0.0.1, since a browser sends such a body without a CORS
preflight.
v0.2.0
Added
-
Hardware tests.
pytest -m intel_gpuandpytest -m intel_npurun on a real Intel GPU and
NPU; they were empty markers before. They cover generation, compile-cache reuse, probing,
streaming,serveover the real runtime, GPU logits against CPU, a real GGUF import (GPU
generation, and the NPU refusal the registry names), and, opt-in with
YONGA_TEST_VLM_ARTIFACT, multi-turn VLM chat.tools/make_test_artifact.py --npubuilds the
small real LLM the NPU generation tests need. -
NPU profiles carry
MAX_PROMPT_LEN=1024/MIN_RESPONSE_LEN=128, so the manifest records the
static-shape budget an NPU artifact was probed with. -
ProbeStatus.DEFERRED, for a raw compile the pipeline was asked to settle. -
Managed export environments. When the installed packages cannot meet an architecture's exporter
requirements -- Qwen3.5 needstransformers==5.2.*--preparebuilds an isolated environment
under the cache directory and runs the export in it, instead of refusing and leaving the user to
reshape their own environment. It mirrors the packages yonga runs with and changes only the failing
pin; torch and torchvision come from PyTorch's CPU index (1.44 GiB against a 6.2 GB CUDA install).
It is confirmed in the plan, built and verified at preflight before the model download, never used
half-built, reused by later jobs, and recorded in the manifest asexport.environment_key.
yonga env list|removemanage them anddoctorreports their disk use.--no-managed-envrestores
the refusal;--ignore-exporter-requirementsexports in the current environment and never builds
one. Registry data can choose only a version for a package yonga lists, never a package of its own. -
Ten verified architecture families. Registry rules with tested CPU, GPU and NPU claims for
qwen3, qwen2, llama, gemma3-text, gemma2, olmo2, granite, smollm3, lfm2 and qwen2.5-vl (NPU
refused), each measured on the reference host with a small model. phi3 is recorded as not yet
exportable with optimum-intel 2.2.0. -
optimum-intel's own Transformers ranges are honoured before the download. Where the registry
pins nothing, the range optimum-intel's export config declares for the model type (for example
<=5.0for Gemma and Qwen2.5-VL) is read in a child process, cached, and satisfied with a
managed export environment instead of failing after the download. -
export.native_implementationlets the registry declare that Transformers implements an
architecture itself. A repository's own modelling code (Phi-4-mini'smodeling_phi3.py) is then
reported as not run, and remote code stays off. -
Direct GGUF import. A GGUF-only repository with no usable declared derivative is imported
with OpenVINO GenAI's own GGUF reader when a new registryggufrule verifies the file's
architecture (llama, qwen2, qwen3) and every tensor type in it (F32, F16, Q4_0, Q4_K, Q6_K, Q8_0).
Both are read from the file header over HTTP range requests before anything is downloaded, so a
Q4_K_Mfile carryingQ5_0tensors is passed over with the reason in the plan. Only the chosen
file is fetched. The import runs in a child process from a symlink in staging, never writing to
the Hub cache, and the result is validated, repaired and probed like any artifact. It is named by
the file's quantization (gguf-q4_0-gpu).--gguf-filepicks the file;
eval --allow-different-sourcescompares a GGUF import with the original weights. A new issue
rule classifies the NPU's refusal of GGUF-built graphs (npu.pipeline.static_reshape_failed). -
yonga eval CANDIDATE --reference REFmeasures what a quantization cost. It compares the
two artifacts' next-token distributions, teacher-forced, over a shipped corpus of original text
in eight domains (including Python and Turkish), and reports mean KL divergence, top-1 agreement
and perplexity, in total and per passage. It runs the IR directly throughopenvino.Core, on the
CPU at f32 by default, with one model loaded at a time. It refuses artifacts of different weights
or vocabularies. Newcpu-int8-asym-refandcpu-fp16-refprofiles prepare references and are
never a device's default. On Qwen3.5-9B,npu-int4-sym-cwdiverges from INT8 2.5 times as much as
gpu-int4-asym-g128(KL 0.197 vs 0.080, perplexity +22 % vs +7 %): seePROFILE_COMPARISON.md
section 8. -
cpu-int4-asym-g128-r08profile, opt-in: INT4 with 20 % of layers kept at INT8. On
Qwen3-0.6B it cuts the divergence from FP16 by 31 % (KL 0.245 vs 0.356) for 10.5 % more disk.
It is never selected automatically. -
Signed remote registry bundles:
yonga registry update [--url | --file],status,reset.
A bundle is verified against Ed25519 keys the user trusts (yonga ships none) and refused when
expired, older than the installed one, or naming a path outside it. Every repair rule must name a
built-in implementation. The installed bundle is re-verified on every load and ignored if it
stops verifying.cryptographyis now a dependency.tools/registry_bundle.pycreates keys and
signed bundles. -
yonga publish ARTIFACT --repo NS/NAMEuploads a prepared artifact to the Hugging Face Hub
in one commit, with a model card generated from the manifest. It refuses before uploading unless
the destination is writable, both source licenses permit redistribution (read fresh from the Hub;
unrecognised ones need--license-reviewed), the artifact is ready and every file still matches
its digest. The card says the artifact is a quantized conversion and not a fine-tune, discloses
repairs, and claims only devices whose probe passed. Private by default;--dry-runwrites the
card locally; local paths, login name and hostname are anonymised. -
yonga serve MODEL_OR_ARTIFACT [--host] [--port] [--device], an OpenAI-compatible HTTP API:
GET /health,GET /v1/models,POST /v1/chat/completionswith SSE streaming. It serves the same
sessionchatopens, loads the model at startup, binds127.0.0.1by default, answers one
request at a time, stops a generation whose streaming client disconnects, hides reasoning while
streaming too, and refuses parameters it cannot honour with a 400.YONGA_API_KEYenables bearer
authentication. Needs the newserveextra (FastAPI, uvicorn). -
VLM generation is stateless where OpenVINO GenAI allows it. GenAI 2026.4's
VLMPipeline
accepts aChatHistory, so the whole conversation now travels with every turn, as on the LLM
path. Builds without that overload still use the hiddenstart_chat()session. -
Reuse of a completed export after a failed job. A successful export now writes a checkpoint
(fingerprint of every input, plus each file's size, mtime and small-file digest) into staging. When
a later stage fails, an identicalprepareadopts the retained export, restoring repaired files
from their backups and rejecting anything that changed, and skips the download and conversion.
The plan says so and the manifest recordsexport.reused_from_job;--fresh-exportdeclines. An
export interrupted part-way still restarts: optimum-intel cannot resume one. -
yonga search QUERY [--limit N] [--device D]searches the Hub in one request and runs each
result through the registry's identification, marking GGUF-header identities. Compatibility is
shown per architecture next to a format hint, because an MLX or AWQ repository can report a
supported architecture it cannot be converted from. Works without OpenVINO installed. -
yonga logs JOB_ID --bundlewrites a bug-report archive: job log and state, artifact manifest
and validation record, and a host report, redacted again and with the home directory, hostname and
login name anonymised.yonga logsand the bundle now share one job resolver. -
probesays before the wait when a pipeline build will compile more than 2 GiB of weights. -
docs/PROFILE_COMPARISON.md— the same Qwen3.5-9B weights quantized twice
(gpu-int4-asym-g128andnpu-int4-sym-cw), both converted and benchmarked on this
host's GPU, with artifact sizes, conversion duration, cold-compile and warm-load
times and per-prompt TTFT/tok/s/TPOT. It measures size and speed only and says so:
output quality was not evaluated, and the two profiles differ in pipeline properties
as well as in quantization, so the timings are not a clean comparison of the two
quantizations. -
AppImage. Every release carries
yonga-X.Y.Z-x86_64.AppImage: yonga, a relocatable CPython
3.14 and theopenvinoandserveextras in one 135 MB file, for x86_64 Linux with glibc 2.28 or
newer and nothing else installed.tools/build_appimage.shbuilds it with every download pinned
and checksummed, chooses wheels for glibc 2.28 rather than the build host, and fails if a bundled
library needs anything newer. CI builds and smoke-tests it on every push. Verified on the
reference host from the image:doctor,benchon the GPU and the NPU, a GGUF import, a full
conversion through managed environments, andserve. -
Conversion without an export toolchain. When no optimum-intel sits beside yonga -- a plain
pip install yonga, or the AppImage --preparebuilds a managed export environment with the
toolchain yonga is tested with instead of failing at the export. Once it exists, the Transformers
range optimum-intel declares for the model type is read inside it, and an environment within the
range replaces one outside it, still before the download (Gemma-3:<=5.0). Inside an AppImage,
environments are made from a copy of the bundled interpreter kept under the environments
directory, because a virtual environment made from the image's own mount would break when it
exits. -
PyPI publishing through trusted publishing, from
release.yml, running only while the
repository is public, with a manual run to publish an earlier tag (...