Repository navigation
Replies: 2 comments
Testing external backends: Donato's Strix Halo toolboxesThis branch adds **external backends**: a JSON manifest describes an out-of-tree engine, and `lemond` discovers it, launches it as a subprocess, health-checks it, and forwards its OpenAI routes to it. No C++ rebuild for a new engine. This walkthrough is the OCI-container use case from #3585. You run one of Donato's What you are exercising: manifest discovery, platform (OS x accelerator) selection, container launch and teardown, health probing, and route proxying.
Prerequisites
1. Get and build the branchgit clone https://github.com/abn/lemonade.git lemonade
cd lemonade
git checkout feat/external-backend-system
./setup.sh
cmake --build --preset default --target lemond lemonade2. Create the scratch dirs and the manifestThe manifest must exist before the server starts. Put it under export TEST=/tmp/lemonade-toolbox-test
mkdir -p "$TEST"/{cache,config,xdg/lemonade/backends,hf}Write the file below to {
"recipe": "toolbox-llamacpp",
"display_name": "kyuz0 Strix Halo llama.cpp toolbox",
"api_contract_version": "1",
"capabilities": ["chat_completion", "completion"],
"health_probe": {
"type": "http", "endpoint": "/health", "expected_status": 200,
"timeout_seconds": 300, "poll_interval_ms": 500
},
"platforms": {
"linux": {
"vulkan": {
"command": "podman",
"args": [
"run", "--rm", "--init", "--name", "lemonade-{recipe}-{port}",
"--device", "/dev/dri", "--group-add", "video",
"--security-opt", "seccomp=unconfined",
"--security-opt", "label=disable",
"-v", "{hf_cache}:/models:ro", "--network", "host",
"docker.io/kyuz0/amd-strix-halo-toolboxes:vulkan-radv",
"llama-server", "-m", "/models/{checkpoint_relative:main}",
"--host", "127.0.0.1", "--port", "{port}", "-c", "{ctx_size}"
],
"stop_command": "podman",
"stop_command_args": ["rm", "-f", "lemonade-{recipe}-{port}"]
},
"rocm": {
"command": "podman",
"args": [
"run", "--rm", "--init", "--name", "lemonade-{recipe}-{port}",
"--device", "/dev/dri", "--device", "/dev/kfd",
"--group-add", "video", "--group-add", "render",
"--security-opt", "seccomp=unconfined",
"--security-opt", "label=disable",
"-v", "{hf_cache}:/models:ro", "--network", "host",
"docker.io/kyuz0/amd-strix-halo-toolboxes:rocm-10.0",
"llama-server", "-m", "/models/{checkpoint_relative:main}",
"--host", "127.0.0.1", "--port", "{port}", "-c", "{ctx_size}"
],
"stop_command": "podman",
"stop_command_args": ["rm", "-f", "lemonade-{recipe}-{port}"]
}
}
}
}If the image does not accept The image is named explicitly because these toolboxes exist for Strix Halo (gfx1151) only. When a recipe must name a different image or artifact per hardware arch, a platform block accepts an 3. Start an isolated serverUse a scratch directory and a non-default port so you do not touch any Lemonade you already run. XDG_CONFIG_HOME="$TEST/xdg" HF_HOME="$TEST/hf" \
./build/lemond "$TEST/cache" "$TEST/config" --port 13350 &Confirm it can see the recipe: curl -s http://127.0.0.1:13350/api/v1/system-info \
| python3 -c 'import sys,json; print("toolbox-llamacpp" in json.load(sys.stdin)["recipes"])'4. Register a small model, load, and chatCustom models must be named ./build/lemonade --no-discovery --port 13350 pull user.toolbox-270m \
--recipe toolbox-llamacpp \
--checkpoint main unsloth/gemma-3-270m-it-GGUF:gemma-3-270m-it-UD-IQ2_M.gguf
./build/lemonade --no-discovery --port 13350 load user.toolbox-270m
curl -s http://127.0.0.1:13350/api/v1/chat/completions \
-H 'Content-Type: application/json' \
-d '{"model":"user.toolbox-270m","messages":[{"role":"user","content":"In one word, say OK."}],"max_tokens":8}'A short 5. Confirm lifecyclepodman ps --filter name=lemonade- # the engine container is running
./build/lemonade --no-discovery --port 13350 unload
podman ps -a --filter name=lemonade- # empty: the container was removed
kill %1 # stop the test server6. Pin the image by digest (recommended)Tags move; digests do not. What to reportPlease note what worked and what did not, with the OS, kernel, GPU ( |
|
Update: This is an archive for anyone who wants to build on it, or parts of it. Other in-flight RFCs have momentum and priority. I'm not maintaining or pushing this beyond this point - fork and take what you want. Happy to answer questions if someone picks it up. |
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Proposal size: Major feature (spans multiple well-scoped PRs).
Prototype: https://github.com/abn/lemonade/tree/feat/external-backend-system
This RFC covers how an out-of-tree inference engine is described, discovered, launched, and managed by
lemond. The companion Sandbox RFC owns process confinement, and the boundary between the two is stated in section 12. Section 21 states the scope of the first implementation and what it leaves out.1. Why this exists
Lemonade ships a fixed set of backends. If you want to run an engine Lemonade does not build, your options today are to rebuild the C++ server or to wait for a release. Both are wrong for the cases people actually have: a community
llama.cppbuild with extra kernels, a vendor nightly, a container image pinned to a digest, an internal engine that only exists in your network.This RFC makes the engine describable. A small JSON manifest says which executable to run and what it can do.
lemondreads the manifest at runtime, starts the engine as a subprocess, waits for it to answer a health probe, and forwards the fixed Lemonade routes to it. Nothing in the C++ tree changes, so a new engine arrives as a file, not a rebuild.The design keeps one hard line: a manifest describes an engine, it does not program the server. There is no scripting and no request rewriting. If an engine needs more than a manifest can say, it graduates to a built-in backend, and that path stays open.
2. User stories
As an operator, I want to register a third-party or containerized inference engine (custom
llama.cppbuild, vLLM, an NPU runner) withlemondas a declarative manifest, so I can run the engine that fits my hardware without rebuilding the C++ server.As an engine or fork developer, I want to publish my engine, or a fork of an existing one, as a self-describing manifest, so it is discoverable, installable, and reviewable in Lemonade.
As a user with a fork or repackaged build of an already-built-in backend, I want to register it as a
variant_ofthat reuses the built-in's handling, so it runs without re-declaring argv or touching C++.3. High-level design
An out-of-tree engine has one of two shapes:
lemondorchestrates the engine directly from the manifest'scommandandargs. This covers most engines, including containerizedpodmananddockerrecipes.variant_of. The recipe is a binary drop-in for a named built-in.lemondasks the built-in for the argv it would have used and runs that against the external binary, so the manifest declares binary provenance and deltas only.Both run the engine as an out-of-process subprocess. In-process
dlopenplugins are permanently out of scope.A backend never reads
/sysor/procto learn about its host. Centralized hardware and model facts are resolved bylemondand handed to the engine as launch-time tokens. A queryable fact channel (an engine asking for facts at runtime) belongs to a future adapter RFC.An engine that outgrows both shapes graduates to a built-in
WrappedServer. An out-of-process adapter tier is a possible expansion and would be its own RFC.4. The manifest contract
The manifest is validated against the published JSON Schema, pinned by a lock file. Its JSON Schema major is
api_contract_version, so there is no second version number to keep in sync.Validation is strict, and the parser rejects rather than ignores. A manifest is rejected, with a message naming the offending field, when it:
recipe,display_name,api_contract_version,capabilities,platforms);additionalProperties: false);capabilities,capability_enable_args, orendpoints;api_contract_versionthan the running server supports;extendsandvariant_of;variant_ofbinary pinned by hash without asha256.Silent acceptance of an unknown field or capability is a contract violation, not a tolerance.
4.1 Identity
recipe: stable id,^[a-z0-9][a-z0-9_-]{1,63}$. It must not collide with a built-in recipe or with another external recipe found at a higher-priority path.display_name: human-readable name.api_contract_version:"1"for this major.extends(optional): declarative inheritance, see section 6.variant_of(optional): procedural inheritance, see section 7.4.2 Execution block
Each platform block carries
command,args, an optionalstop_commandandstop_command_args, an optionalworking_dir, and anenvobject. Avariant_ofblock also carriesbinary,argv_extra, andreserved_args, and may override the top-levelsource,sha256, andversion_policy. Tokens are resolved when the process is spawned, not when the manifest is read.Normative token vocabulary (v1). The vocabulary is fixed for the contract major and additive across majors. Unknown tokens fail at discovery.
{port},{host},{pid},{log_level},{recipe},{model_name}{checkpoint:NAME},{checkpoint_relative:NAME},{resolved_path},{model_relative_path},{model_dir},{exe_dir},{hf_cache},{cache_dir}{rocm_arch},{cuda_arch},{arch_alias},{target_device},{hip_visible_devices},{cuda_visible_devices},{rocr_visible_devices},{ggml_vk_visible_devices},{ze_affinity_mask}{ctx_size},{batch_size},{ubatch_size},{threads},{cache_type_k},{cache_type_v}{custom:NAME},{custom:NAME:-DEFAULT},{custom_args}{env:NAME},{env:NAME:-DEFAULT}Resolution and injection rules.
execvp; there is no shell string forcommandorargs.-is rejected regardless of argv position, unless it is a validated negative number (-0,-12,-1.5), which the existing option parser already accepts. This covers positional tokens such as{model_name}as well as option values.{custom:NAME}resolves any key present inrecipe_options, including one inherited throughextends.custom_optionsonly controls CLI and config exposure; it is not a whitelist for{custom:}resolution.{custom_args}expands a user string or JSON array into separate argv elements using the existing shell-quote-aware parser, then validates each againstreserved_args. An exact match or a leading<flag>=is rejected.{env:NAME}reads the parent environment through a default-deny allowlist.LEMONADE_*names and secret-shaped names (*API_KEY*,*TOKEN*,*SECRET*,*PASS*,*AUTH*) are never resolvable, so a secret cannot be placed on a command line. The allowlist is implementation-defined and documented. The v1 default is empty plus credential-free cache and CA variables; proxy URL variables are excluded because their values may embed credentials.env_allowlist. The former chooses which host values a manifest may interpolate; the latter chooses which parent variables a child inherits and is enforced by the Sandbox RFC.stop_commandandstop_command_argsare executed as an argv vector. A manifest that needs a container runtime suppliespodmanordockeras thecommand, not a shell pipeline.4.3 Lifecycle fields
health_probe:typeishttp,tcp, orprocess.httpprobesendpointforexpected_status;tcpchecks that{port}is connectable;processtreats the child surviving past spawn as ready.timeout_secondsandpoll_interval_msbound the probe. A child that exits during the probe fails fast instead of waiting out the timeout.requested_ports: how many ports the engine needs, default 1. v1 defines the single-port contract ({port}); multi-port is a reserved extension and any value other than 1 is rejected.reserved_args: flags a user's{custom_args}must not override. Per-blockreserved_argsare unioned with the top-level list, with the block taking precedence.model_management:lemond_managed(default) orself_managed. In v1,self_managedmeanslemonddoes not pre-download model weights for the recipe. The model-inventory and readiness RPC is deferred (section 19), soself_managedimplies nothing else.4.4 Governance
slot_policyisstandard,exclusive_npu,coexist_by_type, orunmetered, and maps directly onto the existing router slot policies.default_acceleratorselects the platform block when the request names no device.4.5 Extension fields
These were kept from the prototype as additive surface. They do not open routes.
endpoints: remaps a declared capability's nonstandard engine path onto Lemonade's fixed route, for engines whose internal path differs from the OpenAI shape. The value is an engine-side path.custom_options: declarative CLI and config knobs the recipe exposes, surfaced in the CLI and config file.downsize_endpoint: engine-side path invoked on soft-idle downsize.4.6 Arch overrides and aliases
The platform matrix is host OS x accelerator. Hardware arch is a further axis, handled two ways so nothing is duplicated.
archinside a platform block is a map of arch pattern to a partial block. The block stays the common config and each entry names only its deltas. Keys are exact arch strings (gfx1151,sm_90) or globs (gfx115*); exact wins, otherwise the first matching glob. Merging is field-wise: a present field replaces the base field andenvmerges key-wise. The only pattern syntax is a glob, so this stays declarative.arch_aliasesat the top maps an arch (or glob) to a short alias. The{arch_alias}token resolves through it, so a name, tag, or path that varies by arch is written once.Selection order for a load: OS, then accelerator (model
device, elsedefault_accelerator, else detection), then the matching arch override. The same merge runs at install time, so the artifact the CLI fetches and the argvlemondlaunches cannot disagree. When an artifact exists only inside an arch override, install needs the arch:lemonade backends install-external --arch gfx1151.5. Capabilities
Membership in the capability set is a canonical registry. A capability maps to exactly one deployment mode and one or more fixed Lemonade routes. A manifest never opens a route: an operation outside the declared set is answered with the existing
unsupported_operationenvelope.Manifest capability names are not the same as the server's deployment-mode labels. The mapping is explicit and normative:
chat_completionPOST /v1/chat/completionscompletionPOST /v1/completionsembeddingsPOST /v1/embeddingsrerankingPOST /v1/rerank(plus/reranking,/rerankeraliases)transcriptionPOST /v1/audio/transcriptionsimagePOST /v1/images/generations,/edits,/variationsttsPOST /v1/audio/speechresponsesPOST /v1/responsesstreaming_transcriptionvariant_ofonly)/realtimeclassificationPOST /v1/classifyaudio_generationPOST /v1/audio/generationsmodel_3dPOST /v1/3d/generationsslotsGET /v1/slots,POST /v1/slots/{id}tokenizePOST /v1/tokenizeThe seven core capabilities are the closed set defined by this specification. The five extensions are additive and kept from the prototype. New capability kinds are reviewed core changes; a manifest cannot introduce one.
streaming_transcriptionisvariant_of-only in v1. The Realtime WS route needs protocol framing that manifest passthrough cannot express, and protocol translation is an explicit non-goal.capability_enable_argsattaches argv when the capability's deployment mode is the one the model loads in. A model deploys in exactly one mode, so every declared capability that shares that mode is enabled together. A manifest that needs finer activation than that does not fit this surface.Gating is per route, not per interface.
chat_completionandcompletionshare thechatmode but are distinct routes, andWrappedServerbundles both inICompletionServer, so the external path consults the declared capability set at the route throughhas_capability, not throughdynamic_cast.6.
extendsextendsinherits declarative data only. It merges the base manifest'srecipe_optionsandcustom_optionsinto the child, with the child's keys taking precedence where they overlap. It never supplies execution: the child still provides its own platform blocks. Capabilities and platform blocks are not inherited. A cycle is rejected. If the base cannot be resolved, the child is dropped with a recorded reason.extendsandvariant_ofare mutually exclusive.This is narrower than inheritance across the whole manifest. If you want a fork that reuses a built-in's argv construction, that is
variant_of; if you want shared default option values, that isextends.7.
variant_ofA
variant_ofrecipe's binary is wire-compatible with a named built-in.lemondbuilds the argv it would have handed the built-in and runs that against the external binary. The manifest reduces to:source(optional),sha256,version_policy;binary,argv_extra, andreserved_argsdeltas.The value must name a built-in backend, and specifically one that exposes a launch plan. The supported values today are
llamacpp,whispercpp, andsd-cpp. Any other value is rejected at load with a message naming them: an unknown recipe, an external recipe, or a built-in that has not yet grown a launch plan.version_policydefaults topinned, which requiressha256.roll_forwardis an explicit opt-out for authors tracking a moving artifact.sourceis fetched and hash-verified by the CLI;lemondnever downloads executables. Avariant_ofwithoutsourceuses a user-provided binary at the declaredbinarypath.source,sha256, andversion_policymay also be set inside a platform block, overriding the top-level values for that block. That is what lets a project that publishes one artifact per platform or accelerator be described in a single manifest. At install time the CLI selects the host OS block and fetches its effective source, verifying its effective hash unless its effective policy isroll_forward. If the manifest offers more than one artifact for the host OS, the CLI requires--acceleratorand names the choices rather than guessing.This is distinct from the
*_binconfig override._binis a global, mutable, single-slot change with no versioning and no separate lifecycle.variant_ofis per-recipe, pin-verified, and separately supervised.Reuse needs a seam. Built-in argv construction is inline in each
WrappedServersubclass'sload(). A launch-plan step lets a backend produce(executable, argv, environment, working_dir)without spawning;variant_ofthen re-targets the executable and appendsargv_extra.llamacpp,whispercpp, andsd-cppexpose that step; the other built-ins return a clear error until they do.variant_ofis about backend type, not runtime. It answers one question: which built-in's argument handling does this binary follow? It says nothing about how the binary is obtained or what launches it. A container image is a runtime choice that can hold any backend, so a containerized fork ofllamacppis described as a passthrough manifest (the runtime owns the image and the argv), not asvariant_of. Combining the two, reusing a built-in's argv while wrapping it in a container runtime, is a provisioning concern and is not part of this surface.8. Discovery and integrity
Descriptor directories, highest priority first:
$XDG_CONFIG_HOME/lemonade/backends/or%APPDATA%\Lemonade\backends\.<cache>/lemonade/backends/or%USERPROFILE%\.cache\lemonade\backends\./usr/share/lemonade-server/backends/,/usr/local/share/lemonade-server/backends/,/Library/Application Support/Lemonade/backends/,/etc/lemonade/backends/, or%ProgramData%\Lemonade\backends\.A descriptor in a higher-priority directory shadows the same recipe lower down. A recipe that collides with a built-in is rejected rather than shadowed.
Validation on POSIX rejects symlinked descriptor files, files with any group or world write bit (
mode & 0022), user files not owned bygeteuid(), system files not owned by root, and ancestor directories that are group or world writable without the sticky bit. On Windows it rejects owner SIDs outside the process user or administrators and DACLs that grant write toEveryoneorUsers.lemondnever downloads executables or model weights. Container images should be pinned by digest, because a mutable tag can change under you. This is a recommended best practice, not an enforced rule: the image reference lives inside the free-formcommandandargs, whichlemonddoes not parse. Enforcing digest pinning is a future enhancement that needs a first-class container field (image, runtime, devices, mounts) the server turns into runtime argv (section 19).Descriptors that fail validation are rejected with a clear message and recorded in the registry's rejected list. They are never silently loaded. In this build the rejected list is not yet logged or exposed over an endpoint, which makes a misconfigured manifest harder to diagnose than it should be (section 21).
lemondscans the discovery directories when it first needs an external recipe, and/system-infoenumerates them on first use. Once an external backend has been created, a background watcher re-scans every 2 seconds, so manifest edits are picked up without a restart.9. Execution and lifecycle
At load time,
lemondselects a platform block, resolves tokens, spawns the process, probes for readiness, and only then marks the model ready. On unload it runs the stop command and stops the child.Platform selection, in order:
devicerecipe option, when the manifest has a block with that exact key.default_accelerator, when the manifest declares one and the block exists.metalon macOS, otherwisecudawhen a CUDA architecture is detected, thenrocm, thenvulkan,gpu, andcpu.If the manifest has no block for the host OS, the load fails with a clear error.
For a passthrough recipe,
commandandargsare resolved into argv and env. For avariant_ofrecipe,lemondcalls the built-in's launch plan, substitutes<cache>/external/<recipe>/<binary>for the executable, appendsargv_extra, and merges the plan's environment with the manifest'senv. A missing binary fails the load and namesinstall-external.The health probe runs with the configured timeout and poll interval. A child that dies during the probe fails fast. If the probe type is
http, an after-ready watchdog keeps checking the endpoint.health_probe: processtreats survival past spawn as ready, which is the weakest option and only fits engines that expose no port.On unload, the stop command runs first (as an argv vector, with tokens resolved from the load-time snapshot), then
lemondstops the child. On soft-idle downsize,lemondcalls the manifest'sdownsize_endpointif it declares one.10. Management and consent
Models reference a recipe through
user_models.json(recipeplusrecipe_options). The management surface islemonade backends(list, install, uninstall) and the external subcommands:These commands operate locally and do not require the server or an admin API key. Before doing anything, they print the manifest's provenance (display name,
variant_of,source,sha256,version_policy, and declared capabilities) and this disclosure:Confirmation is interactive.
--yesskips the prompt. On a non-interactive terminal without--yes, the command refuses rather than assuming consent.install-externaldownloads and verifies avariant_ofbinary. It picks the artifact for the host platform; if the manifest offers more than one artifact for this host OS, it requires--acceleratorand names the choices. For a passthrough manifest with nosource, there is nothing to download, and the command reports that the binary is user-provided.uninstall-externalremoves the manifest and the install directory, and refuses to remove a system descriptor (remove those with the package manager).lemonade backends --alllists discovered recipes, and external recipes appear inGET /v1/system-infowith"is_external": true.When the Sandbox RFC adds the grant block as an additive contract change, grants and the re-consent-on-widening rule move there and are surfaced through the same consent flow.
11. Versioning and compatibility
Schema major equals
api_contract_version. The server accepts manifests at or below its supported major and rejects a newer one with a clear message rather than partially honoring it. Evolution is additive: per-major schema files, a lock file pinning each released major, load-time shims that upgrade older documents to the latest internal form, and retention of every prior major. Because the versioned schema isadditionalProperties: false, the only additive-within-a-major surfaces are the explicitly openextensionsbag andrecipe_options. A new first-class field, capability name, or token requires a new major. The v1 schema isreleased: falseuntil it is published; editing it before then requires a reviewed lock refresh.12. Security boundaries
Two things sit on opposite sides of a line, and conflating them is the mistake this section exists to prevent.
What this RFC does. It validates descriptors and their file permissions, restricts tokens to a closed vocabulary, gates
{env:}through a default-deny allowlist, rejects leading-dash token values, enforcesreserved_args, and keepslemondfrom downloading executables or model weights. The CLI hash-verifies the binaries it does download.What this RFC does not do. It does not confine the child process. There is no filesystem confinement, no ambient-environment scrubbing, no egress policy, and no sandbox grant block. An external backend runs as a subprocess with
lemond's own privileges, so a manifest can run any command thelemonduser can run. Treat every manifest as trusted code. Those protections belong to the companion Sandbox RFC, which also ownssandbox statusand grant consent.One gap inside the boundary: manifest
envkeys are literal and are not filtered against loader variables such asLD_PRELOADorLD_LIBRARY_PATH. The allowlist governs the{env:}token, not theenvblock's keys. Until the Sandbox RFC lands, a manifest that you choose to install can set any environment value it likes on its own child.13. Worked examples
The three worked examples below follow.
13.1 Passthrough: a user-installed Vulkan engine
{ "recipe": "llamacpp-vulkan-custom", "display_name": "llama.cpp Server (Vulkan Custom)", "api_contract_version": "1", "capabilities": ["chat_completion", "completion"], "slot_policy": "standard", "health_probe": { "type": "http", "endpoint": "/health", "expected_status": 200, "timeout_seconds": 60, "poll_interval_ms": 200 }, "platforms": { "linux": { "vulkan": { "command": "/opt/lemonade/bin/llamacpp/vulkan/llama-server", "args": [ "-m", "{checkpoint:main}", "--host", "{host}", "--port", "{port}", "-c", "{ctx_size}", "-t", "{threads}" ], "stop_command": "kill", "stop_command_args": ["-9", "{pid}"], "env": { "GGML_VK_VISIBLE_DEVICES": "{custom:vk_device:-0}" } } } } }You install and manage the binary yourself.
lemondruns it with the tokens resolved to the current model and port. There is nothing to download and nosourcefield.13.2 Passthrough: a digest-pinned container
The containerized
dflash-rocmexample (abridged here to the parts that matter) runs a container withpodman, mounts the model directory read-only, and pins the image by digest:{ "recipe": "dflash-rocm", "api_contract_version": "1", "capabilities": ["chat_completion", "completion"], "health_probe": {"type": "http", "endpoint": "/health", "timeout_seconds": 300, "poll_interval_ms": 500}, "platforms": { "linux": { "rocm": { "command": "podman", "args": [ "run", "--rm", "--name", "lemonade-{recipe}-{port}", "--device", "/dev/kfd", "--device", "/dev/dri", "--group-add", "video", "--group-add", "render", "-v", "{model_dir}:/models:ro", "--network", "host", "ghcr.io/luce-org/lucebox-hub:rocm-7.2@sha256:9f3a...c2", "/opt/lucebox-hub/server/build/dflash_server", "/models/{checkpoint_relative:main}", "--host", "{host}", "--port", "{port}" ], "stop_command": "podman", "stop_command_args": ["rm", "-f", "-t", "0", "lemonade-{recipe}-{port}"] } } } }The stop command is not a shell pipeline, it is
podman rm. The image digest is the pin; a mutable tag would not be. Nothing enforces the digest: it is a convention the author follows and a reviewer can check by reading the args.13.3
variant_of: allama.cppfork or nightly{ "recipe": "llamacpp-rocm-nightly", "display_name": "llama.cpp ROCm Nightly (community fork)", "api_contract_version": "1", "variant_of": "llamacpp", "capabilities": ["chat_completion", "completion", "embeddings"], "source": "https://example.invalid/llamacpp-rocm/nightly/llama-server-linux-rocm.tar.gz", "version_policy": "pinned", "sha256": "sha256:0123456789abcdef0123456789abcdef0123456789abcdef0123456789abcdef", "platforms": { "linux": { "rocm": { "binary": "llama-server", "argv_extra": ["--gpu-layers", "{custom:gpu_layers:-0}"], "reserved_args": ["--port", "--host", "-m"] } } } }install-externalfetches the archive, verifies the hash, extracts it under<cache>/external/<recipe>/, andlemondlaunches that binary with the built-in's argv.argv_extraadds only the fork's own flags, andreserved_argsstops a model option from overriding the flagslemondalready builds.13.4
variant_of: one artifact per platformA fork that publishes a separate archive for each accelerator sets the provenance inside each platform block instead of at the top level. The top-level
sourcestays unset, and every block carries its own URL and hash:{ "recipe": "llamacpp-per-platform", "display_name": "llama.cpp Fork (one artifact per accelerator)", "api_contract_version": "1", "variant_of": "llamacpp", "capabilities": ["chat_completion", "completion"], "platforms": { "linux": { "rocm": { "binary": "llama-server", "source": "https://example.invalid/llamacpp-fork/linux-rocm.tar.gz", "sha256": "sha256:1111111111111111111111111111111111111111111111111111111111111111", "argv_extra": ["--gpu-layers", "999"], "reserved_args": ["--port", "--host", "-m"] }, "cuda": { "binary": "llama-server", "source": "https://example.invalid/llamacpp-fork/linux-cuda.tar.gz", "sha256": "sha256:2222222222222222222222222222222222222222222222222222222222222222" }, "cpu": { "binary": "llama-server", "source": "https://example.invalid/llamacpp-fork/linux-cpu.tar.gz", "sha256": "sha256:3333333333333333333333333333333333333333333333333333333333333333" } }, "windows": { "cuda": { "binary": "llama-server.exe", "source": "https://example.invalid/llamacpp-fork/windows-cuda.zip", "sha256": "sha256:4444444444444444444444444444444444444444444444444444444444444444" } } } }On a Linux host with both a ROCm and a CUDA GPU,
lemonade backends install-external llamacpp-per-platformcannot choose for you and fails with the available accelerators;--accelerator rocm(orcuda, orcpu) selects one. On a single-artifact host the choice is unambiguous.lemondthen runs the block that matches its own platform selection, so the installed artifact and the launched argv agree.13.5 Arch overrides and aliases
Two DRY shapes: one for a name that changes by arch, one for structure that changes by arch.
Image name by alias. One block, one alias table, no duplicated block:
{ "recipe": "toolbox-arch-alias", "arch_aliases": { "gfx1151": "strix-halo", "gfx1150": "strix-point" }, "api_contract_version": "1", "capabilities": ["chat_completion"], "platforms": { "linux": { "vulkan": { "command": "podman", "args": ["run", "--rm", "docker.io/kyuz0/amd-{arch_alias}-toolboxes:vulkan-radv", "llama-server", "-m", "/models/{checkpoint_relative:main}", "--port", "{port}"] } } } }On a gfx1151 host the image resolves to
...amd-strix-halo-toolboxes...; on gfx1150 it becomes...amd-strix-point-toolboxes.... The naming rule lives inarch_aliasesonce. If the arch is unmapped,{arch_alias}fails loudly instead of substituting an empty string.Per-arch artifacts by override. The base block holds the shared config; each arch names only its artifact and deltas:
{ "recipe": "llamacpp-arch-overrides", "variant_of": "llamacpp", "arch_aliases": { "gfx1151": "strix-halo" }, "api_contract_version": "1", "capabilities": ["chat_completion"], "platforms": { "linux": { "rocm": { "binary": "llama-server", "argv_extra": ["--device-name", "{arch_alias}"], "arch": { "gfx1151": { "source": "https://example.invalid/gfx1151.tar.gz", "sha256": "sha256:5555555555555555555555555555555555555555555555555555555555555555", "argv_extra": ["--mtp"] }, "gfx115*": { "source": "https://example.invalid/gfx115x.tar.gz", "sha256": "sha256:6666666666666666666666666666666666666666666666666666666666666666" } } } } } }install-external llamacpp-arch-overrides --arch gfx1151fetches the gfx1151 artifact;lemondmerges the same entry when it launches. Thegfx115*entry covers every other gfx115x arch from one declaration. Both examples are also shipped under the example manifests.14. Scenarios
These are the concrete cases those pieces enable. Each one maps to fields defined in section 4.
A community
llama.cppbuild on Vulkan. You compiledllama-serveryourself for a GPU Lemonade does not cover. You write a passthrough manifest with alinux.vulkanblock, pointcommandat the binary, and declarechat_completionandcompletion. Drop it in~/.config/lemonade/backends/. It appears inlemonade backends --all, and once a model is registered against the recipe, requests flow. No rebuild.A containerized engine pinned to a digest. You want an engine you cannot build locally, and you want it to run identically every time. A passthrough manifest with
podman runand an image digest gives you that.{model_dir}mounts the model,{checkpoint_relative:main}points the container at it, and{recipe}and{port}keep the container name unique.lemondhealth-probes the port and later runspodman rm.A fork of a built-in. Your
whisper.cppfork adds a flag. You do not want to describe the whole argv; you want the built-in's careful argument handling against your binary. Avariant_of: "whispercpp"manifest withsource,sha256,binary, andargv_extradoes that. The CLI verifies the download, andlemondonly ever launches what the CLI installed.An engine that manages its own weights. An internal engine pulls its own model files out of band, so Lemonade must not try to download them. You set
model_management: "self_managed"and register the model;lemondskips the pre-download and the completeness check. A deeper readiness handshake is deferred, so the engine still has to answer the health probe.An engine that serves more than chat. A local server exposes embeddings, reranking, and a classify route. You declare all of them, and
lemondgates each route on the capability you declared. A request to a route you did not declare comes back asunsupported_operation, not a confusing proxy error.15. Maintenance plan
CI. JSON-Schema validation of manifests, the schema-lock test, registry unit tests (ownership and permissions, discovery priority,
extendsmerging, token resolution, version rejection), and a live server integration suite on Linux x86_64 with a fake engine. Container smoke tests run wherepodmanis available.Human. macOS and Windows path validation runs through the existing cross-platform CI. New-engine recipes are community-submitted and triaged; they are never merged by automation alone. Adding a capability name or a token is a reviewed contract change (section 5).
Long term. The capability set and token vocabulary are additive and reviewed. Every new schema major is frozen in the lock file and kept.
16. Risks
reserved_args, and graduation to a built-in keep the manifest from becoming a program.{env:}allowlist, no binary downloads bylemond, and hash pinning all raise the cost. They do not confine the process, which is the Sandbox RFC's job.variant_ofargv reuse is a real refactor across built-in backends. It lands incrementally; the initial implementation exposes launch plans forllamacpp,whispercpp, andsd-cpp.variant_ofplus the declarative surface answer the "declarative cannot cover existing backends" concern. 1:1 idempotency with built-in backends stays an explicit non-goal.17. Breaking changes
None for existing users. The change is additive. One behavior change: descriptor files in the search paths are now validated strictly, so a non-compliant descriptor is rejected with a clear error instead of being silently loaded.
18. Non-goals
19. Deferred items
sandbox status(Sandbox RFC).self_managedsupport: the model-inventory RPC and readiness handshake.requested_portssemantics beyond the single-port v1 contract.variant_ofargv in a container runtime. Today provisioning is a single HTTPS download with a pinned hash; anything richer needs its own RFC.20. Design decisions
{env:VAR}is allowlist-only and default-deny.chat_completionandcompletionare distinct routes; the bundled interface is too coarse.extensionsbag allows in-major additions.streaming_transcriptionisvariant_of-only in v1.21. Scope of the first implementation
This RFC lands as a series of additive changes. Built-in backends and existing configs are untouched.
The first implementation covers:
api_contract_version, unknown field, unknown capability, unknown token, provenance conditional, andextends/variant_ofexclusivity rejection.{custom_args}expansion.responses.variant_ofthrough the launch-plan seam, provided byllamacpp,whispercpp, andsd-cpp, with per-platform and per-archsource/sha256/version_policyoverrides,arch_aliaseswith the{arch_alias}token, CLI-side hash-verified install, and--accelerator/--archselection when a recipe publishes one artifact per platform or arch.extendsmergingrecipe_optionsandcustom_optionswith cycle detection.self_managedskipping the pre-download and completeness check.capability_enable_argsapplied per deployment mode, anddownsize_endpointinvoked on soft-idle downsize.recipe_optionsandcustom_optionsregistered so{custom:}and{custom_args}resolve from model options./system-infolisting, the discovery watcher, andlemonade backends install-external/uninstall-externalwith a provenance and no-sandbox disclosure.Not in the first implementation (none of these contradict the schema):
variant_ofcovers only built-ins that expose a launch plan. The others return a clear error until they do.requested_ports > 1is rejected; v1 is single-port.self_managedreadiness RPC is deferred. The flag only suppresses auto-download and completeness checks.imagecapability.22. Proposals this RFC resolves
Two open proposals shaped this surface. Both are served, in part, by what is described here, and each is explicitly bounded.
#3583, "RFC: user-defined list of backend forks." The proposal asks for a named list of alternate builds of an existing backend and a way to pick one per model.
variant_ofis the mechanism. Each fork is a small manifest that names the built-in it is wire-compatible with (variant_of: llamacpp), where the binary comes from (source,sha256,version_policy), and only the argv deltas (argv_extra,reserved_args). A model picks a fork by naming its recipe, andlemondreuses the built-in's argv construction, so adding a fork needs no C++ and no rebuild. Two differences from the exact CLI in the proposal are deliberate: selection is by recipe at model registration rather than a--backendflag onload, and every fork is pin-verified and separately supervised. The*_binconfig override the proposal currently relies on is unaffected by this RFC.#3585, "Support for OCI backends." A containerized engine is a passthrough manifest:
commandispodmanordocker,argsarerun ... image@sha256:..., and the runtime pulls and caches the image itself. The digest is the pin, the runtime is the isolation boundary, and thedflash-rocmexample is exactly this shape. A container-only engine such as HaloGen fits as a passthrough recipe with its own argv. A container image that happens to be a fork of a built-in, such as a llama.cpp or whisper container, also fits as passthrough, not asvariant_of, for the reason in section 7:variant_ofis about backend type, not runtime.What this RFC does not add for #3585, and why:
The split to hold onto:
variant_ofanswers which built-in's argument handling a binary follows, and runtime answers how that binary is obtained and launched. Keeping them separate is what lets one mechanism describe a host fork and the other describe a container without either pretending to be the other.23. References
All reactions