Releases: emergenthq-net/gpu-broker
Releases · emergenthq-net/gpu-broker
Release list
gpu-broker 0.5.0
Added
gpu-broker connectagainst a broker with cloud failover keeps Claude Code's and Codex's own
provider keys and adds the broker key asx-gpu-broker-key; other clients keep the broker key
alone.gpu-broker clientsshows which keys each client sends. Connect asks
GET /v1/upstreams/passthrough, which a client key may read (only whether failover is on and
which APIs pass a client's key through);/v1/upstreamsstays main-token only.- Failover: unstreamed requests go upstream as streams and are reassembled into the provider's
exact unstreamed answer, so a down provider fails over afterfirst_byte_sand a long answer
is never cut off; a provider that refuses streams is asked again unstreamed. - Cloud first, local on failure (
upstreams:, docs/failover.md):claude-*/gpt-*go to the
real provider with the client's own key and fall back to a local model on an outage, a hang,
5xx/529 or used-up quota or credit; client errors are returned as-is. A circuit breaker per
provider (and per key for quota) skips a failing provider at once and probes it back.
Answers carryx-gpu-broker-served-byandx-gpu-broker-fallback; a dashboard card shows
the breakers and failover events. The broker credential may be sent asx-gpu-broker-key.
A client's own key goes only to providers withpass_client_key(the official hosts by default). - An MCP server (the
mcpextra): tools to list models, generate images and video, edit
images, animate an image, make 3D splats, poll and fetch jobs, ask the local LLM and read the
GPU. Streamable HTTP at/mcpbehind the chat-route credentials (a client key sees only its
own jobs), andgpu-broker mcpover stdio. Long jobs return their id aftermcp.wait_s;
small images come back inline.gpu-broker connectregisters it with Claude Code, Claude
Desktop and Codex (claude-code-mcp,claude-desktop,codex-mcp)./healthreportsmcp.
A client key sees only jobs it owns (by key id) and sends files as base64 unless
mcp.client_url_inputs; tools take only models that can run now. - A video's length is the request param
num_frames;framesis only the list of input views
(a number there is refused with a pointer tonum_frames). comfy.auth_env: a ComfyUI behind a Bearer token.connectwrites symlinked configs through, never replaces a user's owngpu-brokerMCP
server, warns when a running Claude Code drops its entry, and lists the files it changed.- The GPU thread's queue is
fairby default: interactive jobs first, then the requester that
has used the least expected GPU time, so one busy client no longer holds everyone else's jobs
behind its backlog.- A call waiting for a pool slot no longer blocks the jobs behind it.
scheduler.policy: fifo(envBROKER_SCHEDULER_POLICY) restores strict arrival order.- Background jobs age into the interactive class after
scheduler.max_wait_s. - A switch never jumps a call waiting for a slot on the resident model for longer than
scheduler.evict_wait_sunless it is of a higher class. /v1/jobstakesx-priority; without it, a job's class followsdefaults.background_requesters.
gpu-broker replay <events.jsonl>: replays a broker's event log through the queue rules in
simulated time and reports waits per requester, jobs/hour and residency churn, beside what the
broker did. Offline; the yardstick for scheduler changes.
Fixed
- A job's new state and its
job.<state>event are written together, so a client that sees a job
finish in/v1/jobsalways finds itsjob.done/job.failedin/v1/events. - MCP
initializereports the installed gpu-broker version inserverInfo(it was empty).
Changed
- With
fair, a/v1/jobsLLM call from a requester not indefaults.background_requesters
(and withoutx-priority: background) is interactive, so it may use the slots kept by
reserved_interactive; before, only chat routes classified their calls. - With
fair, a requester indefaults.background_requesterscan only lower its class with
x-priority(orinteractivein a job), not raise it.scheduler.may_claim_interactive
lists who may claim interactive (default: everyone not inbackground_requesters).fifo
keeps the old rule. - While quiesced (
/v1/admin/quiesce), new jobs, chats, embeddings and sessions get HTTP 503
withRetry-After(intervals.quiesced_retry_s, default 5) andx-should-retry: true,
instead of being queued and then failed as orphans by the restart. - A restart re-queues jobs the previous process had queued but not started, in their original
order and with their staged input files (eventjob.requeued).- Jobs that had started, and direct chats, are still failed as orphaned.
- A re-queued job fails with "input files missing" if its files are gone, and with
"model '' no longer in catalog" if its model left the catalog; a re-queued interactive
session fails with "session expired by restart".
gpu-broker 0.4.0
Added
gpu-broker setup: one command from install to a running broker, asking no questions.- Finds the GPU, and llama.cpp, vLLM, Ollama and ComfyUI on their usual ports, with the
systemd units or Docker containers that run them. - Writes the config and catalog for them (the starter files if it finds nothing) and a new
API token inbroker.env(mode 600). Existing files and the token are kept. - Installs and starts the
gpu-brokerservice, never as root: a user service (lingering)
forsystemctl --usermodel servers; else, with root or passwordless sudo, a system
service run as the installation's owner or a dedicatedgpu-brokeraccount, with a
visudo-checked sudoers rule for exactly the catalog's units. It refuses code anyone else
could change. Otherwise it prints theservecommand, and runs it in the foreground at a
terminal. - Runs
check, waits for/health, and opens the dashboard (not over SSH). --dry-run,--yes,--dir. Seedocs/setup.md.
- Finds the GPU, and llama.cpp, vLLM, Ollama and ComfyUI on their usual ports, with the
- Catalog:
health_pathfor LLM servers without/health(Ollama:/api/version). - README: a TLDR block first, and the logo (light and dark).
gpu-broker init: writes a starterconfig.yamlandcatalog.yamlto/etc/gpu-broker
(or--dir) and creates the folders they name.- Existing files are kept unless
--force.
- Existing files are kept unless
- Docs pages:
docs/catalog.md,docs/exec-recipes.md,docs/api.md,docs/hardware.md,
docs/threat-model.md, andAGENTS.mdfor automated reviewers.
Changed
- README: leads with the one-person case; installs from PyPI; reference material moved to
docs/. - Internal: modules that mixed two responsibilities are split, with no behaviour change.
catalogschema(entry types and validation) out ofcatalog.settingsschema(the dataclasses) out ofsettings.admission(job submit) out ofbroker.drivers.systemdanddrivers.dockerout ofdrivers.local.demo.assembleout ofdemo.run.- The old modules re-export the moved names, except the two concrete drivers: import
SystemdDriverandDockerDriverfromgpu_broker.drivers.systemdandgpu_broker.drivers.docker.
gpu-broker 0.3.2
Fixed
- The sdist is self-testing: unpack it, install it with
[dev], runpytest.MANIFEST.in
shipstests/(withtests/helpers.pyand the fixtures),host/,examples/andscripts/.
The amdgpu fixtures'/proc/<pid>/fdsymlinks, which an sdist drops because they dangle,
are listed intests/fixtures/amdgpu/fd-links.txtand recreated in a temp dir at test time.
CI checks this on every change.
gpu-broker 0.3.1
Changed
- Tests and examples use neutral container ids and paths, and the demo's leak check reads its
denylist fromGPU_BROKER_LEAK_DENYLISTinstead of spelling it out. New
scripts/leak_scan.pyscans the tree and the built sdist and wheel against that denylist;
CI (jobleak-scan) and the release workflow run it before anything is published.
gpu-broker 0.3.0
Added
gpu-broker demo: the real broker, API and dashboard on a simulated GPU, with a simulated
team sending chats, images and videos. No graphics card, model servers or downloads needed;
it opens the dashboard already signed in (#token=link). See README, Try it.- Drop-in replacement for the OpenAI and Anthropic APIs:
POST /v1/messages(Anthropic
Messages, plain and streaming, tools, images, thinking),POST /v1/embeddings(routed to a
caps: [embed]model),x-api-keyauth besideAuthorization: Bearer, andmodel_map
(BROKER_MODEL_MAP) mapping hosted names such asgpt-*andclaude-*to catalog models.
See docs/drop-in.md. - Input images: jobs may carry
image/end_image(base64 or a data: URL), or
image_url/end_image_urlwhenimages.allow_urlsis on (off by default; fetches are
limited to public addresses unlessurl_allow_networksallows more). Catalog entries say
which images a model takes (images:), and substitution only picks models that take them. - Image-to-video and image-edit templates: Wan 2.2 14B i2v, Wan 2.2 5B with a start image,
MiniMax first/last frame, HunyuanVideo 1.5 i2v, Qwen-Image edit and FLUX.2 Klein edit. - Exec runner (
runner: exec): a model can be a program run from a recipe file
(recipes_dir) on the GPU's machine, with no shell, fixed paths and validated parameters
(exec.params,exec.choices).examples/recipes/sharp.recipeturns one image into a 3D
Gaussian splat. Recipeoutputsmay list several globs. Prune units for old job folders are
inexamples/systemd/. - Per-model request defaults in the catalog (
defaults:): request > catalog entry > template. - The dashboard can run image-taking models with an uploaded image, and its model index says
who has the GPU: the resident model, the job it is lent to, or nobody. - AMD GPUs: the amdgpu driver's sysfs files and
/proc/*/fdinfo, no ROCm tools needed
(gpu.vendor,gpu.index; on Proxmoxhost/gpu-broker-gpu). See README, Hardware. /v1/metricssamples carryprocs_unreadable(per-process memory could not be read: it
needs root orCAP_SYS_PTRACE), and/v1/gpunames the GPU reader in use (probe).
Changed
- GPU sample field
sm_mhzis nowclock_mhz(the vendor-neutral name). - Only VRAM used and total are required in a GPU reading; utilisation, power, temperature and
clock may benull(unknown) in/v1/gpuand/v1/metrics. servestarts when no GPU can be read yet, and keeps retrying, instead of refusing to start.
gpu.vendor: autois chosen in the background, withintimeouts.gpu_query_sof wall-clock
time;/v1/gpuanswers{"state": "probing"}meanwhile and never waits onnvidia-smi.- Proxmox host script:
GPU_BUDGET_S(default 8) bounds a wholegpureading, the vendor
choice included;NV_TIMEOUT_Snow bounds eachnvidia-smicall only.
Deprecated
sm_mhzin/v1/metricssamples: kept as an alias ofclock_mhzfor this release only;
it is removed in the next one.
gpu-broker 0.2.0
First public release.
One GPU, many models: an HTTP broker that queues requests and swaps LLM servers and ComfyUI workloads in and out of a single GPU on demand.
- OpenAI-compatible chat endpoint with streaming; interactive chats skip the background queue
- Parallel LLM slots, with a reserved slot for interactive users
- Automatic model switching, substitution and on-demand downloads from a YAML catalog
- Drivers for systemd, Docker and Proxmox hosts
- Built-in dashboard with live GPU metrics and a click-to-use model index
- Bundled ComfyUI graphs for Qwen-Image, Chroma, FLUX.2 Klein, Wan 2.2, HunyuanVideo 1.5, MiniMax H3 and LTX 2.5
See the README for install and configuration.