Skip to content

Releases: emergenthq-net/gpu-broker

gpu-broker 0.5.0

Choose a tag to compare

@aphilips aphilips released this 05 Oct 01:02
c27b18d

Added

  • gpu-broker connect against a broker with cloud failover keeps Claude Code's and Codex's own
    provider keys and adds the broker key as x-gpu-broker-key; other clients keep the broker key
    alone. gpu-broker clients shows which keys each client sends. Connect asks
    GET /v1/upstreams/passthrough, which a client key may read (only whether failover is on and
    which APIs pass a client's key through); /v1/upstreams stays main-token only.
  • Failover: unstreamed requests go upstream as streams and are reassembled into the provider's
    exact unstreamed answer, so a down provider fails over after first_byte_s and a long answer
    is never cut off; a provider that refuses streams is asked again unstreamed.
  • Cloud first, local on failure (upstreams:, docs/failover.md): claude-* / gpt-* go to the
    real provider with the client's own key and fall back to a local model on an outage, a hang,
    5xx/529 or used-up quota or credit; client errors are returned as-is. A circuit breaker per
    provider (and per key for quota) skips a failing provider at once and probes it back.
    Answers carry x-gpu-broker-served-by and x-gpu-broker-fallback; a dashboard card shows
    the breakers and failover events. The broker credential may be sent as x-gpu-broker-key.
    A client's own key goes only to providers with pass_client_key (the official hosts by default).
  • An MCP server (the mcp extra): tools to list models, generate images and video, edit
    images, animate an image, make 3D splats, poll and fetch jobs, ask the local LLM and read the
    GPU. Streamable HTTP at /mcp behind the chat-route credentials (a client key sees only its
    own jobs), and gpu-broker mcp over stdio. Long jobs return their id after mcp.wait_s;
    small images come back inline. gpu-broker connect registers it with Claude Code, Claude
    Desktop and Codex (claude-code-mcp, claude-desktop, codex-mcp). /health reports mcp.
    A client key sees only jobs it owns (by key id) and sends files as base64 unless
    mcp.client_url_inputs; tools take only models that can run now.
  • A video's length is the request param num_frames; frames is only the list of input views
    (a number there is refused with a pointer to num_frames).
  • comfy.auth_env: a ComfyUI behind a Bearer token.
  • connect writes symlinked configs through, never replaces a user's own gpu-broker MCP
    server, warns when a running Claude Code drops its entry, and lists the files it changed.
  • The GPU thread's queue is fair by default: interactive jobs first, then the requester that
    has used the least expected GPU time, so one busy client no longer holds everyone else's jobs
    behind its backlog.
    • A call waiting for a pool slot no longer blocks the jobs behind it.
    • scheduler.policy: fifo (env BROKER_SCHEDULER_POLICY) restores strict arrival order.
    • Background jobs age into the interactive class after scheduler.max_wait_s.
    • A switch never jumps a call waiting for a slot on the resident model for longer than
      scheduler.evict_wait_s unless it is of a higher class.
    • /v1/jobs takes x-priority; without it, a job's class follows defaults.background_requesters.
  • gpu-broker replay <events.jsonl>: replays a broker's event log through the queue rules in
    simulated time and reports waits per requester, jobs/hour and residency churn, beside what the
    broker did. Offline; the yardstick for scheduler changes.

Fixed

  • A job's new state and its job.<state> event are written together, so a client that sees a job
    finish in /v1/jobs always finds its job.done / job.failed in /v1/events.
  • MCP initialize reports the installed gpu-broker version in serverInfo (it was empty).

Changed

  • With fair, a /v1/jobs LLM call from a requester not in defaults.background_requesters
    (and without x-priority: background) is interactive, so it may use the slots kept by
    reserved_interactive; before, only chat routes classified their calls.
  • With fair, a requester in defaults.background_requesters can only lower its class with
    x-priority (or interactive in a job), not raise it. scheduler.may_claim_interactive
    lists who may claim interactive (default: everyone not in background_requesters). fifo
    keeps the old rule.
  • While quiesced (/v1/admin/quiesce), new jobs, chats, embeddings and sessions get HTTP 503
    with Retry-After (intervals.quiesced_retry_s, default 5) and x-should-retry: true,
    instead of being queued and then failed as orphans by the restart.
  • A restart re-queues jobs the previous process had queued but not started, in their original
    order and with their staged input files (event job.requeued).
    • Jobs that had started, and direct chats, are still failed as orphaned.
    • A re-queued job fails with "input files missing" if its files are gone, and with
      "model '' no longer in catalog" if its model left the catalog; a re-queued interactive
      session fails with "session expired by restart".

gpu-broker 0.4.0

Choose a tag to compare

@aphilips aphilips released this 03 Oct 18:06
783e01c

Added

  • gpu-broker setup: one command from install to a running broker, asking no questions.
    • Finds the GPU, and llama.cpp, vLLM, Ollama and ComfyUI on their usual ports, with the
      systemd units or Docker containers that run them.
    • Writes the config and catalog for them (the starter files if it finds nothing) and a new
      API token in broker.env (mode 600). Existing files and the token are kept.
    • Installs and starts the gpu-broker service, never as root: a user service (lingering)
      for systemctl --user model servers; else, with root or passwordless sudo, a system
      service run as the installation's owner or a dedicated gpu-broker account, with a
      visudo-checked sudoers rule for exactly the catalog's units. It refuses code anyone else
      could change. Otherwise it prints the serve command, and runs it in the foreground at a
      terminal.
    • Runs check, waits for /health, and opens the dashboard (not over SSH).
    • --dry-run, --yes, --dir. See docs/setup.md.
  • Catalog: health_path for LLM servers without /health (Ollama: /api/version).
  • README: a TLDR block first, and the logo (light and dark).
  • gpu-broker init: writes a starter config.yaml and catalog.yaml to /etc/gpu-broker
    (or --dir) and creates the folders they name.
    • Existing files are kept unless --force.
  • Docs pages: docs/catalog.md, docs/exec-recipes.md, docs/api.md, docs/hardware.md,
    docs/threat-model.md, and AGENTS.md for automated reviewers.

Changed

  • README: leads with the one-person case; installs from PyPI; reference material moved to docs/.
  • Internal: modules that mixed two responsibilities are split, with no behaviour change.
    • catalogschema (entry types and validation) out of catalog.
    • settingsschema (the dataclasses) out of settings.
    • admission (job submit) out of broker.
    • drivers.systemd and drivers.docker out of drivers.local.
    • demo.assemble out of demo.run.
    • The old modules re-export the moved names, except the two concrete drivers: import
      SystemdDriver and DockerDriver from gpu_broker.drivers.systemd and gpu_broker.drivers.docker.

gpu-broker 0.3.2

Choose a tag to compare

@aphilips aphilips released this 03 Oct 14:29
28c9d76

Fixed

  • The sdist is self-testing: unpack it, install it with [dev], run pytest. MANIFEST.in
    ships tests/ (with tests/helpers.py and the fixtures), host/, examples/ and scripts/.
    The amdgpu fixtures' /proc/<pid>/fd symlinks, which an sdist drops because they dangle,
    are listed in tests/fixtures/amdgpu/fd-links.txt and recreated in a temp dir at test time.
    CI checks this on every change.

gpu-broker 0.3.1

Choose a tag to compare

@aphilips aphilips released this 03 Oct 14:11
42dad52

Changed

  • Tests and examples use neutral container ids and paths, and the demo's leak check reads its
    denylist from GPU_BROKER_LEAK_DENYLIST instead of spelling it out. New
    scripts/leak_scan.py scans the tree and the built sdist and wheel against that denylist;
    CI (job leak-scan) and the release workflow run it before anything is published.

gpu-broker 0.3.0

Choose a tag to compare

@aphilips aphilips released this 03 Oct 12:13
8e430c3

Added

  • gpu-broker demo: the real broker, API and dashboard on a simulated GPU, with a simulated
    team sending chats, images and videos. No graphics card, model servers or downloads needed;
    it opens the dashboard already signed in (#token= link). See README, Try it.
  • Drop-in replacement for the OpenAI and Anthropic APIs: POST /v1/messages (Anthropic
    Messages, plain and streaming, tools, images, thinking), POST /v1/embeddings (routed to a
    caps: [embed] model), x-api-key auth beside Authorization: Bearer, and model_map
    (BROKER_MODEL_MAP) mapping hosted names such as gpt-* and claude-* to catalog models.
    See docs/drop-in.md.
  • Input images: jobs may carry image / end_image (base64 or a data: URL), or
    image_url / end_image_url when images.allow_urls is on (off by default; fetches are
    limited to public addresses unless url_allow_networks allows more). Catalog entries say
    which images a model takes (images:), and substitution only picks models that take them.
  • Image-to-video and image-edit templates: Wan 2.2 14B i2v, Wan 2.2 5B with a start image,
    MiniMax first/last frame, HunyuanVideo 1.5 i2v, Qwen-Image edit and FLUX.2 Klein edit.
  • Exec runner (runner: exec): a model can be a program run from a recipe file
    (recipes_dir) on the GPU's machine, with no shell, fixed paths and validated parameters
    (exec.params, exec.choices). examples/recipes/sharp.recipe turns one image into a 3D
    Gaussian splat. Recipe outputs may list several globs. Prune units for old job folders are
    in examples/systemd/.
  • Per-model request defaults in the catalog (defaults:): request > catalog entry > template.
  • The dashboard can run image-taking models with an uploaded image, and its model index says
    who has the GPU: the resident model, the job it is lent to, or nobody.
  • AMD GPUs: the amdgpu driver's sysfs files and /proc/*/fdinfo, no ROCm tools needed
    (gpu.vendor, gpu.index; on Proxmox host/gpu-broker-gpu). See README, Hardware.
  • /v1/metrics samples carry procs_unreadable (per-process memory could not be read: it
    needs root or CAP_SYS_PTRACE), and /v1/gpu names the GPU reader in use (probe).

Changed

  • GPU sample field sm_mhz is now clock_mhz (the vendor-neutral name).
  • Only VRAM used and total are required in a GPU reading; utilisation, power, temperature and
    clock may be null (unknown) in /v1/gpu and /v1/metrics.
  • serve starts when no GPU can be read yet, and keeps retrying, instead of refusing to start.
    gpu.vendor: auto is chosen in the background, within timeouts.gpu_query_s of wall-clock
    time; /v1/gpu answers {"state": "probing"} meanwhile and never waits on nvidia-smi.
  • Proxmox host script: GPU_BUDGET_S (default 8) bounds a whole gpu reading, the vendor
    choice included; NV_TIMEOUT_S now bounds each nvidia-smi call only.

Deprecated

  • sm_mhz in /v1/metrics samples: kept as an alias of clock_mhz for this release only;
    it is removed in the next one.

gpu-broker 0.2.0

Choose a tag to compare

@aphilips aphilips released this 01 Oct 02:17

First public release.

One GPU, many models: an HTTP broker that queues requests and swaps LLM servers and ComfyUI workloads in and out of a single GPU on demand.

  • OpenAI-compatible chat endpoint with streaming; interactive chats skip the background queue
  • Parallel LLM slots, with a reserved slot for interactive users
  • Automatic model switching, substitution and on-demand downloads from a YAML catalog
  • Drivers for systemd, Docker and Proxmox hosts
  • Built-in dashboard with live GPU metrics and a click-to-use model index
  • Bundled ComfyUI graphs for Qwen-Image, Chroma, FLUX.2 Klein, Wan 2.2, HunyuanVideo 1.5, MiniMax H3 and LTX 2.5

See the README for install and configuration.