Skip to content

gpu-broker 0.3.0

Choose a tag to compare

@aphilips aphilips released this 03 Oct 12:13
8e430c3

Added

  • gpu-broker demo: the real broker, API and dashboard on a simulated GPU, with a simulated
    team sending chats, images and videos. No graphics card, model servers or downloads needed;
    it opens the dashboard already signed in (#token= link). See README, Try it.
  • Drop-in replacement for the OpenAI and Anthropic APIs: POST /v1/messages (Anthropic
    Messages, plain and streaming, tools, images, thinking), POST /v1/embeddings (routed to a
    caps: [embed] model), x-api-key auth beside Authorization: Bearer, and model_map
    (BROKER_MODEL_MAP) mapping hosted names such as gpt-* and claude-* to catalog models.
    See docs/drop-in.md.
  • Input images: jobs may carry image / end_image (base64 or a data: URL), or
    image_url / end_image_url when images.allow_urls is on (off by default; fetches are
    limited to public addresses unless url_allow_networks allows more). Catalog entries say
    which images a model takes (images:), and substitution only picks models that take them.
  • Image-to-video and image-edit templates: Wan 2.2 14B i2v, Wan 2.2 5B with a start image,
    MiniMax first/last frame, HunyuanVideo 1.5 i2v, Qwen-Image edit and FLUX.2 Klein edit.
  • Exec runner (runner: exec): a model can be a program run from a recipe file
    (recipes_dir) on the GPU's machine, with no shell, fixed paths and validated parameters
    (exec.params, exec.choices). examples/recipes/sharp.recipe turns one image into a 3D
    Gaussian splat. Recipe outputs may list several globs. Prune units for old job folders are
    in examples/systemd/.
  • Per-model request defaults in the catalog (defaults:): request > catalog entry > template.
  • The dashboard can run image-taking models with an uploaded image, and its model index says
    who has the GPU: the resident model, the job it is lent to, or nobody.
  • AMD GPUs: the amdgpu driver's sysfs files and /proc/*/fdinfo, no ROCm tools needed
    (gpu.vendor, gpu.index; on Proxmox host/gpu-broker-gpu). See README, Hardware.
  • /v1/metrics samples carry procs_unreadable (per-process memory could not be read: it
    needs root or CAP_SYS_PTRACE), and /v1/gpu names the GPU reader in use (probe).

Changed

  • GPU sample field sm_mhz is now clock_mhz (the vendor-neutral name).
  • Only VRAM used and total are required in a GPU reading; utilisation, power, temperature and
    clock may be null (unknown) in /v1/gpu and /v1/metrics.
  • serve starts when no GPU can be read yet, and keeps retrying, instead of refusing to start.
    gpu.vendor: auto is chosen in the background, within timeouts.gpu_query_s of wall-clock
    time; /v1/gpu answers {"state": "probing"} meanwhile and never waits on nvidia-smi.
  • Proxmox host script: GPU_BUDGET_S (default 8) bounds a whole gpu reading, the vendor
    choice included; NV_TIMEOUT_S now bounds each nvidia-smi call only.

Deprecated

  • sm_mhz in /v1/metrics samples: kept as an alias of clock_mhz for this release only;
    it is removed in the next one.