Repository navigation
gpu-broker 0.2.0
First public release.
One GPU, many models: an HTTP broker that queues requests and swaps LLM servers and ComfyUI workloads in and out of a single GPU on demand.
- OpenAI-compatible chat endpoint with streaming; interactive chats skip the background queue
- Parallel LLM slots, with a reserved slot for interactive users
- Automatic model switching, substitution and on-demand downloads from a YAML catalog
- Drivers for systemd, Docker and Proxmox hosts
- Built-in dashboard with live GPU metrics and a click-to-use model index
- Bundled ComfyUI graphs for Qwen-Image, Chroma, FLUX.2 Klein, Wan 2.2, HunyuanVideo 1.5, MiniMax H3 and LTX 2.5
See the README for install and configuration.