Skip to content

Releases: redhat-et/comfyui-on-openshift

v0.2.0 — The demo, the bill, the tiers

Choose a tag to compare

@joyful-ii-V-I joyful-ii-V-I released this 03 Sep 17:14

The release where the pool becomes something you can watch, price, and route. v0.1.0 proved the platform on a laptop; v0.2.0 puts the proof in front of people who don't run make test.

See it run

make demo-local. The real pool — gateway, Redis queue, workers — on the machine in front of you, rendering on the laptop's own GPU, zero dollars of cloud. Every click is scripted, so the take reruns identically on any machine with a GPU. Two recorded takes and a 9-second autoplay teaser live in docs/pitch/, beside a six-slide deck that renders directly on GitHub (slide PNGs in the folder README, click-through PDF, live HTML — regenerated by one command, docs/pitch/export.py).

Missing models, one button. ComfyUI-Manager names what's missing in the canvas; the gateway's Install to pool button pulls it server-side — allowlisted source, .safetensors verified at byte nine, fetched once for everyone — because the workers have no internet by design. And Manager itself now installs as the pip module ComfyUI pins behind --enable-manager (a custom_nodes checkout of Manager 4.x cannot import); lint enforces the form, and the nightly boots the ENABLE_MANAGER=true worker variant.

Read the bill

GET /api/showback?format=focus emits FOCUS 1.2 — the FinOps Foundation's billing schema — priced by GPU_HOURLY_RATE, with a billed_jobs denominator per user. scripts/savings-report.py turns a month of showback into the per-team savings number the README's cost model predicts. Finance gets a CSV their tooling already reads, not a JSON blob and an apology.

Pick the card

VRAM-tier routing, laptop half (roadmap N2): a job runs on the card class it declared, and on nothing else. GPU_TIERS on the gateway, WORKER_TIER on the worker, per-tier KEDA scaling and machine pools via #TIER manifest blocks in setup.sh. The default tier deliberately rides the bare comfy:queue — a pool upgraded mid-flight drains its backlog instead of stranding it. Proven end-to-end by check-58 without a cluster.

The paper trail

  • docs/12: the first cluster day as a checklist, costed at ~$50.
  • docs/13: the vLLM-Omni relationship — throughput tier and flexibility tier, the showback-driven promotion rule, and the bridge's boundary in writing.
  • docs/14: the market case, sourced — the category's enterprise flip, ComfyUI's own intake curve, the compliance-captive NDA segment, and the agentic second market (Comfy MCP went official in June; multi-tenant agent access is the roadmap's highest-value open item).
  • docs/09 restructured to the transfer's shape: definition of done, receipts ledger, concerns on the record, and a section of the things a receiving team would otherwise learn slowly.

Known limits

  • The tier pools' cluster half (real 0→1→0 scaling per card class) rides the next cluster day, on the docs/12 checklist.
  • make demo-local time-shares one GPU: pods are simulated, silicon is not; SSO, KEDA and scale-to-zero still need the cluster.
  • The Omni bridge remains a coarse prompt→output call, image-only today — the graph never leaves ComfyUI.
  • CI proves everything that needs no cluster; cluster-applied policy remains schema-validated, not exercised, until the validation day.

v0.1.0 — One GPU pool, a whole team

Choose a tag to compare

@joyful-ii-V-I joyful-ii-V-I released this 02 Sep 08:44
97e388d

First tagged release. Stock ComfyUI on OpenShift for a team: one autoscaled GPU pool behind cluster SSO, with a fair queue, per-user showback and quota, and GPU workers that scale to zero.

What is in it

Two configurations on one cluster. A single-user pod (make up) and the multi-user gateway + Redis queue + KEDA-autoscaled worker pool (make enterprise), sharing one GPU operator and one set of volumes. ROSA (Red Hat OpenShift Service on AWS) is the first-class target; any OpenShift 4.x cluster works with PLATFORM=openshift.

What the platform does, wired up. oauth-proxy for identity, a namespace default-deny NetworkPolicy, a least-privilege Redis ACL user for the workers, no ServiceAccount token in any pod that does not need one, arbitrary-UID images, and GPU pods with no Service and no Route.

What this repo adds. Fair queueing across submitters; a narrow, bounded retry that never replays a workflow ComfyUI already has; a reaper that recovers from worker death without losing or duplicating a job; an ownership fence that is a compare-and-set; per-user output workspaces with path confinement; GPU-second showback and an optional per-user quota that fails open; a scheduled warm floor (WARM_WORKERS) so a team iterating during working hours never waits out a cold start.

Proof that runs on a laptop. make test: 40 shell unit tests, 210 pytest tests, and 378 end-to-end assertions against a real Redis, the real gateway and worker, and a stub ComfyUI, in about a minute. No cluster, no GPU, no AWS account. The nightly builds both GPU images, boots them as an arbitrary UID, and scans them.

Since the audit sweep (#8, #12)

  • Gateway: outputs, job state, cancel, progress stream and showback are scoped to the submitter under oauth; stats and metrics cached instead of scanning the keyspace; sockets bounded; requeue-then-reap no longer fails a live retry.
  • Worker: a timed-out job interrupts ComfyUI; the executing claim is atomic; previews, cross-workspace moves and unvalidated subfolders are refused; explicit Redis timeouts and an agent liveness probe.
  • Images: torch's vendored CUDA libraries resolve on UBI9's lib/lib64 split (the first nightly caught libcudnn.so.9 missing at boot); base packages updated; pip removed from the final image.
  • Docs: README leads with the number, a thirty-second quickstart and an honest comparison table; the monthly cost table reproduces from the repo's own rates.

Known limits

  • Cold start is 8–17 minutes from zero unless WARM_WORKERS or SCALE_TO_ZERO=false keeps a card warm.
  • Custom nodes are baked at image build; runtime installs do not persist and pip is not in the image.
  • The gateway is where a finished workflow runs, not a canvas; authoring stays in ComfyUI.
  • CI proves everything that needs no cluster. NetworkPolicy, the Redis ACL as applied, the KEDA cron floor and the probes are schema-validated and lint-checked, not exercised on a cluster by CI.

Full audit and plan: see the PR descriptions for #8 and #12.