Releases: channel-ai/comfyui-api
Release list
v1.15.14
fix(cache): hardlink cached inputs so ComfyUI 0.28 accepts them (1.15.13)
fix a hardlink cache issue
v1.15.12
Two-commit hardening of the wedged-job watchdog, driven by production incidents on the LTX deployment (6495a152e2):
Observed failure: a job wedges (ComfyUI threads stuck in uninterruptible NFS page-fault waits — GPU 0%, VRAM held), /ready 503s, the watchdog logs "exiting to trigger redeploy" — and then the process stays alive (one replica: 17 hours). The container never recycles; the fleet accumulates zombie workers. Separately observed: the watchdog firing hours late because cgroup memory-throttle starves the event loop.
3024bce (1.15.11, previously unpushed): SIGKILL the ComfyUI process group before exiting — releasing its memory is what un-throttles the cgroup enough for the proxy to die.
this commit (1.15.12): replace process.exit(1) with self-SIGKILL. exit still needs the throttled JS runtime to be scheduled, and killing ComfyUI doesn't always release memory (pages pinned in unkillable hard-mount NFS waits). Kernel signal delivery needs no JS. The logger is flushed first so the fatal line survives.
Note for ops: an in-process watchdog can still fire late under heavy throttle — an AutoDL-side health check on /ready with auto-restart remains the backstop for fully-frozen replicas (one was observed where even fork hung).
🤖 Generated with Claude Code
fix(comfy): exclude warmup prompt from self-kill watchdog (1.15.10)
Follow-up to #3 (watchdog, merged at 1.15.9). 1.15.10.
Problem
warmupComfyUI() POSTs to the api's own /prompt, so the boot warmup ran through the watchdog added in #3. A cold-boot warmup is a full LTX generation loading ~40GB of models from network storage (especially under --highvram) and can legitimately exceed JOB_TIMEOUT_SECONDS (600s). The watchdog then process.exit(1)'d mid-warmup → container redeployed → reloaded 40GB → stalled again: a fleet-wide restart loop, every replica stuck at /ready 503, never reaching fully ready.
Fix
- The warmup POST is tagged
x-comfyui-api-warmup: true. runPromptAndGetOutputstakes awatchdogMsparam (defaults toconfig.jobTimeoutMs).- The
/prompthandler passeswatchdogMs=0when the warmup header is present.
The watchdog now guards real user jobs only. A hung warmup just leaves the container not-ready (pre-#3 behavior) instead of looping the whole fleet.
Tests
tsc clean; existing 7 poller+watchdog unit tests pass (armJobWatchdog(0) no-op path already covered).
🤖 Generated with Claude Code
v1.15.9: self-kill watchdog for wedged jobs (
ComfyUI can deadlock mid-job — observed on LTX, the tiled VideoVAE decode stalls on the first tile during dynamic-VRAM staging (GPU idle, uninterruptible via /interrupt; not length-dependent — 5s clips hang, 65s clips complete). The job then never completes or errors, so runPromptAndGetOutputs never returns, inFlight stays pinned, and with MAX_QUEUE_DEPTH=1 the replica is stuck at /ready 503 until manually restarted.
v1.15.8
Merge pull request #2 from channel-ai/fix/history-poller-stop fix(comfy): make HistoryEndpointPoller.stop() actually stop
1.15.5
hostname in response and enforce queue size
1.15.4
1.15.3
1.15.2
add a metadata field in api response for all non-file value