Skip to content

fix: Resolve first-run CUDA OOM and optimize VRAM usage during multi-GPU cloning

Choose a tag to compare

@FearL0rd FearL0rd released this 06 Feb 17:30
· 16 commits to main since this release

Critical memory optimization for Parallel Anything node to prevent CUDA out-of-memory
errors on first execution and improve stability across multi-GPU setups.

Problem:

  • First-run OOM errors occurred due to CUDA memory fragmentation and holding
    duplicate model copies (CPU + GPU) simultaneously during cloning
  • No handling for devices that failed during cloning due to insufficient VRAM
  • Original model restoration caused temporary 2x VRAM usage spikes

Solution:

  • Skip cloning when target device matches source device (use reference)
  • Incremental state_dict loading to minimize peak memory usage
  • Aggressive ComfyUI model cache unloading before cloning operations
  • Progressive cleanup between device clones with synchronized CUDA operations
  • Graceful degradation: skip OOM devices and redistribute workload

Technical Changes:

  • Added aggressive_cleanup() helper for forced GC and CUDA cache clearing
  • Modified safe_model_clone() with incremental parameter loading and
    early-exit if model already on target device
  • Updated setup_parallel() to use reference counting for original device
  • Implemented successful_devices tracking for OOM fallback handling
  • Added periodic torch.cuda.empty_cache() during large model transfers
  • Enhanced error handling with device-level try/except blocks
  • Fixed state dict parsing for nested modules in incremental loading

Memory optimizations:

  • Unload ComfyUI models (unload_all_models) before cloning
  • Delete state_dict entries immediately after transfer
  • Force CPU transition only when necessary (not when reusing original)
  • Synchronize CUDA devices before/after memory operations

Breaking Changes: None
Backwards Compatible: Yes