fix: Resolve first-run CUDA OOM and optimize VRAM usage during multi-GPU cloning
Critical memory optimization for Parallel Anything node to prevent CUDA out-of-memory
errors on first execution and improve stability across multi-GPU setups.
Problem:
- First-run OOM errors occurred due to CUDA memory fragmentation and holding
duplicate model copies (CPU + GPU) simultaneously during cloning - No handling for devices that failed during cloning due to insufficient VRAM
- Original model restoration caused temporary 2x VRAM usage spikes
Solution:
- Skip cloning when target device matches source device (use reference)
- Incremental state_dict loading to minimize peak memory usage
- Aggressive ComfyUI model cache unloading before cloning operations
- Progressive cleanup between device clones with synchronized CUDA operations
- Graceful degradation: skip OOM devices and redistribute workload
Technical Changes:
- Added aggressive_cleanup() helper for forced GC and CUDA cache clearing
- Modified safe_model_clone() with incremental parameter loading and
early-exit if model already on target device - Updated setup_parallel() to use reference counting for original device
- Implemented successful_devices tracking for OOM fallback handling
- Added periodic torch.cuda.empty_cache() during large model transfers
- Enhanced error handling with device-level try/except blocks
- Fixed state dict parsing for nested modules in incremental loading
Memory optimizations:
- Unload ComfyUI models (unload_all_models) before cloning
- Delete state_dict entries immediately after transfer
- Force CPU transition only when necessary (not when reusing original)
- Synchronize CUDA devices before/after memory operations
Breaking Changes: None
Backwards Compatible: Yes