Replies: 2 comments 3 replies
|
Partial modules are transferred to and offloaded from VRAM during computation as usual. Before that PR, some devices were forced to run text encoders on the CPU by default. It just made them work like the main diffusion/flow model, which can run on the GPU even when the entire model cannot fit into VRAM. |
|
The text encoder does not have to fit in your 12 GB. With dynamic VRAM the weights stay in system RAM and are streamed to the GPU as each layer is needed, the same way the diffusion model is handled. So only the compute moves to the GPU, which is much faster than running the encoder on the CPU. This matches what chaObserv said. The change is in ComfyUI/comfy/model_management.py Lines 1219 to 1230 in b87fe48 if args.gpu_only:
return get_torch_device()
if comfy.memory_management.aimdo_enabled:
return get_torch_device()
elif vram_state in (VRAMState.HIGH_VRAM, VRAMState.NORMAL_VRAM):
if should_use_fp16(prioritize_performance=False):
return get_torch_device()
else:
return torch.device("cpu")
else:
return torch.device("cpu")Before commit d39cdfd (#16374), dynamic VRAM was OR'd with the HIGH/NORMAL VRAM check, so it still went through the The encoder is also not loaded onto the GPU up front. In ComfyUI/comfy/model_management.py Lines 1232 to 1234 in b87fe48 --gpu-only is set. The weights start in RAM and are pulled to the GPU on demand.
Without dynamic VRAM the old behaviour is unchanged, so the CPU is still used in low VRAM modes. (I looked this up in the code with help from an AI assistant; the links are to the lines I checked.) |
Uh oh!
There was an error while loading. Please reload this page.
For people with systems that have inadequate VRAM the solution has been to have the text encoder stay in main system RAM and use the CPU, right?
On a system with 64 GB of system RAM and a 12 GB GPU, how does it make sense to put the text encoder on the GPU?
All reactions