Skip to content

v0.34.0 — HiResFix round-2 perf

Choose a tag to compare

@Deaththegrim Deaththegrim released this 07 May 00:11
· 71 commits to main since this release

v0.34.0 — HiResFix round-2 perf

Three more wins on top of v0.33.0, same review pattern: move expensive things out of the inner loop, switch to the safer implementation when the size warrants it.

OPT1 — ControlNet loaded once per HiResFix run

The control-net was being read from disk inside _hires_one_iteration — with iterations=5 + use_controlnet=True that's five redundant loads. Now loaded once in _apply_hires_fix and passed through. The hint image still resamples per iteration since the target size changes; only the model load is hoisted.

OPT3 — Auto-tile VAE decode for big latents

New _smart_vae_decode helper auto-promotes "true" → tiled when the latent's longest dim exceeds 192 (≈ 1536px image). Saves you from picking "true (tiled)" manually for HiResFix outputs that would OOM the non-tiled path. SDXL's standard 128-latent (1024px) stays on the fast path; HiResFix at 2× or higher promotes itself.

Inverse pairs with auto_tile=True on _vae_encode for the pixel-mode encode-back step. The Anima sampler's _vae_decode gets the same heuristic. The final decode in _apply_hires_fix uses smart decode, so HiResFix never returns a half-decoded OOM.

"true (tiled)" and "false" are honoured verbatim — promotion only applies to plain "true".

OPT4 — Pixel upscale model stays GPU-resident across iterations

Previously the upscale model did .to(device) at the top of every iteration and .to(vae_offload_device()) at the bottom — for iterations=5 + mode=both that's 5 device round-trips. Now loaded onto device once, kept resident through all iterations, and offloaded in a try/finally after the loop. Upscale tensors themselves still flow through the standard tiled-scale path.

OPT6 (verified — no change needed)

comfy.utils.common_upscale already handles 5D Qwen-Image latents via the orig_shape-reshape pattern (utils.py line 1032), so the Anima HiResFix's _interpolation_upscale works correctly on flow-matching models with 5D layout. No code change.

Skipped: OPT5

Considered short-circuiting when total_scale ≈ 1.0 but rejected: scale=1.0 + denoise<1.0 is a valid refinement-only HiResFix run, not a no-op.

Backward compat

All changes are internal — no socket / widget / workflow JSON shape changes. Drop-in safe with v0.33.x workflows.

Install / upgrade

cd ComfyUI/custom_nodes/ComfyUI-GrimmRibbity
git pull
# restart ComfyUI