Skip to content

Overview

eduardoabreu81 edited this page Oct 8, 2026 · 2 revisions

Overview

Every option tested so far, rated from 1 to 5 for quality and for speed, as a quick map. The ratings come from watching and listening to the clips and from the measured times on an NVIDIA A40 (48 GB) and an RTX 4090 (24 GB); the linked pages have the numbers, the clips and the settings.

Rating
🟦🟦🟦🟦🟦 5 - the best tested
🟩🟩🟩🟩 4 - good
🟨🟨🟨 3 - acceptable
🟧🟧 2 - weak
🟥 1 - do not use

Speed is relative within each table. Ratings are from individual tests, not benchmarks.

Checkpoints

Option Quality Speed Notes
W4A8 (Kijai) - recommended 🟦🟦🟦🟦🟦 🟦🟦🟦🟦🟦 INT8 quality from an 11.7 GiB file; 20% faster than INT8 on 24 GB cards
INT8 ConvRot (Comfy-Org) 🟦🟦🟦🟦🟦 🟨🟨🟨 The reference; does not fit in 24 GB with the rest, so Forge swaps it there
INT4BQ (tsolful) 🟨🟨🟨 🟩🟩🟩🟩 Fine picture and sound, follows the actions less precisely
INT4 ConvRot (Merserk) 🟥 🟩🟩🟩🟩 Broken: the bird of the first test vanished
GGUF Q4_K (unsloth) 🟩🟩🟩🟩 🟧🟧 Close to W4A8, a quieter glass clink in the bar scene; the slowest format
FastH3 8-step V2 + its VSA 🟩🟩🟩🟩 🟦🟦🟦🟦🟦 Distilled, 8 steps; louder, the regular model sits more naturally in its surroundings
H3 Eros Max beta5 (community, adult-oriented) 🟩🟩🟩🟩 🟩🟩🟩🟩 Turbo merged in, 8 steps without a LoRA

Details: Models, Format Comparisons

Text encoders and VAEs

Option Quality Speed Notes
Text encoder INT8 (Comfy-Org) 🟦🟦🟦🟦🟦 🟨🟨🟨 Speed here means system RAM: about 13 GB more than INT4
Text encoder INT4 (Merserk) - recommended 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 Almost the same clips as INT8
Text encoder NVFP4 AWQ 🟥 n/a Refused: prompts come out wrong in Forge
Video VAE INT8 (Kijai) 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 Matched the fp16 VAE at 47.9 dB PSNR
taeh3 live preview (TAESD) 🟩🟩🟩🟩 🟩🟩🟩🟩 Clear from about step 10 of 20

Details: Models

Sampling

Option Quality Speed Notes
h3 preset: Res Multistep, Simple, 20 steps, Shift 12 🟦🟦🟦🟦🟦 🟧🟧 The reference
Turbo LoRA: Res Multistep, Simple, 8 steps, Shift 6 🟩🟩🟩🟩 🟦🟦🟦🟦🟦 Music from the first second and bolder camera moves, a softer picture
Turbo LoRA: ER SDE, Beta, 9 steps, Shift 6 🟩🟩🟩🟩 🟩🟩🟩🟩 12-17% slower than the 8-step turbo; the best expressions in an anime test
Audio shift 6 instead of 3 🟩🟩🟩🟩 not rated Another take of the same sound: a different rhythm, a little louder
Below 544 pixels on the short side 🟧🟧 not rated The picture degrades; with a turbo LoRA it distorts

Details: Speed Options, Speed Comparisons

Attention and LoRAs

Option Quality Speed Notes
PyTorch SDPA (default) 🟦🟦🟦🟦🟦 🟨🟨🟨 The reference
ck attention (--use-ck-attention) 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 3-10% faster with W4A8 and INT8, the same picture
Sparse Attention, default range 0.15-0.85 🟨🟨🟨 🟩🟩🟩🟩 About 25% faster; turns and spins may morph
Sparse Attention from 0.50, Extra Tokens 256 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 About 15% faster and keeps rotations
ck attention + Sparse Attention from 0.50 🟦🟦🟦🟦🟦 🟦🟦🟦🟦🟦 17% faster, the same picture
Turbo LoRA merged (Forge's default) 🟦🟦🟦🟦🟦 🟦🟦🟦🟦🟦 No cost per step; needs RAM while patching
Turbo LoRA on the fly (Automatic fp16 LoRA) 🟩🟩🟩🟩 🟩🟩🟩🟩 10% slower per step (Forge's own path: 21%); no extra RAM
Acc 8-Step LoRA (Euler, 8 steps) 🟦🟦🟦🟦🟦 🟦🟦🟦🟦🟦 About 40% faster than 20 steps with the same look; combines with other LoRAs

Details: Speed Options, Speed Comparisons

16 GB cards (32 GB of RAM)

Option Quality Speed Notes
W4A8 + INT4, --cuda-malloc 🟦🟦🟦🟦🟦 🟦🟦🟦🟦🟦 The whole checkpoint on the card for short clips: 5.7 s per step at 640×384
W4A8 + INT4, --cuda-malloc + Never OOM 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 7.5 s per step; for larger clips
GGUF Q4_K + INT4, Never OOM 🟩🟩🟩🟩 🟨🟨🟨 9.9 s per step at 640×384; ran 8 seconds at 960×544
INT8 + INT8 text encoder not rated n/a Does not fit in 32 GB of RAM

Details: 16 GB Cards

Pictures

Option Quality Speed Notes
First frame (img2img) 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 About 15% more time
First and last frame 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 About 21% more time
Ref2VA, 1 reference picture 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 Identity carried over very well; the setting is followed more loosely; +8%
Ref2VA, 3 reference pictures 🟦🟦🟦🟦🟦 🟨🟨🟨 Person, place and object all kept; +29%
Ref2VA, 9 reference pictures 🟩🟩🟩🟩 🟧🟧 Every subject in the clip, the place rebuilt from another angle
Ref2VA turbo LoRA, 4 steps 🟧🟧 🟦🟦🟦🟦🟦 Picture holds, the sound is poor
Ref2VA checkpoint without pictures 🟨🟨🟨 🟨🟨🟨 Works, but nothing ties it to a reference
Singularity v1.3 (Ref2VA fine-tune) 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 The most natural skin in our close-ups

Details: First and Last Frame, Reference Pictures

Videos, sound and control

Option Quality Speed Notes
A voice as the guide audio 🟦🟦🟦🟦🟦 🟩🟩🟩🟩 The voice is followed, lips in sync
Canny control 🟩🟩🟩🟩 🟨🟨🟨 Keeps the scene and every step; new person and season
Inpainting with a still mask 🟩🟩🟩🟩 🟩🟩🟩🟩 Only the masked part changes, the source's sound kept
Pose control with a character sheet 🟨🟨🟨 🟩🟩🟩🟩 Head turns follow; a hidden leg can be lost; describe the person, not the moves
Character Swap LoRA, 5 s 🟩🟩🟩🟩 🟨🟨🟨 Scene and position kept; expressions and fine hand details are not
Character Swap LoRA, 15 s 🟥 🟧🟧 The swap does not hold that long
Motion-only reference (person in negative) 🟦🟦🟦🟦🟦 🟨🟨🟨 Face and motion kept, on a street too; needs a clean mask of the person

Details: Reference Videos and Sound, Motion Control, Character Swap

The short version

  • 48 GB card: INT8 or W4A8 checkpoint, either text encoder, the h3 preset; turn on Sparse Attention from 0.50 for clips of 5 seconds and more.
  • 24 GB card: W4A8 checkpoint + INT4 text encoder, about 35 GB of system RAM.
  • 16 GB card, 32 GB of RAM: the same files, Forge launched with --cuda-malloc, Never OOM for larger clips (16 GB Cards).
  • Drafts: the turbo LoRA at 8 steps, Shift 6, from 544p up; final clips at 20 steps.
  • A person or an object you already have a picture of: a Ref2VA checkpoint with reference pictures.
  • Faster without losing quality: the Acc 8-Step LoRA for your checkpoint, Euler, 8 steps.
  • Follow another video: Motion Control with canny to keep the scene, pose to keep only the motion; Character Swap to replace a person.

Clone this wiki locally