Repository navigation
Overview
eduardoabreu81 edited this page Oct 8, 2026
·
2 revisions
Every option tested so far, rated from 1 to 5 for quality and for speed, as a quick map. The ratings come from watching and listening to the clips and from the measured times on an NVIDIA A40 (48 GB) and an RTX 4090 (24 GB); the linked pages have the numbers, the clips and the settings.
| Rating | |
|---|---|
| 🟦🟦🟦🟦🟦 | 5 - the best tested |
| 🟩🟩🟩🟩 | 4 - good |
| 🟨🟨🟨 | 3 - acceptable |
| 🟧🟧 | 2 - weak |
| 🟥 | 1 - do not use |
Speed is relative within each table. Ratings are from individual tests, not benchmarks.
| Option | Quality | Speed | Notes |
|---|---|---|---|
| W4A8 (Kijai) - recommended | 🟦🟦🟦🟦🟦 | 🟦🟦🟦🟦🟦 | INT8 quality from an 11.7 GiB file; 20% faster than INT8 on 24 GB cards |
| INT8 ConvRot (Comfy-Org) | 🟦🟦🟦🟦🟦 | 🟨🟨🟨 | The reference; does not fit in 24 GB with the rest, so Forge swaps it there |
| INT4BQ (tsolful) | 🟨🟨🟨 | 🟩🟩🟩🟩 | Fine picture and sound, follows the actions less precisely |
| INT4 ConvRot (Merserk) | 🟥 | 🟩🟩🟩🟩 | Broken: the bird of the first test vanished |
| GGUF Q4_K (unsloth) | 🟩🟩🟩🟩 | 🟧🟧 | Close to W4A8, a quieter glass clink in the bar scene; the slowest format |
| FastH3 8-step V2 + its VSA | 🟩🟩🟩🟩 | 🟦🟦🟦🟦🟦 | Distilled, 8 steps; louder, the regular model sits more naturally in its surroundings |
| H3 Eros Max beta5 (community, adult-oriented) | 🟩🟩🟩🟩 | 🟩🟩🟩🟩 | Turbo merged in, 8 steps without a LoRA |
Details: Models, Format Comparisons
| Option | Quality | Speed | Notes |
|---|---|---|---|
| Text encoder INT8 (Comfy-Org) | 🟦🟦🟦🟦🟦 | 🟨🟨🟨 | Speed here means system RAM: about 13 GB more than INT4 |
| Text encoder INT4 (Merserk) - recommended | 🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | Almost the same clips as INT8 |
| Text encoder NVFP4 AWQ | 🟥 | n/a | Refused: prompts come out wrong in Forge |
| Video VAE INT8 (Kijai) | 🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | Matched the fp16 VAE at 47.9 dB PSNR |
| taeh3 live preview (TAESD) | 🟩🟩🟩🟩 | 🟩🟩🟩🟩 | Clear from about step 10 of 20 |
Details: Models
| Option | Quality | Speed | Notes |
|---|---|---|---|
| h3 preset: Res Multistep, Simple, 20 steps, Shift 12 | 🟦🟦🟦🟦🟦 | 🟧🟧 | The reference |
| Turbo LoRA: Res Multistep, Simple, 8 steps, Shift 6 | 🟩🟩🟩🟩 | 🟦🟦🟦🟦🟦 | Music from the first second and bolder camera moves, a softer picture |
| Turbo LoRA: ER SDE, Beta, 9 steps, Shift 6 | 🟩🟩🟩🟩 | 🟩🟩🟩🟩 | 12-17% slower than the 8-step turbo; the best expressions in an anime test |
| Audio shift 6 instead of 3 | 🟩🟩🟩🟩 | not rated | Another take of the same sound: a different rhythm, a little louder |
| Below 544 pixels on the short side | 🟧🟧 | not rated | The picture degrades; with a turbo LoRA it distorts |
Details: Speed Options, Speed Comparisons
| Option | Quality | Speed | Notes |
|---|---|---|---|
| PyTorch SDPA (default) | 🟦🟦🟦🟦🟦 | 🟨🟨🟨 | The reference |
ck attention (--use-ck-attention) |
🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | 3-10% faster with W4A8 and INT8, the same picture |
| Sparse Attention, default range 0.15-0.85 | 🟨🟨🟨 | 🟩🟩🟩🟩 | About 25% faster; turns and spins may morph |
| Sparse Attention from 0.50, Extra Tokens 256 | 🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | About 15% faster and keeps rotations |
| ck attention + Sparse Attention from 0.50 | 🟦🟦🟦🟦🟦 | 🟦🟦🟦🟦🟦 | 17% faster, the same picture |
| Turbo LoRA merged (Forge's default) | 🟦🟦🟦🟦🟦 | 🟦🟦🟦🟦🟦 | No cost per step; needs RAM while patching |
| Turbo LoRA on the fly (Automatic fp16 LoRA) | 🟩🟩🟩🟩 | 🟩🟩🟩🟩 | 10% slower per step (Forge's own path: 21%); no extra RAM |
| Acc 8-Step LoRA (Euler, 8 steps) | 🟦🟦🟦🟦🟦 | 🟦🟦🟦🟦🟦 | About 40% faster than 20 steps with the same look; combines with other LoRAs |
Details: Speed Options, Speed Comparisons
| Option | Quality | Speed | Notes |
|---|---|---|---|
W4A8 + INT4, --cuda-malloc
|
🟦🟦🟦🟦🟦 | 🟦🟦🟦🟦🟦 | The whole checkpoint on the card for short clips: 5.7 s per step at 640×384 |
W4A8 + INT4, --cuda-malloc + Never OOM |
🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | 7.5 s per step; for larger clips |
| GGUF Q4_K + INT4, Never OOM | 🟩🟩🟩🟩 | 🟨🟨🟨 | 9.9 s per step at 640×384; ran 8 seconds at 960×544 |
| INT8 + INT8 text encoder | not rated | n/a | Does not fit in 32 GB of RAM |
Details: 16 GB Cards
| Option | Quality | Speed | Notes |
|---|---|---|---|
| First frame (img2img) | 🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | About 15% more time |
| First and last frame | 🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | About 21% more time |
| Ref2VA, 1 reference picture | 🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | Identity carried over very well; the setting is followed more loosely; +8% |
| Ref2VA, 3 reference pictures | 🟦🟦🟦🟦🟦 | 🟨🟨🟨 | Person, place and object all kept; +29% |
| Ref2VA, 9 reference pictures | 🟩🟩🟩🟩 | 🟧🟧 | Every subject in the clip, the place rebuilt from another angle |
| Ref2VA turbo LoRA, 4 steps | 🟧🟧 | 🟦🟦🟦🟦🟦 | Picture holds, the sound is poor |
| Ref2VA checkpoint without pictures | 🟨🟨🟨 | 🟨🟨🟨 | Works, but nothing ties it to a reference |
| Singularity v1.3 (Ref2VA fine-tune) | 🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | The most natural skin in our close-ups |
Details: First and Last Frame, Reference Pictures
| Option | Quality | Speed | Notes |
|---|---|---|---|
| A voice as the guide audio | 🟦🟦🟦🟦🟦 | 🟩🟩🟩🟩 | The voice is followed, lips in sync |
| Canny control | 🟩🟩🟩🟩 | 🟨🟨🟨 | Keeps the scene and every step; new person and season |
| Inpainting with a still mask | 🟩🟩🟩🟩 | 🟩🟩🟩🟩 | Only the masked part changes, the source's sound kept |
| Pose control with a character sheet | 🟨🟨🟨 | 🟩🟩🟩🟩 | Head turns follow; a hidden leg can be lost; describe the person, not the moves |
| Character Swap LoRA, 5 s | 🟩🟩🟩🟩 | 🟨🟨🟨 | Scene and position kept; expressions and fine hand details are not |
| Character Swap LoRA, 15 s | 🟥 | 🟧🟧 | The swap does not hold that long |
| Motion-only reference (person in negative) | 🟦🟦🟦🟦🟦 | 🟨🟨🟨 | Face and motion kept, on a street too; needs a clean mask of the person |
Details: Reference Videos and Sound, Motion Control, Character Swap
- 48 GB card: INT8 or W4A8 checkpoint, either text encoder, the h3 preset; turn on Sparse Attention from 0.50 for clips of 5 seconds and more.
- 24 GB card: W4A8 checkpoint + INT4 text encoder, about 35 GB of system RAM.
-
16 GB card, 32 GB of RAM: the same files, Forge launched with
--cuda-malloc, Never OOM for larger clips (16 GB Cards). - Drafts: the turbo LoRA at 8 steps, Shift 6, from 544p up; final clips at 20 steps.
- A person or an object you already have a picture of: a Ref2VA checkpoint with reference pictures.
- Faster without losing quality: the Acc 8-Step LoRA for your checkpoint, Euler, 8 steps.
- Follow another video: Motion Control with canny to keep the scene, pose to keep only the motion; Character Swap to replace a person.
MiniMax H3 for Forge Neo
Using it
- Getting Started
- Settings and Controls
- Writing Prompts
- First and Last Frame
- Reference Pictures
- Reference Videos and Sound
- Motion Control
- Character Swap
- 16 GB Cards
- Speed Options
- Examples
- Bloopers
- All Generations
Comparisons
Reference