Skip to content
eduardoabreu81 edited this page Oct 8, 2026 · 5 revisions

Models

H3 needs four parts: the checkpoint (the diffusion model), the text encoder (Qwen3-VL 32B), the video VAE and the audio VAE. This page lists the files that were checked, where they go and what is not supported yet.

The extension recognizes files by their tensor layout and quantization, not by their name. A file called H3 on Civitai or Hugging Face may still be another architecture or format; if the extension cannot use a file, it says so instead of generating garbage.

Folders

Part Forge folder Selected in
Checkpoint models/Stable-diffusion the checkpoint selector
Text encoder models/text_encoder VAE / Text Encoder
Video VAE, audio VAE models/VAE VAE / Text Encoder
LoRAs models/Lora the prompt, <lora:name:weight>
Fun ControlNet models/ControlNet Control in the MiniMax H3 panel
Preview decoder (taeh3) models/VAE-taesd downloaded automatically

Select exactly one text encoder, one video VAE and one audio VAE. The Components section of the MiniMax H3 panel shows what was recognized.

Recommended set

The tested starting point, from Comfy-Org/MiniMax-H3 (revision e5eb578):

Part File Size
Checkpoint minimax_h3_fl2va_pruned_int8_convrot 19.5 GiB
Text encoder qwen3vl_32b_minimax_h3_int8_convrot 25.3 GiB
Video VAE minimax_h3_video_vae_fp16 4.9 GiB
Audio VAE minimax_h3_audio_vae_fp32 0.6 GiB

The text encoder includes the vision weights that first and last frame and reference pictures use.

Smaller set for 24 GB and 16 GB cards

Swap the checkpoint and the text encoder; the VAEs stay the same. Tested on an RTX 4090 and an A40 on 5 October 2026, and on a 16 GB RTX 2000 Ada with 32 GB of RAM on 6 October (16 GB Cards).

Part File Size
Checkpoint minimax_h3_fl2va_pruned_w4a8_mixed (Kijai) 11.7 GiB
Text encoder qwen3vl_32b_minimax_h3_int4_convrot (Merserk) 13.9 GiB
  • The W4A8 checkpoint stores 4-bit weights and computes with 8-bit activations. Its clips are as good as the INT8 checkpoint's; with the same seed they differ more than any other swap, because the checkpoint decides the picture.
  • The INT4 text encoder gives almost the same clips as the INT8 one: same framing, same action. One cartoon with a chain of four actions (a cat ends up wearing a flower pot as a hat) went wrong with the INT8 checkpoint and the INT4 text encoder - a second pot appeared - and right with the W4A8 checkpoint and the same text encoder.
  • On the RTX 4090 this pair used about 35 GB of system RAM and ran about 20% faster than the INT8 checkpoint, which does not fit in 24 GB with the rest. See Performance and Memory.
  • The same text encoder is also on Civitai (model 2830162, minimaxH3INT4Convrot_qwen3vl32bInt4): it is the same file.
  • The same scenes generated in each format, with sound: Format Comparisons.

Checkpoints

File Size Status
minimax_h3_fl2va_pruned_int8_convrot (Comfy-Org) 19.5 GiB Recommended. Used for every standard example.
minimax_h3_fl2va_pruned_w6a8 (Comfy-Org) 14.9 GiB Works. Bird test 64.4 s against 53.5 s.
minimax_h3_fl2va_pruned_fp8_scaled (Comfy-Org) 19.5 GiB Works. Slower on the A40 (99.6 s), which has no FP8 tensor cores; may do better on newer cards.
minimax_h3_fl2va_int8_convrot, full, not pruned (Comfy-Org) 31.7 GiB Recognized, not generated with: it needs more system RAM than the test machine had.
fastvideo_fasth3_8step_v2_pruned_int8_convrot (FastVideo) 22.1 GB Works at 8 steps, Shift 10. See FastH3.
minimax_h3_fl2va_pruned_w4a8_mixed (Kijai) 11.7 GiB Works. Recommended for 24 GB cards: same quality as INT8, same speed on the A40, faster on the RTX 4090.
minimax_h3_fl2va_pruned_INT4BQ (tsolful, INT4 and INT8 layers) 14.8 GiB Works, acceptable. Pictures and sound are fine, but it follows the actions less precisely: a skater rode toward the camera instead of being tracked, Holi powder was not thrown up. W4A8 is smaller and better.
minimax_h3_fl2va_pruned_int4_convrot (Merserk; also Civitai model 2830162) 10.6 GiB Loads, not recommended. About 25% faster per step, but the quality breaks down: the bird of the first test vanished from the clip. Its INT4 text encoder is good.
minimax_h3_fl2va_pruned-Q4_K.gguf (unsloth) 10.6 GiB Works. Close to W4A8 to watch and listen to, but the slowest format (Forge dequantizes it at every step), with the RAM of the W4A8 set. Other GGUF builds from Q2_K to Q8_0 load the same way.
GGUF IQ1/IQ2/IQ3/IQ4 builds Refused: Forge Neo cannot dequantize these types.
H3 Eros Max beta5, INT8 and W4A8 (Civitai) 13 GiB (W4A8) Works. See community checkpoints.
minimax_h3_ref2va_pruned_int8_convrot (Comfy-Org) 19.5 GiB Works: the reference pictures mode, up to 9 pictures. Its tensors are identical to FL2VA, so the extension tells them apart by ref2va in the file name.
Minimax-h3_Singularity_ref2va_Pruned_v1.3_int8 (WarmBloodAban) Works, a Ref2VA fine-tune; the most natural skin in our close-ups. See Examples.
minimax_h3_ref2va_pruned_w4a8_mixed (Kijai) 11.0 GiB Works, the same quality and speed as the INT8 Ref2VA on the A40.

"Pruned" checkpoints store the per-block modulation layers in a compact shared form; they are smaller and need less RAM than the full ones. The extension reads both forms.

Text encoders

File Size Status
qwen3vl_32b_minimax_h3_int8_convrot (Comfy-Org) 25.3 GiB Recommended. Used for every example.
qwen3vl_32b_minimax_h3_int4_convrot (Merserk) 13.9 GiB Works. Almost the same clips as the INT8 text encoder, about 11 GiB smaller. Recommended for 24 GB cards.
qwen3vl_32b_minimax_h3_nvfp4_awq (Comfy-Org) Refused: Forge Neo loads it without a warning, but its prompts come out wrong (a bird prompt gave a dog).
GGUF text encoders Refused: llama.cpp keeps Qwen3-VL's vision part in a separate file, which first and last frame need.

The text encoder only runs while the prompt is encoded, so its precision does not change the speed of a clip; it changes the system RAM. Forge keeps it in RAM with the checkpoint. Loading it only while the prompt is encoded is on the roadmap.

VAEs

File Size Status
minimax_h3_video_vae_fp16 (Comfy-Org) 4.9 GiB Recommended.
minimax_h3_audio_vae_fp32 (Comfy-Org) 0.6 GiB Recommended.
Original MiniMax FL2VA video VAE and audio VAE 9.7 / 0.6 GiB Work, with output identical to the Comfy-Org VAEs. Both are called model.safetensors at the source: give them different names.
minimax_h3_video_vae_int8_convrot (Kijai) Works. Only the decoder's large linear layers are INT8. On a test clip it matched the fp16 VAE at 47.9 dB PSNR, visually the same.

LoRAs

File Status
minimax_h3_fl2v_turbo_8step_v1.0_comfyui_bf16 (Comfy-Org, 1.8 GiB) Works. 208 keys, none skipped.
minimax_h3_ref2v_turbo_4step_v0.1_comfyui_bf16 (Comfy-Org, 1.8 GiB), for Ref2VA Works, but the sound is poor at its 4 steps.
larryvrh turbo LoRAs (4-step and v4) Not supported yet: keys without the diffusion_model. prefix, and per-block adaln_proj layers that pruned checkpoints do not have.
MiniMax-H3-FL2VA-Acc-8Step_pruned_comfy and MiniMax-H3-Ref2VA-Acc-8Step_pruned_comfy (alibaba-pai, converted by Kijai, 1.7 GB each) Works, 8 steps with Euler. See Acc 8-Step LoRAs.
minimax_h3_ref2v_turbo_8step_v1.0_768p_comfyui_bf16 (lightx2v), for Ref2VA Works at 8 steps, Res Multistep.
h3_character_swap_pro4500_1000 (akatz-ai) Works, with the Acc or turbo LoRA. See Character Swap.
h3-realism-people-t2v-i2v-r2v (fal), trigger r34l1sm Works, with the turbo or Acc LoRA.
INT8 repacks of LoRAs Not supported yet.

See Speed Options for the settings and the RAM note.

Fun ControlNet

File Size Status
minimax_h3_fun_controlnet_union_2.0_pruned_int8_convrot (alibaba-pai, Kijai's INT8) 4.5 GB Works, used for every control example. See Motion Control.
minimax_h3_fun_controlnet_union_2.0_pruned_bf16 8.4 GB The same model in bf16.
The first union model (minimax_h3_fun_controlnet_union_pruned_*) Older; use 2.0.

Community checkpoints

Community fine-tunes work when they keep H3's architecture and use a quantization Forge Neo understands. The extension reads the quantization from the file itself, including the _quantization_metadata that some converters write, and ignores extra keys it does not need.

H3 Eros Max beta5

A community fine-tune of H3 on Civitai, with the turbo distillation merged in. It is an adult-oriented model; here it was tested only with ordinary scenes.

Files tested The INT8 file (Star Converter format) and the 13 GiB file labelled "fp8", which is actually a W4A8 format (4-bit weights with a codebook). Both load as they are.
With The standard INT8 text encoder and the Comfy-Org VAEs
Settings Res Multistep, Simple, 8 steps, CFG 1, Shift 12, no LoRA
Time About 140 s for 73 frames at 1152×768; 777 s for 362 frames at 576×1024
Examples Five clips

Adult checkpoints on Civitai need an account to download.

Checking another file

  1. Put it in the right folder and refresh the model list.
  2. Select it. If the MiniMax H3 panel does not appear for a checkpoint, the extension did not recognize it as H3.
  3. Generate a short test (the bird from Getting Started). A clear error in the result area means the format is not supported; garbage output means the format loads but is read wrongly, which is worth an issue with the file's link.

License

MiniMax H3 has its own license. Commercial use of the outputs needs a commercial license from MiniMax; see the official ComfyUI guide and the model repository. Community fine-tunes may add their own terms.

Clone this wiki locally