MiniMax-H3 generates video and audio in the same diffusion process β not a render followed by a dubbing pass, but two streams denoised across one shared set of sampling steps. Synchronization comes out of generation itself rather than from aligning tracks afterwards.
This plugin brings the full H3 runtime inside the ComfyUI process and covers all three task paths: text to video+audio (T2VA), keyframe driven generation (FL2VA β supply only a first frame and it is image-to-video), and ordered multimodal references (Ref2VA). With INT8 weights and layerwise offload it runs on a single 24GB GPU.
RunningHub is a MiniMax partner; this plugin is developed and maintained by RunningHub.
- Three tasks in one plugin β text to video+audio (T2VA), keyframe driven generation (FL2VA: first frame, last frame, or both), and ordered image/audio/video references (Ref2VA).
- Image-to-video is FL2VA with a first frame only β the supplied image becomes the literal frame 0 of the output.
- Joint AV sampling β a dual-sigma rectified-flow sampler drives video and audio with independent shift schedules, keeping audio locked to motion.
- 2.5Γ less denoising work with
res_multistepβ a second-order exponential integrator reaches the quality of 50-step Euler in about 21 sigma points (20 DiT calls instead of 49). Same weights, no requantisation;eulerstays the default. End-to-end gain depends on the workload β text encoding, DiT load and VAE decode are not affected. - INT8 or BF16 β single-file INT8 checkpoints or the upstream sharded BF16 release. The loaders never switch between them silently.
- 24GB-class single GPU β automatic layerwise DiT offload, adaLN precompute with weight release (~40% of DiT weights dropped afterwards), shape-aware activation reserve, and residency leasing between runs.
- Typed component contracts β every loader output carries release and component fingerprints, so mixing FL2VA and Ref2VA parts, or swapping a checkpoint mid-graph, fails closed instead of rendering something wrong.
- Optional approximate acceleration β velocity-cache or Cache-DiT, both off by default.
cd ComfyUI/custom_nodes
git clone https://github.com/HM-RunningHub/ComfyUI_RH_MinMaxH3.git
pip install -r ComfyUI_RH_MinMaxH3/requirements.txtRestart ComfyUI afterwards β node definitions are read once at start-up.
Requirements
- ComfyUI 0.27+ (0.28+ recommended)
- A CUDA build of PyTorch matching your ComfyUI, plus Triton and
comfy-kitchen ffmpegandffprobeonPATHfor Ref2VA video and audio referencestransformers>=4.57.0,<=5.8.1(Qwen3-VL support)
MiniMax-H3 is a large model. INT8 reduces storage and transfer cost but does not make it small β expect substantial host RAM and fast storage on top of VRAM.
Weights live in two places, and both are required:
| Location | Holds | Why it is needed |
|---|---|---|
ComfyUI/models/MiniMax-H3/ |
Flat single-file converted weights | The tensors |
ComfyUI/models/diffusers/MiniMax-H3/ |
Upstream sharded release | config.json, source/config.json, tokenizer, preprocessor_config.json |
The flat root carries weights only, with no sidecar files. Component type
and partition are decided entirely by the filename; the architecture is read
from the sharded release that the model_root widget points at.
ComfyUI/
βββ models/
βββ MiniMax-H3/ # converted single-file weights
β βββ MiniMax-H3-FL2VA-int8_convrot.safetensors # DiT, FL2VA partition (T2VA + FL2VA)
β βββ MiniMax-H3-Ref2VA-int8_convrot.safetensors # DiT, Ref2VA partition
β βββ qwen3-vl-32b-int8_convrot.safetensors # Qwen3-VL text/multimodal encoder
β βββ MiniMax-H3-video_vae.safetensors # 24-channel video VAE
β βββ MiniMax-H3-audio_vae.safetensors # 32-channel audio VAE
β
βββ diffusers/
βββ MiniMax-H3/ # upstream sharded release
βββ FL2VA/
β βββ transformer/ # BF16 DiT + config.json
β βββ text_encoder/ # Qwen3-VL + tokenizer/processor
β βββ video_vae/
β βββ audio_vae/
βββ Ref2VA/
βββ ... # same component layout
# Converted single-file weights
hf download Gluttony10/MiniMax-H3-INT8-CONVROT --local-dir ComfyUI/models/MiniMax-H3
# Sharded release β running INT8 only needs its configs, tokenizer and processor (~66 MB, see note below)
hf download MiniMaxAI/MiniMax-H3 --local-dir ComfyUI/models/diffusers/MiniMax-H3 --include "FL2VA/*" "Ref2VA/*" --exclude "*.safetensors"pip install modelscope
# Converted single-file weights
modelscope download --model Gluttony10/MiniMax-H3-INT8-CONVROT --local_dir ComfyUI/models/MiniMax-H3
# Sharded release β running INT8 only needs its configs, tokenizer and processor (~66 MB, see note below)
modelscope download --model MiniMax/MiniMax-H3 --local_dir ComfyUI/models/diffusers/MiniMax-H3 --include "FL2VA/*" "Ref2VA/*" --exclude "*.safetensors"| Model | Link | Description |
|---|---|---|
| INT8-CONVROT weights | HuggingFace Β· ModelScope | Converted single-file DiT / text-encoder / VAE weights β models/MiniMax-H3/ |
| MiniMax-H3 release | HuggingFace Β· ModelScope | Sharded components and their configs β models/diffusers/MiniMax-H3/ |
The upstream release supplies
config.json,source/config.json, the tokenizer andpreprocessor_config.json. The converted weights alone are not enough to load the model.
When you only run the INT8 single-file weights, the only parts of the release
you actually need are its non-weight files β per-component configs, tokenizer
and processor, about 66 MB in total; that is what --exclude "*.safetensors"
above is for. Note that the BF16 sharded entries (dropdown items without a
.safetensors suffix) will still be listed because their configs exist, but
selecting one fails since the weight shards are missing. To use BF16, re-run
the download without --exclude (the FL2VA + Ref2VA subtrees total ~268 GB).
| Selection | Loader dropdown value | Notes |
|---|---|---|
| DiT INT8 | MiniMax-H3-FL2VA-int8_convrot.safetensors / MiniMax-H3-Ref2VA-int8_convrot.safetensors |
Smallest footprint; the partition is proven by the filename |
| DiT BF16 (sharded) | MiniMax-H3-FL2VA / MiniMax-H3-Ref2VA |
Highest fidelity; layerwise offload keeps it viable on 24GB |
| Text encoder INT8 | qwen3-vl-32b-int8_convrot.safetensors |
Roughly 26GB versus 62GB for BF16 |
| Text encoder BF16 | qwen3-vl-32b |
Sharded component directory |
| Video / Audio VAE | MiniMax-H3-video_vae.safetensors / MiniMax-H3-audio_vae.safetensors |
Always FP32 weights, selected independently |
Each loader dropdown lists only its own component type, filtered by task partition β an FL2VA node never offers Ref2VA weights.
Every task wires three loaders (DiT, Qwen3-VL, dual VAE) into a target, a conditioning/encode step, an empty AV latent, the dual-sigma sampler, and the AV decode.
Ready-to-load graphs live in examples/workflows/:
| Task | Workflow | Condition material |
|---|---|---|
| T2VA | t2va.json |
none |
| FL2VA | fl2va_first_frame.json |
first frame β this is image-to-video |
| FL2VA | fl2va_last_frame.json |
last frame |
| FL2VA | fl2va_first_last_frame.json |
first + last frame |
| Ref2VA | ref2va_image.json |
one reference image |
| Ref2VA | ref2va_image_audio.json |
image β audio, ordered chain |
| Ref2VA | ref2va_video_audio.json |
video carrying an audio track |
Shared defaults: explicit 832Γ480, five seconds, 50 sigma points, shifts
12/3, accel=off, denoise_video=true. Replace the placeholder media
filenames with files already uploaded to ComfyUI input/.
Keyframes and references are different mechanisms. An FL2VA keyframe
occupies a real frame position in the output (index 0 or the final frame) and
must match the target canvas. A Ref2VA reference has no frame position, may use
its own resolution, and steers identity rather than becoming a frame.
Ref2VA reference order is significant. Chain each reference node's
references output into the next one; reordering the chain changes the
multimodal prompt and the condition rows.
All nodes register under the RunningHub/MiniMax H3/* category.
| Node | Purpose |
|---|---|
RHMiniMaxH3DirectModelLoader |
T2VA DiT |
RHMiniMaxH3DirectTextEncoderLoader |
T2VA Qwen3-VL |
RHMiniMaxH3DirectVAELoader |
T2VA dual VAE (video_vae_path + audio_vae_path) |
RHMiniMaxH3FL2VAModelLoader / β¦TextEncoderLoader / β¦VAELoader |
FL2VA partition |
RHMiniMaxH3Ref2VAModelLoader / β¦TextEncoderLoader / β¦VAELoader |
Ref2VA partition |
| Node | Purpose |
|---|---|
RHMiniMaxH3T2VATarget / RHMiniMaxH3T2VATextEncode |
Text-only target and prompt encode |
RHMiniMaxH3FL2VAFirstFrameCondition |
First frame, optional last frame |
RHMiniMaxH3FL2VALastFrameCondition |
Last frame only |
RHMiniMaxH3FL2VATarget / RHMiniMaxH3FL2VAEncode |
Keyframe target and encode |
RHMiniMaxH3Ref2VAImageReference / β¦AudioReference / β¦VideoReference |
Ordered reference chain |
RHMiniMaxH3Ref2VATarget / RHMiniMaxH3Ref2VAEncode |
Reference target and encode |
| Node | Purpose |
|---|---|
RHMiniMaxH3EmptyAVLatent |
Allocate the joint AV latent from a target |
RHMiniMaxH3SeparateAVLatent / RHMiniMaxH3CombineAVLatent |
Split or rejoin the video and audio streams |
RHMiniMaxH3EncodeVideoAVLatent |
Encode existing frames into an AV latent (video-to-audio) |
RHMiniMaxH3FrameRate |
Experimental frame-rate conditioning |
RHMiniMaxH3DualSigmaSampler |
Joint video + audio sampling. sampler_mode selects euler (default, 50 sigma points) or res_multistep (second order, ~21 points). Full parameter guide: docs/sampling.md |
RHMiniMaxH3DecodeAV |
Decode to IMAGE frames and AUDIO |
Feature flags live in minimax_h3_nodes/runtime/h3_settings.py and can each be
turned off independently for rollback.
- Memory β
ENABLE_DIT_LAYERWISE_OFFLOAD(auto),DIT_LAYERWISE_PREFETCH,OPT_DYNAMIC_ACTIVATION_RESERVE,OPT_RESIDENCY_LEASE+RESIDENCY_POLICY. - Hot path β
OPT_ADALN_PRECOMPUTE/OPT_ADALN_RELEASE_WEIGHTS,OPT_SDPA_PRECOMPUTED_BOUNDS,OPT_PREPARED_STRUCTURE,OPT_INPLACE_EULER_UPDATE,OPT_FUSED_QK_ROPE. - Acceleration (approximate, off by default) β set
accelon the sampler tominimax-h3-velocity-cache-v1orminimax-h3-cache-v1. These are not ground-truth paths; the sidecar records which one ran. - Observability β
OPT_TELEMETRYrecords stage timings, per-step P50/P95 and peak VRAM;OPT_WRITE_SIDECARwrites a JSON sidecar beside each render.
- Two all-in-one generation nodes.
Video Gen (Text / Keyframes)covers t2va, fl2va and v2a β the task is inferred from which inputs you wire β andVideo Gen (References)covers ref2va with the reference limits from the official manual (β€9 images, β€3 videos, β€3 audios, β€12 total, audio never alone). They orchestrate the same single-purpose nodes internally, so every fingerprint check and cache key behaves as on the granular path. Each takes optionalh3_model,h3_text_encoderandh3_vae_bundleinputs, so you can override any one component while the rest load frommodel_root. - Loaders merged from nine nodes to six. The FL2VA and Ref2VA loader pairs differed only by a task constant, and the text encoder and the dual VAE ship byte-identical weights under both partitions β so those two now expose no partition widget at all, and only the DiT keeps one. The superseded six stay registered and keep resolving in saved workflows; they are just hidden from the node search and the slot-drag menus.
- adaLN is now quantized.
adaln_projwas excluded fromint8_convroton the assumption that a runtime precompute would release those weights instead. That path never engages under INT8 β the modulation cache refuses to take over Comfy cast-weights linears, which INT8 weights always are β so 24.35 GiB of BF16 adaLN stayed resident for the whole run, neither quantized nor freed. Quantizing it brings the Ref2VA DiT from 43.77 GiB to 31.65 GiB. Verified per layer: adaLN's own reconstruction error is 0.85%, below the 0.907% mean across all 252 quantized linears. - Configurable slow-step guard. The sampler aborts after two consecutive
denoise steps exceed a threshold, to catch a wedged run. The threshold was
fixed at 75s and the size exemption keyed off
width >= 1344, which let a 1280Γ736 15-second job β 2.7Γ the work of the calibrated baseline β fail on every single run. The exemption now compareswidth Γ height Γ framesagainst that baseline, andstep_abort_secondson the sampler config makes the threshold adjustable (0 disables it) for long jobs or cards that have to offload weights.
-
res_multistepsampler mode. The video and audio streams run on different shift schedules, so each is now integrated with a second-order exponential integrator on its own schedule. About 21 sigma points (20 DiT calls) match the quality of 50-step Euler, with no visible difference on the same seed. Measured on a single GPU at 832Γ480/125f, both runs warm:denoise loop end to end euler, 50 points (49 DiT calls)248.7s 549s res_multistep, 21 points (20 DiT calls)101.2s 406s 2.46Γ 1.35Γ The denoise loop is about 45% of wall-clock time at this size; the rest is text encoding, DiT load and VAE decode, which this change does not touch. Larger canvases spend proportionally more time denoising, so the end-to-end gain there is higher. Select it with
sampler_modeon the sampler and setsigma_pointsto 21;eulerremains the default so existing workflows are untouched. The mode forcesaccel=offβ the velocity-cache and Cache-DiT profiles are calibrated for 50 steps and would over-skip at 20. -
Video VAE allocation elisions. Q/K norm skips a redundant fp32 round trip when the norm has no affine parameters (CUDA already accumulates in fp32, so the half-precision result is bit-identical); gated FFN, scaled residuals and
norm_silubecame in-place; causal temporal padding is now a singleF.padinstead ofzeros_like+cat. Output is bit-identical β verified end-to-end at PSNR = inf on the same seed. -
Minimum output duration lowered from 5s to 4s. The widget default stays at 5.0, so existing workflows are unaffected.
- Per-type loader COMBOs; the dual VAE loader takes
video_vae_pathandaudio_vae_pathseparately. - Flat single-file weights root at
ComfyUI/models/MiniMax-H3/, classified by filename with no sidecar required. - Example workflows for all task types under
examples/workflows/.
Apache License 2.0 β see LICENSE.
This plugin contains adaptations of the MiniMax-H3 runtime; provenance and the list of adapted components are recorded in NOTICE.md. No model weights are bundled. Checkpoint files remain subject to the licence and confidentiality terms that apply to them β review those terms before redistributing the weights or using them commercially.
- RunningHub China
- RunningHub International
- MiniMax on GitHub
- ComfyUI
- MiniMax-H3-INT8-CONVROT on HuggingFace
- MiniMax-H3-INT8-CONVROT on ModelScope
- MiniMax-H3 release on HuggingFace
- MiniMax-H3 release on ModelScope
- Apache License 2.0
Built on the MiniMax-H3 joint audio-video diffusion model by MiniMax (GitHub). The native runtime here is adapted from the H3 source package released by MiniMax under Apache License 2.0; see NOTICE.md for the baseline snapshot and the list of changes.
Built to run inside ComfyUI, whose upstream code and conventions parts of this plugin draw on. ComfyUI is licensed GPL-3.0 and is a runtime dependency here; its source is not bundled.
Packaged for ComfyUI by RunningHub.
