Hi! I've been playing with the new MiniMax H3 model and ran into an issue
with the audio output. Wanted to report it in case others hit the same thing.
What happens
The video comes out perfect, but the audio sounds like the microphone is
blown out — very loud with harsh crackling/popping, kind of like when you
scream into a cheap mic. The content itself is correct (the right sounds
happen at the right time), it just sounds distorted.
How to reproduce
- Official MiniMax H3 image-to-video template workflow, unchanged
- The smallest/pruned model set (nvfp4 qwen3vl, INT8 diffusion model, fp16 video VAE,
fp32 audio VAE)
- Default sampler settings from the template
- Happens more often with loud/intense audio scenes (explosions, impacts,
crowds). Changing the seed sometimes avoids it.
My setup
- ComfyUI 0.30.1 (Desktop)
- Windows
- AMD Radeon RX 7800 XT, 16 GB VRAM (ROCm 7.14)
- 64 GB RAM
- PyTorch 2.12.0+rocm7.14.0
Video generation works great otherwise — really impressed with the model,
just this one audio issue. Thanks!
Hi! I've been playing with the new MiniMax H3 model and ran into an issue
with the audio output. Wanted to report it in case others hit the same thing.
What happens
The video comes out perfect, but the audio sounds like the microphone is
blown out — very loud with harsh crackling/popping, kind of like when you
scream into a cheap mic. The content itself is correct (the right sounds
happen at the right time), it just sounds distorted.
How to reproduce
fp32 audio VAE)
crowds). Changing the seed sometimes avoids it.
My setup
Video generation works great otherwise — really impressed with the model,
just this one audio issue. Thanks!