Repository navigation
Format Comparisons
The same prompt and seed generated with different checkpoint and text encoder files, to see what each format changes. Each MP4 plays the formats one after another, each with its own sound, with the format's name in the corner; the GIF previews put two of them side by side. For a quick summary see the Overview; times and memory per format are in Performance and Memory.
In short: W4A8 matches INT8, the INT4 text encoder is close to lossless, GGUF Q4_K is close to W4A8 but the slowest, and INT4BQ follows the actions less precisely. Changing the checkpoint changes a clip much more than changing the text encoder.
Made with the smaller set for 24 GB cards - Kijai's W4A8 checkpoint and the INT4 text encoder - on one RTX 4090 with version 0.5.0, h3 preset (Res Multistep, Simple, 20 steps, CFG 1, Shift 12), no Never OOM. The prompts follow MiniMax's prompt format. The same scenes with other files are compared in Performance and Memory.
| Model | Size | Frames | Sampler | Steps | Shift | Seed | Time |
|---|---|---|---|---|---|---|---|
| W4A8 + INT4 text encoder | 640×384 | 124 (5.17 s) | Res Multistep | 20 | 12 | 101 | 84.3 s |
Prompt
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a low tracking shot at a puppy's height frames a golden retriever puppy on a sunny backyard lawn as soap bubbles drift past. The camera tracks the puppy at fast speed as it leaps, snaps at the bubbles and spins around, while two children in the background blow more bubbles and laugh.\n\noverall_soundscape: The children laugh and the puppy barks excitedly over a light breeze and birdsong. Its paws thump and rustle across the grass.\n\nnon_diegetic_music: N/A
| Model | Size | Frames | Sampler | Steps | Shift | Seed | Time |
|---|---|---|---|---|---|---|---|
| W4A8 + INT4 text encoder | 576×768 | 243 (10.13 s) | Res Multistep | 20 | 12 | 202 | 363.5 s |
Prompt
integrated_multimodal_description: [Shot 1] Live-action, cinematic, a medium shot frames two Brazilian women in their thirties at a small table in a busy boteco in Rio de Janeiro at night, cold beer glasses between them and a football match playing on a TV behind them. The woman with curly hair in a red and black Flamengo shirt, with a warm, lively voice (S1), leans forward laughing and says: <d>[Portuguese] Não acredito que você ainda torce pro Fluminense!</d> The camera pushes in with small amplitude at slow speed as the woman with straight black hair in a maroon, green and white Fluminense shirt, with a low, teasing voice (S2), raises her glass with a smirk and answers: <d>[Portuguese] Pelo menos o meu time sabe jogar bola.</d> Both burst out laughing and clink their glasses.\n\noverall_soundscape: Crowd chatter and clinking glasses fill the bar while the football commentary from the TV plays underneath. The two women laugh loudly and their glasses clink together at the end.\n\nnon_diegetic_music: N/A
| Model | Size | Frames | Sampler | Steps | Shift | Seed | Time |
|---|---|---|---|---|---|---|---|
| W4A8 + INT4 text encoder | 960×544 | 243 (10.13 s) | Res Multistep | 20 | 12 | 303 | 427.1 s |
Prompt
integrated_multimodal_description: [Shot 1] Live-action, cinematic, an aerial drone shot glides slowly over a small coastal fishing village during a festival at dusk, colorful boats rocking in the harbor and strings of lights over the narrow streets. The camera trucks right with large amplitude at slow speed, revealing a crowd dancing in the main square while a brass band plays festive music on a small stage. Fireworks then burst over the sea in red, gold and green, reflected in the dark water, as the camera tilts up with small amplitude.\n\noverall_soundscape: The crowd cheers and chatters in the square, waves lap against the moored boats, and fireworks boom and crackle over the sea.\n\nnon_diegetic_music: N/A
The three scenes above were also generated with the other tested files, same seed, on the same RTX 4090. Each MP4 plays every format one after another, with its own sound and its name in the corner. Times, memory and speed per format are in Performance and Memory.
- Changing the checkpoint (W4A8, INT8, GGUF) changes the clip the most; changing only the text encoder (INT8 + INT4 against INT8 + INT8) gives nearly the same picture.
- GGUF Q4_K against W4A8: both good to watch and listen to; in the bar scene the clink of the glasses at the end is louder with W4A8.
- The bar scene shows why every person needs their own description. The first prompt gave the Flamengo shirt to the first woman and only a hairstyle to the second: both came out in Flamengo shirts, with the same curly hair. The corrected prompt, below, gives the second woman "a maroon, green and white Fluminense shirt", and every format then drew the two shirts.
The first bar prompt (both women in Flamengo shirts)
The same as the corrected prompt above, with the second woman described only as "the woman with straight black hair, with a low, teasing voice (S2)".
Six prompts, each generated with the same seed in six checkpoint and text encoder combinations, on one NVIDIA A40 with version 0.5.0, h3 preset (Res Multistep, Simple, 20 steps, CFG 1, Shift 12), 124 frames (5.17 s). The preview puts the INT8 checkpoint (left or top) and the W4A8 one (right or bottom) together; the MP4 plays all six formats one after another, each with its own sound, with its name in the corner:
- INT8 checkpoint + INT8 text encoder (the reference)
- W4A8 (Kijai) + INT8
- INT4BQ (tsolful) + INT8
- INT8 + INT4 text encoder (Merserk)
- INT4BQ + INT4
- W4A8 + INT4
The verdict per format is in Models and Performance and Memory. In short: W4A8 matches INT8, the INT4 text encoder is close to lossless, and INT4BQ follows the actions in the prompt less precisely.
| Size | Frames | Sampler | Steps | Shift | Seed |
|---|---|---|---|---|---|
| 576×768 | 124 (5.17 s) | Res Multistep | 20 | 12 | 11 |
Speech, steam, a busy kitchen. The INT8 and W4A8 clips came out nearly in sync with each other, and INT4BQ kept good lip sync while changing the shot slightly.
Prompt
A middle-aged Italian chef with flour on his apron stands in a busy restaurant kitchen, looks into the camera and says in Italian with a warm, slightly hoarse voice: "Il segreto è la pazienza, e un po' di burro." He laughs and tosses a handful of fresh basil into a sizzling pan. Steam, clattering pans and kitchen chatter in the background. Handheld documentary look, no music, no subtitles.
| Size | Frames | Sampler | Steps | Shift | Seed |
|---|---|---|---|---|---|
| 576×768 | 124 (5.17 s) | Res Multistep | 20 | 12 | 22 |
Skin, wrinkles, freckles and hair in golden light. Acceptable in every format.
Prompt
Extreme close-up portrait of an elderly woman with deep wrinkles and freckles, silver hair moving in a light breeze, golden hour sunlight on her skin. She slowly turns her head toward the camera, blinks and gives a small, knowing smile. Shallow depth of field, fine skin texture, soft wind and distant birds. No music, no subtitles.
| Size | Frames | Sampler | Steps | Shift | Seed |
|---|---|---|---|---|---|
| 640×384 | 124 (5.17 s) | Res Multistep | 20 | 12 | 33 |
Fast motion and a tracking camera. INT4BQ with the INT8 text encoder made him ride toward the camera instead of being tracked; INT4BQ with the INT4 text encoder added a body turn the prompt did not ask for, and still landed it.
Prompt
A skateboarder in a red hoodie races down a steep San Francisco street, carves between parked cars and lands a kickflip over a manhole cover, the camera tracking alongside at speed. Bright midday sun, hard shadows, wheels roaring on asphalt, a passer-by shouts "whoa!". No music, no subtitles.
| Size | Frames | Sampler | Steps | Shift | Seed |
|---|---|---|---|---|---|
| 640×384 | 124 (5.17 s) | Res Multistep | 20 | 12 | 44 |
Very saturated colors. INT4BQ with the INT8 text encoder did not throw the powder up. The W4A8 soundtracks came out about 6 LU quieter than INT8.
Prompt
Holi festival in India: a crowd throws clouds of vivid magenta, yellow, cyan and green powder into the air, everyone laughing and dancing, colors drifting slowly across the frame in the sunlight. Very saturated colors, slow-motion feeling, joyful shouting and drums. No subtitles.
| Size | Frames | Sampler | Steps | Shift | Seed |
|---|---|---|---|---|---|
| 640×384 | 124 (5.17 s) | Res Multistep | 20 | 12 | 55 |
Low light, neon reflections and rain. Acceptable in every format.
Prompt
Night in a narrow Tokyo alley in heavy rain: neon signs in pink and blue reflect on the wet pavement, a woman with a transparent umbrella walks away from the camera and stops under a flickering sign. Low light, film grain, raindrops hitting the umbrella and puddles, a distant train. No music, no subtitles.
| Size | Frames | Sampler | Steps | Shift | Seed |
|---|---|---|---|---|---|
| 640×384 | 124 (5.17 s) | Res Multistep | 20 | 12 | 66 |
A chain of four actions, the hardest prompt of the set. INT4BQ with the INT8 text encoder ended with the cat inside the pot; INT8 with the INT4 text encoder dropped it in the pot and put a second pot on its head; INT4BQ with the INT4 text encoder showed the pot above the cat and lost it; W4A8 with the INT4 text encoder wore the pot as a hat.
Prompt
A 2D hand-drawn cartoon in a bright, flat style: a round orange cat chases a paper airplane across a rooftop, trips, rolls into a flower pot and pops out wearing it as a hat, looking proud. Bold black outlines, simple shapes, playful pizzicato music and cartoon sound effects. No text.
MiniMax H3 for Forge Neo
Using it
- Getting Started
- Settings and Controls
- Writing Prompts
- First and Last Frame
- Reference Pictures
- Reference Videos and Sound
- Motion Control
- Character Swap
- 16 GB Cards
- Speed Options
- Examples
- Bloopers
- All Generations
Comparisons
Reference











