Repository navigation
[Feedback & Benchmark] Generating 300s with MiniMax Music 3 on M4 Max (OOM Workaround, Progress Bar Issue & Language Bias) #894
Keylin2935
started this conversation in
Show and tell
Replies: 0 comments
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Hi everyone,
I’ve been heavily testing the
mlx-community/MiniMax-Music3-mxfp8model on Apple Silicon, pushing it to its architectural limits (up to 300 seconds). I wanted to share my benchmark times, a workaround to prevent OOM crashes, and some severe limitations I encountered regarding prompt adherence and API feedback.💻 Hardware & Setup
mlx-community/MiniMax-Music3-mxfp8⏱️ Benchmark Results & Generation Time
Here are the exact metrics for a single 270-second generation:
During my first attempts to generate full tracks, macOS would trigger an OOM kill after ~11 minutes due to KV cache saturation.
The fix: I had to free up over 50 GB of space on my internal SSD (to allow macOS to heavily use Swap) and implement a background monitoring thread in my Python script to actively force garbage collection every 5 seconds:
With this running, the MLX active GPU memory stayed stable instead of infinitely leaking, allowing the 28-minute generation to successfully complete on the M4 Max.
🐛 2. The tqdm Progress Bar Issue
If you wrap the generator in tqdm (for res in tqdm(model.generate(...)):), the progress bar does not update step-by-step.
For my 270s generation, the bar sat at 0% (0/30) for 28 minutes, and only jumped to 3% (1/30) at the very last second before completing. It seems the underlying implementation holds the generation entirely on the Metal GPU and doesn't yield properly to Python until massive chunks (or the whole generation) are done.
📉 3. Qualitative Issues at Long Durations (Language Bias)
While the M4 Max survived the workload, the model's coherence completely broke down at 270+ seconds:
Cultural / Language Overriding: My prompt explicitly asked for "Melancholic chamber pop, dream pop, cinematic, Taylor Swift style". However, because my lyrics were in French, the model completely ignored the English musical prompt and defaulted to generating traditional French folk/accordion music. The cultural bias of the language heavily overrides the genre prompt.
Ignored Formatting: It attempts to sing stage directions placed in brackets/parentheses instead of using them as structural cues.
Conclusion:
Generating 270s-300s in one shot works natively on a 64GB M4 Max if you aggressively clear the cache in a background thread and have SSD Swap available. Expect an RTF around 6.4x. However, due to the complete lack of real-time progress feedback and the model's tendency to hallucinate genres based on the language of the lyrics, I highly recommend users stick to 45s-60s chunks to iterate faster.
Hope this data helps anyone trying to push the duration limits on Mac!
All reactions