Skip to content

Stable Audio Open Small — CoreML

Choose a tag to compare

@john-rocky john-rocky released this 04 Apr 05:09
· 152 commits to master since this release
b45b00a

CoreML conversion of stabilityai/stable-audio-open-small — text-to-music generation (497M params).

Generates up to 11.9 seconds of stereo 44.1kHz audio from text prompts.

Models

File Size Description
StableAudioT5Encoder 94 MB T5-base text encoder (INT8)
StableAudioNumberEmbedder 367 KB Seconds conditioning (FP16)
StableAudioDiT 292 MB Diffusion transformer (INT8, use cpuAndGPU)
StableAudioDiT_FP32 1.2 GB Diffusion transformer (FP32 compute, use cpuOnly, best quality)
StableAudioVAEDecoder 138 MB Oobleck stereo decoder (FP16)

Usage

See StableAudioDemo sample app and convert_stable_audio.py.

Conversion notes

  • DiT FP16 weights cause NaN in attention on iOS GPU → use INT8 or FP32 compute
  • T5 INT8 may produce occasional NaN → sanitize before DiT input
  • DiT FP32 requires cpuOnly (GPU background permission restriction on iOS)