Skip to content

v0.3.7 — beta

Pre-release
Pre-release

Choose a tag to compare

@OuincheWinch OuincheWinch released this 08 Oct 14:42

Tested on the owner's own machine, from the installed build, before publication.

Phases 1-5: Full optimization pipeline shipped

TeaCache step-skipping (Phase 1)

4-21% speedup on FLUX.2, Krea2, Z-Image. Auto-disabled for FLUX.2 img2img where KV cache applies.

Disk-backed prompt cache (Phase 2)

SQLite + .npy serialization — survives restarts, ~9% cold-start improvement.

Krea2 TE q4@load (Phase 3)

Eliminates 7.5 GB bf16 spike, saves ~52 s cold start. Removed 150+ lines of disk-cache/WeightApplier patch logic.

mx.compile + 2-step warmup (Phase 4)

11-20% speedup after 2 warmup steps on FLUX.2, Krea2, Z-Image.

Quantization sweep (Phase 5)

q4 confirmed optimal; q3/q2 catastrophic quality loss on FLUX.2/Krea2 (PSNR ~10dB). Z-Image ignores quant param.

Krea2: 8 steps now default

4-step + distillation LoRA = 8 steps compute equivalent. 8-step is now default for quality.

Arena Benchmark: 40/40 scenes @ 768×512

Model Steps TeaCache Avg Time
FLUX.2-klein 4B 4 0.15 43s
FLUX.2-klein 9B 4 0.15 95s
Krea2 Turbo 4 0.20 256s
Z-Image Turbo 8 0.12 152s

Assets verified byte-for-byte against local.

Checksums

ALL_SHA256SUMS published beside assets. Verify with shasum -a 256 DiffusionBear-0.3.7-arm64.zip.