Skip to content

Splash 1.1.0

Latest

Choose a tag to compare

@jianc99 jianc99 released this 26 Sep 12:51
· 45 commits to main since this release
3e1f9ec

Splash 1.1.0 adds upstream GGUF and MLX models, Unsloth 1–8-bit quantizations, and Prism ML Bonsai support, with faster inference and more flexible memory management on Apple silicon.

Models

  • Load supported models directly from Hugging Face: Qwen3.8-27B and Qwen3.6-35B-A3B, from Unsloth GGUF or MLX affine 4-bit checkpoints. Splash automatically selects the matching DFlash2 draft.
  • Run Unsloth quantizations from UD-IQ1_S through Q8_0, including mixed-precision UD variants. UD-Q8_K_XL and BF16 target weights are not supported.
  • Run Prism ML Ternary Bonsai 2 (PQ2_0) with input rotation and vision support. Smaller quantizations make these models usable on 24 GB Macs.
  • Use the target model's own tokenizer and chat template. Vision loads from the same model source; --language-only skips it.
  • Prepare target and draft weights once, with bounded temporary memory, and reuse the disk-backed cache on later starts. Existing Splash model packages remain supported.

Performance and accuracy

On the same Unsloth UD-Q4_K_M files, Splash measured 2.5–3.2× llama.cpp's decode throughput on 35B and 4.5–5.3× on 27B across M5 Pro and M3 Max. Against llama.cpp with MTP on 27B, the measured advantage was 2.7–4.6×.

Over 16,384 positions of prose, code and chat, Splash and llama.cpp ranked the same token first at 99.30–99.45% of positions on 27B and 97.83–98.14% on 35B with BF16 KV. These are next-token agreement measurements, not task-accuracy scores. See the comparison and test conditions.

This release also improves mixed greedy/sampling batches, avoids repeated request and cache-admission work, and retains active model weights in memory to reduce page faults.

Memory and reliability

  • Add optional BF16 KV cache with --kv-format bf16; INT8 remains the default.
  • Add an optional SSD cache tier with --max-cache-disk, supporting both KV formats without further quantization. Rolling checkpoints can replace older disk copies when the tier fills. This cache is temporary and does not persist across server restarts.
  • Keep final logits in FP32 and improve Metal synchronization, memory accounting, model loading, and generation error handling.
  • Fix tool argument handling, structured output edge cases, image processing, and browser chat storage.

Clients and APIs

  • Add Pi coding agent support and fix OpenCode 2 provider registration.
  • Expose effective context limits in model discovery and per-request timing information for compatible clients.
  • Improve Responses API errors, LAN hostname configuration, and browser chat over HTTP.

Install or upgrade

Requires Apple M3 or newer and macOS 26.4+. The prebuilt package includes Python and Metal kernels.

brew update
brew install incoai/tap/splash  # or: brew upgrade incoai/tap/splash
splash serve --model unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M

Stop a running Splash server before upgrading, then restart it. Existing model downloads and agent sessions are preserved. Upstream models need disk space for the original weights, their prepared copy, and the draft; later starts reuse the prepared copy.

Thanks to everyone who contributed code, testing, and feedback. Community contributions are welcome!

Full changelog