Skip to content

Releases: paperniuk/splash

Splash 1.1.0 for M1/M2 (1.1.0-m1)

Choose a tag to compare

@paperniuk paperniuk released this 27 Sep 18:44

Prebuilt Splash 1.1.0 for M1/M2 Macs, with custom Metal kernels. No compiling needed.

curl -fsSL https://github.com/paperniuk/splash/releases/download/1.1.0-m1/install-m1.sh | sh
splash-m1 serve --model incoai/Qwen3.6-35B-A3B-Splash

Already on 1.0.2-m1.x? Run the same curl line; it installs next to the old version and switches splash-m1 to the new one.

New in 1.1.0-m1: everything from upstream Splash 1.1.0 (https://github.com/incoai/splash/releases/tag/1.1.0), ported to M1/M2:

  • GGUF models: Unsloth quantizations of Qwen3.8-27B and Qwen3.6-35B-A3B, and Prism ML Ternary Bonsai 2 (PQ2_0), load straight from Hugging Face. Upstream runs them on Apple's MPP matrix library, which hangs the GPU on the M1, so this build has its own GGUF kernels for Apple7/8 GPUs.
  • Vision on M1: the vision encoder also ran on MPP; it now has its own kernels too. A 1024x1024 image encodes in 0.87 s on an M1 Max.
  • Long prompts in chunks: on M1 a single full-size prefill command runs for about 12 s at 8K context and about a minute at 170K, and macOS kills commands that long (ImpactingInteractivity), which ended long agent sessions with an engine failure. This build splits prompt processing into GPU commands of about 5 s each, with no measurable cost (8K prompt: 60.8 s before, 60.9 s after). A 7-hour OpenCode session reached 229,790 tokens of context without an abort; 1.0.2-m1 failed at 176K. --prefill-mode full restores upstream's behaviour.
  • Memory: optional --kv-format bf16, the SSD cache tier --max-cache-disk, and upstream's memory accounting fixes. Model weights stay resident in GPU memory on M1, which fixes the cached 0 problem on 32 GB Macs (issue #3).

Bonsai for 16 GB Macs: prism-ml/Ternary-Bonsai-2-27B-gguf:PQ2_0 is a 27B model in ~7 GB. With --language-only Splash needs about 10.2 GB for it, weights, draft and buffers included. Capped to an 11 GB budget it starts with a 28K context and decodes at 49-58 tok/s on short prompts. A stock 16 GB Mac gives the GPU about 10.7 GB, which leaves Splash just short, so raise the GPU limit first:

sudo sysctl iogpu.wired_limit_mb=12288   # resets at reboot
splash-m1 serve --model prism-ml/Ternary-Bonsai-2-27B-gguf:PQ2_0 --language-only

This is measured with a capped budget on a 64 GB Mac; I have no 16 GB Mac to confirm, so reports are very welcome. 24 GB Macs need no sysctl.

On an M1 Max 64 GB (npanj's five-prompt benchmark, decode tok/s):

Model tok/s Time to first token, 8K prompt
incoai/Qwen3.6-35B-A3B-Splash 145.3 (1.0.2-m1.1: 139.0) not measured
incoai/Qwen3.8-27B-Splash 38.1 (1.0.2-m1.1: 37.7) 61 s
unsloth/Qwen3.8-27B-GGUF:UD-Q4_K_M 31.6 66 s
prism-ml/Ternary-Bonsai-2-27B-gguf:PQ2_0 27.6 63 s

Quality: on 54 physics and arithmetic word problems every model above scores 54/54. The vision encoder matches upstream's fp32 reference fixtures (worst row cosine 0.9999 on 27B, 0.997 on 35B) and describes test images correctly through serve. make test-engine-metal passes (31 tests, shader validation on).

  • Needs: Apple Silicon Mac, macOS 26.4+. The Splash packages want 36 GB of unified memory or more; smaller GGUF quantizations fit 24 GB, and Bonsai fits 16 GB as above.
  • Bundled: engine binary, Metal shader library, and its own Python. No Xcode, Homebrew, or pip install.
  • Side by side: installs splash-m1 and leaves Homebrew's splash alone. Models live in the Hugging Face cache and are shared.
  • Not tested on M1 yet: MLX checkpoints (mlx-community/...-4bit) run through the same kernels as the Splash packages but have not been tried end to end. Bonsai PTQ1_0 is not supported (upstream neither).

This is an unofficial community port. Splash, its models and draft models are built by Inco (https://github.com/incoai/splash); official Splash runs on M3 and newer. Tested on an M1 Max; M1 Pro/Ultra and M2 use the same code path, and M1 Ultra and M2 Max owners have reported the 1.0 port working (incoai#131). Reports welcome.

Built from commit 88411cb (branch apple7-m1-kernels-1.1) on macOS 26.7 with Apple clang 21.0.0. Installer based on npanj's install-q8.sh.

File SHA-256
splash-1.1.0-m1-arm64-macos26.tar.gz 12b34731ea9e909259ea803c1a26ddbde82747beb8537133e41bdcc948877c73
install-m1.sh 7a85a0856f7dad65e8d7cc3e241ce23337af6d4713277827696468c56b1a0894
engine splash 58abcf6d2df91827370e4216696a5ef598378ac1e1c3231eee08f9d034f0ea14
splash.metallib 1b64015a3e75c50e0a48b8c192ad2ef18ccc74f10616b211052de4706de3827b

Splash 1.0.2-m1.1 (prebuilt, M1/M2)

Choose a tag to compare

@paperniuk paperniuk released this 25 Sep 21:02

Prebuilt Splash for M1/M2 Macs, with custom Metal kernels. No compiling needed.

curl -fsSL https://github.com/paperniuk/splash/releases/download/1.0.2-m1.1/install-m1.sh | sh
splash-m1 serve --model incoai/Qwen3.6-35B-A3B-Splash

Already on 1.0.2-m1? Run the same curl line; it installs next to the old version and switches splash-m1 to the new one.

New in 1.0.2-m1.1: the last parts that still used Apple's MPP matrix library on the M1, attention and the MoE expert layers, now run on register-matrix kernels written for Apple7/8 GPUs. On an M1 Max 64 GB, compared with 1.0.2-m1:

  • Qwen3.6-35B-A3B: 98.9 -> 143.8 tok/s on npanj's five-prompt benchmark (1.45x). Time to first token for a 38K-token prompt: 106 s -> 72 s.
  • Qwen3.8-27B: short prompts unchanged (~40 tok/s); at 51K context decode goes 20 -> 25 tok/s, and a 38K-token prompt starts 47 s sooner (359 s -> 312 s).
  • Attention per layer: 1.85-1.95x faster for both the decode (verify) step and prompt processing, from 8K up to 176K context.
  • Server defaults for agents: --max-new-tokens 65536 (was 32768) and --request-timeout 10000 s (was 1800), so long reasoning turns are not cut off.

Quality: on 54 physics and arithmetic word problems both models score 54/54 with Splash's own kernels and with these. Next-token top-1 agreement with Splash's own kernels is 99.84% (27B) and 98.99% (35B, whose expert routing is more sensitive to rounding); every mismatch checked was a near-tie. make test-engine passes, including new tests that compare the register tiles with the MPP tiles and a CPU reference.

  • Needs: Apple Silicon Mac, macOS 26.4+. Official Splash asks for 36 GB of unified memory (48 GB+ recommended); Splash sizes the context to the memory it gets.
  • Bundled: engine binary, Metal shader library, and its own Python. No Xcode, Homebrew, or pip install.
  • Side by side: installs splash-m1 and leaves Homebrew's splash alone. Models live in the Hugging Face cache and are shared, so nothing is downloaded twice.
  • Models: incoai/Qwen3.8-27B-Splash (dense 27B) and incoai/Qwen3.6-35B-A3B-Splash (MoE 35B).

This is an unofficial community port. Splash, its models and draft models are built by Inco (https://github.com/incoai/splash); official Splash runs on M3 and newer. Tested on an M1 Max; M1 Pro/Ultra and M2 use the same code path but are untested. Reports welcome.

Built from commit 5967821 (branch apple7-m1-kernels) on macOS 26.7 with Apple clang 21.0.0. Installer based on npanj's install-q8.sh.

File SHA-256
splash-1.0.2-m1.1-arm64-macos26.tar.gz b76bd0ce1b2e9d108232f21321b7574c6b71090185727f4b0ee1def6cb281752
install-m1.sh 1e0e6bb362b3cde918744f455c093aec5a994ca9878a40cbb0330eb2f73e86f8
engine splash 11ea81c9c30d40525dc36f6bd83a06da07ee5858ed3b824cdf005f8a10042bb3
splash.metallib dd807d1bb9ca1d3c08c4fe866ddd583ce60aa1a5e6aa0b73311942a9f729696d

Splash 1.0.2-m1 (prebuilt, M1/M2)

Choose a tag to compare

@paperniuk paperniuk released this 24 Sep 07:44

Prebuilt Splash for M1/M2 Macs, with custom Metal kernels. No compiling needed.

curl -fsSL https://github.com/paperniuk/splash/releases/download/1.0.2-m1/install-m1.sh | sh
splash-m1 serve --model incoai/Qwen3.8-27B-Splash
  • Needs: Apple Silicon Mac, macOS 26.4+. Official Splash asks for 36 GB of unified memory (48 GB+ recommended); Splash sizes the context to the memory it gets.
  • Bundled: engine binary, Metal shader library, and its own Python. No Xcode, Homebrew, or pip install.
  • Side by side: installs splash-m1 and leaves Homebrew's splash alone. Models live in the Hugging Face cache and are shared, so nothing is downloaded twice.
  • Models: incoai/Qwen3.8-27B-Splash (dense 27B) and incoai/Qwen3.6-35B-A3B-Splash (MoE 35B).

Stock Splash only runs on M3 and newer. This build adds M1/M2 (Apple7/8) support with kernels written for those GPUs. On an M1 Max, compared with Splash's own kernels on the same machine: Qwen3.8-27B decode is about 2.1x faster (18.9 -> 39.1 tok/s average on npanj's five-prompt benchmark) and prefill 2.3-2.7x faster. The 35B MoE runs at 94.7 tok/s on the same benchmark; its expert layers still use the stock kernels. Details and methodology: incoai#131

Tested on an M1 Max; M1 Pro/Ultra and M2 use the same code path but are untested. Reports welcome.

Built from commit d2f902e (branch apple7-m1-kernels) on macOS 26.7 with Apple clang 21.0.0. Installer based on npanj's install-q8.sh.

File SHA-256
splash-1.0.2-m1-arm64-macos26.tar.gz 3c561140ccec7a32593fce252d9ffc567f509fba1e10d8c95d90923bee8f7af5
install-m1.sh 99fe7e57742ecc846b2a3a7cb7ee864d1bee8c21fc3e496af2fc8c470389f15c
engine splash 4765d5cfe9b1ef52e049131dd524c306cb3bf2516accf3fa60653ffe0d2de97c
splash.metallib f88ad0442ce8116970d760a8fe691320320d962c319aca96ea93f73e8e872db6