Skip to content

Splish 1.0

Choose a tag to compare

@publicExcess publicExcess released this 27 Sep 06:05
· 31 commits to m5 since this release

Splish 1.0 is an unofficial fork of Inco's Splash 1.1.0, tuned for Apple M5 GPUs and measured on a 40-core M5 Max. It is not affiliated with Inco.

Against Splash 1.1.0 as shipped: about 1.25× at one request (+11% to +35%) and up to 1.5× at 2–4 requests on Qwen3.8-27B-family models. Qwen3.6-35B-A3B is up to +22% faster, and whole-file code edits a further +24–42% with the copy rule. Quality is unchanged (95/95). Figures are for 4-bit affine models unless marked GGUF.

git clone https://github.com/publicExcess/splish.git && cd splish
make -j4
./splish serve --model mlx-community/Qwen3.8-27B-4bit

What's in it:

  • New verify kernels for the M5 tensor units (SplitSums32, lighter barriers).
  • Kernel choices measured for the 40-core M5 Max, loaded from a file (tuning/, SPLASH_KERNEL_CHOICES). ./splish picks them automatically.
  • A copy rule for coding agents: verbatim continuations are drafted exactly.
  • Faster GGUF decode (G4a, bit-identical).
  • Benchmark and correctness tools (dev/m5/), and dev/m5/report.py for performance reports.

Full results, method, what did not work and the ideas left to try are in the README. The next release (v1.1) adds an auto-tuner for other M5 chips.