Skip to content

Releases: publicExcess/splish

Splish 1.1

Choose a tag to compare

@publicExcess publicExcess released this 29 Sep 20:56

Everything here is measured on the 40-core M5 Max unless it says otherwise.

New

  • Swift-1.5 / Qwen3.8-27B: the v12 kernel choices. One request, greedy: v1.0 (v8) → v12 is +4.1%
    (+3.1..+5.1, 8 prompts × 2, one session). The DFlash draft's projections at 2–4 requests had no
    choices; serving gains are C=2 +2.7%, C=3 +3.7%, C=4 +6.0%. Seeded samples differ from v1.0 at the
    same seed (the draft proposes differently); greedy text is unchanged.
  • Qwen3.6-35B-A3B: its own tuned table (v2). One request: +18% steady, greedy +10.0% (+8.1..+11.9);
    2–4 requests +2.6 / +1.8 / 0%. Quality 95/95.
  • GGUF: the draft is tuned too (q80-v2, kquant-v2). step_bench v1 → v2: Q8_0 −2.8 / −2.8 / −4.9 /
    −6.6% step time at 1–4 requests; Q4_K_M −4.5 / −3.9% at 3–4.
  • 20-core M5 Pro support (from @mikebuckets171, splish#1): a Qwen3.8-27B table measured on the M5 Pro
    (step −6.8% at 1 request, −6.8 to −8.4% at 2), picked automatically; the Qwen3.6-35B-A3B case uses the
    first M5 Max table, the one measured there.

Fixed

  • Tokenizer: transformers' Qwen2 tokenizer dropped the \p{M} rule from the package's pre-tokenizer,
    so text with combining marks (Hindi, Thai, Arabic, …) was split into too many tokens. The server now
    restores the package's own pre-tokenizer (Hindi sample 33 → 21 tokens).
  • Images: an image more elongated than 200:1 (e.g. an accidental 8192 × 17 screen selection) is now centred on
    white instead of rejected; a rejected image stayed in the client's history and failed every later turn.
  • step_bench: intervals use Student-t (1.96 × SE was ~1.6× too narrow at 4 rounds). Thanks
    @mikebuckets171.

Docs

  • A shorter README: what Splish is for, headline results, quality (engine equivalence), quick start and settings. Detailed measurements by version moved to RESULTS.md.
  • The G4a, concurrent-encoder and repeat statements now say where the differences are statistically significant.
  • kernel_bench covers six more shapes, and sizes its buffers as the engine does (staged GGUF plans at 24 rows).

Not in this release

  • The auto-tuner for other M5 chips, announced for v1.1 in the v1.0 notes, moves to a later release.

Splish 1.0

Choose a tag to compare

@publicExcess publicExcess released this 27 Sep 06:05

Splish 1.0 is an unofficial fork of Inco's Splash 1.1.0, tuned for Apple M5 GPUs and measured on a 40-core M5 Max. It is not affiliated with Inco.

Against Splash 1.1.0 as shipped: about 1.25× at one request (+11% to +35%) and up to 1.5× at 2–4 requests on Qwen3.8-27B-family models. Qwen3.6-35B-A3B is up to +22% faster, and whole-file code edits a further +24–42% with the copy rule. Quality is unchanged (95/95). Figures are for 4-bit affine models unless marked GGUF.

git clone https://github.com/publicExcess/splish.git && cd splish
make -j4
./splish serve --model mlx-community/Qwen3.8-27B-4bit

What's in it:

  • New verify kernels for the M5 tensor units (SplitSums32, lighter barriers).
  • Kernel choices measured for the 40-core M5 Max, loaded from a file (tuning/, SPLASH_KERNEL_CHOICES). ./splish picks them automatically.
  • A copy rule for coding agents: verbatim continuations are drafted exactly.
  • Faster GGUF decode (G4a, bit-identical).
  • Benchmark and correctness tools (dev/m5/), and dev/m5/report.py for performance reports.

Full results, method, what did not work and the ideas left to try are in the README. The next release (v1.1) adds an auto-tuner for other M5 chips.