Skip to content

fastkernel 1.1.0

Choose a tag to compare

@abhishekgahlot2 abhishekgahlot2 released this 06 Oct 22:39
· 14 commits to main since this release

The fastest inference engine for Qwen3.8-27B and Qwen3.6-35B-A3B on Apple Silicon. Prebuilt, no Xcode needed.

New in 1.1.0: built on Splash 1.3.0. On an M5 Max it writes answers 1.11× faster than Splash 1.3.0 (2026-10-07). Compared with fastkernel 1.0.0, it reuses its prompt cache better.

Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.

  1. Download fastkernel-1.1.0-macos-arm64.tar.gz below.
  2. tar -xzf fastkernel-1.1.0-macos-arm64.tar.gz && cd fastkernel
  3. Got it through a browser or AirDrop? Run /usr/bin/xattr -dr com.apple.quarantine . once.
  4. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash
  5. Chat at http://127.0.0.1:8000, or run ./splash opencode (or claude, codex, hermes) in a second terminal.

24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.

The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).

sha256 {{SHA256}}