Repository navigation
fastkernel 1.1.0
The fastest inference engine for Qwen3.8-27B and Qwen3.6-35B-A3B on Apple Silicon. Prebuilt, no Xcode needed.
New in 1.1.0: built on Splash 1.3.0. On an M5 Max it writes answers 1.11× faster than Splash 1.3.0 (2026-10-07). Compared with fastkernel 1.0.0, it reuses its prompt cache better.
Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.
- Download
fastkernel-1.1.0-macos-arm64.tar.gzbelow. tar -xzf fastkernel-1.1.0-macos-arm64.tar.gz && cd fastkernel- Got it through a browser or AirDrop? Run
/usr/bin/xattr -dr com.apple.quarantine .once. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash- Chat at http://127.0.0.1:8000, or run
./splash opencode(orclaude,codex,hermes) in a second terminal.
24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.
The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).
sha256 {{SHA256}}