Skip to content

fastkernel 1.1.1

Choose a tag to compare

@abhishekgahlot2 abhishekgahlot2 released this 07 Oct 07:06
· 13 commits to main since this release

fastkernel 1.1.1: robustness fixes from code review. Decode speed and outputs are unchanged.

  • A GPU command that fails partway through being built now waits for the work it already sent to the GPU.
  • GPU errors during cleanup now mark the engine unhealthy.
  • On GGUF models, the smaller draft vocabulary falls back to the full one instead of stopping the engine at startup.
  • Memory accounting for the draft head is corrected.

Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.

  1. Download fastkernel-1.1.1-macos-arm64.tar.gz below.
  2. tar -xzf fastkernel-1.1.1-macos-arm64.tar.gz && cd fastkernel
  3. Got it through a browser or AirDrop? Run /usr/bin/xattr -dr com.apple.quarantine . once.
  4. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash
  5. Chat at http://127.0.0.1:8000, or run ./splash opencode (or claude, codex, hermes) in a second terminal.

24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.

The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).

sha256 f1666c24110874fd5d1a4a3459dd3a8169d87d1669534e94452c1f26346bc529