Repository navigation
Releases: loopai-hq/pulsar
Release list
Pulsar 1.1.6
Faster long prompts: 6.9% faster writing on a 32K-token agent task (Qwen3.8-27B), 3.6% (Qwen3.6-35B-A3B).
- Prompts are read 1.7% faster.
- With the Neural Engine off, answers match 1.1.5 on 12 of 12 Qwen3.8-27B prompts.
- The Neural Engine split calibrates more reliably at startup; the first start after installing spends about 6 s on it.
./pulsar --versionprints "Pulsar 1.1.6".
Measured on an M5 Max, 2026-10-10. The long-prompt kernels apply on the 40-core M5 Max; other Macs run 1.1.5's.
Download pulsar-1.1.6-macos-arm64.tar.gz (sha256 d739a01a888d204b1e273df831c8ee74b4cf769cc6c3e5504e7e28105c8f6b1f) and follow the Quick start in the README. This release ships as a prebuilt download; the source in this repository is 1.1.4.
Pulsar 1.1.5
Qwen3.8-27B on an M5 Max: 153 tok/s writing code, 344 tok/s editing it (up to 477 tok/s).
- 1.44× faster than lithos-metal 0.1.2 (official release) across 10 prompts; faster on every prompt.
- Same answers as 1.1.4, with faster gate/up kernels and sturdier GDN buffer checks.
- Reads the guesser word list straight from a model package that declares it.
./pulsar --versionprints "Pulsar 1.1.5".
Download pulsar-1.1.5-macos-arm64.tar.gz (sha256 6a01d6c44b95f0f7121f8a0a8f68290e30fb859d01d1edf5ef6be668dd576cae) and follow the Quick start in the README. This release ships as a prebuilt download; the source in this repository is 1.1.4.
Pulsar 1.1.4
Pulsar 1.1.4: fastkernel is now Pulsar. Same engine, same speed and the same outputs as fastkernel 1.1.3.
- New name. The commands are now
./pulsar serveand./pulsar opencode, the engine isengine/pulsar, and the
download ispulsar-1.1.4-macos-arm64.tar.gz. TheSPLASH_*settings keep their names. Links to
github.com/loopai-hq/fastkernel lead here. - Same engine. The GPU kernels are byte-identical to 1.1.3's, and every output test matches 1.1.3.
Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.
- Download
pulsar-1.1.4-macos-arm64.tar.gzbelow. tar -xzf pulsar-1.1.4-macos-arm64.tar.gz && cd pulsar- Got it through a browser or AirDrop? Run
/usr/bin/xattr -dr com.apple.quarantine .once. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./pulsar serve --model incoai/Qwen3.8-27B-Splash- Chat at http://127.0.0.1:8000, or run
./pulsar opencode(orclaude,codex,hermes) in a second terminal.
24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.
The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).
sha256 baef6123d2f8095a44d1d382a6d61a52c5936cd7f0e9515b24f8a3c050bdd8f8
fastkernel 1.1.3
fastkernel 1.1.3: 1% decode speed improvement over v1.1.1, measured in a paired serving test (96 pairs, identical outputs). Outputs are byte-identical to v1.1.1.
- A shorter sampling search. For sampled answers with top-k at most 32 and no min-p, the top-k and top-p search
reads the 32 most likely tokens of each vocabulary shard instead of the whole vocabulary. It picks the same tokens. - Less state written by the linear-attention (GDN) layers. The checking pass no longer stores the layers' running
state. When a step keeps fewer than all eight guessed tokens, the commit replays the kept ones instead, which skips a
151 MB write on Qwen3.8-27B. - A deferred state commit. With one request, each step's GDN state commit runs inside the next step's pass instead
of as separate GPU work.
All three are on by default. SPLASH_SAMPLER_TOPK32=0, SPLASH_GDN_SCAN_NOSTORE=0 and SPLASH_GDN_DEFER=0 turn
them off (docs/SWITCHES.md).
Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.
- Download
fastkernel-1.1.3-macos-arm64.tar.gzbelow. tar -xzf fastkernel-1.1.3-macos-arm64.tar.gz && cd fastkernel- Got it through a browser or AirDrop? Run
/usr/bin/xattr -dr com.apple.quarantine .once. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash- Chat at http://127.0.0.1:8000, or run
./splash opencode(orclaude,codex,hermes) in a second terminal.
24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.
The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).
sha256 82065032637c798711dff5736227a8abc6f1dc45de6aa808edff509d0f2e77c4
fastkernel 1.1.1
fastkernel 1.1.1: robustness fixes from code review. Decode speed and outputs are unchanged.
- A GPU command that fails partway through being built now waits for the work it already sent to the GPU.
- GPU errors during cleanup now mark the engine unhealthy.
- On GGUF models, the smaller draft vocabulary falls back to the full one instead of stopping the engine at startup.
- Memory accounting for the draft head is corrected.
Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.
- Download
fastkernel-1.1.1-macos-arm64.tar.gzbelow. tar -xzf fastkernel-1.1.1-macos-arm64.tar.gz && cd fastkernel- Got it through a browser or AirDrop? Run
/usr/bin/xattr -dr com.apple.quarantine .once. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash- Chat at http://127.0.0.1:8000, or run
./splash opencode(orclaude,codex,hermes) in a second terminal.
24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.
The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).
sha256 f1666c24110874fd5d1a4a3459dd3a8169d87d1669534e94452c1f26346bc529
fastkernel 1.1.0
The fastest inference engine for Qwen3.8-27B and Qwen3.6-35B-A3B on Apple Silicon. Prebuilt, no Xcode needed.
New in 1.1.0: built on Splash 1.3.0. On an M5 Max it writes answers 1.11× faster than Splash 1.3.0 (2026-10-07). Compared with fastkernel 1.0.0, it reuses its prompt cache better.
Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.
- Download
fastkernel-1.1.0-macos-arm64.tar.gzbelow. tar -xzf fastkernel-1.1.0-macos-arm64.tar.gz && cd fastkernel- Got it through a browser or AirDrop? Run
/usr/bin/xattr -dr com.apple.quarantine .once. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash- Chat at http://127.0.0.1:8000, or run
./splash opencode(orclaude,codex,hermes) in a second terminal.
24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.
The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).
sha256 {{SHA256}}