Skip to content

Releases: loopai-hq/pulsar

Pulsar 1.1.6

Choose a tag to compare

@abhishekgahlot2 abhishekgahlot2 released this 10 Oct 09:48

Faster long prompts: 6.9% faster writing on a 32K-token agent task (Qwen3.8-27B), 3.6% (Qwen3.6-35B-A3B).

  • Prompts are read 1.7% faster.
  • With the Neural Engine off, answers match 1.1.5 on 12 of 12 Qwen3.8-27B prompts.
  • The Neural Engine split calibrates more reliably at startup; the first start after installing spends about 6 s on it.
  • ./pulsar --version prints "Pulsar 1.1.6".

Measured on an M5 Max, 2026-10-10. The long-prompt kernels apply on the 40-core M5 Max; other Macs run 1.1.5's.

Download pulsar-1.1.6-macos-arm64.tar.gz (sha256 d739a01a888d204b1e273df831c8ee74b4cf769cc6c3e5504e7e28105c8f6b1f) and follow the Quick start in the README. This release ships as a prebuilt download; the source in this repository is 1.1.4.

Pulsar 1.1.5

Choose a tag to compare

@abhishekgahlot2 abhishekgahlot2 released this 09 Oct 21:54

Qwen3.8-27B on an M5 Max: 153 tok/s writing code, 344 tok/s editing it (up to 477 tok/s).

  • 1.44× faster than lithos-metal 0.1.2 (official release) across 10 prompts; faster on every prompt.
  • Same answers as 1.1.4, with faster gate/up kernels and sturdier GDN buffer checks.
  • Reads the guesser word list straight from a model package that declares it.
  • ./pulsar --version prints "Pulsar 1.1.5".

Download pulsar-1.1.5-macos-arm64.tar.gz (sha256 6a01d6c44b95f0f7121f8a0a8f68290e30fb859d01d1edf5ef6be668dd576cae) and follow the Quick start in the README. This release ships as a prebuilt download; the source in this repository is 1.1.4.

Pulsar 1.1.4

Choose a tag to compare

@abhishekgahlot2 abhishekgahlot2 released this 09 Oct 08:26

Pulsar 1.1.4: fastkernel is now Pulsar. Same engine, same speed and the same outputs as fastkernel 1.1.3.

  • New name. The commands are now ./pulsar serve and ./pulsar opencode, the engine is engine/pulsar, and the
    download is pulsar-1.1.4-macos-arm64.tar.gz. The SPLASH_* settings keep their names. Links to
    github.com/loopai-hq/fastkernel lead here.
  • Same engine. The GPU kernels are byte-identical to 1.1.3's, and every output test matches 1.1.3.

Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.

  1. Download pulsar-1.1.4-macos-arm64.tar.gz below.
  2. tar -xzf pulsar-1.1.4-macos-arm64.tar.gz && cd pulsar
  3. Got it through a browser or AirDrop? Run /usr/bin/xattr -dr com.apple.quarantine . once.
  4. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./pulsar serve --model incoai/Qwen3.8-27B-Splash
  5. Chat at http://127.0.0.1:8000, or run ./pulsar opencode (or claude, codex, hermes) in a second terminal.

24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.

The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).

sha256 baef6123d2f8095a44d1d382a6d61a52c5936cd7f0e9515b24f8a3c050bdd8f8

fastkernel 1.1.3

Choose a tag to compare

@abhishekgahlot2 abhishekgahlot2 released this 07 Oct 22:28

fastkernel 1.1.3: 1% decode speed improvement over v1.1.1, measured in a paired serving test (96 pairs, identical outputs). Outputs are byte-identical to v1.1.1.

  • A shorter sampling search. For sampled answers with top-k at most 32 and no min-p, the top-k and top-p search
    reads the 32 most likely tokens of each vocabulary shard instead of the whole vocabulary. It picks the same tokens.
  • Less state written by the linear-attention (GDN) layers. The checking pass no longer stores the layers' running
    state. When a step keeps fewer than all eight guessed tokens, the commit replays the kept ones instead, which skips a
    151 MB write on Qwen3.8-27B.
  • A deferred state commit. With one request, each step's GDN state commit runs inside the next step's pass instead
    of as separate GPU work.

All three are on by default. SPLASH_SAMPLER_TOPK32=0, SPLASH_GDN_SCAN_NOSTORE=0 and SPLASH_GDN_DEFER=0 turn
them off (docs/SWITCHES.md).

Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.

  1. Download fastkernel-1.1.3-macos-arm64.tar.gz below.
  2. tar -xzf fastkernel-1.1.3-macos-arm64.tar.gz && cd fastkernel
  3. Got it through a browser or AirDrop? Run /usr/bin/xattr -dr com.apple.quarantine . once.
  4. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash
  5. Chat at http://127.0.0.1:8000, or run ./splash opencode (or claude, codex, hermes) in a second terminal.

24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.

The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).

sha256 82065032637c798711dff5736227a8abc6f1dc45de6aa808edff509d0f2e77c4

fastkernel 1.1.1

Choose a tag to compare

@abhishekgahlot2 abhishekgahlot2 released this 07 Oct 07:06

fastkernel 1.1.1: robustness fixes from code review. Decode speed and outputs are unchanged.

  • A GPU command that fails partway through being built now waits for the work it already sent to the GPU.
  • GPU errors during cleanup now mark the engine unhealthy.
  • On GGUF models, the smaller draft vocabulary falls back to the full one instead of stopping the engine at startup.
  • Memory accounting for the draft head is corrected.

Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.

  1. Download fastkernel-1.1.1-macos-arm64.tar.gz below.
  2. tar -xzf fastkernel-1.1.1-macos-arm64.tar.gz && cd fastkernel
  3. Got it through a browser or AirDrop? Run /usr/bin/xattr -dr com.apple.quarantine . once.
  4. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash
  5. Chat at http://127.0.0.1:8000, or run ./splash opencode (or claude, codex, hermes) in a second terminal.

24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.

The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).

sha256 f1666c24110874fd5d1a4a3459dd3a8169d87d1669534e94452c1f26346bc529

fastkernel 1.1.0

Choose a tag to compare

@abhishekgahlot2 abhishekgahlot2 released this 06 Oct 22:39

The fastest inference engine for Qwen3.8-27B and Qwen3.6-35B-A3B on Apple Silicon. Prebuilt, no Xcode needed.

New in 1.1.0: built on Splash 1.3.0. On an M5 Max it writes answers 1.11× faster than Splash 1.3.0 (2026-10-07). Compared with fastkernel 1.0.0, it reuses its prompt cache better.

Needs: an M3 or newer Mac, macOS 26.4 or later, Python 3.12 to 3.14, and internet on the first run.

  1. Download fastkernel-1.1.0-macos-arm64.tar.gz below.
  2. tar -xzf fastkernel-1.1.0-macos-arm64.tar.gz && cd fastkernel
  3. Got it through a browser or AirDrop? Run /usr/bin/xattr -dr com.apple.quarantine . once.
  4. SPLASH_DRAFT_HEAD_IDS=$PWD/data/head-ranked.u32 ./splash serve --model incoai/Qwen3.8-27B-Splash
  5. Chat at http://127.0.0.1:8000, or run ./splash opencode (or claude, codex, hermes) in a second terminal.

24 GB Mac: first run sudo sysctl iogpu.wired_limit_mb=20480, and put SPLASH_TEXT_ONLY=1 in front of step 4.
Qwen3.6-35B-A3B: use incoai/Qwen3.6-35B-A3B-Splash in step 4.

The first run sets up its Python packages (20 seconds) and downloads the model (17.4 GB).

sha256 {{SHA256}}