Skip to content

v0.2.0

Latest

Choose a tag to compare

@Nehanth Nehanth released this 07 Sep 18:32
· 5 commits to main since this release
18824b4

Qwen 3.8 27B split across a MacBook and an iPhone in browser tabs, 10.7 tok/s on the same Wi‑Fi. Recorded September 7, 2026 (attached video).

Since v0.1: prefill GEMM (1.6x prompt processing), multi-token-prediction speculative decoding with bit-identical output, adaptive draft depth, sliced and striped hidden-state transport, linear room topology, 2048-token context with a clean stop, room emulator (npm run e2e), f16 rounding fix, and the bench log with every measurement. See CHANGELOG.md.