ThunderMLX v0.3.0 — lossless high-context decode speedup
Single-pass exact sparse-selection kernel — bit-exact (60/60 parity), identical model outputs, selection cost 18.7ms→4.0ms/token @200k (4.7x). Real agent decode with native tools: 50k 22.8→24.4 (+7%), 80k 21.5→23.9 (+11%), 150k 17.7→22.4 (+26%), 200k 15.6→21.5 (+38%). Chat/prefill/TTFT/cache unchanged by design and measurement. Field-validated in a live 178k-token agent session (7,374-token write at 20.6 tok/s flat, 63 tool calls, zero retries). Enable: MLX_M3_MSA_SELECT_V2=1. Full methodology + honest observations: ops/fable_lab/RESULTS.md.