0.2.0 — the MXFP4 native-byte lane
The e2m1 codebook is a table swap of the NF4 mainloop, so the MXFP4 lane ships in this package: grouped GEMM on a checkpoint's exact released bytes with per-tensor sha256 provenance, the pipelined hot/cold decode engine, QLoRA with recompute-in-backward over frozen native bytes, and the executable verify_provenance receipt.
Stamped receipts (in docs/mxfp4/): gpt-oss-120b fused-native serve at reference perplexity (the +9.4% NF4-requant tax deleted); 120b QLoRA at 9.82 GB peak VRAM with 144/144 file == loaded == post-train hashes. Kimi K3 gather scaffolding included, K2-verified — every K3-specific number waits for the 2026-07-27 weights drop and the per-model oracle STOP gate.
Wheel ships: nf4_grouped, nf4_pack_ref, host_gather, mxfp4_pack_ref, mxfp4_grouped, mxfp4_loader, mxfp4_pipelined, mxfp4_qlora, mxfp4_native_load, moonshot_gather, verify_provenance, plus the run/gate harnesses.
Hardened over five review rounds before merge (10 findings: wheel self-consistency, padding safety on all three forward paths, sharded-checkpoint provenance, snapshot resolution, marker-guard fixed-string matching).