int8 vector search for Arm that reaches SDOT and SMMLA. 9.1x faster than the fastest working FAISS int8 mode at slightly better recall, with every number reproducible in CI on free Arm runners.
-
Updated
Aug 5, 2026 - Python
int8 vector search for Arm that reaches SDOT and SMMLA. 9.1x faster than the fastest working FAISS int8 mode at slightly better recall, with every number reproducible in CI on free Arm runners.
2.12x faster IQ4_XS prefill on Arm Neoverse: the missing smmla repack kernel for llama.cpp. IQ4_XS is the smallest standard 4-bit GGUF and had an accelerated matmul path on Intel AMX but none on Arm. Measured stock-vs-patched on free Neoverse N2 CI.
Add a description, image, and links to the i8mm topic page so that developers can more easily learn about it.
To associate your repository with the i8mm topic, visit your repo's landing page and select "manage topics."