This release contains many improvements for ARM CPUs. The D part of Gourdon’s prime counting algorithm has been vectorized using ARM NEON, speeding up the D algorithm by about 10% on Apple Silicon CPUs. The AC algorithm has been vectorized using ARM SVE, and the ARM SVE implementation of the D algorithm has also been improved.
ChangeLog
D_arm_neon.hpp: ARM NEON filtering for Gourdon’s D algorithm.D.cpp: Refactor runtime SIMD dispatching.AC.cpp: Refactor runtime SIMD dispatching.- Fix potential MinGW GCC 16 assertion warning in debug mode.
AC_libdivide.hpp: Factor out inner-loop arithmetic usingsum_pi_libdivide().AC_default.hpp: Factor out inner-loop arithmetic usingsum_pi().- Remove unnecessary integer literal suffixes.
sieve/count_simd.hpp: New ARM NEON sieve count kernel, for ARM CPUs without ARM SVE.D_arm_sve.hpp: Use manual ARM SVE vectorization instead of relying on auto-vectorization.D_arm_sve.hpp: Improve CPU pipelining.D_avx512.hpp: Improve CPU pipelining.D_default.hpp: Improve CPU pipelining.AC_arm_sve.hpp: Add ARM SVE implementation of AC algorithm.AC.cpp: Add ARM SVE runtime dispatch.fast_div.hpp: Add ARM SVE integer division functions.test/fast_div.cpp: Add ARM SVE tests.- Add
AGENTS.mdfile for AI agents. - Update to libprimesieve 12.16.