Skip to content

primecount-8.8

Latest

Choose a tag to compare

@kimwalisch kimwalisch released this 28 Sep 12:23

This release contains many improvements for ARM CPUs. The D part of Gourdon’s prime counting algorithm has been vectorized using ARM NEON, speeding up the D algorithm by about 10% on Apple Silicon CPUs. The AC algorithm has been vectorized using ARM SVE, and the ARM SVE implementation of the D algorithm has also been improved.

ChangeLog

  • D_arm_neon.hpp: ARM NEON filtering for Gourdon’s D algorithm.
  • D.cpp: Refactor runtime SIMD dispatching.
  • AC.cpp: Refactor runtime SIMD dispatching.
  • Fix potential MinGW GCC 16 assertion warning in debug mode.
  • AC_libdivide.hpp: Factor out inner-loop arithmetic using sum_pi_libdivide().
  • AC_default.hpp: Factor out inner-loop arithmetic using sum_pi().
  • Remove unnecessary integer literal suffixes.
  • sieve/count_simd.hpp: New ARM NEON sieve count kernel, for ARM CPUs without ARM SVE.
  • D_arm_sve.hpp: Use manual ARM SVE vectorization instead of relying on auto-vectorization.
  • D_arm_sve.hpp: Improve CPU pipelining.
  • D_avx512.hpp: Improve CPU pipelining.
  • D_default.hpp: Improve CPU pipelining.
  • AC_arm_sve.hpp: Add ARM SVE implementation of AC algorithm.
  • AC.cpp: Add ARM SVE runtime dispatch.
  • fast_div.hpp: Add ARM SVE integer division functions.
  • test/fast_div.cpp: Add ARM SVE tests.
  • Add AGENTS.md file for AI agents.
  • Update to libprimesieve 12.16.