Skip to content

fearless_simd-v1.0.0

Latest

Choose a tag to compare

@LaurenzV LaurenzV released this 21 Sep 19:38
dcbd824

Crates.io | Docs

Fearless SIMD 1.0 is here! 🎉

This release has an MSRV of 1.89.

Added

  • Added the optional #[simd] attribute in fearless_simd_macros 0.1.0 to annotate SIMD-generic functions. The first argument can be a SIMD token, vector, mask, reference, or user-defined wrapper implementing ExtractToken. The core library does not depend on the macro crate. (#347, #383 by @Shnatsel)
  • Added Native<S> to SimdFloatElement and SimdIntElement to select native-width vectors in scalar-generic code. (#378 by @Shnatsel)
  • Added the public ExtractToken trait for SIMD tokens, vectors, masks, references, and user-defined wrappers.
  • Added reverse for all SIMD vector and mask types. (#356 by @Shnatsel)
  • Added rotate_elements_left and rotate_elements_right to mask types. Rotations wrap the offset, matching the existing non-mask vector operations. (#360 by @Shnatsel)
  • Added lane-wise saturating_add and saturating_sub for all integer vector types and backends. (#352 by @Shnatsel)
  • Added lane-wise count_ones and count_zeros operations for all integer vector types and backends. The result has the same vector type as the input, including for signed integers. (#345 by @Shnatsel)
  • Added reduce_min, reduce_max, reduce_min_precise, and reduce_max_precise for all non-mask vector types. The precise floating-point variants ignore quiet NaNs, returning NaN only when all lanes are quiet NaNs; signaling NaNs have implementation-defined behavior. (#358 by @Shnatsel)
  • Added reduce_sum and reduce_product for all non-mask vector types. Integer arithmetic wraps. Floating-point reductions use a fixed order that produces identical results across backends for a given vector type and lane count, except for NaN bit patterns. Summation limits each lane’s contribution to at most log2(N) roundings for N lanes. (#357, #361 by @Shnatsel)
  • Added mul_add_precise and mul_sub_precise for floating-point vectors. They guarantee the infinite-precision product-plus-add rounded once, including on SIMD levels without hardware fused multiply-add instructions. They are not susceptible to the bug in Rust’s standard library, std::simd, and musl libc that causes incorrect rounding for subnormal results. SSE4.2 gets SIMD emulation of these operations for better performance. (#323, #324 by @Shnatsel)
  • Documented the storage representation of the SIMD vector types. The documented representation will not change without a semver major version change. (#330 by @danderson)
  • Added TryFrom bounds to SimdIntElement, allowing attempted conversion from all primitive integer types. (#335 by @danderson, @Shnatsel)
  • Added a security policy. Starting with v1.0, the latest Fearless SIMD version for each MSRV will receive security backports for at least three years after that Rust version was released. (#367 by @Shnatsel, @DJMcNab)

Changed

  • Breaking change: witness() has moved from SimdBase and SimdMask to their new ExtractToken supertrait and has been renamed to token().
  • Breaking change: SimdBase::N and SimdMask::N have been renamed to LEN, matching the std::simd naming. (#366 by @Shnatsel)
  • Breaking change: SimdBase::as_array now borrows the vector and returns an array reference, while owned extraction has moved to to_array. The old as_array_ref and as_array_mut methods have been replaced by as_array and as_mut_array, matching the std::simd API. (#351 by @Shnatsel)
  • Breaking change: abs has moved from SimdFloat to SimdBase and is now available on integer vectors. Signed integers use wrapping absolute value, leaving the minimum representable value unchanged; unsigned integers are unchanged. (#371 by @Shnatsel)
  • Simd::vectorize now marks its inner wrappers #[inline], allowing inlining into callers with compatible target features. (#347 by @Shnatsel)
  • Conversions between 64-bit integers and floating-point values have been optimized on x86, particularly for 256-bit AVX2 vectors. (#348 by @Shnatsel)
  • x86 code generation has been improved for 8-bit integer multiplication, 8-bit per-lane left shifts on AVX-512, and 8-bit and 16-bit integer unzip on SSE4.2 and AVX2. (#350 by @Shnatsel)
  • swizzle_dyn and swizzle_dyn_precise have been optimized on AVX2. On WebAssembly with relaxed-simd enabled, 128-bit swizzle_dyn now uses the relaxed swizzle instruction. (#322, #362 by @Shnatsel)
  • SimdMask::to_bitmask has been optimized on NEON, including dedicated implementations for wide masks. (#344 by @Shnatsel)
  • Mask reductions (any_true, all_true, any_false, and all_false) on masks wider than a native register now combine vector halves before reducing, avoiding branches and repeated scalar extraction. (#343 by @Dr-Emann)
  • Level::is_fallback is now marked inline, allowing callers to eliminate the function call when the result is a compile-time constant. (#336 by @Dr-Emann)

Fixed

  • Fixed a possible dispatch panic with custom x86 target-feature configurations by including adx in the AVX-512 feature checks that control AVX2 dispatch availability. (#377 by @Shnatsel)
  • Fixed fract on NEON for large finite values and infinities. Large finite values now return zero, and infinities return NaN, matching the other backends. (#365 by @Shnatsel)
  • Fixed the sign of zero returned by mul_sub on NEON for some combinations of signed inputs. (#323 by @Shnatsel)
  • Hardened the hidden kernel! implementation helpers so callers cannot bypass SIMD token and target-feature checks by invoking them directly. (#363 by @Shnatsel)

Full Changelog: v0.7.0...fearless_simd-v1.0.0