Fearless SIMD 1.0 is here! 🎉
This release has an MSRV of 1.89.
Added
- Added the optional
#[simd]attribute infearless_simd_macros0.1.0 to annotate SIMD-generic functions. The first argument can be a SIMD token, vector, mask, reference, or user-defined wrapper implementingExtractToken. The core library does not depend on the macro crate. (#347, #383 by @Shnatsel) - Added
Native<S>toSimdFloatElementandSimdIntElementto select native-width vectors in scalar-generic code. (#378 by @Shnatsel) - Added the public
ExtractTokentrait for SIMD tokens, vectors, masks, references, and user-defined wrappers. - Added
reversefor all SIMD vector and mask types. (#356 by @Shnatsel) - Added
rotate_elements_leftandrotate_elements_rightto mask types. Rotations wrap the offset, matching the existing non-mask vector operations. (#360 by @Shnatsel) - Added lane-wise
saturating_addandsaturating_subfor all integer vector types and backends. (#352 by @Shnatsel) - Added lane-wise
count_onesandcount_zerosoperations for all integer vector types and backends. The result has the same vector type as the input, including for signed integers. (#345 by @Shnatsel) - Added
reduce_min,reduce_max,reduce_min_precise, andreduce_max_precisefor all non-mask vector types. The precise floating-point variants ignore quiet NaNs, returning NaN only when all lanes are quiet NaNs; signaling NaNs have implementation-defined behavior. (#358 by @Shnatsel) - Added
reduce_sumandreduce_productfor all non-mask vector types. Integer arithmetic wraps. Floating-point reductions use a fixed order that produces identical results across backends for a given vector type and lane count, except for NaN bit patterns. Summation limits each lane’s contribution to at mostlog2(N)roundings forNlanes. (#357, #361 by @Shnatsel) - Added
mul_add_preciseandmul_sub_precisefor floating-point vectors. They guarantee the infinite-precision product-plus-add rounded once, including on SIMD levels without hardware fused multiply-add instructions. They are not susceptible to the bug in Rust’s standard library,std::simd, and musl libc that causes incorrect rounding for subnormal results. SSE4.2 gets SIMD emulation of these operations for better performance. (#323, #324 by @Shnatsel) - Documented the storage representation of the SIMD vector types. The documented representation will not change without a semver major version change. (#330 by @danderson)
- Added
TryFrombounds toSimdIntElement, allowing attempted conversion from all primitive integer types. (#335 by @danderson, @Shnatsel) - Added a security policy. Starting with v1.0, the latest Fearless SIMD version for each MSRV will receive security backports for at least three years after that Rust version was released. (#367 by @Shnatsel, @DJMcNab)
Changed
- Breaking change:
witness()has moved fromSimdBaseandSimdMaskto their newExtractTokensupertrait and has been renamed totoken(). - Breaking change:
SimdBase::NandSimdMask::Nhave been renamed toLEN, matching thestd::simdnaming. (#366 by @Shnatsel) - Breaking change:
SimdBase::as_arraynow borrows the vector and returns an array reference, while owned extraction has moved toto_array. The oldas_array_refandas_array_mutmethods have been replaced byas_arrayandas_mut_array, matching thestd::simdAPI. (#351 by @Shnatsel) - Breaking change:
abshas moved fromSimdFloattoSimdBaseand is now available on integer vectors. Signed integers use wrapping absolute value, leaving the minimum representable value unchanged; unsigned integers are unchanged. (#371 by @Shnatsel) Simd::vectorizenow marks its inner wrappers#[inline], allowing inlining into callers with compatible target features. (#347 by @Shnatsel)- Conversions between 64-bit integers and floating-point values have been optimized on x86, particularly for 256-bit AVX2 vectors. (#348 by @Shnatsel)
- x86 code generation has been improved for 8-bit integer multiplication, 8-bit per-lane left shifts on AVX-512, and 8-bit and 16-bit integer
unzipon SSE4.2 and AVX2. (#350 by @Shnatsel) swizzle_dynandswizzle_dyn_precisehave been optimized on AVX2. On WebAssembly withrelaxed-simdenabled, 128-bitswizzle_dynnow uses the relaxed swizzle instruction. (#322, #362 by @Shnatsel)SimdMask::to_bitmaskhas been optimized on NEON, including dedicated implementations for wide masks. (#344 by @Shnatsel)- Mask reductions (
any_true,all_true,any_false, andall_false) on masks wider than a native register now combine vector halves before reducing, avoiding branches and repeated scalar extraction. (#343 by @Dr-Emann) Level::is_fallbackis now marked inline, allowing callers to eliminate the function call when the result is a compile-time constant. (#336 by @Dr-Emann)
Fixed
- Fixed a possible dispatch panic with custom x86 target-feature configurations by including
adxin the AVX-512 feature checks that control AVX2 dispatch availability. (#377 by @Shnatsel) - Fixed
fracton NEON for large finite values and infinities. Large finite values now return zero, and infinities return NaN, matching the other backends. (#365 by @Shnatsel) - Fixed the sign of zero returned by
mul_subon NEON for some combinations of signed inputs. (#323 by @Shnatsel) - Hardened the hidden
kernel!implementation helpers so callers cannot bypass SIMD token and target-feature checks by invoking them directly. (#363 by @Shnatsel)
Full Changelog: v0.7.0...fearless_simd-v1.0.0