Skip to content

v0.7.0

Choose a tag to compare

@LaurenzV LaurenzV released this 11 Aug 18:25
· 57 commits to main since this release
57ec21a

Crates.io | Docs

This release has an MSRV of 1.89.

Added

  • Added i64x2, i64x4, i64x8, u64x2, u64x4, and u64x8 vector types, the native-width i64s and u64s associated types, and 64-bit integer operations across all backends. (#253, #310 by @Shnatsel)
  • Added an Sse2 level. This is the new baseline for i686-* and x86_64-* targets, replacing Fallback. It is detected at runtime on Tier-2 i586-* targets. (#270 by @Shnatsel)
  • Added shift_elements_left, shift_elements_right, rotate_elements_left, and rotate_elements_right to non-mask vectors. Shifts accept a padding element and fill the entire vector when the offset is at least its lane count; rotations wrap the offset. (#274 by @Shnatsel)
  • Added full-vector swizzle_dyn and swizzle_dyn_precise byte swizzles. swizzle_dyn permits implementation-defined results for out-of-range indices, while swizzle_dyn_precise always returns zero for them. (#276, #304 by @Shnatsel)
  • Added the SimdElement::BITS constant, exposing the bit width of a vector's lane type to generic code. (#296 by @danderson)
  • Added trait bounds on SimdElement, and introduced the SimdIntElement and SimdFloatElement subtraits. These allow generic code to access many math and utility operations on the elements of SIMD vector types. (#302 by @danderson)
  • Added the SimdWiden and SimdNarrow traits, providing widening, narrowing, and saturating narrowing operations for all integer and floating-point vector types. (#300 by @Shnatsel)
  • Four-way interleaved load and store operations are now exposed on the vector types they operate on and are available to generic code through SimdInterleaved. (#321 by @Shnatsel)

Changed

  • Breaking change: the load_interleaved_128_* and store_interleaved_128_* methods, which exchanged one 512-bit vector, have been replaced by load_four_interleaved_* and store_four_interleaved_*. The new methods exchange an array of four 128-bit vectors directly and are available for all non-mask scalar types. (#298 by @Shnatsel)
  • Breaking change: new_unchecked() on SIMD level tokens such as Avx2 has been renamed to assume_supported(). It is safe to call from contexts containing the appropriate #[target_feature] annotations. Functions without those annotations can still call it inside an unsafe block. (#293 by @Shnatsel)
  • Breaking change: the fxsr CPU feature is now required for all x86 SIMD levels. It is present in hardware on all SIMD-capable CPUs, but can be disabled in some emulators combined with a custom Rust target specification. (#270 by @Shnatsel)
  • Breaking change: operations shared by integer and floating-point vectors have moved from SimdInt/SimdFloat to SimdBase, allowing code generic over any non-mask vector to use arithmetic, comparisons, zip/unzip, and interleave/deinterleave operations. (#308 by @Shnatsel)
  • Breaking change: min, max, min_precise, and max_precise have moved from SimdInt/SimdFloat to SimdBase. (#313 by @Shnatsel)
  • Breaking change: the 204 vector-specific Simd array conversion methods have been replaced by the generic SimdBase::load_array, load_array_ref, as_array, as_array_ref, as_array_mut, and store_array methods. Masks continue to use SimdMask::from_slice and store_slice. (#292 by @Shnatsel)
  • On x86_64 targets with static SSE2 support, Level::baseline() now returns Sse2 instead of Fallback. (#270 by @Shnatsel)
  • The scalar Fallback backend and Level::Fallback variant are no longer compiled when the target has a better ambient SIMD baseline, such as SSE2 on x86 or NEON on AArch64. The force_support_fallback feature continues to make them available for testing. (#320 by @Shnatsel)
  • Runtime CPU feature detection performed by Level::new() is now cached on x86. (#278 by @Shnatsel)
  • Integer shifts by an amount greater than or equal to the element width are now explicitly documented as platform-dependent. Scalar fallback shifts use wrapping shift amounts instead of potentially panicking in debug builds. (#283 by @Shnatsel)
  • Full-vector 8-bit shifts on x86 have been optimized, including a 2.4× faster left-shift formulation and faster signed and unsigned right shifts. (#291 by @Shnatsel)
  • Native-width non-mask vector types now share u8s as their byte representation, enabling Bytes::bitcast between arbitrary lane types in code generic over Simd. (#284 by @Shnatsel)
  • Generic bounds now encode existing relationships between masks, vectors, blocks, elements, and split/combined vector types. (#285 by @Shnatsel)
  • SimdBase::Array now guarantees Copy, Debug, by-value IntoIterator, AsRef, AsMut, and conversion from its vector type. (#285 by @Shnatsel)
  • Simd and SimdBase now require Debug. (#309 by @Shnatsel)
  • The Simd::vectorize documentation now explains when to use it and includes an end-to-end example. (#312 by @Shnatsel)
  • Generated code and metadata have been substantially reduced, cutting x86 build time by roughly one third. (#292, #317, #318 by @Shnatsel)

Removed

  • Breaking change: removed the low-level reinterpret_f32_*, reinterpret_f64_*, reinterpret_i32_*, reinterpret_u32_*, reinterpret_u8_*, cvt_to_bytes_*, and cvt_from_bytes_* methods. Use Bytes::bitcast for arbitrary same-width bit reinterpretation, or Bytes::to_bytes and Bytes::from_bytes for direct byte-vector conversions. (#284 by @Shnatsel)
  • Breaking change: removed the WithSimd trait, which only delegated to the dispatch! macro. Use dispatch! directly instead. (#306 by @Shnatsel)

Fixed

  • Integer negation in the scalar fallback now wraps for the minimum signed value, matching SIMD backends instead of potentially panicking in debug builds. (#253 by @Shnatsel)
  • Fixed x86 8-bit left shifts: overflowing u8 lanes now wrap instead of saturating, and i8 lanes now match Rust's signed shift semantics. (#288, #290 by @danderson)

Full Changelog: v0.6.0...v0.7.0