Skip to content

kithara dsp

Pavel Litvinenko edited this page Sep 27, 2026 · 1 revision

kithara-dsp

Documentation reviewed from source revision b693d34ed. This records the documented contract at that revision; it is not a new runtime validation. API and usage · All crates.

Vector DSP kernels over planar f32 slices, plus the one import path for firewheel's fade, mix and smoothing helpers.

Purpose

Layout and sanitize loops used to live inline in kithara-signal and kithara-decode, on the scalar path or on fast-interleave. kithara-dsp owns them as free functions whose backend the build target picks, so a hot loop is vectorised once and every consumer gets the same bits.

Public surface

  • deinterleave_variable repeats the signature of fast_interleave's function for f32 and returns (). It writes the same bits (pinned against fast_interleave for 1–8 × 1–8 channels, with NaN payloads, infinities and subnormals in the data). Surplus output channels stay untouched, surplus input channels are ignored.
  • interleave_channel_major and deinterleave_channel_major address every plane of one strided channel-major slice (channel c fills slice[c * stride..(c + 1) * stride]), so any channel count runs without a slice of plane references and without allocating. They equal the fast_interleave functions over chunks_exact(stride); samples past the last whole plane are ignored or stay untouched.
  • Deliberate difference: where fast_interleave panics on a short slice, these handle the whole frames every used slice holds (the common prefix) and leave the rest untouched.
  • sanitize is the only function that rewrites values: NaN, ±infinity, subnormals and −0.0 become +0.0, bit-identical to kithara_signal::sanitize_sample (pinned by an exhaustive 2³² bit-pattern test in tests/crates/dsp). No other kernel sanitizes implicitly.
  • Nothing allocates or panics: the crate is #![forbid(unsafe_code)] and denies clippy::indexing_slicing, clippy::arithmetic_side_effects and clippy::panic.

Backends

Backends are private modules chosen by cfg; callers never name one.

Channels macOS, iOS everything else wasm32
1 copy_from_slice copy_from_slice copy_from_slice
2 vDSP ztoc/ctoz through kithara-apple (the only unsafe) fearless_simd at the best level the CPU reports fearless_simd, +simd128 fixed at compile time
3+ scalar strided copy, four frames per iteration same same
sanitize fearless_simd fearless_simd fearless_simd

On x86 the SIMD level is detected once per process; elsewhere it is fixed at compile time. The crate's unit tests run every pair backend (portable at the native level, portable at Level::fallback(), Accelerate on Apple) against a scalar oracle for every size and offset, specials included.

Why the strided copy moves four frames per iteration: a loop that moves one sample per iteration runs at half speed whenever it straddles a 4096-byte page, and opt-level = "z" neither unrolls nor aligns it. Four samples per iteration amortize that fetch and roughly halve the cost on every target.

Why not BLAS for 3+ channels: Accelerate's cblas_scopy with a stride of 2 or more quiets a signaling NaN (sets bit 22), so it is not a bit-exact copy. At stride 1, and in vDSP ztoc/ctoz, bits are preserved. The strided path stays a plain scalar copy on every target.

wasm +simd128 minimum browsers

Safari / iOS 16.4, Chrome 91, Firefox 89. .cargo/config.toml sets +simd128 for every wasm32 build; the kithara-ffi web build passes --enable-simd to wasm-opt.

Integration

kithara-signal interleaves and deinterleaves its buffers through the layout functions; its pooled planar buffer goes through the channel-major ones, so no channel count allocates. kithara-decode sanitizes resampled planes.

Facade over firewheel

kithara_dsp::fade::FadeCurve and kithara_dsp::param::* (SmoothedParam, SmootherConfig, SmoothingFilter, SmoothingFilterCoeff, Mix, MixDSP and their constants) re-export firewheel's types as the one import path. The arch check firewheel_dsp_facade denies importing these modules from firewheel directly; node, event and buffer APIs stay direct firewheel imports. A re-export can later become a local type of the same name without touching consumers.

Crossfade gains in kithara-host (crossfader_gain) and kithara-play (CrossfadeSettings::gains) speak FadeCurve, so the host crossfader and the play crossfade share firewheel's curve, including its exact 0/1 ends below 1e-5 and above 0.99999.

no_panic spike (outcome)

#[no_panic] on the portable kernels was evaluated and rejected: the attribute moves the kernel body into a closure that does not inherit the fearless_simd target-feature context, so on x86_64 the AVX2 arm stops inlining its intrinsics. The panic guarantee is carried by the restriction lints above instead.

Build notes

  • [profile.dev.package.kithara-dsp] opt-level = 2: the kernels sit on kithara-decode's dev path, and opt-0 fearless_simd turns every lane op into a call.
  • fearless_simd is in hakari final-excludes so the dev-only force_support_fallback feature does not leak into release builds.

Clone this wiki locally