Skip to content

kithara beat

Pavel Litvinenko edited this page Sep 28, 2026 · 2 revisions

kithara-beat

Documentation reviewed from source revision 4a470dd2c. This records the documented contract at that revision; it is not a new runtime validation. API and usage · All crates.

Contract

  • Input: whole-track mono f32 PCM at 22 050 Hz. Decode, downmix, and resample are the caller's job — this crate has no decoder or resampler dependency and does no I/O.
  • Output: RawBeats { beats, downbeats } — pooled positions in seconds, sorted, deduplicated, every downbeat snapped to the nearest beat. Grid cleanup and source-frame conversion belong to the consumer (kithara-analysis).
  • Models load from (mel, beat) ONNX bytes via BeatThis::builder() — the caller decides embed vs file vs download; one embed-*-model feature exposes the bundled bytes. The same builder takes the BeatConfig the peak picker runs with and a caller-owned PoolRegion<S> where S: HasPool<f32>. Tensor inputs, copied runtime outputs, and inference scratch all reuse that facade's registered sample pool.

Backends

nn is the neural pipeline below; one embed-*-model feature adds the bundled weights and implies it. dsp is SpectralBeats: its spectrum and autocorrelation come from kithara-dsp, no rten and no model data, so it compiles for wasm32, and it reports beats and never downbeats. Neither feature implies the other.

SpectralBeats::new takes the Tempo it searches, a band and a prior inside it in BPM, 48..=185 around 120 by default; anything else is TempoError. The neural pipeline has no tempo policy. An FFT the backend cannot set up is BeatDetectError::Init.

Pipeline

  1. mel.rs — ONNX mel model: input "audio_pcm" [1, N] -> output "mel_spectrogram" [1, T, 128], hop 441 -> 50 fps. Using the ONNX mel rather than hand-rolled DSP is what guarantees numerical parity with the training pipeline.
  2. inference.rs — chunked beat model (input "spectrogram"): 1500-frame windows starting at -6, stride 1488, 6 border frames trimmed per side, the last start pulled back to align with the spectrogram end. Chunks run in reverse order so earlier chunks overwrite later ones in overlaps — that is the keep_first rule. Outputs are read as beat / downbeat with beat_logits / downbeat_logits as export-name fallbacks.
  3. postprocess.rs — minimal peak picking, no DBN, driven by BeatConfig: a frame is a peak when its logit clears peak_threshold and it is the max of the 2 * peak_half_width + 1 frames around it; peaks at most dedup_width frames apart merge to their running mean; each downbeat then snaps to the nearest beat.

Two groups of numbers, with different owners:

  • Chunk geometry (chunk 1500, border 6, stride 1488, hop 441, 50 fps) stays constant. It is the segmentation the model was trained on, not a knob — a different chunking runs the model outside its training regime.
  • Picking policy (BeatConfig: threshold 0, half-width 3 = +-60 ms, dedup width 1) is the caller's. Its defaults are the frozen parity values: they define numerical parity with the Python reference, the golden fixtures are held to them, and a consumer that moves them must fold the new values into its analysis cache key. The crate does not do that for it.

Validation

cargo test -p kithara-beat --features embed-small-model is the scoped probe for the golden parity test; just test leaves that feature off. The goldens were recorded from the small model, so the test is gated on it alone, and asserts F-measure >= 0.99 at the +-70 ms MIR window. It takes a couple of minutes: whole-track inference, rten* built at opt-level = 3.

cargo test -p kithara-beat --features dsp is the scoped probe for the signal-processing backend. tests/degara.rs scores it against the recorded reference; its unit tests generate their own signal and still mean something when a golden cannot be regenerated.

tests/fixtures/README.md holds every fixture's provenance and parity criterion.

Clone this wiki locally