Releases: SAY-5/spoofline
Release list
v5.1.0
The browser demo now runs on the package's own code path, and the README's figures are gated on committed runs. web/scripts/export.py scores its reference through spoofline.scoring.RunScorer and exports both graphs through spoofline.export.export_stream; it had kept a second copy of the ONNX wrapper and imported score_single_clip, which the v4 series removed, so the documented way to rebuild the page had been broken at HEAD. The page shows what the repository computes: the logistic fusion beside the weighted sum in the calibration table and in the clip lab, and the stream each fusion decision is attributed to, with the hero naming the corpus, the seed, the held out pair and the training commit, and saying that the catch strips replay exported logits rather than scoring in the tab.
docs/runs/ holds the JSON of the demo run and of the full profile sweep, and tests/test_readme_numbers.py re-renders every block and table the README quotes from them, so a stale figure fails the suite instead of drifting; the reduced profile variance table is the one block left without an artifact, and the README says so above it. CI runs the page as well: npm ci, npm run typecheck, npm run selfcheck and the production bundle, with ruff now linting web/scripts, and the self-check makes 374 assertions over 25 clips because both fusion probabilities, both decisions and the attribution have to match PyTorch.
spoofline sweep --pairs default restricts the sweep to the profile's held out pair, and three seeds of it at the full profile put the headline in context: fused unseen precision 0.972 (std 0.030, bootstrap interval 0.941 to 1.000), so the demo run's 0.941 sits at the bottom of the interval, while video alone holds precision 1.000 in all three runs. spoofline robustness takes its seed, held out families and split membership from the run's results.json and splits.json instead of recomputing them from the profile, so a run trained with another seed can be post-processed. Splits.counts() reports the clips dropped for carrying a held out family, so the four split sizes and the dropped count add up to the corpus, and the report orders its own rows so a run rendered from results.json prints what the run printed. The corpus cache compares the whole corpus shape rather than the seed and the clip count alone, so a corpus rendered with other frame or audio settings is refused instead of reused. On Linux torch and torchaudio resolve from the PyTorch CPU index, so a CPU only test run no longer installs the CUDA toolkit, and CI installs with uv sync --locked.
Accessibility and prose: one live region for the calibration table instead of one per cell, a plain text hero heading, the draggable target named in the chart label, larger small type, and a citation for the CNN-LSTM arrangement in the architecture note. Four new test modules cover the README figures against the committed runs, the browser demo's export path, reading a finished run's seed and splits back, and committed artifacts carrying no machine paths: 172 tests at this commit.
v5.0.0
spoofline export --onnx writes both streams as ONNX with the normaliser folded into the graph and a dynamic batch and step axis, and refuses to finish unless every clip's ONNX Runtime logit is within 1e-4 of the PyTorch logit; on the demo run the largest difference is 3.81e-06 over 32 clips for both streams.
spoofline model-card renders a card from the last evaluation run with the data note, the splits, every threshold and its score formula, seen and unseen metrics for all four detectors, the robustness table and the limitations, and the card of the demo run is committed as docs/MODEL_CARD.md.
spoofline score now takes any number of clips and with --json emits one document carrying schema_version, the run directory and every logit, probability, flag, both decisions and the triggering stream per clip.
spoofline bench reports per clip p50 and p95 CPU latency for each stream and end to end on both engines: PyTorch 10.07 ms, 2.54 ms and 17.04 ms at p50, ONNX Runtime 4.50 ms, 2.05 ms and 8.44 ms, on one thread while another heavy job shared the machine.
The test suite is 151 tests, and two separate full demo runs produced byte identical summary blocks apart from their timings.
v4.0.0
spoofline robustness degrades every bona fide test clip of a finished run with eight benign perturbations at 30 severities in total, JPEG quality, video noise, brightness and contrast drift, frame dropout, audio noise, resampling, mild reverb and short audio dropouts, and reports the false alarm rate of each stream and both fusions.
On the demo run the clean false alarm rate of the weighted sum is 0.116, and benign audio changes break it: 40 dB SNR noise raises it to 0.849, a 12 kHz resample to 1.000 and a 0.1 second reverb to 0.709.
Video degradations mostly pass through the weighted sum, which puts only 0.05 on video, with JPEG quality 30 at 0.174 and brightness drift 0.35 at 0.244, while video noise and frame dropout change nothing.
The new abstain option skips clips whose calibrated stream probabilities disagree by more than a margin, and at margin 0.9 it reaches unseen precision 1.000 at coverage 0.821 while lowering seen precision to 0.881 at coverage 0.487.
The honest conclusion is that the calibrated precision holds for clean capture only, and a deployment would need benign channel augmentation and recalibration.
v3.0.0
Adds a learned logistic fusion over both calibrated stream probabilities plus their absolute disagreement, fitted on the calibration split only, with its operating point chosen by the same precision constrained threshold search at 0.95 as every other detector.
Every fusion decision is attributed to the stream that triggered it by silencing one stream at a time, printed as triggered_by by spoofline score and tabulated against the attacked modality in the pipeline summary.
Over the 48 runs of the reduced profile sweep the logistic fusion reaches unseen precision 0.946 against 0.944 for the weighted sum, 0.934 for the OR rule and 0.916 for the AND rule.
It does not measurably narrow the gap to the best single stream: per run that gap is minus 0.036 for logistic and minus 0.038 for the weighted sum, and the paired difference in unseen precision is plus 0.002 with a 95 percent interval from minus 0.011 to plus 0.017.
The weighted sum remains the primary decision and the logistic fusion is reported beside it.
v2.0.0
spoofline sweep repeats the whole train, calibrate, fuse and evaluate protocol for all 16 leave two families out splits, each pairing one held out video family with one held out audio family, across several seeds.
It reports the mean, the sample standard deviation and a percentile bootstrap 95 percent interval of the mean for precision, recall, F1, EER and AUC per detector, and caches every run's raw logits so the evaluation is recomputed without retraining.
The variance table in the README comes from a documented reduced profile of 960 clips with 8 frames and 1 second of audio and 8 video and 6 audio epochs, which finished 48 runs over 3 seeds in 1125 seconds on a 10 core CPU.
Over those runs fused precision on unseen families averages 0.944 with a standard deviation of 0.048, against 0.985 on seen families.
Fused precision still sits 0.038 below the better single stream of the same run on average, while fused recall on unseen families is 0.566 against 0.412 for video and 0.268 for audio.
v1.0.0
Baseline release of spoofline, a two stream audio and video spoof detector built on one CNN LSTM per modality.
Each stream gets a Platt calibration fitted on a held out calibration split and a threshold chosen at a target precision of 0.95, and the two calibrated scores are fused by a weighted sum with AND and OR rules kept for reference.
Evaluation follows a leave one attack family out protocol with identity disjoint train, calibration and test pools, and reports precision, recall, F1, EER and AUC on seen and unseen families.
The corpus comes from a deterministic generator with eight real signal transformation attack families, because the public spoofing datasets need signed licences.
On the single measured run the fused detector reaches unseen family precision 0.941 at recall 0.600, below the video stream's 1.000 precision but well ahead on recall, F1 and AUC.