Skip to content

Benchmarks

angelatgithub edited this page Sep 20, 2026 · 2 revisions

Benchmarks

Every number here is measured, reproducible, and gated on correctness — including the cells where a mojo kernel is not the fastest option. This page consolidates the per-package measured tables; each package README in the repo carries its own full table with per-cell latencies.

Methodology

  • Correctness gate before timing. No timing pass runs until the kernel's output has been asserted against the real reference package at that kernel's documented tolerance (several are bit-identical). A fast wrong answer is a failed benchmark, not a fast one.
  • Median of 5 runs for every cell (cold-start cells: median of 5 fresh processes, or ×3 patterns for search workloads). Absolute latencies shift with machine load; speedup ratios within each table are internally consistent.
  • Cold and warm reported separately. Cold = first call in a fresh process (dlopen, runtime init, allocator warm-up). Warm = steady state. The consolidated table below cites warm numbers; notable cold regressions are flagged in the lowlights.
  • Seeded, synthetic corpora. All inputs are generated locally from fixed seeds (Zipf-ish vocabularies, typo-seeded patterns, seeded MO coefficients, seeded soundings), so runs are reproducible and no upstream dataset is redistributed.
  • Same machine, pinned environment. Unless a package README says otherwise: Apple M4 Max, macOS 26.6.2 arm64, Python 3.12.14, NumPy 2.5.3, Node v23.10.0, Mojo 1.1.0, measured 2026-09-19. Reference packages are pinned PyPI/npm releases (each README names its oracle version).
  • Reproduce: pixi run bench, pixi run bench-cclib, pixi run bench-fuse for the three flagship kernels; benchmarks/bench_<name>.py / .mjs for the rest. The README tables are regenerated from these scripts, never hand-edited.

Consolidated measured table

Warm steady state, peak-and-range per package. "Parity" is the measured agreement against the reference oracle, asserted on both the native and forced-fallback backends.

Package Accelerates Measured speedup (warm) Parity
bm25-mojo rank_bm25 116×–8,769× (1k–100k docs) max abs diff 3.6e-15
cclib-mojo cclib method.volume 72× vs fastest shipping backend; 1,898–7,033× vs PyQuante1 path max abs diff 1.5e-12
pypdf-filters-mojo pypdf filters 809–1,776× LZW; 12–15× PNG/TIFF predictor end-to-end identical decoded bytes
pykalman-mojo pykalman 5–827× filter; 3.9–589× smooth within 1e-10
elephant-mojo elephant surrogates 218–342× dither; 2× p-value spectrum bit-exact spectrum; dither ≤1e-3
uproot-mojo uproot 172.6× vector<string> walk; 9.4× cold file→array byte-for-byte
metpy-mojo MetPy cape_cin 156.6× (625-col grid); 78× single column within documented tolerances
natural-mojo natural edit distances 3.7–64.5× (8–1,024 code units) bit-identical distances
dynesty-mojo dynesty 4.3–51.9× (defaults → bootstrap=0 vectorized) posterior/evidence parity, asserted
jsonpath-mojo jsonpath_ng 1.7–49.2× (shrinks with document size) same matches, asserted
croniter-mojo croniter 15.3–46.0× per get_next call same fire times, asserted
obspy-mojo ObsPy Konno–Ohmachi 3.9–6.7× steady-state; up to 45× with window build rtol 1e-9, atol 1e-12 (float64)
nx-mojo networkx 19.9–40× betweenness; 1.2–2.1× Dijkstra asserted vs networkx, both backends
fuse-mojo Fuse.js 7.1.0 8.8–38.7× warm; 2.9–13.4× cold bit-identical (0 diff)
ruptures-mojo ruptures 2.2–48× (Dynp 35–45×, PELT 10.5–48×, Binseg 2.2–2.9×) same breakpoints, asserted
ase-mojo ASE neighbor list 20.3–28.3× native (8.5–12.2× NumPy fallback) identical (i, j, S) sets
jsonschema-mojo jsonschema 2.1–3.1× small docs; 21–28× large nested same verdicts and error surfaces
sacrebleu-mojo sacrebleu 14.1–15.2× chrF; 2.9–3.1× BLEU end-to-end 0.0 measured diff (gate 1e-9)
toml-mojo tomllib / tomlkit 1.8× vs tomllib; ~8× vs tomlkit (64 KB–1 MB) same parsed objects, asserted
langdetect-mojo langdetect 4.3–6.9× per detection same detections, asserted
vader-mojo vaderSentiment 4.1× warm same scores, asserted
ckmeans-mojo simple-statistics ckmeans 1.6–3.4× exact cluster assignments
minisearch-mojo MiniSearch fuzzy/prefix 0.8–3.1× by query mix identical results, asserted
bio-mojo Biopython Bio.SeqIO 1.4–2.3× (GenBank 2.1–2.3×) record-for-record, asserted
ta-mojo pandas-ta-classic 1.8–93× (indicator- and length-dependent) bit-exact native; battery max 8.5e-14
jmespath-mojo jmespath 0.01–1.1× — serialization-bound exact incl. exception classes (5,000+ cells)

Honest lowlights

We publish the cells where the kernel loses, because the parity harness reports them whether we like the answer or not:

  • jmespath-mojo is mostly ≤1× on its own benchmark. The kernel's raw evaluation is several times faster than the reference interpreter, but marshal.dumps serialization across the FFI boundary eats the win on typical queries (best cell: 1.10× on a 5-condition filter; worst: 0.01×). It ships because parity and the fallback are exact — not because it wins. This is what an honest 1.1× looks like.
  • minisearch-mojo dips to 0.8–0.9× on exact-match and small fuzzy mixes at 10k–50k docs; it only earns its keep on prefix/autoSuggest mixes (up to 3.1× at 50k).
  • Cold starts can cost more than they save on small work: ta-mojo's fastest indicator (EMA) measures 0.4× cold; vader-mojo is 0.85× cold (import + first call); uproot-mojo warm steady-state full open+read is 0.9× — its win is the cold path and the basket walk.
  • bm25-mojo vs bm25s (numba incumbent): bm25-mojo v2 wins every cell of the 9-cell grid — 1.10×–1.95× warm steady-state (100k/20t re-verified at 1.12× over 11 interleaved samples) — plus 10–1000× on cold starts (AOT vs numba's ~4 s JIT). The v2 design: idf-weighted CSR baked at index time, query batching, no-rezero accumulation (float32 storage was measured and rejected: 94× over the 1e-8 parity gate, so float64 stays and is bit-faithful to rank_bm25 at 3.6e-15). Full autopsy and table in Kernel: bm25 and benchmarks/AUTOPSY-bm25s.md in the repo.
  • Scaling caveats: pykalman-mojo's 827× is a small-system number (n_s=2); at n_s=20 the win is 5×. jsonpath-mojo slides from 49× to 1.7× as documents grow. sacrebleu-mojo's BLEU is bounded by the shared Python tokenizer (~3× end-to-end; the kernel's n-gram statistics alone are ~20×).

If a workload of yours shows a regression we didn't measure, that's a bug report we want: open an issue with the corpus shape.


Kernel catalog · How It Works · Apache-2.0, © 2026 Algenta

mojo-kernels — clean-room Mojo kernels as drop-in accelerators

Start

Understand

Contribute

Project

Apache-2.0 · © 2026 Algenta

Clone this wiki locally