Skip to content

How It Works

angelatgithub edited this page Sep 19, 2026 · 1 revision

How It Works

Every package in this repo is built by the same factory. Understand one and you understand all 26.

kernels/<name>/src/<name>.mojo     clean-room Mojo kernel, exported batch C ABI
        │  build.sh: mojo build --emit shared-lib
        ▼
lib<name>mojo.{dylib,so}           AOT-compiled native library
        │  per-platform packaging + repair
        ▼
platform wheel / npm platform pkg  self-contained: Mojo runtime vendored in
        │  ctypes (Python) / koffi (Node) loader, ABI handshake
        ▼
thin wrapper, same API as the reference library
        │  on any load/ABI failure
        ▼
vendored pure-language fallback    silently correct, everywhere

Stage 1 — the kernel: mojo build --emit shared-lib

Each kernel is a clean-room Mojo implementation of a published algorithm (textbook BM25, Bitap/Shift-And, contracted Cartesian Gaussians, …), compiled ahead-of-time into a plain shared library. The kernel exposes a batch-shaped C ABI — a small set of entry points along the lines of <name>mojo_abi_version / create / score / destroy — so FFI cost is paid per call, not per item: one call scores a whole corpus, evaluates a whole 3-D grid, or smooths a whole spectrum. This shape is the single most important design rule; it is what makes the speedups survive the language boundary. AOT compilation also means no JIT warmup and no numba-style first-call stall — cold-start numbers are measured and published separately from warm ones (see Benchmarks).

Stage 2 — per-platform, self-contained packaging

The shared library is useless if it can't run on a machine that has never seen Mojo. So every package is built per platform and repaired into self-containment:

  • Python: hatchling build hooks vendor the kernel into the wheel; delocate (macOS) / auditwheel repair + patchelf (Linux) copy the Mojo runtime libraries into the wheel and rewrite load paths to be wheel-relative. Result: py3-none-macosx_*_arm64 and py3-none-manylinux_*_x86_64 wheels with no absolute rpaths and nothing to compile. No sdist — a source tarball cannot rebuild the native library.
  • TypeScript: an npm workspace per kernel — @…-mojo/core (thin wrapper + koffi loader + vendored fallback) plus @…-mojo/<platform>-<arch> packages pulled in via optionalDependencies, packed and repaired the same way.

(Redistribution terms for Modular's runtime binaries should be confirmed with Modular before any public release — this is tracked in each package README and in the FAQ.)

Stage 3 — the loader: ctypes / koffi + ABI handshake

The wrapper is deliberately thin. At import it resolves the native library in a fixed order — $<NAME>_MOJO_NATIVE_LIB → the bundled wheel/platform library → the repo development build output — loads it with ctypes (Python) or koffi (Node), and checks <name>mojo_abi_version() against its own expected version before any native call. A version mismatch, a missing symbol, a wrong-architecture library: all fall back cleanly. The ABI is versioned precisely so a stale library can never produce silently wrong results; changing a signature means bumping the version and teaching the loader to refuse mismatches.

Stage 4 — the fallback: correct everywhere, by construction

Every package vendors a pure-language reference implementation (pure Python/NumPy, or the reference library's own Apache-2.0 build for the TypeScript kernels). The fallback is selected automatically on any load failure and can be forced with <NAME>_MOJO_DISABLE_NATIVE=1. Crucially, the fallback is asserted against the same oracle at the same tolerance as the native backend — the differential suite runs twice per kernel, once per backend. The fallback is never a stub; on several packages it is itself faster than the library being replaced (e.g. jsonschema-mojo's fallback is ~3× the oracle on small documents).

Why no Windows native — and why that's fine

There is no Mojo toolchain for Windows today, so there are no Windows native builds. This is not a gap in the support matrix so much as the reason the architecture exists: on Windows, the fallback IS the product. A Windows user gets the same package name, the same API, and silently correct results from the vendored implementation, with zero native code involved. When a Windows Mojo toolchain arrives, a win_amd64 wheel becomes a packaging exercise, not a port.

Warm vs. cold, separated on purpose

"How fast is it?" is two questions. Cold = first call in a fresh process: dlopen, runtime init, allocator warm-up — paid once. Warm = steady state after that. Kernels that win warm can lose cold (and vice versa: vader-mojo is 4.1× warm, 0.85× cold), so every benchmark reports both, and the consolidated table cites warm numbers with cold called out where it matters. Because the kernels are AOT-compiled, their cold start is small and bounded — there is no JIT compilation in anyone's latency budget.

Determinism

Same input → same output, on both backends: kernels are single-threaded by default unless documented otherwise (fuse-mojo's thread fan-out owns disjoint document ranges and assembles results in fixed order), benchmarks use seeded corpora, and scoring arithmetic follows the reference's operation order. This is what makes the parity gates — and the speedup claims — meaningful.


Next: Benchmarks · Writing a Kernel · FAQ

Apache-2.0, © 2026 Algenta

mojo-kernels — clean-room Mojo kernels as drop-in accelerators

Start

Understand

Contribute

Project

Apache-2.0 · © 2026 Algenta

Clone this wiki locally