Skip to content
angelatgithub edited this page Sep 19, 2026 · 1 revision

FAQ

Why not contribute these kernels upstream directly?

Because the upstreams can't take them. A Mojo kernel means a compiled shared library, a per-platform packaging pipeline, and a new language toolchain in the build — most pure-Python/JS libraries have a hard policy against exactly that, and they should: their universal installability is their feature. The drop-in package is the compromise that respects it: users opt in with one changed import line, upstream keeps its API and its purity, and the reference package is used only as a test oracle. Where an official extension point exists, we use it instead — the networkx backend-dispatch track and the pypdf optional-extra pitch are on the Roadmap.

What about Windows?

There is no Mojo toolchain for Windows today, so there are no Windows native builds — and that's a designed-for case, not a gap. Every package vendors a pure-language fallback that is differential-tested against the same oracle at the same tolerance as the native backend. On Windows the fallback is the product: same package name, same API, silently correct results. When a Windows toolchain exists, a win_amd64 wheel is a packaging task, not a port. (For Python, a py3-none-any fallback wheel can be built explicitly, e.g. CCLIB_MOJO_ALLOW_PURE_WHEEL=1.)

Is redistributing the Mojo runtime in your wheels allowed?

The wheels are self-contained: delocate (macOS) / auditwheel repair / patchelf (Linux) vendor the Mojo runtime libraries and rewrite load paths to be package-relative, so nothing on the end user's machine needs to come from Modular. The package READMEs carry the standing note that redistribution terms for Modular's runtime binaries should be confirmed with Modular before any public release — we track this openly rather than assuming. If you redistribute our wheels inside your own product, the same diligence applies to you.

Why doesn't fuse-mojo support extended search (useExtendedSearch, $and/$or)?

Deliberately. Extended search is a second query language layered on the Bitap core (exact-match ', prefix ^, negation !, OR |, space-AND, plus logical query objects). Bit-exact parity there means reimplementing that language's every corner — the honest scope decision was to make the 90% option surface bit-exact and throw UnsupportedOptionError for the rest, on both backends (the vendored fallback is Fuse.js's basic build, which has the same restriction, so behavior is identical either way). The full supported/unsupported matrix is on Kernel: Fuse. If you need extended search, that's an upvote we want: open an issue.

Do the kernels use the GPU?

No. Every kernel on main is CPU: AOT-compiled SIMD (float64, reference operation order), plus deterministic thread fan-out where documented (fuse-mojo's pthread shim owns disjoint document ranges and assembles in fixed order). Several kernels are even deliberately single-threaded today (gaussgrid: determinism first; multithreading is documented future work). Nothing in the repo claims GPU support; if a future kernel takes that path it will ship with the same parity gates.

Do I need to install Mojo (or anything from Modular)?

No. pip install <name>-mojo / npm install @<name>-mojo/core gets you a self-contained per-platform binary with the runtime vendored in, or the vendored pure-language fallback on platforms without a build. The Mojo toolchain exists only in the development repo, pinned by pixi.lock, for people building kernels from source.

Why is jmespath-mojo slower in its own benchmark — is something broken?

Nothing is broken; that's what an honest boundary cost looks like. The kernel's raw evaluation is several times faster than the reference interpreter, but crossing the FFI boundary requires marshal.dumps serialization of the document, and on typical queries that tax exceeds the win (0.01–1.1×; best cell 1.10× on heavy filters). It ships because its parity is exact (5,000+ differential cells, exact exception classes) and its fallback is correct everywhere — and because publishing the autopsy is more useful than hiding it. If your workload runs many heavy queries against large documents, the native backend earns its keep; for light lookups, pip jmespath stays ahead, and we say so in the README.

Why should I trust the speedup numbers?

Because they're built to be falsifiable: a correctness gate asserts parity against the real reference package before every timing pass; corpora are seeded and synthetic; every cell is a median of 5; cold and warm are separated; the machine/environment block is published; and the scripts (pixi run bench*, benchmarks/bench_<name>.*) regenerate the tables — hand-edits aren't allowed. Most importantly, the losing cells are published too (jmespath 0.01×, minisearch 0.8×, ta 0.4× cold, bm25s beating bm25-mojo warm). Full methodology and the lowlights: Benchmarks.


More questions: GitHub Issues · How It Works · Apache-2.0, © 2026 Algenta

mojo-kernels — clean-room Mojo kernels as drop-in accelerators

Start

Understand

Contribute

Project

Apache-2.0 · © 2026 Algenta

Clone this wiki locally