Skip to content

erlang_wasm 0.3.0

Choose a tag to compare

@benoitc benoitc released this 18 Sep 22:38
· 223 commits to main since this release

0.3.0

This release is about running other people's code: safely, and fast enough to
be worth doing.

Three parts, meant to be used together.

  • A worker runs one untrusted request at a time, in its own process, with
    its own deadline and its own limits.
  • Snapshots let a language like Python start once, when the worker starts,
    instead of starting again on every request.
  • The compiled tier now works for a worker like that. It did not before.

Together they take a CPython request from about a minute to 35 ms.

Run untrusted code, one request at a time

script_worker is the worker. It knows about modules, imports, deadlines and
output limits. It knows nothing about WASI, or JSON, or what your guest calls
its entry point. That part is an adapter: one module per language.

{ok, _} = worker_reaper:start_link(#{scratch => "/var/tmp/w"}),
{ok, W} = script_worker:start_link(my_adapter, #{root => scratch}),
{ok, R} = script_worker:run(W, Request).

Three languages come with adapters already. js_worker and python_worker
take a function written by whoever is sending the request. Lua is
lua_reactor_adapter.

Every language needs its own limits, and an adapter will never raise one for
you. Python will not even start until you raise several of them. The Python
guide lists them.

Read next: workers to run one,
the adapter contract to write one, and
JavaScript, Python or
Lua for a language.

Smaller things: max_output_bytes now accepts separate bounds for stdout and
stderr. t:wasm:extern/0 names the type extern/2 returns.

Start an interpreter once, not once per request

Starting CPython takes about a minute and a half. Doing that per request is not
an option, and keeping one interpreter alive across requests leaks one caller's
state into the next.

So capture it once, and give every request a fresh copy:

{ok, Image} = wasm:snapshot(Init),
{ok, Fresh} = wasm:restore(Image, FreshImports, #{}).

The copy is genuinely fresh. Globals, memory and tables come from the image,
but the imports are the ones you pass in now, so one request cannot reach
another's files or sockets.

The instance you capture has to be created with snapshotable => true, and a
restore refuses an image that does not match the module it is handed.
wasm:save_snapshot/2 and load_snapshot/2 put an image on disk.
max_snapshot_bytes caps what one node keeps in memory.

Read next: snapshots.

Compiling hot code, and why it helps now

Turn it on with compile => true and fuel => infinity. Those two go
together: leaving a fuel limit in place quietly keeps you on the interpreter.
Point code_cache_dir at a directory you own, and a restart reuses what was
compiled last time. That is minutes of work turned into seconds.

What changed:

  • A fresh instance uses compiled code immediately. It used to wait for a
    function to be called 32 times. A worker that builds a new instance per
    request almost never got there, so 31 requests in 32 ran interpreted next to
    compiled code that was sitting right there.
  • Restoring a snapshot is three times faster. It used to write out the
    module's initial data and then blank it again, even though the image was
    about to overwrite all of it. A CPython request went from 64 ms to 35 ms.
  • A compile can be given a memory cap, and can be interrupted.
    compile_max_heap_words caps a single compile. compile_budget_heap_words
    caps the whole machine: divide it by the cap and that is how many compiles
    run at once. A guest that does not get a slot keeps interpreting and tries
    again later. Both are off unless you turn them on.
  • The compiled-code cache is checked, not trusted. It verifies the
    directory and every file it reads, and quietly recompiles if anything looks
    wrong. It will not read through a symlink or out of a world-writable
    directory.

Read next: the compiled tier.

If requests are slower than you expect, set a heap floor

A restored instance holds almost nothing on the Erlang heap, so the runtime
gives its process a tiny one and then collects garbage hundreds of times during
a single call.

runner_min_heap_words fixes it. The right value depends on the guest:
200,000 for QuickJS and Lua, 1,000,000 for CPython. Going higher than that
makes things worse, not better. capture_min_heap_words does the same for the
snapshot.

Read next: tuning.

Breaking

script_worker used to be the QuickJS worker. It is called qjs_worker now
and behaves exactly as it did. The old name now belongs to the
language-neutral worker described above.