Skip to content

toks 0.3.0

Choose a tag to compare

@Tom-A-Lynch Tom-A-Lynch released this 05 Oct 11:06
· 176 commits to master since this release

toks 0.3.0

toks is Actual Computer's tokenizer: give it the tokenizer.json a model ships with and it returns exactly the ids
Hugging Face tokenizers 0.23.2 returns, from a small C library with hand-written asm kernels for arm64 (NEON) and
x86-64 (AVX2), no runtime and no dependencies, plus a Python package with an hf-style Tokenizer API. This is the
first public release; the release commit is d56a5c1 of the history this repository was cut from (the bundles'
MANIFEST names it).

What's new since 0.2.0

  • Vocabulary lookups. toks_token_to_id(ctx, bytes, len) gives the id a token's bytes decode from (added tokens
    included, as hf's token_to_id; the key is the bytes, not hf's vocabulary spelling: " hello", not "Ġhello"),
    and toks_id_flags(ctx, id) says whether an id is an added token, a special one or one raw byte. Python:
    token_to_id (str in hf's spelling, or bytes), id_to_token, token_bytes, id_flags.
  • Sizing the output. toks_encode_bound(ctx, len) is the most ids toks_encode can return for any text of len
    bytes under any flags, so an output buffer is allocated once and never retried.
  • Streaming decode without a hold limit. toks_stream_hold gives a stream caller memory for a byte-fallback run,
    so no valid run is refused (0.2's 44-byte hold returned TOKS_E_LIMIT on long runs of rare 4-byte characters).
  • The default scratch remembers. Flags 0 now means a 2 MiB piece cache and a 4 MiB segment memo: text a worker
    has seen before (a re-sent conversation) is answered without encoding it again. TOKS_SCRATCH_MEMO_MIB(0) turns
    the memo off, TOKS_SCRATCH_MEMO_MIB(n) resizes it; the output is identical at every size.
  • ABI minor 3 (TOKS_ABI_MINOR); the reserved precompiled-image surface left the public header.

Exact

Every result is checked against hf tokenizers 0.23.2 (Kimi K3 against its own reference), at the release commit:

  • full case sets, nightly run on d56a5c1: x86-64 and arm64, auto and scalar tiers, 21 suites per job, 3,964,440
    compared and 0 failed in each of the four jobs; python wheels (linux x86-64) 113 suites, 1,303,805 compared, 0
    failed; bpe 4,626,610 and normalizer 313,357 compared, 0 failed.
  • the release machines (gb10c: NVIDIA GB10, neon and scalar; tr9970x: AMD Threadripper 9970X, avx2 and scalar;
    m2ultra2: Apple M2 Ultra, neon): every target's parity sample on every tier, 217,388-221,783 cases per target
    (encode in every mode, pieces, decode, stream decode) and Kimi K3's 166,284 texts, 0 diffs; make test on both
    tiers; every test program under ASan + UBSan and test_par under TSan (gb10c); the wheels for CPython 3.10-3.14
    built and tested on each; Windows x86-64 build and tests (aimax395).
  • the receipts: docs/release/0.3.md and docs/release/rc/d56a5c10fe19/.

Fast

Measured at commit 245cc5c, the release candidate before the release commit; the release commit differs from it
in src/core/unigram.c only (an id-buffer bound in the Unigram model, which none of the table's eleven tokenizers
uses: all are BPE) and two comments in include/toks.h. The
table is re-measured at the release commit after the release.

  • one core, fresh text: 13-151x faster than hf tokenizers in every cell of the speed table on the asm tiers (88
    cells per machine on gb10c, tr9970x and m2ultra2, 4 KiB chunks and whole-corpus calls, all exact), and faster than
    tiktoken in every cell tiktoken can run (65 / 65 per machine).
  • against gigatoken, the fastest tokenizer we know of: faster cold in 85 / 83 / 84 of 85 cells (gb10c / tr9970x /
    m2ultra2), in the pass state in 83 / 82 / 84, and on warm replays in 43 / 49 / 50.
  • against tok v1, the asm tokenizer toks replaces: faster in all 264 cell x state medians on gb10c and tr9970x.
  • the receipts: docs/bench/e2e.md (generated from the raw logs in docs/bench/raw/*-245cc5c00541*.log) and
    docs/release/rc/245cc5c00541/.

Known limits

The release report's UNMET table (docs/release/0.3.md) has every item that is not met, with its evidence:

  • the speed table at the release commit (above: measured at 245cc5c);
  • the gigatoken gate (paired runs with intervals per cell) on the new default scratch;
  • the proof package (Eva on the JSON readers and the compiler, every CBMC harness);
  • the fuzzing budget (24 cpu-hours per ISA per entry point; 6.8-7.2 reached);
  • the integration package run against an inference engine;
  • control-token isolation (no reserved control id reachable from plain text), not implemented;
  • the Python binding's per-target case-set parity on the release machines (it runs in CI's linux wheels jobs);
  • deferred past this release: the memo's ids-only record format, and a dense pair-lookup table.

Where toks is not ahead yet, by cell: README.md, "Where toks is not ahead yet".

Pinning (docs/release.md, "How a consumer pins toks")

toks 0.3.0 commit 89adbba2a2b1c014dd549f2545dbf8cfbb23ef81
linux-x86_64 sha256 1d295ada9b0075d82fc1d2b80d3e467bad9796d834960f754ea7944a2024b8f3
linux-arm64  sha256 b3fcebf4a22292d6b671cf96e8ec63aee7225b323dff6302ee9a5983bb7344cc
macos-arm64  sha256 761080bf6acda4168c7859e0235861d20422516d934dc217f067cd9739ca58ae

The bundles' MANIFEST names d56a5c10fe19fd9870642674845655e6bfc14d56, the release commit's id in the private
history; the tag's commit carries the same source (docs/release/public.md). Integration flags: one scratch per worker
thread, kept across requests, initialized with flags 0; encode calls with flags 0. Every file's sha256: SHA256SUMS.

From the release commit to the tag only docs and the tools that write them changed (src/, include/, python/
and the Makefile are the release commit's): tools/bench/e2e_table.py (the speed table's load column shows every
run's loads, so a re-measured cell shows both runs), tools/release/rc_report.py (--speed-note: the report says
where its speed table was measured) and tools/release/rc.sh (the default path of the assembler tok v1's cells
need on the x86-64 bench machine).

The wheels are attached here and not published to PyPI.