Skip to content

toks 0.3.1

Choose a tag to compare

@Tom-A-Lynch Tom-A-Lynch released this 06 Oct 19:51
· 131 commits to master since this release
ac5940e

toks 0.3.1

toks is Actual Computer's tokenizer: give it the tokenizer.json a model ships with and it returns exactly the ids
Hugging Face tokenizers 0.23.2 returns, from a small C library with hand-written asm kernels for arm64 (NEON) and
x86-64 (AVX2), no runtime and no dependencies, plus a Python package with an hf-style Tokenizer API. 0.3.1 is the
first release cut from this repository's own history; the release commit is ac5940ee712957c7279fd2365dd9f3f260af002c.

What's new since 0.3.0

  • The piece dictionary on the BPE miss path. K6 (the BPE of one piece) seats a dictionary of 131,072 frequent
    English pieces in the words table's free ways at load, so a first-sight piece that is in it is answered without
    running the merges: English prose and code cold +3-10% on the release machines, multilingual and CJK 0..+3%, no
    cell slower; load time +48 ms on the largest vocabulary (docs/kernels.md §6, raw logs under docs/bench/raw/).
  • The test method of the stall bar rests on millisecond baselines (best of five batches per size), so a shared
    machine cannot move its verdict.
  • Documentation: the README's release note and voice, the contribution policy (CONTRIBUTING.md: issues, bug
    reports, receipts, ideas and pull requests are welcome; every pull request that merges is run by an Actual
    Computer engineer, who takes contributors' changes in with authorship kept), the losses table beside the speed
    table (docs/bench/losses.md), the K6 grid experiment (measured, not taken; docs/kernels.md §5.3), the SMT
    receipts for toks_par on the 9970X (docs/bench/par.md), the fuzz campaign's fourth finding (docs/fuzz.md).
  • Tooling: the release candidate's wheels step reads its own case sets; bench and gate tooling under tools/bench.

The C ABI is 0.3.0's (TOKS_ABI_MAJOR 0, TOKS_ABI_MINOR 3); the wheel's version follows the header.

Exact

Every result is checked against hf tokenizers 0.23.2 (Kimi K3 against its own reference):

  • the release commit's CI: make test on both tiers on linux x86-64 and arm64, macOS arm64 and Windows x86-64
    (the Actions runs on ac5940ee712957c7279fd2365dd9f3f260af002c);
  • the nightly full-parity run dispatched on the release commit: run 37519924259: x86_64-scalar: 21 suites, 3964440 compared, 0 failed; arm64-auto: 21 suites, 3964440 compared, 0 failed; x86_64-auto: 21 suites, 3964440 compared, 0 failed; arm64-scalar: 21 suites, 3964440 compared, 0 failed (4 jobs x 21 suites, 3,964,440 cases each, 0 failed).
  • the 0.3.0 release report (docs/release/0.3.md) for the release machines' parity samples, sanitizer runs and
    wheels at 0.3.0; 0.3.1's source differs from 0.3.0's in the dictionary (K6's miss path, src/core/bpe_build.c,
    src/gen/dict.c) and the Unigram bound; the dictionary's every entry is certified against the BPE reference at
    load (check_tables).

Fast

The speed table (docs/bench/e2e.md) stands at commit 245cc5c, the 0.3.0 release candidate: 13-151x faster than
hf tokenizers in every cell on the asm tiers, 2-24x faster than tiktoken in every cell it can run, ahead of
gigatoken cold in 252 of 255 cells. 0.3.1's dictionary moves the English prose and code cold cells up by 3-10%
(docs/kernels.md §6); the whole table is re-measured at the next cut.

Where the artifacts were built

  • linux arm64: the GB10 release machine (gb10c, docs/machines.md), tools/release.sh at the release commit:
    make test on both tiers first, then the bundle and the five wheels (their tests on each CPython).
  • macOS arm64: the M2 Ultra release machine (m2ultra2), the same way.
  • linux x86-64: the x86-64 release machine was unreachable at the cut, so this set comes from a clean CI runner with
    the same pinned toolchain (clang 21.1.8, glibc 2.34; .github/workflows/release-artifacts.yml, run
    37521052367), tools/release.sh the same way, from master at 361883a3945e1bb6fb159d7f1a4468fc08bc3ac0:
    the release commit plus that workflow file, the library source identical. Its MANIFEST names that commit.

Known limits

The 0.3.0 report's UNMET table (docs/release/0.3.md) still applies: the fuzzing budget (24 cpu-hours per ISA per
entry point; 7 reached and counting), the proof package, the gigatoken gate on the new default scratch, the
integration package, control-token isolation, the Python binding's per-target parity on the release machines, and
the speed table at the release commit.

Pinning (docs/release.md, "How a consumer pins toks")

toks 0.3.1 commit ac5940ee712957c7279fd2365dd9f3f260af002c
linux-x86_64 sha256 4e5a5d468aa89eaa987ae73637c981a65fc7e829fce7658bd9fbe656a502f90a
linux-arm64  sha256 62179e2b3eac17a05afdef7aec8eb0e85504f05c949f62d0eff49dd0afb40843
macos-arm64  sha256 97bfa1db1272b80cbcd893a8ea026a4113b0a066b64c2d6a101152772ca09357

Integration flags: one scratch per worker thread, kept across requests, initialized with flags 0; encode calls with
flags 0. Every file's sha256: SHA256SUMS. The wheels are attached here and not published to PyPI.