Repository navigation
Releases: actual-computer/toks
Release list
toks 0.3.2
toks 0.3.2
toks is Actual Computer's tokenizer: give it the tokenizer.json a model ships with and it returns exactly the ids
Hugging Face tokenizers 0.23.2 returns, from a small C library with hand-written asm kernels for arm64 (NEON) and
x86-64 (AVX2), no runtime and no dependencies, plus a Python package with an hf-style Tokenizer API. 0.3.2 is the
first release under the Apache License, Version 2.0; the release commit is 55a5230b08a75916a2b36f92c320f057376833ea.
License
From this release on, toks is open source under the Apache License, Version 2.0
(SPDX Apache-2.0): use it, modify it, vendor it and ship it inside anything, closed products included, at any
scale, keeping the LICENSE and NOTICE files with your copies. 0.3.0 and 0.3.1 shipped under the Business Source
License 1.1 and their tags keep that text. LICENSING.md is the plain-language page. The wheels declare
License-Expression: Apache-2.0 and carry LICENSE, NOTICE, LICENSING.md and THIRD_PARTY_NOTICES.md under
licenses/; the bundles carry the same four files, each named in their MANIFEST with its sha256.
What's new since 0.3.1
- ABI 0.4: the hf and tiktoken primitives (#13).
toks_template(the post-processor's single-sequence
template),toks_added(the added tokens in id order, with their lstrip / rstrip / single_word / normalized
flags),TOKS_NO_TRUNCATE/TOKS_NO_PAD(encode without the file's truncation and padding) and
TOKS_DECODE_RAW(the bytes the ids spell, never repaired);toks_inforeports the file's truncation, padding
and the template's shape, and an abi 0.3 caller'stoks_infosize is still answered.TOKS_ABI_MINOR3 -> 4;
TOKS_ABI_MAJORstays 0. - Python: hf's
Encodingand tiktoken'sEncodingontoks.Tokenizer, over abi 0.4 (#34), and
Tokenizer.encode_bound(n), the C library'stoks_encode_bound(#20). - Exactness: a SentencePiece pending unk stays pending across byte-fallback chars, as hf's
merge_wordkeeps
it (#36);config.crefuses what hf refuses before any model reads the file (#22); the last id, 2^21 - 2, loads
(#24). - Memory geometry:
make test-guardputs every table and every scratch region on its own pages, and a fault on
a guard page says where (#14, #38, #39, #40, #47); no writable data in the C core (#24);toks_plat_arena's first
touch is a write, sotoks_par's scratches keep their first 2 MiB frame huge on a kernel that splits the huge
zero page (#15, #29); huge pages for a default scratch, measured (#46). - The stall screen (#35):
test_stallnames suspects and fails growth at x8 then x12 from 128 KiB to 2 MiB;
test_par's parallel floor (#7). - Measured and not taken, with their raw logs in the tree: the memo's keyed check and bytes-then-check records
(#28, #41), the low-id pair grid on avx2 (#6), K6's round one in two passes (#23). - Documentation and tooling: the gigatoken gate's picture on the default scratch, its m6 paragraph generated
from the probe's log (#45, #49); receipts that name their commits (#30); the README's tiktoken range by the hf
clause's rule (#42, #43); two more x86-64 machines indocs/machines.md(#27); the external pull request policy
decided by repository permission (#33); the release-artifacts workflow (#26); the API audit's severity-3 tests
(#21).
Exact
Every result is checked against hf tokenizers 0.23.2 (Kimi K3 against its own reference):
- the release commit's CI:
make teston both tiers on linux x86-64 and arm64, macOS arm64 and Windows x86-64
(the Actions runs on55a5230b08a75916a2b36f92c320f057376833ea: test.yml run 37710610274 for linux x86-64
and arm64 on both tiers, windows.yml run 37710610279, macos.yml run 37710610281); - the nightly full-parity run dispatched on the release commit: run 37710682088: x86_64-scalar: 21 suites, 3964440
compared, 0 failed; arm64-auto: 21 suites, 3964440 compared, 0 failed; x86_64-auto: 21 suites, 3964440 compared, 0
failed; arm64-scalar: 21 suites, 3964440 compared, 0 failed
(4 jobs x 21 suites, 3,964,440 cases each, 0 failed); its norm job: 313,357 compared, 0 failed; its bpe job:
4 suites, 4,626,610 compared, 0 failed. - the 0.3.0 release report (
docs/release/0.3.md) for the release machines' parity samples, sanitizer runs and
wheels at 0.3.0. The C core and header of 0.3.2 differ from 0.3.1's in 24 files undersrc/andinclude/
(732 insertions, 173 deletions): the ABI 0.4 primitives (#13),spm_c.c's pending unk (#36),config.c's
refusals (#22), no writable data and the last id's load (#24), the guard build's table and scratch pointers (#14)
and the arena's first touch (#15); the asm kernels are unchanged.
Fast
The speed table (docs/bench/e2e.md) stands at commit 245cc5c, the 0.3.0 release candidate: 13-151x faster than
hf tokenizers in every cell on the asm tiers, 4-23x faster than tiktoken in every cell it can run, ahead of
gigatoken cold in 252 of 255 cells. 0.3.2 is a license and patch cut and does not re-measure it; the table is
re-measured at a later cut.
Where the artifacts were built
Each set is tools/release.sh at the release commit on the machine named (docs/machines.md), from a tree synced by
tools/remote.sh: make test on both tiers first (every suite 0 failures), then the bundle, then the five wheels
with their tests on each CPython (cp313 runs the whole Python suite).
- linux arm64: the GB10 release machine (
gb10c), load 0.07 at the start; cp313: 324 passed, 46 skipped. - macOS arm64: the M2 Ultra release machine (
m2ultra2), load 1.2; cp313: 313 passed, 46 skipped. - linux x86-64: the 9970X release machine (
tr9970x), reachable again for this cut, load 2.7, the build pinned to
taskset -c 0-7(untimed); cp313: 314 passed, 46 skipped. The bundle'slibtoks.sofloor is GLIBC_2.34 (its
MANIFEST'sglibcline); the wheels are manylinux2014.
Known limits
The 0.3.0 report's UNMET table (docs/release/0.3.md) still applies: the fuzzing budget (24 cpu-hours per ISA per
entry point), the proof package, the gigatoken gate on the new default scratch, the integration package,
control-token isolation, the Python binding's per-target parity on the release machines, and the speed table at
the release commit.
Pinning (docs/release.md, "How a consumer pins toks")
toks 0.3.2 commit 55a5230b08a75916a2b36f92c320f057376833ea
linux-x86_64 sha256 e0780282bd8caabcee355e12404d2d3ead3ed5a52442fb73eb88ee2ba4b23b76
linux-arm64 sha256 d273392667d6d80f66cf61a61e3129a7dfca16176b0be990752a0d08e4fd2fb8
macos-arm64 sha256 410f4981ba76a1b473640d0fcd9bb63b457b841abf4c0f2027826badc61fe82b
Integration flags: one scratch per worker thread, kept across requests, initialized with flags 0; encode calls with
flags 0. Every file's sha256: SHA256SUMS. The wheels are attached here and not published to PyPI.
toks 0.3.1
toks 0.3.1
toks is Actual Computer's tokenizer: give it the tokenizer.json a model ships with and it returns exactly the ids
Hugging Face tokenizers 0.23.2 returns, from a small C library with hand-written asm kernels for arm64 (NEON) and
x86-64 (AVX2), no runtime and no dependencies, plus a Python package with an hf-style Tokenizer API. 0.3.1 is the
first release cut from this repository's own history; the release commit is ac5940ee712957c7279fd2365dd9f3f260af002c.
What's new since 0.3.0
- The piece dictionary on the BPE miss path. K6 (the BPE of one piece) seats a dictionary of 131,072 frequent
English pieces in the words table's free ways at load, so a first-sight piece that is in it is answered without
running the merges: English prose and code cold +3-10% on the release machines, multilingual and CJK 0..+3%, no
cell slower; load time +48 ms on the largest vocabulary (docs/kernels.md§6, raw logs underdocs/bench/raw/). - The test method of the stall bar rests on millisecond baselines (best of five batches per size), so a shared
machine cannot move its verdict. - Documentation: the README's release note and voice, the contribution policy (CONTRIBUTING.md: issues, bug
reports, receipts, ideas and pull requests are welcome; every pull request that merges is run by an Actual
Computer engineer, who takes contributors' changes in with authorship kept), the losses table beside the speed
table (docs/bench/losses.md), the K6 grid experiment (measured, not taken;docs/kernels.md§5.3), the SMT
receipts fortoks_paron the 9970X (docs/bench/par.md), the fuzz campaign's fourth finding (docs/fuzz.md). - Tooling: the release candidate's wheels step reads its own case sets; bench and gate tooling under
tools/bench.
The C ABI is 0.3.0's (TOKS_ABI_MAJOR 0, TOKS_ABI_MINOR 3); the wheel's version follows the header.
Exact
Every result is checked against hf tokenizers 0.23.2 (Kimi K3 against its own reference):
- the release commit's CI:
make teston both tiers on linux x86-64 and arm64, macOS arm64 and Windows x86-64
(the Actions runs onac5940ee712957c7279fd2365dd9f3f260af002c); - the nightly full-parity run dispatched on the release commit: run 37519924259: x86_64-scalar: 21 suites, 3964440 compared, 0 failed; arm64-auto: 21 suites, 3964440 compared, 0 failed; x86_64-auto: 21 suites, 3964440 compared, 0 failed; arm64-scalar: 21 suites, 3964440 compared, 0 failed (4 jobs x 21 suites, 3,964,440 cases each, 0 failed).
- the 0.3.0 release report (
docs/release/0.3.md) for the release machines' parity samples, sanitizer runs and
wheels at 0.3.0; 0.3.1's source differs from 0.3.0's in the dictionary (K6's miss path,src/core/bpe_build.c,
src/gen/dict.c) and the Unigram bound; the dictionary's every entry is certified against the BPE reference at
load (check_tables).
Fast
The speed table (docs/bench/e2e.md) stands at commit 245cc5c, the 0.3.0 release candidate: 13-151x faster than
hf tokenizers in every cell on the asm tiers, 2-24x faster than tiktoken in every cell it can run, ahead of
gigatoken cold in 252 of 255 cells. 0.3.1's dictionary moves the English prose and code cold cells up by 3-10%
(docs/kernels.md §6); the whole table is re-measured at the next cut.
Where the artifacts were built
- linux arm64: the GB10 release machine (
gb10c, docs/machines.md),tools/release.shat the release commit:
make teston both tiers first, then the bundle and the five wheels (their tests on each CPython). - macOS arm64: the M2 Ultra release machine (
m2ultra2), the same way. - linux x86-64: the x86-64 release machine was unreachable at the cut, so this set comes from a clean CI runner with
the same pinned toolchain (clang 21.1.8, glibc 2.34;.github/workflows/release-artifacts.yml, run
37521052367),tools/release.shthe same way, from master at361883a3945e1bb6fb159d7f1a4468fc08bc3ac0:
the release commit plus that workflow file, the library source identical. ItsMANIFESTnames that commit.
Known limits
The 0.3.0 report's UNMET table (docs/release/0.3.md) still applies: the fuzzing budget (24 cpu-hours per ISA per
entry point; 7 reached and counting), the proof package, the gigatoken gate on the new default scratch, the
integration package, control-token isolation, the Python binding's per-target parity on the release machines, and
the speed table at the release commit.
Pinning (docs/release.md, "How a consumer pins toks")
toks 0.3.1 commit ac5940ee712957c7279fd2365dd9f3f260af002c
linux-x86_64 sha256 4e5a5d468aa89eaa987ae73637c981a65fc7e829fce7658bd9fbe656a502f90a
linux-arm64 sha256 62179e2b3eac17a05afdef7aec8eb0e85504f05c949f62d0eff49dd0afb40843
macos-arm64 sha256 97bfa1db1272b80cbcd893a8ea026a4113b0a066b64c2d6a101152772ca09357
Integration flags: one scratch per worker thread, kept across requests, initialized with flags 0; encode calls with
flags 0. Every file's sha256: SHA256SUMS. The wheels are attached here and not published to PyPI.
toks 0.3.0
toks 0.3.0
toks is Actual Computer's tokenizer: give it the tokenizer.json a model ships with and it returns exactly the ids
Hugging Face tokenizers 0.23.2 returns, from a small C library with hand-written asm kernels for arm64 (NEON) and
x86-64 (AVX2), no runtime and no dependencies, plus a Python package with an hf-style Tokenizer API. This is the
first public release; the release commit is d56a5c1 of the history this repository was cut from (the bundles'
MANIFEST names it).
What's new since 0.2.0
- Vocabulary lookups.
toks_token_to_id(ctx, bytes, len)gives the id a token's bytes decode from (added tokens
included, as hf'stoken_to_id; the key is the bytes, not hf's vocabulary spelling:" hello", not"Ġhello"),
andtoks_id_flags(ctx, id)says whether an id is an added token, a special one or one raw byte. Python:
token_to_id(str in hf's spelling, or bytes),id_to_token,token_bytes,id_flags. - Sizing the output.
toks_encode_bound(ctx, len)is the most idstoks_encodecan return for any text oflen
bytes under any flags, so an output buffer is allocated once and never retried. - Streaming decode without a hold limit.
toks_stream_holdgives a stream caller memory for a byte-fallback run,
so no valid run is refused (0.2's 44-byte hold returnedTOKS_E_LIMITon long runs of rare 4-byte characters). - The default scratch remembers. Flags 0 now means a 2 MiB piece cache and a 4 MiB segment memo: text a worker
has seen before (a re-sent conversation) is answered without encoding it again.TOKS_SCRATCH_MEMO_MIB(0)turns
the memo off,TOKS_SCRATCH_MEMO_MIB(n)resizes it; the output is identical at every size. - ABI minor 3 (
TOKS_ABI_MINOR); the reserved precompiled-image surface left the public header.
Exact
Every result is checked against hf tokenizers 0.23.2 (Kimi K3 against its own reference), at the release commit:
- full case sets, nightly run on
d56a5c1: x86-64 and arm64, auto and scalar tiers, 21 suites per job, 3,964,440
compared and 0 failed in each of the four jobs; python wheels (linux x86-64) 113 suites, 1,303,805 compared, 0
failed; bpe 4,626,610 and normalizer 313,357 compared, 0 failed. - the release machines (gb10c: NVIDIA GB10, neon and scalar; tr9970x: AMD Threadripper 9970X, avx2 and scalar;
m2ultra2: Apple M2 Ultra, neon): every target's parity sample on every tier, 217,388-221,783 cases per target
(encode in every mode, pieces, decode, stream decode) and Kimi K3's 166,284 texts, 0 diffs;make teston both
tiers; every test program under ASan + UBSan andtest_parunder TSan (gb10c); the wheels for CPython 3.10-3.14
built and tested on each; Windows x86-64 build and tests (aimax395). - the receipts:
docs/release/0.3.mdanddocs/release/rc/d56a5c10fe19/.
Fast
Measured at commit 245cc5c, the release candidate before the release commit; the release commit differs from it
in src/core/unigram.c only (an id-buffer bound in the Unigram model, which none of the table's eleven tokenizers
uses: all are BPE) and two comments in include/toks.h. The
table is re-measured at the release commit after the release.
- one core, fresh text: 13-151x faster than hf tokenizers in every cell of the speed table on the asm tiers (88
cells per machine on gb10c, tr9970x and m2ultra2, 4 KiB chunks and whole-corpus calls, all exact), and faster than
tiktoken in every cell tiktoken can run (65 / 65 per machine). - against gigatoken, the fastest tokenizer we know of: faster cold in 85 / 83 / 84 of 85 cells (gb10c / tr9970x /
m2ultra2), in the pass state in 83 / 82 / 84, and on warm replays in 43 / 49 / 50. - against tok v1, the asm tokenizer toks replaces: faster in all 264 cell x state medians on gb10c and tr9970x.
- the receipts:
docs/bench/e2e.md(generated from the raw logs indocs/bench/raw/*-245cc5c00541*.log) and
docs/release/rc/245cc5c00541/.
Known limits
The release report's UNMET table (docs/release/0.3.md) has every item that is not met, with its evidence:
- the speed table at the release commit (above: measured at
245cc5c); - the gigatoken gate (paired runs with intervals per cell) on the new default scratch;
- the proof package (Eva on the JSON readers and the compiler, every CBMC harness);
- the fuzzing budget (24 cpu-hours per ISA per entry point; 6.8-7.2 reached);
- the integration package run against an inference engine;
- control-token isolation (no reserved control id reachable from plain text), not implemented;
- the Python binding's per-target case-set parity on the release machines (it runs in CI's linux wheels jobs);
- deferred past this release: the memo's ids-only record format, and a dense pair-lookup table.
Where toks is not ahead yet, by cell: README.md, "Where toks is not ahead yet".
Pinning (docs/release.md, "How a consumer pins toks")
toks 0.3.0 commit 89adbba2a2b1c014dd549f2545dbf8cfbb23ef81
linux-x86_64 sha256 1d295ada9b0075d82fc1d2b80d3e467bad9796d834960f754ea7944a2024b8f3
linux-arm64 sha256 b3fcebf4a22292d6b671cf96e8ec63aee7225b323dff6302ee9a5983bb7344cc
macos-arm64 sha256 761080bf6acda4168c7859e0235861d20422516d934dc217f067cd9739ca58ae
The bundles' MANIFEST names d56a5c10fe19fd9870642674845655e6bfc14d56, the release commit's id in the private
history; the tag's commit carries the same source (docs/release/public.md). Integration flags: one scratch per worker
thread, kept across requests, initialized with flags 0; encode calls with flags 0. Every file's sha256: SHA256SUMS.
From the release commit to the tag only docs and the tools that write them changed (src/, include/, python/
and the Makefile are the release commit's): tools/bench/e2e_table.py (the speed table's load column shows every
run's loads, so a re-measured cell shows both runs), tools/release/rc_report.py (--speed-note: the report says
where its speed table was measured) and tools/release/rc.sh (the default path of the assembler tok v1's cells
need on the x86-64 bench machine).
The wheels are attached here and not published to PyPI.