Strata v0.1.27
Answers in Chinese, Japanese or Korean are written faster, RTX 20 cards are supported, and five fixes.
Faster CJK answers (#137). The draft layer (MTP speculative decoding) can only propose tokens from its subset of
the vocabulary, and the shipped subset had 27 of the vocabulary's 55,328 Chinese-character tokens. So an answer in
Chinese got almost no drafts. The subset now includes all Chinese, Japanese and Korean tokens:
- On this PC (RTX 5070, Q2_0, greedy), answers in Chinese are 15-38% faster: a concept explanation 55 -> 76 tokens/s,
a translation 67 -> 81, code with Chinese comments 68 -> 79. An answer in Japanese: 63 -> 72. - On the Coder, the Chinese answers are 3-8% faster.
- The larger draft head costs English answers 1-2% (the same text, the same drafts accepted). It takes about 110 MiB
more VRAM, which came out of the reserve, not the expert cache (the same number of cached experts on 12 GB). - Setup replaces the old subset on existing installs at the next start. A subset you made yourself is kept.
tools/draft_vocab.pybuilds and inspects subsets.
Thanks @demon851113 and @qwased for the measurements.
RTX 20 (Turing) cards run Strata (#87, thanks @hireymage).
- The ready-made engine now includes code for RTX 20 cards (sm_75), and setup accepts them.
- On these cards, the tensor-core prompt kernels are replaced by their portable versions.
- The ready-made image encoder has no RTX 20 code yet: with images on the GPU, setup compiles it (it offers to install
the build tools). - It was tested by the contributor on an RTX 2070. We have no RTX 20 card to test it ourselves, so please report how
it runs.
Fixes:
- AMD (#157): the HIP build no longer needs the CUDA Toolkit's headers. Thanks @Dasug for the report and
@samuelishida for the fix (#160). - Setup on some Ryzen PCs (#159): Windows can report "no AVX2" on CPUs that have it (a Ryzen 9 3950X). Setup now
asks the CPU itself when Windows says no. - Adding a GPU with
--gpus(#128): when a chosen card is of a generation the installed engine has no code for,
setup compiles the engine for all the chosen cards, at setup and at start. Before, the start crashed with "no kernel
image". - Images plus the text
<|image_pad|>(#150): a message that contains that text (an agent reading docs about
the chat format), together with an image, gave an error 400. Only the markers the chat template writes for images
are images now.
Checked before the release:
- Byte-identical to 0.1.26 with the old draft subset on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, the Coder),
including the prompt path's internal state. - With the new subset: the no-repeat and sanity checks on every quant, and needle tests at 8K, 16K and 32K.
- Builds and runs the same on Linux.
- On the RX 7900 XTX (AMD): the HIP test suite and a live server test.
Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.27.
The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.
You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.
Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.