Skip to content

Strata v0.1.31

Choose a tag to compare

@Niko1221 Niko1221 released this 01 Oct 05:13
· 224 commits to main since this release

Unsloth's Q4 model runs on a 64 GB PC (experimental), AMD decode up to 15% faster, models load twice as fast on Windows, and a batch of fixes.

Experimental: Unsloth UD-Q4_K_XL on a normal PC. Strata can now run Unsloth's 4-bit Qwen3.8-Flash-Next
(UD-Q4_K_XL, 111 GB in 4 files, 72 GiB of experts) on a PC whose RAM is smaller than the model:

  • The model's four files are read as they are: a missing file is an error, a layer split across files is fine.
  • Its expert formats (Q4_K, Q5_K, Q5_1, Q8_0) run natively on the GPU and the CPU; the GPU dequantizers match
    llama.cpp's bit for bit, the dot products match ggml's CPU results within float rounding.
  • Three tiers for the experts: the most used in VRAM, the next ones in RAM (--resident-budget-gib N), the rare
    ones read straight from the model files on the SSD, prefetched one layer ahead from the router's prediction.

Measured on an RTX 5070 12 GB with 64 GB RAM, a 40 GiB RAM budget: about 7-8.5 tokens/s for answers, with
correct output. A big-RAM PC holds more in RAM and is faster. Not in setup yet: see docs/UNSLOTH_Q4.md for the
manual import. Ported from eddoursul's fork (#245), with pieces from gopinath87607 (#255) and jagsan-cyber (#247).

Low-RAM mode without the 30-70 GB copy. The low-RAM mode (--mmap-experts) now reads a native pack's experts
straight from the model's GGUF files when the pack has no experts.bin - the same tokens and logits as with the
copy.

Faster:

  • Windows: models without experts.bin load about twice as fast (the expert reader used 4 KB reads; #230,
    reported by activeing123). IQ3_XXS starts in ~30 s instead of ~47 s here.
  • AMD RDNA4: decode +15% on the Radeon AI PRO R9700 and +5-7% on the RX 9070 XT (byte-level intrinsics from
    #262, ttio2tech); the live test on the R9700 went from 67 to 82 tokens/s.
  • The Q2_0 expert kernels (#241, gputier) and the IQ expert kernels (#242, gputier) do less work per token -
    bitwise identical output.
  • RTX 20 (Turing): long prompts ~12% faster - the prompt attention runs on tensor cores (#270, kenh0u; output on
    RTX 20 cards changes slightly, as for RTX 30+ before), and the hyper-connection kernel stages smaller tiles (#258,
    hireymage, bitwise identical).

Fixes:

  • Server: a request that ended late could clear the next request's status and drop its answer (#266, tonykee).
  • Server: a long unbroken run of Chinese/Japanese/Korean text (or a URL) no longer takes seconds to tokenize
    (#268, cha0yang: 1.6 s -> 5 ms for 2,000 characters).
  • Tool calls cut off by the end of the answer are reported as unfinished instead of complete, and their broken
    arguments are not passed on (#211, #231 by alphastorm).
  • Windows: when a verify window stalls, the watchdog now releases the GPU's waits before it ends the engine, so the
    GPU is not left "lost" until a reboot (#267, 1593914054).
  • Linux multi-GPU: the whole expert arena is pinned again (the 8 GiB cap was meant for Windows only; it cost
    prompt speed 3x on a 4090 + 3060; #253, bettercallcaleb). STRATA_ARENA_PIN_GIB=N sets a cap.
  • An MCP server that fails to start always says why.
  • A stall report names the prompt chunk and layer it stopped at (#251).
  • Setup: says why the low-RAM mode uses one GPU, and --low-ram off --yes installs without asking (#250); a
    checkout installs the engine of its own version, pinned model revisions and pinned Python packages (#214,
    alphastorm).
  • A CPU without AVX2 is refused at start with a clear message instead of crashing.
  • The cuBLAS setup failures are named instead of reported as one line (#240, lukmanfauzie).

New, all off unless you turn them on:

  • reasoning_budget_tokens in a request (or the config): caps the thinking, then lets the model answer (#123).
  • STRATA_GR_V3=1: the hyper-connection read in two kernels, +2-4% decode here, another summation order (#186,
    q8atnight).
  • STRATA_ARENA_PIN_GIB=auto (Windows): keeps the sliced pin below the GPU's shared-memory budget, for a PC where
    later allocations fail after pinning (#243, icegita: please try it).
  • "split_skip_if_fits": true (multi-GPU config): no split when the first card holds every expert.
  • STRATA_LOGPOS: per-position log-probabilities for quality tests (#235, enkynakamura).

AMD: setup installs a layer split over several AMD cards (--gpus 1,0); the RX 7800 XT (gfx1101, #254 by
jhohertz) and RX 9060 XT (gfx1200, #256 by Efeisot) are community-validated.

Also: a "Where things are stored" section in the README (#226); a community benchmark guide with RTX 5090
results (#233, hagope) and RTX 3090 / dual 3090 results (#234, mad9home); iq_parity runs from a fresh checkout
(#264, j-luwierski); an experimental build flag admits Volta (#236, cardonja; not in the ready-made engine).

Checked before the release:

  • Every merged change was gated on its own: byte-identical to 0.1.30 with a fixed cache on all four quants on this PC
    (Q2_0, IQ3_XXS, IQ3_S, the Coder), including the prompt path's internal state, in nine rounds; default-settings
    speed A/Bs on Q2_0 and IQ3_S (equal or faster). Two changes were reworked because they cost speed here: the
    Windows pin cap (#243, now opt-in) and a polling wait in the verify window (#267, now the watchdog's job); #257
    was taken out again (not faster on an AVX2-only CPU).

  • Real use at a 64K context on all five models (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, the Coder), through the server:

    • a fact found in a 57,479-token document;
    • the follow-up turn reusing the whole prompt;
    • a tool call;
    • a cancelled long prompt followed by a new request;
    • a sampled answer.

    None of them restarted the engine or ran low on VRAM.

  • The Unsloth Q4 model end to end on this PC (correct answers at 24, 32 and 40 GiB RAM budgets); the new format
    kernels against ggml-cpu and llama.cpp's dequantizers; the split-file loader and pack tests; the low-RAM mode
    reading the GGUF in place gives the same tokens and logits as experts.bin.

  • The server's tests (100) and setup's tests.

  • Linux (WSL, RTX 5070): builds, and Q2_0 is byte-identical to 0.1.30 with a fixed cache (10/10); the default settings pass.

  • AMD (RX 9070 XT and R9700): the HIP build with its tests (all pass but the two that need a model fixture or an
    AVX-512 CPU), the live server test on each card (R9700 82 tokens/s, 9070 XT 34 tokens/s), and the two-card split.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.31.

The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.