Skip to content

Releases: Niko1221/Strata

Strata v0.1.39

Choose a tag to compare

@Niko1221 Niko1221 released this 04 Oct 12:32

Faster decode and long prompts, several requests at once, older and other hardware as experimental opt-ins, the
OpenAI Responses API for Codex, a fix for slower prompts with a RAM budget, and a batch of fixes from your reports.

Faster (the same answers where nothing says otherwise):

  • Decode (#646): the verify pass runs with fewer launches and host round trips (sub-warp expert packing, staged
    inputs, the PLE and MTP steps batched). Measured here on one RTX 5070 against 0.1.38, 10 interleaved pairs, medians:
    Q2_0 +6% (story 68.4 -> 72.7 tok/s, code 75.8 -> 80.0), IQ3_XXS +6% / +2.5% (46.1 -> 48.8, 50.7 -> 51.9); after
    4K and 32K prompts +5% to +7%. The output is byte-identical to 0.1.38's (10/10 on all four quants with a fixed
    cache). The zero-doorbell verify graph itself runs only when every expert of a layer is in VRAM. A 12 GB card never
    gets there, so the gain on big cards was not measured here.
  • Long prompts (#583, on by default): the streamed expert ring is sized in bytes for the pack, and --prefill auto
    picks the largest chunk that keeps it full. Short prompts, and prompts that fit 0.1.38's chunk, keep 0.1.38's ring
    and their exact bits (4K: identical). The gain depends on the free VRAM. On an RTX 5070 with 32K prompts, IQ3_XXS is
    +18.5% with a 1,500-slot expert cache and unchanged with setup's default config. A long prompt's bits change against
    0.1.38, because the experts go through a different mix of cached and streamed groups. Its quality stays in the same
    band, teacher-forced over the next 2,001 tokens against the FP16 prompt path (IQ3_XXS): 8K prompt KL 0.042 (0.1.38:
    0.054), top-1 93.5% (92.2%); 32K prompt KL 0.020 (0.019), top-1 95.3% (95.0%). It doesn't help every pack: the
    Coder at a 32K context reads a 30K prompt about 8% slower (with a 64K context it was about 6% faster).
    STRATA_RING_BYTES=0 restores 0.1.38's ring.
  • Very long contexts: the attention's top-k past the register kernel's reach (262K-524K cells) runs with a
    histogram per warp: a 243K-token prompt +26% on an RTX 3060 (#603).
  • More than one GPU (measured by their authors, not here: we have one GPU): setup now adds
    --remote-expert-opt (#578) to a config with two or more GPUs. It skips the host's work for tokens whose experts
    all run on the helper cards (dual RTX 4090: +63% mixed text, +132% code over the plain helper path), and setup --no-remote-expert-opt leaves it out. A layer split of 3+ GPUs overlaps its stages better (#598, 4x RTX 3060: +10
    to 28% on prompts). STRATA_PREFILL_HELP=1 lets a split's idle card stream a share of a one-chunk prompt's experts
    (#663). It is opt-in because it rounds differently. With one or two GPUs the pipeline runs as before.
  • Also: the Linux expert arena on transparent huge pages (#650; STRATA_NO_LARGEPAGES=1 keeps 4 KB pages), an
    opt-in AVX2 codebook gather for the IQ CPU kernels (STRATA_IQ256_GATHER=1, #622), and thread affinity on PCs with
    more than 64 CPUs (#626).

Several requests at once (#465, opt-in): "parallel": N in strata-<model>.json (or setup --parallel N)
decodes up to N conversations together in the engine. More requests wait for a free slot, and a long prompt gives
way to a short one at a chunk boundary. A request left alone goes back to the one-at-a-time path, and /metrics
shows each slot. Every slot's greedy answer is the same as when it runs alone. On a 12 GB card it cuts waiting time
and costs speed. RTX 5070, Q2_0, four requests at once: the last one starts after 1.8 s instead of 11.2 s, but
together they decode 11% slower (63.1 against 70.7 tok/s). A request alone loses 11% with 2 slots and 22% with 4,
because each slot's session takes 0.56 GiB from the expert cache. So setup recommends it only where the experts
mostly fit in VRAM. A 4-GPU layer split (4x 16 GB, IQ3_S, PR #559) served 8 requests at 360 tok/s in total against
120 for one. docs/BATCHING.md has the numbers and the options.

Older and other hardware, experimental: three opt-in paths that community members wrote and measured on their
own machines. We have none of this hardware. Each is compile-checked and unit-tested here, and the ready-made
engines and their output are unchanged (checked byte-identical).

  • Older NVIDIA cards (Pascal, Volta: P40, P100, GTX 10, V100, Titan V; #395 #600 #540 #655 #627): CUDA 13 cannot
    compile for them, so setup keeps a second engine built with CUDA 12.9 (strata-windows-x64-cuda12.zip on
    Windows, compiled on Linux). It is used only when you choose such a card: a PC with only Pascal / Volta cards, a
    card named with --gpu N / --gpus, or --cuda 12. --cuda 12 is also the way to run with an NVIDIA driver
    older than 580 (528+ on Windows, 525+ on Linux). The choice is kept per model. The Volta prompt attention and the
    BF16 path through FP16 / fp32 are compiled into this engine only. Untested here: every Pascal and Volta card, the
    old drivers, and an RTX 50 card in that engine (it gets a warning; keep it on its own model). docs/OLDER_GPUS.md
    has the reporters' numbers (V100: prompts 1,123-1,251 tok/s, UD-IQ4_XS). RTX 20 owners can try the same FP16
    tensor-core path for the prompt's BF16 products with STRATA_BF16_TC=1 (+15-18% on an RTX 2080 Ti, #655). Its
    sums are not bitwise cuBLAS's.
    Older AMD cards are built by hand: gfx906 (Instinct MI50 / MI60, Radeon VII; -DSTRATA_HIP_GFX906=ON, #638 #677:
    2x MI50, the Coder, 50 tok/s at 4K) and gfx1012 (RX 5500 XT, #442). The RX 6700 XT (gfx1031, #524) goes through
    setup.
  • Intel Arc (#423): maxfridbe's SYCL port of the engine, built from source on Linux with Intel oneAPI:
    ./setup.sh --backend sycl. Reported on the Arc Pro B70 / B50 and the B580 (Coder IQ1_M 70-78 tok/s on a B70) on
    earlier versions. Here it compiles (oneAPI 2026.1) and its kernel tests run on a CPU device. Untested: the 0.1.39
    port on an Arc, Windows (no build path yet; setup points at Linux), WSL2, the A-series and integrated Arc GPUs, and
    images. There is no ready-made Intel engine. docs/INTEL_ARC.md.
  • Older CPUs without AVX2 (#394 #595): AVX-only (Sandy / Ivy Bridge, Xeon E5 v1/v2, Bulldozer) and
    SSE4.2-only CPUs (Nehalem / Westmere). Run setup as usual: it warns and compiles the engine for that CPU
    (STRATA_ISA_FLOOR=avx or none, 10-20 minutes once) instead of stopping. Only the i-quant models run there, and
    the CPU's share is slow (forced on our Ryzen: 11-17 tok/s with the AVX build, 3.5-3.8 with SSE4.2, against 26).
    Untested here on a real old CPU; contributors ran earlier versions of it on Xeon E5-2680 / E5-2687W / X5690.
    docs/INSTALL.md#older-cpus-experimental.

Prompts with a RAM budget are fast again (#577): 0.1.38 decided too early whether to read the experts past the
file cache, and it compared against the size of every model file. On a 96 GB PC with UD-Q4_K_XL and a 72 GiB budget
it chose wrong, so every refill after a prompt read the drive (prompts 15-40% slower than 0.1.34). The choice now
counts only the expert bytes outside the RAM copy, and is made again once that copy is built. The tokens are the same
either way.

The OpenAI Responses API (#451): POST /v1/responses, so Codex CLI works with Strata (tested with Codex 0.160.0,
a tool loop included; later turns reused ~96% of the prompt from the cache). It is stateless, as Codex uses it, and
covers function tools, tool results, reasoning effort, JSON schemas and the streaming events. Not supported:
previous_response_id (nothing is stored), hosted tools and reasoning summaries. The Codex config.toml is in
docs/DETAILS.md.

Security: SECURITY.md says how to report a problem privately (GitHub's private vulnerability reporting) and what
the server exposes: 127.0.0.1 by default, the API key, the Host and Origin checks of 0.1.38, CORS, and the opt-in MCP
tools and request monitor.

Fixes from your reports:

  • A reply stuck on one token is ended (#606): a reply that repeats one token 256 times ends there
    ("repeat_stop_tokens" sets the length, 0 turns it off), and the q8_1 activations stay finite.
  • Starting on a tight card (#620): the native head and the logits are loaded before the expert arena, and a failed
    allocation names the free VRAM.
  • The RAM check before the arena (#633): a container whose memory limit is below the arena is refused with the
    numbers instead of being killed during the load. Less RAM available only warns.
  • A rotational disk (#605): --ple-io direct warns at start, setup keeps the n-gram table in RAM there, and the
    stall report counts disk waits.
  • layer_split in the config (#644) is checked before the start, takes a JSON list, and says the format.
  • Linux CUDA toolkits (#601): STRATA_NVCC picks the toolkit and CUDA_HOME is honoured.
  • Windows Pascal/Volta source builds (#585) link the shared CUDA runtime.
  • The hit rate (#588) names the PCIe share beside it.
  • The server: a malformed tools value is a 400 instead of a dropped connection (#592); a literal <think> in a
    message is read as text (#537); /load and /unload read the request body before replying (#630).
  • Setup: running it again keeps the run config's other keys (#629); --no-browser / "open_browser": false (#609
    #631); --vision-tokens N (#625); --draft-vocab fr, the English/code subset plus French (#597); a calibration on
    Linux HIP is kept for its card (#566); Unsloth's UD-IQ4_XS as an experimental choice (#621); on hybrid CPUs with
    more E-cores than P-cores, setup writes a recommended --pool-workers you can edit (#642).
  • The web page: a Model settings card (#564) and a Conversation cache card in the Monitor tab (#596).
  • Docs: the decode window profiler (#610) and the five reading errors #604 found.

New options (off by default; the default output is unchanged):

  • "effort_position": "end" (#458): a request that only changes the effort reuses the cache...
Read more

Strata v0.1.38

Choose a tag to compare

@Niko1221 Niko1221 released this 03 Oct 01:17

Faster prompts, --kv q4_0 prompts on tensor cores, a 6 GB card that starts, and a large batch of community PRs.

Security: the server without an API key (DNS rebinding, cross-site requests). Without api_key, the server now
answers only requests addressed to a name it knows (localhost, an IP address, the address it listens on and, when it
listens beyond this PC, this PC's name), and a browser page from another site can no longer send it requests (403).
/unload and /load take only JSON: curl -X POST -H "Content-Type: application/json" localhost:8080/unload.
Clients without an Origin header (curl, the OpenAI and Anthropic SDKs, other servers) are not affected, and with an
API key the key decides. Behind a reverse proxy, a tunnel or Docker under another name, without a key: add the
name, "allowed_hosts": ["strata.example.com"] in strata-<model>.json (or STRATA_ALLOWED_HOSTS), or set an API
key. Details: docs/DETAILS.md.

Faster prompts:

  • The prompt path gathers a group of experts in one launch with one wait (#372), reads the first chunk's n-gram rows
    beside layer 0 instead of before it (#374), and runs the DeltaNet recurrence three heads per thread (#413). The
    answers are the same bits as before.
  • Measured on an RTX 5070 12 GB (3 alternating pairs each against 0.1.37): Q2_0 prompts +10% at 4K, +2.7% at 32K,
    +3.0% at 128K; IQ2_XS +1.9% at 32K and +6.4% at 128K (4K: within the run-to-run spread).
  • Decode time per verify round is the same or a little shorter. Since #463 a greedy answer no longer depends on when
    the adaptive tier's expert copies land, so on some long prompts the drafts are accepted a little more or less often
    than before (Q2_0 at 128K: 70.9 instead of 75.1 tokens/s on the measured prompt, IQ2_XS at 32K: 73.0 instead of
    69.7).

--kv q4_0 (#452): prompts read their Q4_0 K/V on tensor cores (RTX 30 and newer; RTX 20 keeps the old kernel
for now): a 32K prompt on Q2_0 in 14.3 s instead of 18.8 s. Against an FP16 K/V reference, over 2,000 teacher-forced
positions after the prompt: at 32K, KL 0.026 and top-1 95.1% (0.031 / 93.6% with the old kernel); at 8K, KL 0.079 and
top-1 89.3% (0.077 / 90.0%). The default --kv int8 is unchanged.

A 6 GB card starts (#496): when the default VRAM reserve leaves the expert cache too little room, the engine
lowers the reserve (down to 300 MiB) until the cache fits, and warns if the card ends nearly full. If even that is
not enough, the start stops with how many MiB are short and what frees them; setup gives the same tip on cards under
8 GB.

More formats and CPUs:

  • Q5_0 experts on the GPU (#473): community Q4_K_M GGUFs that put Q5_0 on the down projections now load. Q5_0
    n-gram tables also read with --ple-io direct.
  • IQ4_XS on AVX-2-only CPUs (#415): the multi-token kernel the other i-quants already had. AVX-512 CPUs are
    unchanged.

Loading and memory:

  • Windows, short on RAM (#357 #362): the experts are read past the file cache when it cannot keep them anyway, at
    start and in the RAM-budget tier (faster starts and decode on 32-64 GB PCs, the same tokens).
  • Linux (#488): STRATA_NO_LARGEPAGES=1 works there too, and a hugetlb pool that is too small is named.

AMD: hipBLASLt tables for gfx1201 and gfx1200 with hipBLASLt 1.2.2 (system ROCm 7.2.x; ~1.5-2x prompt speed there,
#386 #387), and the draft layer's prompt pass runs per group on HIP, which fixes a prompt hang on gfx1201 (#382).

New options (off by default; the default output is unchanged):

  • --peer-device N (#531): a second GPU as an extra expert cache (needs P2P, e.g. NVLink); see docs/SECOND_GPU.md.
  • --adapt-decay F (#407): how fast the adaptive tier forgets (0.7, as before).
  • STRATA_GR_DOWN_MAX4=1 (#443): a smaller decode kernel variant for short windows (the same bits).

Also: a native pack with STRATA_PF_FUSED=1 whose fused kernels don't cover every layer (Unsloth UD-IQ4_XS) no
longer gives garbage or hangs on long prompts; the verify window's PCIe call launches less (#363), the adaptive tier
waits for its expert copies so a decode no longer depends on their timing (#463), a stager race with unpinned blobs is
closed (#385), RTX 20 cards pick a faster top-k at long contexts (#512), and --dump-logits no longer writes a header
without rows (#463).

Fixes from your reports:

  • AMD on a Linux desktop (#560 #516): when the desktop or apps crash once the model is loaded, the expert cache
    has filled the card the desktop needs. Setup and the server now recommend --vram-reserve-mib 3072 there (nothing
    is changed for you).
  • A slow CUDA build (#542): MMQ reads the GPU's limits in a way a mismatched CUDA runtime cannot shift, and the
    engine warns when its runtime is older than the toolkit it was built with.
  • The weight arena does not fit (#486): the error says how much VRAM was free at that moment.
  • Setup and the server (#549 #545 #530 #564): setup --update skips a strata-*.json that is no model config;
    the "exceeds the context" error says how to get past it; a reply that reaches its token limit while still thinking
    is logged with what helps; each start prints the settings it uses.
  • Docs: an OpenCode config example (#543); MODELS.md says plainly that the Coder is the 32 GB fit, weaker outside
    code, and recommends the full models for general use.
  • A parity test waits for its uploads (#548).

Thanks to everyone who sent PRs for this release, and especially to sergqwer and BlueKingMuch for the prompt-path and
decode work.

Checked before the release:

  • The same answers as 0.1.37 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each with a fixed cache,
    and the prompt path's internal state at 4K and 20K tokens (Q2_0, the Coder).
  • Speed: the 5-pair A/B on Q2_0 and IQ3_S: Q2_0 +4.0% / +1.9%, IQ3_S +0.2% / +0.1%, the same expert slots; the
    prompt/decode table above.
  • Kernel parity tests (the new DeltaNet, grouped-expert, top-k and Q4_0 attention tests included), real use at a
    57K-token prompt
    (Q2_0, the Coder), Linux (WSL) Q2_0 identical 10/10.
  • AMD on Windows: the HIP zip builds; the AMD changes are untested on an AMD card here (we have none).
  • After the report fixes were merged: Q2_0 and IQ3_XXS still identical 10/10 and at 4K, and the Q2_0 speed A/B
    +0.7% / +0.6% (the new start-up checks cost nothing).
  • Tests: tools/setup 287, server 197 (the new security tests included).

Updating: run UPDATE.bat (Linux: ./update.sh). Setup installs engine 0.1.38.

The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:

  • strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an
    NVIDIA driver 580 or newer. Contents: strata.exe, strata-vision.exe (the optional image encoder), BUILD.json.
  • strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030), ROCm 10.2.0a20260930
    from AMD's TheRock builds, needs a current AMD driver. Contents: strata.exe, strata-device.exe, the HIP runtime
    next to them, BUILD.json, rocm\ (the ROCm libraries and their licenses).

Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:

Buy Me A Coffee

Strata v0.1.37

Choose a tag to compare

@Niko1221 Niko1221 released this 02 Oct 19:51

A frozen engine no longer hangs the server, Windows AMD counts the desktop's VRAM, and a batch of setup fixes from your reports.

The server notices a silent engine (#481): if the engine prints nothing during a request for engine_silence_s
seconds (default 300, longer while a long prompt is being read; 0 turns it off), or never answers a stop, the server
ends the request with an error and the next request restarts the engine. Before, both sides could wait for each
other forever. The cause of the reported freeze is still being traced with the reporter.

AMD on Windows (#497, PRs #380 and #377 by BlueKingMuch):

  • The engine counts the desktop's VRAM (the WDDM budget), so the expert cache no longer spills into system RAM
    (an RX 6800: 31 -> 42 tokens/s). STRATA_WDDM_BUDGET=0 turns it off.
  • The PCIe probe is timed on the host there (it read impossible values before).

A steadier PCIe probe (#485): the best of four timed bursts, and the share of missed experts read over PCIe
follows the measured speed smoothly below 20 GB/s instead of jumping. Above 20 GB/s nothing changes.

Setup:

  • UD-Q4_K_XL on several GPUs (#498): with enough RAM (~135 GB), setup offers the layer split; --gpus is
    honoured. A start with --gpus on a RAM-budget config no longer fails.
  • Mirrors and manual downloads (#495): HF_ENDPOINT is honoured (e.g. https://hf-mirror.com), also for the
    draft layer's download. Setup prints the folder and file names it expects for files you already have.
  • --vram-reserve-mib N (#493): keeps N MiB of VRAM free for other programs; documented in INSTALL.md.
  • Start errors (#496): when the engine stops before it is ready, the error shows the end of its log.
  • Messages (#505): the start messages match the mode (mapped, resident, RAM budget).

Checked before the release:

  • The same answers as 0.1.36 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each, and the prompt
    path's internal state at 4K and 20K tokens.
  • Speed:
    • the 5-pair A/B on Q2_0 and IQ3_S: +0.9 / -0.1 / +0.9 / +0.3%, the same expert slots;
    • a decode re-check at 32K and 128K (3 + 2 alternating pairs): the same speed.
  • Kernel parity tests, real use at a 57K-token prompt (Q2_0, the Coder), Linux (WSL) Q2_0 identical
    10/10.
  • AMD on Windows: the HIP zip builds. The two PRs are reviewed and Windows-HIP-only, but they are untested on an
    AMD card here (we have none); BlueKingMuch measured them on an RX 6800.
  • Tests: tools/setup 277, server 172.

Updating: run UPDATE.bat (Linux: ./update.sh). Setup installs engine 0.1.37.

The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:

  • strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an
    NVIDIA driver 580 or newer. Contents: strata.exe, strata-vision.exe (the optional image encoder), BUILD.json.
  • strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030), ROCm 10.2.0a20260930
    from AMD's TheRock builds, needs a current AMD driver. Contents: strata.exe, strata-device.exe, the HIP runtime
    next to them, BUILD.json, rocm\ (the ROCm libraries and their licenses).

Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:

Buy Me A Coffee

Strata v0.1.36

Choose a tag to compare

@Niko1221 Niko1221 released this 02 Oct 16:38

Faster: Q2_0 reads prompts 16-22% faster and writes up to 18% faster at long contexts. Plus one-click updates (UPDATE.bat), an opt-in memory of which experts you use, and clearer logs.

Speed (#136), measured on an RTX 5070 12 GB, 0.1.35 -> 0.1.36:

Q2_0 Reads your prompt Writes answers
4K prompt 1,294 -> 1,570 tokens/s (+21%) 89 -> 93.5 (+5%)
32K prompt 2,170 -> 2,653 (+22%) 80 -> 81
128K prompt 2,123 -> 2,468 (+16%) 64.5 -> 76.4 (+18%)
  • New prompt kernels for Q2_0 (RTX 30 and newer): the experts are multiplied by fused int8 tensor-core kernels
    that read each expert where it already is (no copies), with the grouping done on the GPU. They are as accurate as
    before: measured against an FP16 reference over 2,000 teacher-forced tokens after a long document, as close as the
    previous kernels at 8K and closer at 32K (KL 0.009 vs 0.012, perplexity 4.121 vs 4.169, the reference 4.129).
    Answers are not byte-identical to 0.1.35 any more, the same kind of change as a new GPU kernel's rounding;
    STRATA_PF_FUSED=0 in the config's env brings the previous kernels back.
  • Faster long-context decoding on RTX 50 cards: the attention's block selection and the greedy token pick run on
    thread-block clusters. The output is exactly the same (bit-identical tested). Older cards keep the previous
    kernels.
  • Opt-in for the IQ sizes: STRATA_PF_FUSED=1 also runs fused kernels for IQ2_XS / IQ3_XXS / IQ3_S (IQ2_XS
    prompts +12% at 4K, +3% at 32K; the IQ3 sizes about even, so they stay off by default).
  • The approach was inspired by the measurements of another engine for this model,
    flashrt; Strata's kernels are its own.

UPDATE.bat / update.sh (#475): updates Strata without starting the model. In a git clone it pulls the new
files (a zip copy is told to download the new zip), then updates the Python packages, the engine, each model's
settings and the draft vocabulary. Then it stops.

Keep what the expert cache learned (#477, opt-in): with "expert_profile_save": "path\\to\\my-profile.bin" in
strata-<model>.json (or --expert-profile-save), the engine saves which experts your work uses (every 10 minutes
and on exit), and the next start begins from it instead of the shipped profile. Off by default.

Fixes:

  • A prompt cancelled while it is being read (#471): the log and /metrics say how much was read, with the real
    reading speed.
  • "the draft head does not fit" (#474): the engine says what it needs, the free VRAM, and the smaller
    --draft-vocab cyrillic / en options with their sizes; the server shows the hint.

Checked before the release:

  • The same answers as 0.1.35 on IQ3_XXS, IQ3_S and the Coder (10/10 each), and on Q2_0 with STRATA_PF_FUSED=0
    (10/10). Q2_0's default answers changed by design (the quality check above).
  • The new kernels' tests: the fused kernels against an FP32 reference and the previous kernels for every
    format (the same error), the cluster kernels bit-identical to the previous ones over 650 cases up to 262K, the
    existing parity tests.
  • Speed: the 5-pair A/B on Q2_0 and IQ3_S (+1.9 / +0.6 / +0.6 / +0.7%, the same expert slots) and the table above.
  • Real use at a 57K-token prompt through the server on Q2_0 and the Coder; the fused path also through serve
    sessions with cancels, KV streaming and mixed chunk sizes.
  • Linux (WSL, RTX 5070): Q2_0 identical to 0.1.35 with STRATA_PF_FUSED=0 (10/10), and the defaults pass.
  • AMD: the Windows HIP zip builds (AMD keeps its own kernels; the new ones are NVIDIA-only).
  • Tests: the tools' and setup's (261), the server's (158).

Updating: run UPDATE.bat from now on. This time: git pull (or download the new zip), then START-HERE.bat
(Linux: ./setup.sh). Setup installs engine 0.1.36.

The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:

  • strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an
    NVIDIA driver 580 or newer. Contents: strata.exe, strata-vision.exe (the optional image encoder), BUILD.json.
  • strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030), ROCm 10.2.0a20260930
    from AMD's TheRock builds, needs a current AMD driver. Contents: strata.exe, strata-device.exe, the HIP runtime
    next to them, BUILD.json, rocm\ (the ROCm libraries and their licenses).

Strata v0.1.35

Choose a tag to compare

@Niko1221 Niko1221 released this 02 Oct 09:05

Windows AMD no longer crashes on the first prompt, the low-RAM mode works on 32 GB Windows PCs, and a batch of fixes from your reports.

AMD on Windows (#468, #461): the engine now uses the HIP runtime it ships with. Windows looks for a program's DLLs
in its own folder, then in System32, and only then on PATH, where the engine's rocm\bin is. AMD drivers that
install their own amdhip64_7.dll in System32 won, and the bundled libraries crashed on the first prompt (an access
violation on a W7900, hipErrorInvalidDeviceFunction on a 7900 XTX). Setup now puts the bundled runtime next to
strata.exe (every start, so a 0.1.34 install is fixed by running START-HERE.bat), the zip ships it there, and the
engine log names the runtime it loaded.

Low-RAM mode on 32 GB Windows PCs (#467): the resident mode (experts the GPU doesn't hold kept in RAM) never
engaged. Windows counted the files the GPU cache had just read as used memory. The engine now releases them before
it checks, and when the RAM still can't hold every expert, it keeps the hottest ones that fit, instead of falling
back to reading all of them from the SSD.

Fixes:

  • Multi-GPU (#448): a small second card in a layer split capped every prompt chunk at 512 tokens (6x slower
    prompts). The log now names that card, and setup recommends the big card alone (the small one is better as an
    extra expert cache, docs/SECOND_GPU.md).
  • Server:
    • messages sent as a JSON string works, and a malformed list is a 400 (#460);
    • /metrics shows the speculative draft counts per request and in total (#457).
  • Setup:
    • an empty or corrupt config from an earlier install no longer stops a new copy, and configs are written safely
      (#459);
    • a system ROCm without HIP's development files is skipped for the AMD build, and setup uses AMD's wheels instead
      (#446);
    • --gguf-dir says which files it can't use and which it can (#444);
    • --calibrate works with a helper card (--expert-cache-device1), and says so when tuning fails (#447).
  • Docs: the Coder is weaker outside code, including Chinese and other non-English text; use the unpruned sizes for
    those (#438).

Checked before the release:

  • Same answers as 0.1.34 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each with a fixed cache,
    including the prompt path's internal state at 4K and 20K tokens.
  • The same speed: 5-pair A/B on Q2_0 and IQ3_S, -0.2 / -0.7 / +0.1 / -0.3%, with the same expert slots.
  • #467 here:
    • IQ3_XXS in the resident mode: the trim freed 3.5 GiB (51.1 -> 54.5 GiB available), and 38.7 GiB of experts were
      kept in RAM;
    • the same tokens as the mapped mode.
  • Real use at a 57K-token prompt through the server on Q2_0 and the Coder (incl. a cancelled long prompt and the
    next request).
  • Linux (WSL, RTX 5070): Q2_0 byte-identical to 0.1.34 (10/10), and the default settings pass.
  • AMD on Linux (R9700): the HIP build of the release and its tests, all passing except the two that need a model
    fixture or an AVX-512 CPU, as before.
  • AMD on Windows: the zip carries amdhip64_7.dll and amd_comgr.dll next to the exes; setup's copy step is
    tested. The fix itself awaits the reporters' confirmation (no System32 HIP runtime on our test PC).
  • Tests: the tools' and setup's (255, including the golden check that --yes configs are unchanged), the server's
    (149).

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.35.

The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:

  • strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an
    NVIDIA driver 580 or newer. Contents: strata.exe, strata-vision.exe (the optional image encoder), BUILD.json.
  • strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030), ROCm 10.2.0a20260930
    from AMD's TheRock builds, needs a current AMD driver. Contents: strata.exe, strata-device.exe, the HIP runtime
    next to them, BUILD.json, rocm\ (the ROCm libraries and their licenses).

Strata v0.1.34

Choose a tag to compare

@Niko1221 Niko1221 released this 02 Oct 01:21

AMD cards on Windows (new), an MCP server so your AI assistant can set Strata up, a shorter README - and a cancelled request now frees the engine within a second.

AMD on Windows (new, #247 #325): an AMD card on Windows is set up like an NVIDIA one: double-click
START-HERE.bat. On a PC with no NVIDIA card Strata can use, setup picks the AMD card by itself and downloads the
ready-made AMD engine (strata-windows-x64-hip.zip, below). It carries the ROCm libraries it needs, so the AMD
driver (AMD Software: Adrenalin Edition) is the only thing to install. Cards: RX 7900 XT / XTX, RX 7800 XT /
7700 XT, RX 9060 XT, RX 9070 / 9070 XT, Radeon AI PRO R9700, RX 6800 / 6900 series.

  • This is new: #325's author ran this port on an RX 9070 XT. Here, the zip's own runtime finds a Radeon iGPU, and
    a test build of the same code ran a hipBLAS GEMM correctly on it. But the release zip hasn't run a model on a
    discrete card yet. Please report how it runs (docs/AMD_HIP.md, "Reporting a Windows AMD run").
  • For now on Windows: one card per model, no images. Linux AMD is unchanged: multi-GPU, and images through the
    CPU.
  • Thanks to jagsan-cyber (#247), dvasdekis (#325) and araujoluks (#302) for the Windows work this builds on.

Let your AI set it up (#133): paste one line into Claude Code, Cursor, Codex or Copilot (in the README). It
checks the PC, picks the model, installs, starts and tells you how to connect. tools/strata_mcp.py is an MCP server
that lets AI tools install, start, check and stop Strata themselves (docs/MCP_SERVER.md).

A shorter README: what Strata is, how fast it is (now with an AMD PC next to the NVIDIA one), what you need, how
to install, which model to pick. The details moved to docs/INSTALL.md, MODELS.md, TROUBLESHOOTING.md and
HOW_IT_WORKS.md. AMD is no longer called experimental.

Fixes:

  • A request whose client hangs up is cancelled within about a second (#430, #431, jkuepker). Before, a
    non-streamed request ran to max_tokens, and a streamed one read its whole prompt, while the next request waited
    (17-33 s in the reports).
  • The prompt path no longer aborts when llama.cpp's MMQ has no tile that fits a card (J_best=0, #420,
    qni-live). That product takes the FP16 path, and the engine says which type and card. Where everything fits,
    nothing changes.
  • Setup:
    • a resumed download needs disk room only for what is still missing (#425, jctaborda);
    • UD-Q4_K_XL on an AMD card is asked about before its 111 GB download, since it is untested there (#429, jkuepker).
  • Docs: the AMD router's decode gain with a user's repeated measurement next to ours, and their
    STRATA_HIP_WMMA=1 numbers (#432, jkuepker).

Checked before the release:

  • Same answers as 0.1.33 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each with a fixed cache,
    including the prompt path's internal state at 4K and 20K tokens. Checked both after the merge and on the final
    engine.
  • The same speed: 5-pair A/B on Q2_0 and IQ3_S, +1.2 / +0.9 / -0.2 / +0.4%, with the same expert slots.
  • Real use at a 57K-token prompt through the server on Q2_0 and the Coder. A long prompt whose client hung up,
    followed by a short question: the question was answered in 0.6 s (17 s in #431's report).
  • Linux (WSL, RTX 5070): Q2_0 byte-identical to 0.1.33 (10/10), and the default settings pass.
  • AMD on Linux (R9700): the HIP build of the release and its tests, all passing except the two that need a model
    fixture or an AVX-512 CPU, as before. The README's AMD numbers were measured on an RX 9070 XT + Ryzen 9 3900X +
    47 GB RAM with setup's own install.
  • AMD on Windows: the zip built and packaged. Its strata-device with only the zip's libraries lists this PC's
    integrated Radeon and correctly reports that the engine has no code for it (gfx1036).
  • Tests: setup's (132, including the golden check that NVIDIA --yes configs are unchanged), the other tools' and
    the MCP server's (236 together), the server's (146, including client hang-ups, streamed and not).
  • #410: reproduced on IQ3_XXS: --pcie-frac 0 was the missing switch for byte-identical repeats (DETAILS.md).

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.34.

The ready-made Strata engines for Windows, which START-HERE.bat fetches by itself:

  • strata-windows-x64.zip: NVIDIA (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0, needs an
    NVIDIA driver 580 or newer. Contents: strata.exe, strata-vision.exe (the optional image encoder), BUILD.json.
  • strata-windows-x64-hip.zip: AMD (gfx1100, gfx1101, gfx1102, gfx1200, gfx1201, gfx1030), ROCm 10.2.0a20260930
    from AMD's TheRock builds, needs a current AMD driver. Contents: strata.exe, strata-device.exe, BUILD.json,
    rocm\ (the ROCm libraries and their licenses).

Strata v0.1.33

Choose a tag to compare

@Niko1221 Niko1221 released this 01 Oct 21:16

The image encoder runs on every CPU again, and setup recommends instead of forcing: a longer context, more GPUs or a bigger RAM budget you choose is kept, with a note about the risk.

Images (#411, #412): 0.1.32's ready-made strata-vision.exe was built with AVX-512 and OpenMP and stopped at
start on CPUs without AVX-512 (most AMD Ryzen before Zen 4, many Intel desktop CPUs): "illegal instruction" or "the
vision encoder did not start". 0.1.33's is built portable again (AVX2, no OpenMP), like 0.1.31's, and keeps 0.1.32's
RTX 20 code. Running START-HERE.bat once installs it.

Setup recommends, you choose (#406, #403, #364, #384):

  • Context: setup still recommends 128K for IQ3_XXS / IQ3_S on a 64 GB PC, but a longer one you pick (or
    --context 262144) is kept, with a note that it may run out of RAM under load. An update keeps your earlier
    context, and in the low-RAM mode the context no longer counts against RAM.
  • Several GPUs on a low-RAM PC: setup asks whether to use one GPU (the default: experts kept in RAM, steady RAM
    use) or all of them (experts read through the OS file cache: 1.3-1.6x faster in two users' measurements, but RAM
    can fill up during long prompts). --gpus 0,1 is honoured. The engine runs --resident-experts with a layer split
    as --mmap-experts, with a warning, instead of refusing to start.
  • RAM budget (--resident-budget-gib): a budget larger than the free RAM no longer fails the safety check after
    the engine lowers it (#403), and a budget that can't be kept at all is a warning. Setup takes --resident-budget-gib N.
  • Setup asks instead of stopping when a model is above its RAM estimate, or a named card has under 8 GB for a
    split. --yes alone keeps the recommended choices. --kv-streaming on|off|auto is new.
  • With --yes or Enter for every question, setup writes exactly the same config as 0.1.32 (checked on 23
    PC/model combinations).

Fixes:

  • V100 (sm_70) in a layer split: the prompt's attention uses the kernel that card supports again (#371, AxalotLDev).
  • --expert-cache-per-layer starts on native packs and fills every layer (#369, dag08).
  • STRATA_GR_V3=1 (opt-in) keeps the default read on cards where its two-half split fails its checks (Turing; #375,
    kenh0u); its test compares within float rounding (#379, sergqwer).
  • Server: a request the chat template refuses ("No user query found in messages.") gets a 400 error instead of a
    dropped connection (#365); a temporary image file is removed when the encoder fails (#352, Tsury); the image encoder
    can run on its own GPU, "cuda_device" in the vision section (#408, Efs-O).
  • Setup stops when its input ends instead of answering yes, e.g. to installing the build tools (#393, #398 by
    Bortlesboat); a setup test no longer fails on a PC with CUDA installed (#414, rwkeyes).
  • Docs: byte-identical greedy repeats through the server also need --prompt-cache 0 --adapt-swaps 0 (#410).

Checked before the release:

  • Same answers as 0.1.32 on all four quants (Q2_0, IQ3_XXS, IQ3_S, the Coder), 10/10 each with a fixed cache.
    The prompt path's internal state at 4K and 20K tokens is identical too. The default speed is unchanged within noise
    (5-pair A/B, same expert slots). The merged release was checked again on Q2_0 and IQ3_XXS.
  • The image encoder: no AVX-512 instructions and no OpenMP (the same profile as 0.1.31's, which ran on CPUs
    without AVX-512), and the image test passes behind the server (greedy colour and three sampled descriptions).
  • Real use at a 57K-token prompt through the server on Q2_0 and the Coder.
  • Linux (WSL, RTX 5070): Q2_0 byte-identical to 0.1.32 (10/10), and the default settings pass.
  • AMD (R9700): the HIP build and its tests. All pass except the two that need a model fixture or an AVX-512 CPU,
    as in 0.1.32.
  • Tests: setup's tests (120, including a golden check that --yes/Enter configs are byte-identical to 0.1.32's
    on 23 PC/model combinations), the server's (141), the tools', and the parity tests of the changed kernels (QSA
    prompt attention, the hyper-connection read with and without V3, the per-layer expert cache).
  • #403: an engine started with --resident-budget-gib 500 lowers the budget with its margin, starts and answers.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.33.

The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.

Strata v0.1.32

Choose a tag to compare

@Niko1221 Niko1221 released this 01 Oct 17:28

Multi-GPU short prompts fixed, AMD decode +12%, Unsloth's Q4 model in setup with 2.8x faster prompts, and 30+ community changes - with the same answers.

Multi-GPU: short prompts faster again (#340, riverhh76). Since 0.1.30 a layer split lent each card's expert cache to
its prompt path and refilled it after every request, one card after another; 0.1.31's bigger streaming ring made the
loans larger. Short prompts lost up to a third of their speed (2x RTX A5000: -35%). 0.1.32 refills all cards at once,
uses a smaller ring on a split, and lets a card with free VRAM keep its own prompt buffers - with exactly 0.1.31's
output. Measured on an R9700 + RX 9070 XT: 2K prompts 993 -> 1,270 tok/s, 16K 1,852 -> 1,980. STRATA_SPLIT_OWN=1
(opt-in) gives every card its own buffers: 2K 1,470 and 16K 2,230 there, with a slightly different output.
--no-prefill-borrow works on a split again.

AMD:

  • The router is rewritten for RDNA without its serial parts: decode +12% on the R9700 (62 -> 70 tok/s), +4% on the RX
    9070 XT, bit-for-bit the same output. STRATA_HIP_ROUTER_OLD=1 brings the old kernel back.
  • Opt-in STRATA_HIP_WMMA=1 (RX 9070 / R9700): the prompt's attention on RDNA4's matrix cores, +36-51% prompt speed
    on the R9700 (4K/16K) with a slightly different output (#329 by bsorensen110 measured the same speed; credited).
  • Setup can use the CPU image encoder with AMD (#304), and the Monitor shows AMD GPU readings (#301).
  • #337 (bsorensen110): on AMD the prompt's QSA top-k picks its kernel by the blocks a query actually has (the same
    output); its RDNA4 matrix-core block scorer is opt-in (STRATA_SELECT_WMMA=1). #339 (bsorensen110): a gfx1201
    hipBLASLt table for hipBLASLt 1.5 (ROCm nightly; +3.9% prompt on an R9700). #311 (Rafael Grossi): the RX 6800 /
    6900 series (RDNA2, gfx1030) as a community card in setup and the build.

Unsloth UD-Q4_K_XL (experimental):

  • In setup: --family unsloth --model UD-Q4_K_XL downloads the four files (each checked once by size and SHA-256),
    packs them and picks the RAM budget from the PC's RAM.
  • Long prompts 2.8x faster on a 64 GB PC: a 16K prompt reads at 160 tok/s instead of 57 (more experts read from the
    SSD in parallel).
  • Quality against llama.cpp on the same file: the same top token at 99.0% of the positions of a code answer and
    97.5% of a thinking answer (the fork's own numbers: 98-99% and ~96%); perplexities within 3% (docs/UNSLOTH_Q4.md).

Fixes:

  • OrcaRouter IQ3_XXS packs load again (broken since 0.1.25; #326, Gen4536).
  • The draft layer's download checks every tensor's SHA-256 and fetches a corrupt one again; a mirror that ignores
    range requests can no longer write wrong data (#327, lifeidle).
  • During an engine restart a request gets "the engine is starting" (503) instead of an error about a context of 0
    (#344, backstable).
  • Conversation cache (opt-in): a subagent's turns no longer push its parent conversation out (#342, Anhelone).
  • RTX 20 (Turing): the image encoder in the ready-made engine has RTX 20 code, and setup picks the CPU encoder for a
    card it doesn't cover (#331).
  • Setup's --gguf-dir takes any number of files (#305); Pascal/Volta builds through setup with
    STRATA_EXPERIMENTAL_SM60=1 (#295, giostrives); model aliases in the config (#297).

Community changes (all checked byte-identical on four models, or opt-in):

  • Faster, same output: the verify commit overlaps the draft (#284, sergqwer), a finer multi-token hyper-connection
    read chosen per card (#315, BlueKingMuch), the SSD kept awake while PLE rows are read (#317, BlueKingMuch).
    Decode here +2-4% over 0.1.31.
  • Server: /v1/messages/count_tokens, a relative engine path (#278); OPTIONS, opt-in CORS on /v1/* with a trusted-
    origin list, reverse-proxy support (#321, BlitzenCats - without the header bypass the PR had); an opt-in request
    monitor (#332); lazy start and a reliable unload (#333); response_format JSON checks (#334, KadoBOT);
    "anthropic_thinking": "on_request" (opt-in) for clients like Claude Code's helper calls.
  • Opt-in: larger prompt chunks --prefill auto:32768 (#282), BF16 remainder projections (#283), a float64 RoPE table
    (#280), Hadamard-rotated INT8 KV (#293), a BF16 token embedding (#290), an FP8 PLE table (#291), all sergqwer;
    --pool-affinity auto for P/E-core CPUs (#272, praveshkhatana); --coupled-draft (#274).
  • Also: NaN-safe FP16 SwiGLU saturation (#281), --draft-vocab cyrillic (#287), the image encoder's --flash-attn
    (#288), the GPU named at start (#289), STRATA_DUMP_FIRST_LOGITS (#276), experts.bin reuse checks the blobs (#277),
    verify profile fix (#320, xyzzing), retained K/V through the cache gate (#309, chimpera), OrcaRouter Q4_K_S pieces
    (#296, Suoriks).

Checked before the release:

  • Every change gated: byte-identical to 0.1.31 with a fixed cache on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S,
    the Coder), including the prompt path's internal state - each work stream on its own, then the merged release;
    default-settings speed A/Bs (5 pairs) on Q2_0 and IQ3_S: the release decodes +1.5 to +3.8% faster than 0.1.31 with
    the same expert slots. Two changes were reworked because they cost speed or changed defaults: the Unsloth prompt
    kernels are a build option (they took VRAM from every model), and #278's no-thinking default is opt-in.
  • Real use at a 64K context on all five models (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, the Coder), through the server: a fact
    found in a 57,023-token document, the follow-up turn, a tool call, a cancelled long prompt followed by a new
    request, a sampled answer - no engine restart, no low VRAM.
  • Images: the new image encoder (RTX 20 code added) behind the 0.1.32 server with IQ3_XXS: the right answer about a
    picture and different sampled descriptions.
  • The server's tests (117), setup's tests, the pack and draft-layer download tests, the parity tests of the changed
    kernels.
  • Linux (WSL, RTX 5070): builds, and Q2_0 is byte-identical to 0.1.31 (10/10); the default settings pass.
  • AMD (RX 9070 XT and R9700): the HIP build with its tests (all pass but the two that need a model fixture or an
    AVX-512 CPU), the live server test on each card (R9700 92 tokens/s, 9070 XT 36) and on the two-card split (76),
    the split's output identical to 0.1.31 (5+5 starts), the router bitwise identical to the old one; a gfx1030
    build compiles (no RDNA2 card here).
  • The Unsloth model on this PC: setup's flow (mocked end to end), quality against llama.cpp, the prompt speed.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.32.

The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.

Strata v0.1.31

Choose a tag to compare

@Niko1221 Niko1221 released this 01 Oct 05:13

Unsloth's Q4 model runs on a 64 GB PC (experimental), AMD decode up to 15% faster, models load twice as fast on Windows, and a batch of fixes.

Experimental: Unsloth UD-Q4_K_XL on a normal PC. Strata can now run Unsloth's 4-bit Qwen3.8-Flash-Next
(UD-Q4_K_XL, 111 GB in 4 files, 72 GiB of experts) on a PC whose RAM is smaller than the model:

  • The model's four files are read as they are: a missing file is an error, a layer split across files is fine.
  • Its expert formats (Q4_K, Q5_K, Q5_1, Q8_0) run natively on the GPU and the CPU; the GPU dequantizers match
    llama.cpp's bit for bit, the dot products match ggml's CPU results within float rounding.
  • Three tiers for the experts: the most used in VRAM, the next ones in RAM (--resident-budget-gib N), the rare
    ones read straight from the model files on the SSD, prefetched one layer ahead from the router's prediction.

Measured on an RTX 5070 12 GB with 64 GB RAM, a 40 GiB RAM budget: about 7-8.5 tokens/s for answers, with
correct output. A big-RAM PC holds more in RAM and is faster. Not in setup yet: see docs/UNSLOTH_Q4.md for the
manual import. Ported from eddoursul's fork (#245), with pieces from gopinath87607 (#255) and jagsan-cyber (#247).

Low-RAM mode without the 30-70 GB copy. The low-RAM mode (--mmap-experts) now reads a native pack's experts
straight from the model's GGUF files when the pack has no experts.bin - the same tokens and logits as with the
copy.

Faster:

  • Windows: models without experts.bin load about twice as fast (the expert reader used 4 KB reads; #230,
    reported by activeing123). IQ3_XXS starts in ~30 s instead of ~47 s here.
  • AMD RDNA4: decode +15% on the Radeon AI PRO R9700 and +5-7% on the RX 9070 XT (byte-level intrinsics from
    #262, ttio2tech); the live test on the R9700 went from 67 to 82 tokens/s.
  • The Q2_0 expert kernels (#241, gputier) and the IQ expert kernels (#242, gputier) do less work per token -
    bitwise identical output.
  • RTX 20 (Turing): long prompts ~12% faster - the prompt attention runs on tensor cores (#270, kenh0u; output on
    RTX 20 cards changes slightly, as for RTX 30+ before), and the hyper-connection kernel stages smaller tiles (#258,
    hireymage, bitwise identical).

Fixes:

  • Server: a request that ended late could clear the next request's status and drop its answer (#266, tonykee).
  • Server: a long unbroken run of Chinese/Japanese/Korean text (or a URL) no longer takes seconds to tokenize
    (#268, cha0yang: 1.6 s -> 5 ms for 2,000 characters).
  • Tool calls cut off by the end of the answer are reported as unfinished instead of complete, and their broken
    arguments are not passed on (#211, #231 by alphastorm).
  • Windows: when a verify window stalls, the watchdog now releases the GPU's waits before it ends the engine, so the
    GPU is not left "lost" until a reboot (#267, 1593914054).
  • Linux multi-GPU: the whole expert arena is pinned again (the 8 GiB cap was meant for Windows only; it cost
    prompt speed 3x on a 4090 + 3060; #253, bettercallcaleb). STRATA_ARENA_PIN_GIB=N sets a cap.
  • An MCP server that fails to start always says why.
  • A stall report names the prompt chunk and layer it stopped at (#251).
  • Setup: says why the low-RAM mode uses one GPU, and --low-ram off --yes installs without asking (#250); a
    checkout installs the engine of its own version, pinned model revisions and pinned Python packages (#214,
    alphastorm).
  • A CPU without AVX2 is refused at start with a clear message instead of crashing.
  • The cuBLAS setup failures are named instead of reported as one line (#240, lukmanfauzie).

New, all off unless you turn them on:

  • reasoning_budget_tokens in a request (or the config): caps the thinking, then lets the model answer (#123).
  • STRATA_GR_V3=1: the hyper-connection read in two kernels, +2-4% decode here, another summation order (#186,
    q8atnight).
  • STRATA_ARENA_PIN_GIB=auto (Windows): keeps the sliced pin below the GPU's shared-memory budget, for a PC where
    later allocations fail after pinning (#243, icegita: please try it).
  • "split_skip_if_fits": true (multi-GPU config): no split when the first card holds every expert.
  • STRATA_LOGPOS: per-position log-probabilities for quality tests (#235, enkynakamura).

AMD: setup installs a layer split over several AMD cards (--gpus 1,0); the RX 7800 XT (gfx1101, #254 by
jhohertz) and RX 9060 XT (gfx1200, #256 by Efeisot) are community-validated.

Also: a "Where things are stored" section in the README (#226); a community benchmark guide with RTX 5090
results (#233, hagope) and RTX 3090 / dual 3090 results (#234, mad9home); iq_parity runs from a fresh checkout
(#264, j-luwierski); an experimental build flag admits Volta (#236, cardonja; not in the ready-made engine).

Checked before the release:

  • Every merged change was gated on its own: byte-identical to 0.1.30 with a fixed cache on all four quants on this PC
    (Q2_0, IQ3_XXS, IQ3_S, the Coder), including the prompt path's internal state, in nine rounds; default-settings
    speed A/Bs on Q2_0 and IQ3_S (equal or faster). Two changes were reworked because they cost speed here: the
    Windows pin cap (#243, now opt-in) and a polling wait in the verify window (#267, now the watchdog's job); #257
    was taken out again (not faster on an AVX2-only CPU).

  • Real use at a 64K context on all five models (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, the Coder), through the server:

    • a fact found in a 57,479-token document;
    • the follow-up turn reusing the whole prompt;
    • a tool call;
    • a cancelled long prompt followed by a new request;
    • a sampled answer.

    None of them restarted the engine or ran low on VRAM.

  • The Unsloth Q4 model end to end on this PC (correct answers at 24, 32 and 40 GiB RAM budgets); the new format
    kernels against ggml-cpu and llama.cpp's dequantizers; the split-file loader and pack tests; the low-RAM mode
    reading the GGUF in place gives the same tokens and logits as experts.bin.

  • The server's tests (100) and setup's tests.

  • Linux (WSL, RTX 5070): builds, and Q2_0 is byte-identical to 0.1.30 with a fixed cache (10/10); the default settings pass.

  • AMD (RX 9070 XT and R9700): the HIP build with its tests (all pass but the two that need a model fixture or an
    AVX-512 CPU), the live server test on each card (R9700 82 tokens/s, 9070 XT 34 tokens/s), and the two-card split.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.31.

The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.

Strata v0.1.30

Choose a tag to compare

@Niko1221 Niko1221 released this 30 Sep 17:50

Short prompts up to 28% faster (the same answers), multi-GPU and AMD RDNA4 improvements, a server that can give the GPU back, and several opt-in features.

Faster short prompts. Prompts of 1K to 4K tokens read all experts from the stream from 1,024-token chunks
(0.1.29: 2,048): +17-28% prompt speed on this RTX 5070, with the same output (suggested in the issues).
STRATA_PREFILL_STREAM_MIN overrides the threshold.

Multi-GPU (#216, @gopinath87607). In a layer split:

  • Each card now keeps only its own layers' session state, which saved 0.93 GiB on a later card in the author's
    4-GPU test.
  • Each card lends the prompt path the tail of its own expert cache, instead of every card after the first
    holding back 1 GiB for the whole session. --prefill auto can then pick larger chunks (8,192 tokens instead of
    2,048 there).
  • The Monitor's expert cache counts every card (it showed only the first).
  • Snapshots and checkpoints follow the split: a card saves and restores only its own layers.

Single-GPU runs are unchanged (byte-identical).

AMD RDNA4 (#178, @doplxyz). The RX 9070 / 9070 XT and the Radeon AI PRO R9700 (gfx1201) are supported beside
the RX 7900 series: setup picks ROCm and the settings per card, and the engine checks at startup that it was built
for the card it runs on. Validated on doplxyz's machine with both cards. The live server tests pass on each.
Two AMD cards also run as one layer split from a hand-written config (setup installs one AMD card). It gives
exactly the output of one card when the experts sit in the same place; see docs/AMD_HIP.md.

The server can give the GPU back (#208, @bytethecookie).

  • --idle-unload 600 unloads the model after 600 s without requests; the next request loads it again.
  • POST /unload and POST /load do it on demand.
  • --min-free-vram-mib N loads only when that much VRAM is free (else 503), and --before-load "cmd" runs a
    command first (e.g. one that unloads another server's model).

Low-RAM mode, resident variant. When the experts don't fit the RAM but the ones the GPU doesn't hold do,
setup now copies exactly those into RAM at start (--resident-experts) instead of reading them from the file.
The same output. On a cold file cache the prompt went from 150 to 1,197 tok/s and answers were 26% faster.

Opt-in and experimental (off unless you turn them on):

  • Contexts past the trained 262K (#84, @j-luwierski): --rope-scaling yarn|linear and --rope-scale F,
    or setup's --context past 262144. Experimental: at 320K, facts planted at 300K were found at 10%, 50% and
    90% depth. See docs/DETAILS.md for the measured quality.
  • Several conversations kept at once (#189, @jeremiahritchey): --conversation-cache-mib 8192 --conversation-cache-slots 4 parks up to four conversations in RAM, so switching back doesn't re-read the
    prompt. Not yet with a layer split. With the option off, the engine is byte-identical to 0.1.29.
  • Greedy output independent of drafting (#152, reported by @clapbr): STRATA_IQ_MT_MIN=1 in the config's
    env makes the IQ models' CPU experts round the same whatever the verify window held. Costs 1-3% decode on
    IQ3_S (AVX-512); the default is unchanged.
  • Coupled draft sampling: STRATA_SPEC_COUPLED=1 samples drafts and the verifier from one random stream.
    No clear gain in our tests, so it stays off.
  • Docker (#96, @djmaze): docker build -t strata . builds the engine in the image. The server is PID 1 and
    stops cleanly on SIGTERM; KV, GPUs and low-RAM mode are set with environment variables. See the README.
  • Pascal cards (#124, @ruibeikaa): the engine builds and runs on compute capability 6.0 (P100) with exact
    fallbacks for __dp4a and __nanosleep. Not in the ready-made engine; compile it yourself. Not tested here
    (no Pascal card).
  • Shared expert arena (#129, @rhgo1749): Linux: --shared-expert-arena FILE in /dev/shm lets several
    engines on one machine share the RAM copy of the experts.

Checked before the release:

  • Byte-identical to 0.1.29 with a fixed cache on all four quants on this PC (Q2_0, IQ3_XXS, IQ3_S, the Coder),
    including the prompt path's internal state, with the same decode speed.

  • Real use at a 64K context on all five models (Q2_0, IQ2_XS, IQ3_XXS, IQ3_S, the Coder), through the server:

    • a fact found in a 57,566-token document;
    • the follow-up turn reusing the whole prompt;
    • a tool call;
    • a cancelled long prompt followed by a new request;
    • a sampled answer.

    None of them restarted the engine or ran low on VRAM. Prompt speed and free VRAM were the same as 0.1.29.

  • The conversation cache (#189): with it off, the same output as 0.1.29 up to 128K; with it on, the reuse,
    memory-pressure and admission tests pass.

  • Rope scaling (#84): the parity tests, and needles at 300K in a 320K context.

  • The server's tests (80) and the CPU expert parity test (0 failures; with STRATA_IQ_MT_MIN=1 no row differs between
    one token and a group).

  • Linux (WSL, RTX 5070): Q2_0 byte-identical to 0.1.29 with a fixed cache (10/10), and the default settings pass.

  • AMD (RX 9070 XT and R9700): the HIP build with its tests (45 of 47; the 2 others need a model fixture or an
    AVX-512 CPU), the live server test on each card, and the two-card layer split.

Updating: get the latest files (git pull, or download and unzip anywhere), then run START-HERE.bat (Linux:
./setup.sh). Setup installs engine 0.1.30.

The ready-made Strata engine for Windows (RTX 20 / 30 / 40 / 50: sm_75, sm_86, sm_89, sm_120 + PTX), CUDA 13.0.

You don't need to download this yourself: START-HERE.bat fetches it (and NVIDIA's cuBLAS from pip), so no compiler
or CUDA Toolkit is needed, only an NVIDIA driver 580 or newer.

Contents: strata.exe (the engine), strata-vision.exe (the optional image encoder), BUILD.json.