Repository navigation
Faster decode and prompts on NVIDIA, AMD and Intel, plus the fixes since 0.1.42. Decode is up to 38% faster on a 16 GB Radeon, 5-11% on RTX 30-series and Tesla P100 under Linux, up to 19% on four Radeons, and 14-21% on an Arc Pro B70, whose prompts read about twice as fast. Windows streaming from GGUF files decodes 1.35-1.85x faster. Answers are the same as 0.1.42 (byte-identical on every machine we tested) except where a default below says otherwise. Update with UPDATE.bat (Linux: ./update.sh); setup replaces the engine with 0.1.43.
Speed
All numbers are our own interleaved A/Bs against the previous build: 5 pairs (prompts on the B70: 3), medians, the same output tokens in both arms unless a line says so.
| Card | Model | Decode | Prompts |
|---|---|---|---|
| Radeon AI PRO R9700 with a 16 GB card's cache | IQ3_XXS | +21% to +38% | 4K +7.5%, 32K +15.6% |
| Radeon AI PRO R9700 32 GB | Q2_0 / IQ3_XXS | +2% to +10% | +12% to +14% |
| Radeon AI PRO R9700 32 GB | IQ3_S | +6% to +14% | 32K +17%, 64K +27% |
| 4x R9700 (cards at 225 W) | IQ3_S | story +3.8%, code +12.6%, a 4K-token document +18.7% | |
| RTX A4000 (Linux) | Q2_0 | story +6.6%, code +5.5%, a 6K document +5.1%; the new Q3_K kernel adds +2.0-2.7% on top | flat |
| Tesla P100 (Linux) | IQ3_XXS | story +7.0%, code +10.7%, a 6K document +8.0% | flat |
| RTX 5070 (Windows) | Q2_0 / IQ3_XXS | the same as 0.1.42 by default (the two NVIDIA decode changes below are opt-in on Windows) | flat |
| Arc Pro B70 | IQ3_S | +14% to +21% (new expert ranking) | 4K 8.5 -> 3.3 s, 16K 27.4 -> 13.4 s (128K config) |
- AMD Radeon (gfx1200 / gfx1201, RDNA4). The two decode defaults NVIDIA got in 0.1.42 now apply here too, measured on the R9700: the route tail skip (
STRATA_ROUTE_TAIL_SKIP=0turns it off) and the measured PCIe share (STRATA_PCIE_FRAC_DEFAULT=old). Like on NVIDIA they change the answers slightly. The fused prompt kernels are the default for IQ3_S and IQ3_XXS (STRATA_PF_FUSED=0turns them off), and long prompts read in 16K chunks on Linux (STRATA_GFX12_CHUNK=0). Fewer and fused graph nodes in the decode window give another 1.6-2.0%, with the same answers. On a layer split the tail skip stays off from 60% of the experts in VRAM. RDNA3 (gfx11) and older keep their 0.1.42 behaviour: we have no such card to measure on. - 3 or more Radeon cards.
--pipeline-windows 2is on by default for a gfx12 layer split of 3 or more cards in--serve(4x R9700: code +7-12%, 4K prompt +9-11%, 32K +8-19%). The pipelined windows hand the CPU differently sized batches of expert rows, so the CPU-side sums can round differently and answers can differ slightly from the serial order (on 4x R9700 with default settings 10 of 10 test prompts matched the serial order over 160 tokens). With 2 cards it stays opt-in: there it loses about 3% on code.--pipeline-windows 0orSTRATA_PIPELINE_AUTO=0turns it off. - NVIDIA. On Linux the adaptive tier's copies overlap the next decode window (
STRATA_ADAPT_OVERLAP=0turns it off; on Windows it is opt-in with=1, since an RTX 5070 there decoded Q2_0 2.6% slower with it). The Tesla P100 gets its own kernel table (STRATA_SM60_TABLE=0). RTX 30-series cards run the Q2_0 pack's Q3_K weights on the interleaved kernels (STRATA_MMVQ_IL_Q3K=0). On Linux the canonical Q2_0 pack uses the measured PCIe share too (on Windows withSTRATA_PCIE_AUTO_CANON=1). All of these give the same answers. The opt-inSTRATA_ROUTE_RESIDENTruns one warp per token on NVIDIA and AMD (#1737, aly8246; measured on an R9700): +2% decode with a 32 GB cache, +6-8% with a 16 GB card's, same answers as before. - Intel Arc Pro B70 (Battlemage). The prompt path's expert copies run on the card's copy engine (
STRATA_COPY_ENGINE=0), the host KV copy is written 4 bytes at a time, the A770's GEMM staging is skipped, and the prompt loan's experts refill from pinned memory. A new expert ranking (data/expert-profile-sycl.bin, written by setup for cards on the xe driver) holds 71% of a chat's experts instead of 60%. Same answers.
Fixes
-
Windows, GGUF in place: experts are read in parallel again (reported on X). With the model's GGUF shards left mapped in memory, NTFS runs the unbuffered reads of those files one at a time, so a low-memory setup decoded at 2.6 tok/s. The engine now closes every GGUF shard's view once the file tier reads unbuffered.
STRATA_KEEP_MAPPING=1brings the old behaviour back. Measured on an RTX 5070 (Windows, Ryzen 5 7600, 64 GB) with IQ3_XXS read from its GGUF files, 5 interleaved pairs each:Experts held in RAM Story Code Short prompt 22 GiB 18.9 -> 25.5 tok/s (1.35x) 13.5 -> 21.1 (1.56x) 13.3 -> 18.2 (1.37x) 12 GiB 9.6 -> 16.0 (1.67x) 6.5 -> 11.6 (1.78x) 4.6 -> 8.5 (1.85x) -
Dual-socket and multi-node Linux hosts (reported on X). The page-locked expert copy is judged by the emptiest NUMA node too, which only makes it load in steps; it never refuses and never caps. The "not enough RAM" start failure now says how much of the RAM is file cache and prints the fix. New and off by default:
STRATA_DROP_CACHE_ON_EXIT=1makesstrata servegive the model files' cache back when it stops. -
The route tail skip stays off when nearly every expert is in VRAM (#1884, christopherrobertbrooks-tech). On a Tesla V100 32 GB holding 94.8% of the experts 0.1.42 decoded 1-2% slower. The default now stays off from 90% of the experts in the GPU cache (60% on an AMD layer split), and the start-up log says so; those cards get 0.1.41's answers back.
STRATA_ROUTE_TAIL_SKIP=7turns it on anyway;STRATA_TAIL_SKIP_MAX_CACHED=<percent>moves the line. -
Timing numbers in the API (#1869, #1870, sdjger-xiaoniu). A request that went from a batch slot back to the single-request path now reports the whole request, and
timings.prompt_ncounts what the engine actually read, with the rest incache_n. -
A text prompt right after a picture could fault on RX 7000 cards (#1712, mantovaniluca91, who traced it). The prompt chunk's step records and token ids now go up from pinned memory, and a failed upload is reported instead of ignored.
-
Answers of only
!are caught (#879, #1815; the guard is #1892 by lask3802). When a decode window's logits come out all NaN, nothing is sent, the caches are dropped, the prompt is read again from token 0 and the server retries once; the engine logsnon-finite logits (#879). Without a NaN the answers are byte-identical, and it costs at most 0.2% decode.STRATA_NAN_GUARD=0turns it off. The cause on the reporters' machines is still open: we could not make it fire here (1,870 requests on an R9700).
New
BENCH.bat/./bench.sh(python tools/strata_bench.py). Runs a fixed suite against your own install in about 3 to 10 minutes: short chat, 4K and 16K prompts, a five-turn agent session, optional--clients Nand--long. It writes a report in the community-benchmark format, scrubs home folders, names and addresses, and--compareputs two runs side by side. Setup offers it at the end of an interactive install (default no). Seedocs/COMMUNITY_BENCHMARKS.md.
Strata is free and open source. If it runs well on your PC, a coffee keeps the work on it going:
