Repository navigation
Releases: mw00/project-maya
Release list
Project Maya v1.0.33
Several conversations at once on a layer split, a short request no longer waits for a long answer, and fixes from your reports: a start that failed on Windows, a speed that changed between restarts, ZFS, and --calibrate now tunes the disk reads too.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- Several conversations at once on a layer split (#94 by @0xPreDa): with
STRATA_GLM_SEQS=<n>on two GPUs or more, n conversations decode together, each card of the split working on a different one instead of waiting for the others. On 4x RTX 4090, four conversations at once give 225 tok/s in total. It's off by default; one GPU gains nothing from it. - A short request no longer waits for a whole long answer (#93 by @0xPreDa): requests are served in arrival order, and with
STRATA_FAIR_SLICE_S=<s>a long answer that has been decoding for s seconds while another request waits lets it in, then continues from where it was. On 2x V100, a short question asked during a 700-token answer got its first word after 10 s instead of 60. --calibratetunes the disk reads too (#99, from #67): when part of the model is read from the SSD while it answers, the tuning also tries reading each expert in fewer, larger pieces, and keeps a size only when it is more than 3% faster. A Windows laptop on an Intel RST RAID decoded 15% faster with 4 pieces than with the default 8. A Linux NVMe stays fastest at 8, so nothing changes there.- Fixes from your reports (#99):
- A start that failed right after the RAM tier on Windows (
pinned disk staging did not allocate, #67): the small pinned buffers are now set up before the tier, so a cap on pinned memory makes the tier slightly smaller instead of stopping the start. - A speed that changed between restarts (#56): the CPU lane's timing at start is now taken over half a second, and keeps the fastest round, so a burst of other work at that moment no longer makes the whole session slower.
- ZFS: the memory ZFS's cache can give back counts as free RAM when the RAM tier is sized (#89).
- The prefetch says its settings in the engine log when it is on (#6).
- A start that failed right after the RAM tier on Windows (
Checked
On 1x and 2x Tesla V100:
- the build and the GLM parity tests;
- identical greedy and sampled tokens to v1.0.32, and the same answers across a conversation's turns;
- decode the same within run-to-run variation;
- several conversations at once writing the same tokens as one at a time;
- the server end to end and a stress run.
On Windows: the engine compiles. Also: the GitHub checks pass.
Project Maya v1.0.32
A new model to download, Maya-M-Derisked. A split's GPUs now load at once, three GPUs or more no longer pin more RAM than the PC has, and the Ryzen AI Max APUs decode faster.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- Maya-M-Derisked (#97): Maya-M with a directional weight modification by Blackfrost_AI that reduces blanket refusals.
- It is not a new quant: the same IQ2_S files, tensors, MTP draft block and size as Maya-M (116 GB), with some of its weights changed.
- It is experimental, and it lives in a repo of its own.
- Set it up with
./setup.sh --setup --model Maya-M-Derisked(Windows:START-MAYA.bat --setup --model Maya-M-Derisked), or pick it in the setup's model menu. - It reads pictures with Maya's vision files, like the other models.
- A split's GPUs load at once (#80 by @ksanislo): a thread per GPU, instead of one GPU after another.
- On 4x Tesla T4 with Maya-L, start to the first token went from 142 to 96 s.
- The GPUs plan their RAM tiers in turn, so each gets the same tier as before; then they pin and warm them at the same time.
STRATA_GLM_PARALLEL_LOAD=0keeps the old order.
- The CPU lane's timing, kept across starts (#81 by @ksanislo): with
STRATA_GLM_CPU_CAL=<file>, the timing at start (~1.7 s a GPU) is written once and read by later starts. A new build, another card or thread count times again, and a timing that decode finds off is dropped. - Three GPUs or more on Linux no longer freeze the desktop at start (#85, issue #77):
- the later GPUs' prompt buffers are set aside when the RAM tier is sized;
- each later GPU measures the free RAM again;
- a warning says when what is left is short.
- Faster decode on Ryzen AI Max APUs (#95, from the Gorgon Halo work): the decode's dense projections run in one RDNA3 kernel that keeps two blocks' loads in flight, with bit-for-bit the same results.
- Radeon 8065S with Maya-S: decode +3.1%.
- On by default on gfx115x.
STRATA_GLM_MV_RDNA=1turns it on for other gfx11 cards;0turns it off.
- The setup accepts a first shard of metadata only (#85, issue #83): a GGUF whose first shard holds only the tokenizer and settings, as unsloth's do, no longer stops the setup as incomplete.
- For cache studies (by @sociolog):
The README has rows for the new settings.
Checked
On 1x and 2x Tesla V100:
- the build and the GLM parity tests;
- identical greedy tokens to v1.0.31, the parallel and the one-by-one load alike;
- the same RAM tiers as v1.0.31, and decode the same within run-to-run variation;
- the server end to end: answers, pictures, context reloads, a clean stop;
- a stress run.
Maya-M-Derisked was downloaded and checked against its published sha256. Its answers in several languages, greedy and sampled, end on their own, with no loops and no stray characters.
On a Radeon 8065S (Gorgon Halo): the HIP build, and the new kernel bit-exact in all 333 parity cases and in Maya-S's greedy tokens.
On Windows: the engine compiles. Also: the GitHub checks pass.
Project Maya v1.0.31
The expert prefetch can now pay on a RAM-bound split: three new settings choose when its copy starts, how many GPU blocks it takes and which predictions it copies.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
-
The expert prefetch, tunable (#84 by @sociolog):
STRATA_GLM_PREFETCH_Ncopies the next layer's predicted experts that are not in VRAM into its spare slots while a layer computes. As it was, it cost more than it saved on a split whose experts mostly come from RAM. Three settings for it:STRATA_GLM_PREFETCH_AT=fetch: start the copy after the layer's own fetch, when the PCIe link is free (cpu: after its CPU-lane answer);STRATA_GLM_PREFETCH_BLOCKS=<n>: how many GPU blocks the copy takes (half the SMs before);STRATA_GLM_PREFETCH_RANK=<n>: copy only the prediction's first n guesses, the ones that are almost always right.
On an RTX 3090 + 3060 with Maya-M at 21K context, decode (writing the answer) went from 19.64 to 20.22 tok/s with
STRATA_GLM_PREFETCH_N=1 AT=fetch BLOCKS=4 RANK=2, and +4.7% over a 40-turn agent-like conversation.The prefetch stays off by default, and nothing else changes. The README has a row for it.
Checked
On 1x and 2x Tesla V100:
- the build and the GLM parity tests;
- identical greedy tokens with the CPU lane off;
- decode the same within run-to-run variation (2 GPUs: 30.2 / 30.4 against 30.4 / 30.1 tok/s);
- a run with the new settings on, which completes cleanly.
Also: the GitHub checks pass.
Project Maya v1.0.30
AMD GPUs now set up on Windows too, Windows reads its experts from the SSD much faster, and Maya's logs and installer use its own name.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- AMD on Windows, Strix Halo / Gorgon Halo included (experimental) (#69):
START-MAYA.bat --backend hipsets Maya up on Windows 10/11 as./maya.sh --backend hipdoes on Linux. That covers Ryzen AI Max 300 / 400 (Radeon 8050S / 8060S / 8065S) and the RX 7900 XT / XTX and RX 9070 / R9700.- Setup: finds the GPUs, uses AMD's HIP SDK (or offers AMD's ROCm SDK in
.venv), and builds the engine with Visual Studio 2022 or 2026. - The engine: handles a Windows APU's memory and Windows' way of submitting GPU work.
- Checked on a Ryzen AI Max+ PRO 495: setup builds and packs Maya-L, every expert fits on the GPU, and all 15 HIP checks pass.
- Setup: finds the GPUs, uses AMD's HIP SDK (or offers AMD's ROCm SDK in
- Windows reads experts from the SSD much faster (#73 by @klvnblst):
- The model files are unmapped after loading: while they stay mapped, Windows slows every direct read of them. On the RTX 5090 Laptop that reported it, decode is about 3.5x as fast.
- CPUs without hyper-threading leave a third of their cores free for the disk readers.
- The CPU lane times itself on experts read from RAM (#70 by @merbanan): it used to time experts sitting in the CPU's cache and overrate the CPU. +6% decode on a Ryzen 9 7950X with a Tesla V100, and the same plan at every start.
- Maya's own name (#74): the logs say
[maya]andmaya ..., and the installer no longer shows Strata's name. Settings, flags and the API are unchanged, and the README keeps the credit. - Windows setup in the README, step by step (#71 by @handmade0octopus): the prerequisites, copy-paste commands, and page-file advice that matches
--check. - Fix: a restart no longer mistakes the dashboard's saved settings for a model's config.
Checked
On 1x and 2x Tesla V100:
- the builds and parity tests, and identical greedy tokens with the CPU lane off;
- decode (writing the answer) and prefill (reading the prompt) the same within run-to-run variation;
- the server end to end, the setup screen in a real terminal, and the logs saying
[maya].
Also: the engine compiles on Windows, AMD on Windows was run on the author's Ryzen AI Max+ PRO 495, and the GitHub checks pass.
Project Maya v1.0.29
Documentation only: plainer wording about the model files' architecture name.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). Nothing is downloaded again.
What's new
- The README and the changelog describe v1.0.28's change to the model files' architecture name more plainly. The code is unchanged.
Project Maya v1.0.28
The model files now name their architecture with the standard GGUF name, glm5-next.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). Nothing is downloaded again.
What's new
- The model files say
glm5-next. They saidglm5next, an early spelling copied from the file they were made from, which standard GGUF tools don't recognize.- On Hugging Face, each model's first file is replaced by one with only its header changed. Installed models and their packs keep working, and nothing has to be downloaded again.
python tools/gguf_fix_arch.py <the model's first .gguf> --in-placerenames a file you downloaded earlier, rewriting only the header. Maya reads either name.- On an older Maya version, a model downloaded after this change fails its checksum check, so update first.
Checked
- Maya's output is the same with the new files;
- the new files are the old ones byte for byte after the header;
- the GitHub checks.
Project Maya v1.0.27
Maya gets a screen of its own in the terminal, the crash after a long prompt is fixed, a stuck prompt now ends with an error instead of waiting forever, and Maya builds on Windows again.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- A screen of its own in the terminal (#58 by @needmorevram):
- The setup's steps are tabs, the questions are arrow-key menus, and the compile and the download show their progress.
- Maya then runs on the same screen: a loading view while the experts warm up, the Monitor's numbers, and its log.
- Nothing is downloaded for it: its library ships with Maya, byte-identical to PyPI's.
--plain(or a pipe, or a service) keeps the plain text as before.
- No more "illegal memory access" after a long prompt, and no more 0xC0000005 on Windows / AMD (#39 by @boxwrench). It fixes three races in the expert tiers that NVIDIA and AMD share:
- the expert tables' update buffers were rewritten while their copy to the GPU was still in flight;
- a background expert move could land in a VRAM slot a prompt was borrowing;
- a route of NaN scores indexed the expert tables. Now that request fails with an error and the engine keeps running.
- Confirmed on an RTX 3090 + 3060 (#50: the 8K prompt after decoding no longer crashes) and on Windows with an RX 7900 XTX (#53: 5 crashes in 8 sessions before, 0 in 8 now).
- A stuck prompt ends (#40): when a request finishes no prompt layer and no token for 3 minutes, the engine stops with an error that says where it was stuck. On Windows it also writes
strata-stall-<pid>.dmp. The next request starts the engine again. Before, a prompt that stalled with the GPU idle waited forever.STRATA_WATCHDOG_Ssets the time;0turns it off. - Windows builds again (#54 by @jerem91150; #55 by @noahark had the same fix).
- Thinking off stays off (#54): the answer no longer lands inside an empty thinking block. This applies to every GPU.
- AMD:
- a Windows build script (#54);
- RDNA4 prefill (reading the prompt) +8-10% on R9700 / RX 9070 (#38 by @boxwrench).
- One GPU, prefill (reading the prompt) about +1% (#61 by @boxwrench): the last layer's unused outputs are skipped, and the output is the same.
- Robustness (#62 by @needmorevram):
- F16, BF16, Q2_K, Q4_1 and Q5_1 GGUFs read the right embedding rows. The published Maya models were not affected.
- Bad image files are refused instead of ending the engine.
- Switch chats while an answer is being written (#66 by @needmorevram): the other chat opens, and the answer carries on in its own chat.
- A setup for another context keeps your
--calibratetuning (#57). Before, it fell back to the engine's defaults.
Checked
On 1x and 2x Tesla V100, and on Windows:
- the builds and parity tests, including the engine's Windows build;
- identical greedy tokens with the CPU lane off on one and two GPUs, and with the graphs on;
- #50's sequence (every RAM-tier expert on the CPU, then long prompts) on one and two GPUs;
- the watchdog: a forced stall ends with an error and the next request is answered, and it never fires on normal work;
- decode (writing the answer) and prefill (reading the prompt) the same, prefill +1% on one GPU;
- the server end to end, #58's screen in a real terminal, and #66 in a browser against the running model;
- the GitHub checks.
Project Maya v1.0.26
Maya-L runs at 16.4 tokens/s on an RTX 4090 with 192 GB of RAM under Windows 11 - a user's report, now in the README.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). Documentation only; nothing is downloaded again.
What's new
- A new speed report (#63 by @npc97): Maya-L on an RTX 4090 24 GB, a Ryzen 9 9950X3D and 192 GB of DDR5 under Windows 11.
- Every expert fits in VRAM or RAM, so nothing is read from the SSD.
- Decode (writing the answer): 16.4 tokens/s. Prefill (reading the prompt): 1371 tokens/s on an 8K-token prompt.
- It's in the README's user table, and the Windows section now says Maya runs on NVIDIA under Windows, with CUDA 13 working for RTX 20 and newer.
- CPU pinning is Linux-only (#64 by @needmorevram): the README now says
STRATA_GLM_CPU_PINdoes nothing on Windows or with one GPU.
Project Maya v1.0.25
Multi-GPU rigs of three cards or more now use all of the CPU (+19% decode on a 9-GPU split), --calibrate tunes on realistic text without touching your profile, and there are opt-in CUDA graphs for one GPU.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- One CPU pool for three GPUs or more (#60 by @needmorevram):
- A token visits the cards one after another, so giving each card its own small share of the CPU left most of it idle. One pool now serves them all.
- Maya-L on 9 GPUs: decode 25.7 -> 30.7 tokens/s, an 8K prompt 1265 -> 1466 tokens/s.
- One and two GPUs are unchanged.
STRATA_GLM_CPU_SHARED=1shares the pool on two GPUs too, which was faster on a single-CPU 2x V100 box (30.0 vs 28.5 tokens/s).
--calibrateon realistic text (#60):- It now answers fresh prompts for every measurement, so it tunes for real chats rather than three repeated answers.
- It works on a copy of your expert usage profile, so tuning no longer reorders the experts your next start loads first.
- CUDA graphs for one GPU, opt-in (#59 by @handmade0octopus):
STRATA_GLM_KDA_GRAPH=1.- The output is bit-identical.
- RTX 4090D: +2.35% decode. Tesla V100: the same speed.
Checked
On 1x and 2x Tesla V100:
- the builds and parity tests;
- identical greedy tokens at the defaults, and with the graphs on;
- #59's real-model test: 192 bit-for-bit comparisons;
- decode and prefill (reading the prompt) unchanged at the defaults;
- a full
--calibraterun that left the usage file byte-for-byte unchanged; - the GitHub checks.
Project Maya v1.0.24
Better drafts for speculative decode on two GPUs or more: more drafts accepted, never slower.
Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.
What's new
- Better drafts (#48 by @sociolog): when the model's draft block is missing an expert in VRAM, it now fetches the missing ones among its 2 most important experts instead of skipping them all.
- 2x Tesla V100: 82-83% of drafts accepted instead of 76-81%, and decode at 128K context 28.0-28.2 tokens/s instead of 27.0-28.1.
- RTX 3090 + 3060: 19.5 -> 20.6 tokens/s.
- It never measured slower, and the answers are the model's own.
STRATA_GLM_MTP_KEEP=<n>changes how many:0= the old behaviour.
Checked
- On 2x Tesla V100: the build, five draft settings over two rounds, and identical greedy tokens with the CPU lane off.
- The GitHub checks pass.