Skip to content

Releases: mw00/project-maya

Project Maya v1.0.33

Choose a tag to compare

@mw00 mw00 released this 11 Oct 05:03

Several conversations at once on a layer split, a short request no longer waits for a long answer, and fixes from your reports: a start that failed on Windows, a speed that changed between restarts, ZFS, and --calibrate now tunes the disk reads too.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • Several conversations at once on a layer split (#94 by @0xPreDa): with STRATA_GLM_SEQS=<n> on two GPUs or more, n conversations decode together, each card of the split working on a different one instead of waiting for the others. On 4x RTX 4090, four conversations at once give 225 tok/s in total. It's off by default; one GPU gains nothing from it.
  • A short request no longer waits for a whole long answer (#93 by @0xPreDa): requests are served in arrival order, and with STRATA_FAIR_SLICE_S=<s> a long answer that has been decoding for s seconds while another request waits lets it in, then continues from where it was. On 2x V100, a short question asked during a 700-token answer got its first word after 10 s instead of 60.
  • --calibrate tunes the disk reads too (#99, from #67): when part of the model is read from the SSD while it answers, the tuning also tries reading each expert in fewer, larger pieces, and keeps a size only when it is more than 3% faster. A Windows laptop on an Intel RST RAID decoded 15% faster with 4 pieces than with the default 8. A Linux NVMe stays fastest at 8, so nothing changes there.
  • Fixes from your reports (#99):
    • A start that failed right after the RAM tier on Windows (pinned disk staging did not allocate, #67): the small pinned buffers are now set up before the tier, so a cap on pinned memory makes the tier slightly smaller instead of stopping the start.
    • A speed that changed between restarts (#56): the CPU lane's timing at start is now taken over half a second, and keeps the fastest round, so a burst of other work at that moment no longer makes the whole session slower.
    • ZFS: the memory ZFS's cache can give back counts as free RAM when the RAM tier is sized (#89).
    • The prefetch says its settings in the engine log when it is on (#6).

Checked

On 1x and 2x Tesla V100:

  • the build and the GLM parity tests;
  • identical greedy and sampled tokens to v1.0.32, and the same answers across a conversation's turns;
  • decode the same within run-to-run variation;
  • several conversations at once writing the same tokens as one at a time;
  • the server end to end and a stress run.

On Windows: the engine compiles. Also: the GitHub checks pass.

Project Maya v1.0.32

Choose a tag to compare

@mw00 mw00 released this 10 Oct 23:56

A new model to download, Maya-M-Derisked. A split's GPUs now load at once, three GPUs or more no longer pin more RAM than the PC has, and the Ryzen AI Max APUs decode faster.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • Maya-M-Derisked (#97): Maya-M with a directional weight modification by Blackfrost_AI that reduces blanket refusals.
    • It is not a new quant: the same IQ2_S files, tensors, MTP draft block and size as Maya-M (116 GB), with some of its weights changed.
    • It is experimental, and it lives in a repo of its own.
    • Set it up with ./setup.sh --setup --model Maya-M-Derisked (Windows: START-MAYA.bat --setup --model Maya-M-Derisked), or pick it in the setup's model menu.
    • It reads pictures with Maya's vision files, like the other models.
  • A split's GPUs load at once (#80 by @ksanislo): a thread per GPU, instead of one GPU after another.
    • On 4x Tesla T4 with Maya-L, start to the first token went from 142 to 96 s.
    • The GPUs plan their RAM tiers in turn, so each gets the same tier as before; then they pin and warm them at the same time.
    • STRATA_GLM_PARALLEL_LOAD=0 keeps the old order.
  • The CPU lane's timing, kept across starts (#81 by @ksanislo): with STRATA_GLM_CPU_CAL=<file>, the timing at start (~1.7 s a GPU) is written once and read by later starts. A new build, another card or thread count times again, and a timing that decode finds off is dropped.
  • Three GPUs or more on Linux no longer freeze the desktop at start (#85, issue #77):
    • the later GPUs' prompt buffers are set aside when the RAM tier is sized;
    • each later GPU measures the free RAM again;
    • a warning says when what is left is short.
  • Faster decode on Ryzen AI Max APUs (#95, from the Gorgon Halo work): the decode's dense projections run in one RDNA3 kernel that keeps two blocks' loads in flight, with bit-for-bit the same results.
    • Radeon 8065S with Maya-S: decode +3.1%.
    • On by default on gfx115x. STRATA_GLM_MV_RDNA=1 turns it on for other gfx11 cards; 0 turns it off.
  • The setup accepts a first shard of metadata only (#85, issue #83): a GGUF whose first shard holds only the tokenizer and settings, as unsloth's do, no longer stops the setup as incomplete.
  • For cache studies (by @sociolog):
    • STRATA_GLM_VRAM_EVICT=lru refills a layer's spares by evicting its least recently used expert (#78);
    • STRATA_GLM_ROUTE_LOG writes each route's 16 near misses (#90);
    • tools/glm_tier_replay.py replays a route log through a model of the VRAM tier's rules (#91).

The README has rows for the new settings.

Checked

On 1x and 2x Tesla V100:

  • the build and the GLM parity tests;
  • identical greedy tokens to v1.0.31, the parallel and the one-by-one load alike;
  • the same RAM tiers as v1.0.31, and decode the same within run-to-run variation;
  • the server end to end: answers, pictures, context reloads, a clean stop;
  • a stress run.

Maya-M-Derisked was downloaded and checked against its published sha256. Its answers in several languages, greedy and sampled, end on their own, with no loops and no stray characters.

On a Radeon 8065S (Gorgon Halo): the HIP build, and the new kernel bit-exact in all 333 parity cases and in Maya-S's greedy tokens.

On Windows: the engine compiles. Also: the GitHub checks pass.

Project Maya v1.0.31

Choose a tag to compare

@mw00 mw00 released this 10 Oct 11:56

The expert prefetch can now pay on a RAM-bound split: three new settings choose when its copy starts, how many GPU blocks it takes and which predictions it copies.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • The expert prefetch, tunable (#84 by @sociolog): STRATA_GLM_PREFETCH_N copies the next layer's predicted experts that are not in VRAM into its spare slots while a layer computes. As it was, it cost more than it saved on a split whose experts mostly come from RAM. Three settings for it:

    • STRATA_GLM_PREFETCH_AT=fetch: start the copy after the layer's own fetch, when the PCIe link is free (cpu: after its CPU-lane answer);
    • STRATA_GLM_PREFETCH_BLOCKS=<n>: how many GPU blocks the copy takes (half the SMs before);
    • STRATA_GLM_PREFETCH_RANK=<n>: copy only the prediction's first n guesses, the ones that are almost always right.

    On an RTX 3090 + 3060 with Maya-M at 21K context, decode (writing the answer) went from 19.64 to 20.22 tok/s with STRATA_GLM_PREFETCH_N=1 AT=fetch BLOCKS=4 RANK=2, and +4.7% over a 40-turn agent-like conversation.

    The prefetch stays off by default, and nothing else changes. The README has a row for it.

Checked

On 1x and 2x Tesla V100:

  • the build and the GLM parity tests;
  • identical greedy tokens with the CPU lane off;
  • decode the same within run-to-run variation (2 GPUs: 30.2 / 30.4 against 30.4 / 30.1 tok/s);
  • a run with the new settings on, which completes cleanly.

Also: the GitHub checks pass.

Project Maya v1.0.30

Choose a tag to compare

@mw00 mw00 released this 10 Oct 00:41

AMD GPUs now set up on Windows too, Windows reads its experts from the SSD much faster, and Maya's logs and installer use its own name.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • AMD on Windows, Strix Halo / Gorgon Halo included (experimental) (#69): START-MAYA.bat --backend hip sets Maya up on Windows 10/11 as ./maya.sh --backend hip does on Linux. That covers Ryzen AI Max 300 / 400 (Radeon 8050S / 8060S / 8065S) and the RX 7900 XT / XTX and RX 9070 / R9700.
    • Setup: finds the GPUs, uses AMD's HIP SDK (or offers AMD's ROCm SDK in .venv), and builds the engine with Visual Studio 2022 or 2026.
    • The engine: handles a Windows APU's memory and Windows' way of submitting GPU work.
    • Checked on a Ryzen AI Max+ PRO 495: setup builds and packs Maya-L, every expert fits on the GPU, and all 15 HIP checks pass.
  • Windows reads experts from the SSD much faster (#73 by @klvnblst):
    • The model files are unmapped after loading: while they stay mapped, Windows slows every direct read of them. On the RTX 5090 Laptop that reported it, decode is about 3.5x as fast.
    • CPUs without hyper-threading leave a third of their cores free for the disk readers.
  • The CPU lane times itself on experts read from RAM (#70 by @merbanan): it used to time experts sitting in the CPU's cache and overrate the CPU. +6% decode on a Ryzen 9 7950X with a Tesla V100, and the same plan at every start.
  • Maya's own name (#74): the logs say [maya] and maya ..., and the installer no longer shows Strata's name. Settings, flags and the API are unchanged, and the README keeps the credit.
  • Windows setup in the README, step by step (#71 by @handmade0octopus): the prerequisites, copy-paste commands, and page-file advice that matches --check.
  • Fix: a restart no longer mistakes the dashboard's saved settings for a model's config.

Checked

On 1x and 2x Tesla V100:

  • the builds and parity tests, and identical greedy tokens with the CPU lane off;
  • decode (writing the answer) and prefill (reading the prompt) the same within run-to-run variation;
  • the server end to end, the setup screen in a real terminal, and the logs saying [maya].

Also: the engine compiles on Windows, AMD on Windows was run on the author's Ryzen AI Max+ PRO 495, and the GitHub checks pass.

Project Maya v1.0.29

Choose a tag to compare

@mw00 mw00 released this 09 Oct 22:17

Documentation only: plainer wording about the model files' architecture name.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). Nothing is downloaded again.

What's new

  • The README and the changelog describe v1.0.28's change to the model files' architecture name more plainly. The code is unchanged.

Project Maya v1.0.28

Choose a tag to compare

@mw00 mw00 released this 09 Oct 22:00

The model files now name their architecture with the standard GGUF name, glm5-next.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). Nothing is downloaded again.

What's new

  • The model files say glm5-next. They said glm5next, an early spelling copied from the file they were made from, which standard GGUF tools don't recognize.
    • On Hugging Face, each model's first file is replaced by one with only its header changed. Installed models and their packs keep working, and nothing has to be downloaded again.
    • python tools/gguf_fix_arch.py <the model's first .gguf> --in-place renames a file you downloaded earlier, rewriting only the header. Maya reads either name.
    • On an older Maya version, a model downloaded after this change fails its checksum check, so update first.

Checked

  • Maya's output is the same with the new files;
  • the new files are the old ones byte for byte after the header;
  • the GitHub checks.

Project Maya v1.0.27

Choose a tag to compare

@mw00 mw00 released this 09 Oct 19:41

Maya gets a screen of its own in the terminal, the crash after a long prompt is fixed, a stuck prompt now ends with an error instead of waiting forever, and Maya builds on Windows again.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • A screen of its own in the terminal (#58 by @needmorevram):
    • The setup's steps are tabs, the questions are arrow-key menus, and the compile and the download show their progress.
    • Maya then runs on the same screen: a loading view while the experts warm up, the Monitor's numbers, and its log.
    • Nothing is downloaded for it: its library ships with Maya, byte-identical to PyPI's.
    • --plain (or a pipe, or a service) keeps the plain text as before.
  • No more "illegal memory access" after a long prompt, and no more 0xC0000005 on Windows / AMD (#39 by @boxwrench). It fixes three races in the expert tiers that NVIDIA and AMD share:
    • the expert tables' update buffers were rewritten while their copy to the GPU was still in flight;
    • a background expert move could land in a VRAM slot a prompt was borrowing;
    • a route of NaN scores indexed the expert tables. Now that request fails with an error and the engine keeps running.
    • Confirmed on an RTX 3090 + 3060 (#50: the 8K prompt after decoding no longer crashes) and on Windows with an RX 7900 XTX (#53: 5 crashes in 8 sessions before, 0 in 8 now).
  • A stuck prompt ends (#40): when a request finishes no prompt layer and no token for 3 minutes, the engine stops with an error that says where it was stuck. On Windows it also writes strata-stall-<pid>.dmp. The next request starts the engine again. Before, a prompt that stalled with the GPU idle waited forever. STRATA_WATCHDOG_S sets the time; 0 turns it off.
  • Windows builds again (#54 by @jerem91150; #55 by @noahark had the same fix).
  • Thinking off stays off (#54): the answer no longer lands inside an empty thinking block. This applies to every GPU.
  • AMD:
    • a Windows build script (#54);
    • RDNA4 prefill (reading the prompt) +8-10% on R9700 / RX 9070 (#38 by @boxwrench).
  • One GPU, prefill (reading the prompt) about +1% (#61 by @boxwrench): the last layer's unused outputs are skipped, and the output is the same.
  • Robustness (#62 by @needmorevram):
    • F16, BF16, Q2_K, Q4_1 and Q5_1 GGUFs read the right embedding rows. The published Maya models were not affected.
    • Bad image files are refused instead of ending the engine.
  • Switch chats while an answer is being written (#66 by @needmorevram): the other chat opens, and the answer carries on in its own chat.
  • A setup for another context keeps your --calibrate tuning (#57). Before, it fell back to the engine's defaults.

Checked

On 1x and 2x Tesla V100, and on Windows:

  • the builds and parity tests, including the engine's Windows build;
  • identical greedy tokens with the CPU lane off on one and two GPUs, and with the graphs on;
  • #50's sequence (every RAM-tier expert on the CPU, then long prompts) on one and two GPUs;
  • the watchdog: a forced stall ends with an error and the next request is answered, and it never fires on normal work;
  • decode (writing the answer) and prefill (reading the prompt) the same, prefill +1% on one GPU;
  • the server end to end, #58's screen in a real terminal, and #66 in a browser against the running model;
  • the GitHub checks.

Project Maya v1.0.26

Choose a tag to compare

@mw00 mw00 released this 09 Oct 16:42

Maya-L runs at 16.4 tokens/s on an RTX 4090 with 192 GB of RAM under Windows 11 - a user's report, now in the README.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). Documentation only; nothing is downloaded again.

What's new

  • A new speed report (#63 by @npc97): Maya-L on an RTX 4090 24 GB, a Ryzen 9 9950X3D and 192 GB of DDR5 under Windows 11.
    • Every expert fits in VRAM or RAM, so nothing is read from the SSD.
    • Decode (writing the answer): 16.4 tokens/s. Prefill (reading the prompt): 1371 tokens/s on an 8K-token prompt.
    • It's in the README's user table, and the Windows section now says Maya runs on NVIDIA under Windows, with CUDA 13 working for RTX 20 and newer.
  • CPU pinning is Linux-only (#64 by @needmorevram): the README now says STRATA_GLM_CPU_PIN does nothing on Windows or with one GPU.

Project Maya v1.0.25

Choose a tag to compare

@mw00 mw00 released this 09 Oct 16:36

Multi-GPU rigs of three cards or more now use all of the CPU (+19% decode on a 9-GPU split), --calibrate tunes on realistic text without touching your profile, and there are opt-in CUDA graphs for one GPU.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • One CPU pool for three GPUs or more (#60 by @needmorevram):
    • A token visits the cards one after another, so giving each card its own small share of the CPU left most of it idle. One pool now serves them all.
    • Maya-L on 9 GPUs: decode 25.7 -> 30.7 tokens/s, an 8K prompt 1265 -> 1466 tokens/s.
    • One and two GPUs are unchanged. STRATA_GLM_CPU_SHARED=1 shares the pool on two GPUs too, which was faster on a single-CPU 2x V100 box (30.0 vs 28.5 tokens/s).
  • --calibrate on realistic text (#60):
    • It now answers fresh prompts for every measurement, so it tunes for real chats rather than three repeated answers.
    • It works on a copy of your expert usage profile, so tuning no longer reorders the experts your next start loads first.
  • CUDA graphs for one GPU, opt-in (#59 by @handmade0octopus): STRATA_GLM_KDA_GRAPH=1.
    • The output is bit-identical.
    • RTX 4090D: +2.35% decode. Tesla V100: the same speed.

Checked

On 1x and 2x Tesla V100:

  • the builds and parity tests;
  • identical greedy tokens at the defaults, and with the graphs on;
  • #59's real-model test: 192 bit-for-bit comparisons;
  • decode and prefill (reading the prompt) unchanged at the defaults;
  • a full --calibrate run that left the usage file byte-for-byte unchanged;
  • the GitHub checks.

Project Maya v1.0.24

Choose a tag to compare

@mw00 mw00 released this 09 Oct 14:58

Better drafts for speculative decode on two GPUs or more: more drafts accepted, never slower.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • Better drafts (#48 by @sociolog): when the model's draft block is missing an expert in VRAM, it now fetches the missing ones among its 2 most important experts instead of skipping them all.
    • 2x Tesla V100: 82-83% of drafts accepted instead of 76-81%, and decode at 128K context 28.0-28.2 tokens/s instead of 27.0-28.1.
    • RTX 3090 + 3060: 19.5 -> 20.6 tokens/s.
    • It never measured slower, and the answers are the model's own.
    • STRATA_GLM_MTP_KEEP=<n> changes how many: 0 = the old behaviour.

Checked

  • On 2x Tesla V100: the build, five draft settings over two rounds, and identical greedy tokens with the CPU lane off.
  • The GitHub checks pass.