Skip to content

Project Maya v1.0.34

Latest

Choose a tag to compare

@mw00 mw00 released this 11 Oct 13:13
· 1 commit to main since this release

Fixes from your reports: a two-GPU Windows PC whose RAM tier came out smaller in v1.0.33, and builds on CUDA 12.0-12.2. Also two new opt-in settings for machines that read experts from the SSD.

Update: the dashboard's About > Update, or git pull, then ./setup.sh (Windows: START-MAYA.bat). The engine recompiles; nothing is downloaded again.

What's new

  • A split's RAM tiers near Windows' pinning cap (#110, from #105). A two-GPU Windows PC pinned 35.6 GB of its 44 GB RAM budget on v1.0.33, against 42.6 on v1.0.32, so more experts came from the SSD. Now:
    • the GPUs pin their tiers one after another, so only the one that reaches the cap gives way;
    • a tier leaves 512 MB of room under the cap;
    • the expert tables upload through the engine's own pinned buffer, so reaching the cap no longer stops the start with the warmed expert tables did not upload.
  • Windows with two NUMA nodes or more (#110, from #56): the engine log says which processor group and node the engine runs in, and where the RAM tier's pages are. It's there to find why some starts decode slower than others.
  • Builds on CUDA 12.0-12.2 (#107 by @Ajay9o9), the version Ubuntu's nvidia-cuda-toolkit installs.
  • New, opt-in: STRATA_GLM_DISK_READERS=<n> (#108 by @Ajay9o9). Prompts read the experts that come from the SSD with n threads, each expert whole, handed on as soon as it lands. On a DRAM-less NVMe, prompts read 25-50% faster. An NVMe with its own DRAM can lose instead, so measure yours. The answers stay the same.
  • New, opt-in, it changes the answers: STRATA_GLM_TRUNK_KEEP / _MAP (#103 by @sociolog), for a RAM-bound split that would rather be faster. The model's own layers leave out their lowest-ranked experts that aren't in VRAM, as the MTP draft does, layer by layer with the map. On an RTX 3090 + 3060, Maya-M decoded 23% faster with the map in the README's row, KL 0.018 from the full model.

Checked

On 1x and 2x Tesla V100:

  • the build and the GLM parity tests;
  • with every new setting off, identical greedy and sampled tokens to v1.0.33, and the same answers across a conversation's turns;
  • the new settings on: the trunk keep changes the answers, the disk readers don't;
  • the same RAM tier sizes and start time as v1.0.33;
  • decode the same within run-to-run variation;
  • the server end to end.

On Windows: the engine compiles. Also: the GitHub checks pass.