GLM-5.3-Flash on DGX Spark GB10: 0.43 tok/s — glm53 has no CUDA path (make glm53 CUDA=1 is a no-op) #1567
peterMary123
started this conversation in
Show and tell
Replies: 2 comments
|
We are working on adding CUDA support. Some start out with CPU support, and we add the various other types of support later. If you'd like to help us with a PR, you are absolutely welcome. |
0 replies
|
Who cares about cuda, I have older Nvidia Cards and Nvidia just dropped in there new Coda version some important features for the Older Nvidia Cards without any reason to force me to buy newer GPUs, this is absurde. Cuda is dead for me, only Vulkan is the Future. |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
GLM-5.3-Flash on colibri, DGX Spark GB10: 0.43 tok/s — and why
Converted GLM-5.3-Flash (321B-A18B) with tools/convert_glm53.py: 62 shards, 194.7 GB int4 output, 9.2 hours, no errors. Engine built clean on aarch64 (gcc 13, armv8.2-a+dotprod+i8mm, OpenMP). coli info reported ready.
Measured 0.43 tok/s (~135 tokens in 306s, --effort minimal --no-think).
Tuning doesn't help, and the reason is diagnostic. Unconstrained, the engine sized its expert cache from available memory, took ~103 GB, drove the system into swap (15 GiB pinned for 90s). Re-running with --ram 72 (64 GB expert cache) gave 0.44 vs 0.43 tok/s — identical speed on 26 GiB less memory. If cache capacity were the constraint, a 39 GB reduction would have cost something. It didn't, so --cap, --repin and further --ram tuning are all disconnected from the bottleneck.
No GPU path exists for this engine. The glm53 target depends only on$(METAL_OBJ) $ (VK_OBJ) $(VK_SPV) — no CUDA object in its prerequisites. glm53.c has zero #ifdef COLI_CUDA blocks (8 Metal, 2 Vulkan), and there's no backend_cuda_glm53.cu. make glm53 CUDA=1 is accepted but only adds an unused -DCOLI_CUDA define. CUDA backends exist for colibri.c (GLM-5.2), DeepSeek V4 and Inkling only. Toolchain was ready — nvcc 13.0.88, sm_121 gencode present — so the 6,144 CUDA cores sat idle throughout.
--auto-tier silently no-ops here. The launcher only forwards it when arch == "glm"; glm53 takes a separate branch that never reads it. argparse accepts the flag and nothing happens — worth a warning or removal.
For comparison, the same model and box via llama.cpp (glm5next branch, CUDA 13.0, UD-Q2_K_XL): 23.1 tok/s on code, 81% MTP draft acceptance. ~54× faster.
This isn't a criticism of colibri — v1.9.0 targets GLM-5.3-Flash at 25 GB of RAM, and it delivers that. But on a 121 GB machine with a capable GPU, the CPU-only path leaves the hardware unused. A CUDA backend for glm53, or a note in docs/glm53-flash.md that CUDA doesn't apply to this engine, would save people a 9-hour conversion.
Environment: DGX Spark GB10 sm_121, 121 GiB unified, Ubuntu 24.04 aarch64, driver 580.142, CUDA 13.0.88, colibri v1.11.0.
All reactions