TensorFold 0.3.2
TensorFold 0.3.2 adds two things: GLM-5.3-Flash on Mia-AiLab's EXL3 weights, as an experiment, and tensorfold update.
GLM-5.3-Flash on EXL3 (experimental). The CUDA engine now reads Mia-AiLab/GLM-5.3-Flash-EXL3-TR3-4bpw, the checkpoint vLLM serves in Mia-AiLab's DGX Spark recipe, on two Sparks. On those same weights, through the same OpenAI client (decode tok/s, one stream, 64-token replies, median of seeds 1234-1238):
| Code, sampled | Chat, sampled | Code, greedy | Chat, greedy | |
|---|---|---|---|---|
| vLLM, Mia-AiLab's recipe, MTP=3 | 24.5 | 24.3 | 32.2 | 24.7 |
| TensorFold 0.3.2 | 36.4 | 29.7 | 43.8 | 32.9 |
Replies are byte-identical to TensorFold's own serial decoding, as on every other model. The MLX checkpoint (Vontra/GLM-5.3-Flash-MLX-4bit-MTP) is still the faster way to run GLM here, because only the EXL3 checkpoint's routed experts are 4-bit: each Spark reads 10.7 GB a token from it against 5.0. The recipe has the kernels, the checks and what was not checked. The EXL3 format is ExLlamaV3's (MIT).
tensorfold update. tensorfold update installs the newest release. serve also checks GitHub for a newer release once a day in the background, and prints a line if it finds one. It never delays the server. Turn it off with --no-update-check or TENSORFOLD_NO_UPDATE_CHECK=1. From 0.3.1 or earlier, update once by hand:
pip install --upgrade git+https://github.com/ashhart/TensorFold.git@v0.3.2