Repository navigation
Guide GLM 4.7 Flash llama.cpp on 2x RTX 3080
Serve a 31B MoE reasoning model on two consumer 10 GB GPUs at ~120 tok/s with a 64K context — after the vLLM route dies mid-inference and the obvious flags turn out to be wrong on Ampere.
| Model (marketplace ID) | zai-org/GLM-4.7-Flash |
| Weights served |
evilfreelancer/GLM-4.7-Flash-GGUF — IQ4_XS (15.05 GiB) |
| Base model |
zai-org/GLM-4.7-Flash — Glm4MoeLiteForCausalLM, 31B total / ~3B active, 64 routed + 1 shared expert, 47 layers, MLA attention, reasoning |
| Serving engine | llama.cpp llama-server (ghcr.io/ggml-org/llama.cpp:server-cuda), OpenAI-compatible |
| Context window | 65536 tokens (q8_0 KV) split into two 32K slots |
| VRAM | ~18.0 GiB across two cards at 64K (16.19 GiB at 16K) |
| Speed | ~120 tok/s decode, ~1.7k tok/s prefill |
Why llama.cpp and not vLLM? We tried vLLM first and it is worse on every axis here — see the vLLM route. Short version: 4-bit AWQ needs ~18.5 GiB of weights, which does not fit one pair of 10 GB cards with any KV cache left, so it needs all four GPUs at TP=4 — where it manages ~11 tok/s, an eleventh of llama.cpp on half the hardware.
- 2× NVIDIA RTX 3080 (10 GB) dedicated to this model
- NVIDIA driver + NVIDIA Container Toolkit + Docker
- ~16 GB free disk for the weights
- Idle CPU cores — llama.cpp drives the GPUs from CPU threads (see CPU starvation)
One file, no vision projector — this model is text-only.
computing-provider models download evilfreelancer/GLM-4.7-Flash-GGUFThat pulls the whole quant ladder. To take only the file you need:
curl -L -H "Authorization: Bearer $HF_TOKEN" \
-o ~/.swan/models/evilfreelancer/GLM-4.7-Flash-GGUF/GLM-4.7-Flash-IQ4_XS.gguf \
https://huggingface.co/evilfreelancer/GLM-4.7-Flash-GGUF/resolve/main/GLM-4.7-Flash-IQ4_XS.ggufWhy IQ4_XS (15.05 GiB) and not Q4_K_M (17.05 GiB)? Two 10 GB cards give
you 20 GB total. Q4_K_M leaves under 3 GB for KV cache and compute buffers,
which will not hold 64K. IQ4_XS leaves ~4.9 GB, which does.
general.architecture = deepseek2is correct, not a mis-conversion. Every GGUF of this model declares it — evilfreelancer's, unsloth's, and the official ggml-org build. GLM's MoE config uses DeepSeek-style fields (n_routed_experts,n_shared_experts,first_k_dense_replace) and MLA attention, so llama.cpp converts it onto thedeepseek2graph. If you check the GGUF header and seedeepseek2, you have the right file.
MLX builds cannot run here.
lmstudio-community/GLM-4.7-Flash-MLX-*needs Apple Silicon. Neither llama.cpp nor vLLM loads MLX weights.
computing-provider models serve --backend llamacpp \
--weights ~/.swan/models/evilfreelancer/GLM-4.7-Flash-GGUF/GLM-4.7-Flash-IQ4_XS.gguf \
--gpus 0,1 --port 30000 --context-length 65536 --gpu-memory 20000 \
zai-org/GLM-4.7-Flash \
-ngl 99 --tensor-split 1,1 --parallel 2 -fa on -ctk q8_0 -ctv q8_0 \
--jinja --reasoning-format deepseek --reasoning-budget 2048Flags come before the model ID; anything after it is passed through to
llama-server. The equivalent raw docker command:
docker run -d --name swan-cp-zai-org-GLM-4.7-Flash --restart unless-stopped \
--gpus '"device=0,1"' -p 127.0.0.1:30000:8080 --shm-size 4g --ipc host \
-v /path/to/GLM-4.7-Flash-IQ4_XS.gguf:/models/model.gguf:ro \
ghcr.io/ggml-org/llama.cpp:server-cuda \
-m /models/model.gguf --alias zai-org/GLM-4.7-Flash \
--host 0.0.0.0 --port 8080 \
-ngl 99 --tensor-split 1,1 -c 65536 --parallel 2 \
-fa on -ctk q8_0 -ctv q8_0 \
--jinja --reasoning-format deepseek --reasoning-budget 2048--tensor-split 1,1 is fine here — unlike Qwen3.8-27B there is no vision
projector to unbalance the split.
This is the one that took the model off the marketplace. Served with
--parallel 1, the node still advertised the client's default of 10
concurrent requests per model (DefaultModelMax in
internal/computing/concurrency_limiter.go). Ten requests queued onto a
one-slot server, each ran past the 30 s hub timeout, and the self-check's
inference probe timed out twice in a row:
429 provider concurrency limit for model zai-org/GLM-4.7-Flash
502 failed to read stream ... after 300001ms
[selfcheck] failed 2 consecutive probes — deregistered from Swan Inference
Disabled model: zai-org/GLM-4.7-Flash
Concurrency is arrivals × hold time, not arrival rate. A handful of arrivals with five-minute holds reaches ten in flight. Two fixes together:
--parallel 2 # llama.cpp actually serves 2 at oncecurl -X POST .../inference/concurrency/model/zai-org%2FGLM-4.7-Flash \
-H "Authorization: Bearer $(cat $CP_PATH/dashboard.token)" \
-d '{"max":2}'Match the provider limit to the slot count. A fast 429 is far better than a timeout: the hub re-routes instead of waiting.
Re-enabling after an auto-disable also needs the dashboard token:
curl -X POST .../inference/models/zai-org%2FGLM-4.7-Flash/enable \
-H "Authorization: Bearer $(cat $CP_PATH/dashboard.token)"Without --reasoning-budget, this model will spend an entire max_tokens on
thinking and return zero characters of content. At max_tokens: 1500 we got
finish_reason: length, 1500 tokens generated, and an empty content. Even
"name three primary colours" hit a 400-token cap while still reasoning.
--reasoning-budget 2048 bounds it. Also set --reasoning-format deepseek, or
the chain-of-thought is returned as the answer instead of in
message.reasoning_content — customers would get the scratchpad.
GLM-4.7-Flash uses MLA (kv_lora_rank: 512, qk_rope_head_dim: 64), so it
caches 576 elements per token per layer rather than 20 full KV heads. At q8_0
over 47 layers that is roughly 28 KiB/token:
| context | KV cache | fits 2× 10 GB? |
|---|---|---|
| 16K | ~0.45 GiB | yes, lots of room |
| 64K | ~1.8 GiB | yes — use this |
| 96K | ~2.7 GiB | probably, ~1 GB spare |
| 128K | ~3.7 GiB | no, nothing left for compute buffers |
We shipped 16K first by sizing it as ordinary MHA, which was ~10× too pessimistic. Verified at 64K: an 18,035-token prompt recalled a fact buried at its start in 10.9 s end to end.
Same trap as Qwen: mixed or q4_0 KV makes llama.cpp's CUDA flash-attention
fall back to the CPU and prefill collapses. f16 or uniform q8_0.
Docker CSV-parses the --gpus value, so unquoted device=0,1 split into
device=0 and a bare 1 — read as a device count:
docker: Error response from daemon: cannot set both Count and DeviceIDs on device request
Fixed in v0.6.0. On older builds, use the raw docker run above with
--gpus '"device=0,1"' (note the inner quotes).
Worth recording so nobody repeats it.
| cyankiwi AWQ (compressed-tensors) | QuantTrio AWQ (classic) | evilfreelancer IQ4_XS | |
|---|---|---|---|
| Engine | vLLM 0.24.0 | vLLM 0.24.0 | llama.cpp |
| GPUs | 4 (TP=4) | 4 (TP=4) | 2 |
| VRAM | 37.35 GiB | 36.90 GiB | 18.0 GiB |
| Throughput | ~11 tok/s | ~11 tok/s | ~120 tok/s |
| Long generation | EngineDeadError |
survived | survived |
| 3 concurrent | crashed | survived | survived |
Two distinct failures:
--kv-cache-dtype fp8 is unusable on these cards. Triton refuses it:
ValueError: type fp8e4nv not supported in this architecture.
The supported fp8 dtypes are ('fp8e4b15', 'fp8e5')
RTX 3080 is sm_86; fp8 needs sm_89+. A dense Mistral like Cydonia tolerates the flag, so do not copy a working vLLM flag set onto this model — GLM's MoE path hits a kernel that Cydonia's never reaches.
The compressed-tensors pack-quantized path crashes mid-inference. With
cyankiwi/GLM-4.7-Flash-AWQ-4bit, vLLM 0.24.0 loads and generates, then dies:
AttributeError: 'ColumnParallelLinear' object has no attribute 'weight'
EngineDeadError: EngineCore encountered an issue
It survives a short request and dies on a long one or under concurrency — so it
will crash-loop under real traffic. QuantTrio/GLM-4.7-Flash-AWQ (classic
awq/gemm, the awq_marlin path) does not crash, which isolates the bug
to the compressed-tensors path rather than AWQ or the model. It is still ~11
tok/s, so it is not worth using.
curl -s localhost:30000/v1/chat/completions -H 'Content-Type: application/json' -d '{
"model": "zai-org/GLM-4.7-Flash",
"messages": [{"role":"user","content":"What is 17*23? Answer briefly."}],
"max_tokens": 3000
}' | jq '.choices[0] | {finish_reason, content: .message.content, thinking: (.message.reasoning_content|length)}'Expect finish_reason: "stop", content: "17 * 23 is **391**." and a non-zero
thinking length. If content is empty and finish_reason is "length",
raise max_tokens or lower --reasoning-budget.
Check the slots came up as intended:
docker logs swan-cp-zai-org-GLM-4.7-Flash 2>&1 | grep n_ctx_slot
# srv load_model: initializing, n_slots = 2, n_ctx_slot = 32768models serve writes models.json for you and, from v0.6.0, keeps the
Models list in config.toml in step. To do it by hand:
{
"zai-org/GLM-4.7-Flash": {
"endpoint": "http://localhost:30000",
"category": "text-generation",
"gpu_memory": 20000,
"context_length": 65536
}
}Declare 32768 instead if you run --parallel 2 and want the declared window
to match one slot rather than the total — see the
Qwen guide's discussion.
Confirm the hub took it:
curl -s localhost:9087/api/v1/computing/inference/status | jq .registered_models| GPUs | Model | Engine | Context | Speed | Port |
|---|---|---|---|---|---|
| 0–1 | zai-org/GLM-4.7-Flash |
llama.cpp, IQ4_XS | 2 × 32K | ~120 tok/s | 30000 |
| 2–3 | Qwen/Qwen3.8-27B |
llama.cpp, UD-Q4_K_S | 2 × 32K | ~31 tok/s | 30001 |
Three models do not fit: GLM, Qwen and Cydonia need ~45 GiB of weights against
40 GB of VRAM. Give both containers --cpu-shares 262144 on a contended host.
| Symptom | Fix |
|---|---|
cannot set both Count and DeviceIDs |
v0.6.0+, or quote it: --gpus '"device=0,1"'
|
type fp8e4nv not supported in this architecture |
Drop --kv-cache-dtype fp8; Ampere cannot do it |
AttributeError: 'ColumnParallelLinear' object has no attribute 'weight' |
The compressed-tensors AWQ build under vLLM. Use the GGUF |
Empty content, finish_reason: length
|
Raise max_tokens, or lower --reasoning-budget
|
| Chain-of-thought returned as the answer | Add --reasoning-format deepseek
|
429 provider concurrency limit then auto-disable |
Match the provider's per-model limit to --parallel
|
Model shows state: disabled after re-enabling |
Fixed in v0.6.0; before that, trust enabled and registered_models
|
| Prefill collapses to ~50 tok/s | Mixed KV cache types — use uniform q8_0
|
- Qwen3.8-27B (GGUF) on 2× RTX 3080 with llama.cpp — the other half of this box
- Cydonia 24B v4.3 (AWQ) on 4× RTX 3080 — the vLLM path, when a model suits it
- Configuration · Inference Mode · Troubleshooting