Skip to content

Guide GLM 4.7 Flash llama.cpp on 2x RTX 3080

flyworker edited this page Sep 15, 2026 · 1 revision

Guide: GLM-4.7-Flash (GGUF) on 2× RTX 3080 with llama.cpp

Serve a 31B MoE reasoning model on two consumer 10 GB GPUs at ~120 tok/s with a 64K context — after the vLLM route dies mid-inference and the obvious flags turn out to be wrong on Ampere.


What you get

Model (marketplace ID) zai-org/GLM-4.7-Flash
Weights served evilfreelancer/GLM-4.7-Flash-GGUF — IQ4_XS (15.05 GiB)
Base model zai-org/GLM-4.7-Flash — Glm4MoeLiteForCausalLM, 31B total / ~3B active, 64 routed + 1 shared expert, 47 layers, MLA attention, reasoning
Serving engine llama.cpp llama-server (ghcr.io/ggml-org/llama.cpp:server-cuda), OpenAI-compatible
Context window 65536 tokens (q8_0 KV) split into two 32K slots
VRAM ~18.0 GiB across two cards at 64K (16.19 GiB at 16K)
Speed ~120 tok/s decode, ~1.7k tok/s prefill

Why llama.cpp and not vLLM? We tried vLLM first and it is worse on every axis here — see the vLLM route. Short version: 4-bit AWQ needs ~18.5 GiB of weights, which does not fit one pair of 10 GB cards with any KV cache left, so it needs all four GPUs at TP=4 — where it manages ~11 tok/s, an eleventh of llama.cpp on half the hardware.


Hardware & prerequisites

  • 2× NVIDIA RTX 3080 (10 GB) dedicated to this model
  • NVIDIA driver + NVIDIA Container Toolkit + Docker
  • ~16 GB free disk for the weights
  • Idle CPU cores — llama.cpp drives the GPUs from CPU threads (see CPU starvation)

1. Download the weights

One file, no vision projector — this model is text-only.

computing-provider models download evilfreelancer/GLM-4.7-Flash-GGUF

That pulls the whole quant ladder. To take only the file you need:

curl -L -H "Authorization: Bearer $HF_TOKEN" \
  -o ~/.swan/models/evilfreelancer/GLM-4.7-Flash-GGUF/GLM-4.7-Flash-IQ4_XS.gguf \
  https://huggingface.co/evilfreelancer/GLM-4.7-Flash-GGUF/resolve/main/GLM-4.7-Flash-IQ4_XS.gguf

Why IQ4_XS (15.05 GiB) and not Q4_K_M (17.05 GiB)? Two 10 GB cards give you 20 GB total. Q4_K_M leaves under 3 GB for KV cache and compute buffers, which will not hold 64K. IQ4_XS leaves ~4.9 GB, which does.

general.architecture = deepseek2 is correct, not a mis-conversion. Every GGUF of this model declares it — evilfreelancer's, unsloth's, and the official ggml-org build. GLM's MoE config uses DeepSeek-style fields (n_routed_experts, n_shared_experts, first_k_dense_replace) and MLA attention, so llama.cpp converts it onto the deepseek2 graph. If you check the GGUF header and see deepseek2, you have the right file.

MLX builds cannot run here. lmstudio-community/GLM-4.7-Flash-MLX-* needs Apple Silicon. Neither llama.cpp nor vLLM loads MLX weights.


2. Start llama-server on the two GPUs

computing-provider models serve --backend llamacpp \
  --weights ~/.swan/models/evilfreelancer/GLM-4.7-Flash-GGUF/GLM-4.7-Flash-IQ4_XS.gguf \
  --gpus 0,1 --port 30000 --context-length 65536 --gpu-memory 20000 \
  zai-org/GLM-4.7-Flash \
  -ngl 99 --tensor-split 1,1 --parallel 2 -fa on -ctk q8_0 -ctv q8_0 \
  --jinja --reasoning-format deepseek --reasoning-budget 2048

Flags come before the model ID; anything after it is passed through to llama-server. The equivalent raw docker command:

docker run -d --name swan-cp-zai-org-GLM-4.7-Flash --restart unless-stopped \
  --gpus '"device=0,1"' -p 127.0.0.1:30000:8080 --shm-size 4g --ipc host \
  -v /path/to/GLM-4.7-Flash-IQ4_XS.gguf:/models/model.gguf:ro \
  ghcr.io/ggml-org/llama.cpp:server-cuda \
  -m /models/model.gguf --alias zai-org/GLM-4.7-Flash \
  --host 0.0.0.0 --port 8080 \
  -ngl 99 --tensor-split 1,1 -c 65536 --parallel 2 \
  -fa on -ctk q8_0 -ctv q8_0 \
  --jinja --reasoning-format deepseek --reasoning-budget 2048

--tensor-split 1,1 is fine here — unlike Qwen3.8-27B there is no vision projector to unbalance the split.


3. Gotchas that cost us hours

The provider advertises more concurrency than one slot can serve

This is the one that took the model off the marketplace. Served with --parallel 1, the node still advertised the client's default of 10 concurrent requests per model (DefaultModelMax in internal/computing/concurrency_limiter.go). Ten requests queued onto a one-slot server, each ran past the 30 s hub timeout, and the self-check's inference probe timed out twice in a row:

429 provider concurrency limit for model zai-org/GLM-4.7-Flash
502 failed to read stream ... after 300001ms
[selfcheck] failed 2 consecutive probes — deregistered from Swan Inference
Disabled model: zai-org/GLM-4.7-Flash

Concurrency is arrivals × hold time, not arrival rate. A handful of arrivals with five-minute holds reaches ten in flight. Two fixes together:

--parallel 2                      # llama.cpp actually serves 2 at once
curl -X POST .../inference/concurrency/model/zai-org%2FGLM-4.7-Flash \
  -H "Authorization: Bearer $(cat $CP_PATH/dashboard.token)" \
  -d '{"max":2}'

Match the provider limit to the slot count. A fast 429 is far better than a timeout: the hub re-routes instead of waiting.

Re-enabling after an auto-disable also needs the dashboard token:

curl -X POST .../inference/models/zai-org%2FGLM-4.7-Flash/enable \
  -H "Authorization: Bearer $(cat $CP_PATH/dashboard.token)"

A reasoning model with no budget returns empty answers

Without --reasoning-budget, this model will spend an entire max_tokens on thinking and return zero characters of content. At max_tokens: 1500 we got finish_reason: length, 1500 tokens generated, and an empty content. Even "name three primary colours" hit a 400-token cap while still reasoning.

--reasoning-budget 2048 bounds it. Also set --reasoning-format deepseek, or the chain-of-thought is returned as the answer instead of in message.reasoning_content — customers would get the scratchpad.

Context is unusually cheap here — do not size it like ordinary attention

GLM-4.7-Flash uses MLA (kv_lora_rank: 512, qk_rope_head_dim: 64), so it caches 576 elements per token per layer rather than 20 full KV heads. At q8_0 over 47 layers that is roughly 28 KiB/token:

context KV cache fits 2× 10 GB?
16K ~0.45 GiB yes, lots of room
64K ~1.8 GiB yes — use this
96K ~2.7 GiB probably, ~1 GB spare
128K ~3.7 GiB no, nothing left for compute buffers

We shipped 16K first by sizing it as ordinary MHA, which was ~10× too pessimistic. Verified at 64K: an 18,035-token prompt recalled a fact buried at its start in 10.9 s end to end.

Uniform q8_0 KV cache, never mixed

Same trap as Qwen: mixed or q4_0 KV makes llama.cpp's CUDA flash-attention fall back to the CPU and prefill collapses. f16 or uniform q8_0.

models serve --gpus 0,1 was broken before v0.6.0

Docker CSV-parses the --gpus value, so unquoted device=0,1 split into device=0 and a bare 1 — read as a device count:

docker: Error response from daemon: cannot set both Count and DeviceIDs on device request

Fixed in v0.6.0. On older builds, use the raw docker run above with --gpus '"device=0,1"' (note the inner quotes).


The vLLM route, and why we left it

Worth recording so nobody repeats it.

cyankiwi AWQ (compressed-tensors) QuantTrio AWQ (classic) evilfreelancer IQ4_XS
Engine vLLM 0.24.0 vLLM 0.24.0 llama.cpp
GPUs 4 (TP=4) 4 (TP=4) 2
VRAM 37.35 GiB 36.90 GiB 18.0 GiB
Throughput ~11 tok/s ~11 tok/s ~120 tok/s
Long generation EngineDeadError survived survived
3 concurrent crashed survived survived

Two distinct failures:

--kv-cache-dtype fp8 is unusable on these cards. Triton refuses it:

ValueError: type fp8e4nv not supported in this architecture.
The supported fp8 dtypes are ('fp8e4b15', 'fp8e5')

RTX 3080 is sm_86; fp8 needs sm_89+. A dense Mistral like Cydonia tolerates the flag, so do not copy a working vLLM flag set onto this model — GLM's MoE path hits a kernel that Cydonia's never reaches.

The compressed-tensors pack-quantized path crashes mid-inference. With cyankiwi/GLM-4.7-Flash-AWQ-4bit, vLLM 0.24.0 loads and generates, then dies:

AttributeError: 'ColumnParallelLinear' object has no attribute 'weight'
EngineDeadError: EngineCore encountered an issue

It survives a short request and dies on a long one or under concurrency — so it will crash-loop under real traffic. QuantTrio/GLM-4.7-Flash-AWQ (classic awq/gemm, the awq_marlin path) does not crash, which isolates the bug to the compressed-tensors path rather than AWQ or the model. It is still ~11 tok/s, so it is not worth using.


4. Verify it works

curl -s localhost:30000/v1/chat/completions -H 'Content-Type: application/json' -d '{
  "model": "zai-org/GLM-4.7-Flash",
  "messages": [{"role":"user","content":"What is 17*23? Answer briefly."}],
  "max_tokens": 3000
}' | jq '.choices[0] | {finish_reason, content: .message.content, thinking: (.message.reasoning_content|length)}'

Expect finish_reason: "stop", content: "17 * 23 is **391**." and a non-zero thinking length. If content is empty and finish_reason is "length", raise max_tokens or lower --reasoning-budget.

Check the slots came up as intended:

docker logs swan-cp-zai-org-GLM-4.7-Flash 2>&1 | grep n_ctx_slot
# srv load_model: initializing, n_slots = 2, n_ctx_slot = 32768

5. Register it with the computing-provider

models serve writes models.json for you and, from v0.6.0, keeps the Models list in config.toml in step. To do it by hand:

{
  "zai-org/GLM-4.7-Flash": {
    "endpoint": "http://localhost:30000",
    "category": "text-generation",
    "gpu_memory": 20000,
    "context_length": 65536
  }
}

Declare 32768 instead if you run --parallel 2 and want the declared window to match one slot rather than the total — see the Qwen guide's discussion.

Confirm the hub took it:

curl -s localhost:9087/api/v1/computing/inference/status | jq .registered_models

Sharing the box: our 4× RTX 3080 layout

GPUs Model Engine Context Speed Port
0–1 zai-org/GLM-4.7-Flash llama.cpp, IQ4_XS 2 × 32K ~120 tok/s 30000
2–3 Qwen/Qwen3.8-27B llama.cpp, UD-Q4_K_S 2 × 32K ~31 tok/s 30001

Three models do not fit: GLM, Qwen and Cydonia need ~45 GiB of weights against 40 GB of VRAM. Give both containers --cpu-shares 262144 on a contended host.


Troubleshooting

Symptom Fix
cannot set both Count and DeviceIDs v0.6.0+, or quote it: --gpus '"device=0,1"'
type fp8e4nv not supported in this architecture Drop --kv-cache-dtype fp8; Ampere cannot do it
AttributeError: 'ColumnParallelLinear' object has no attribute 'weight' The compressed-tensors AWQ build under vLLM. Use the GGUF
Empty content, finish_reason: length Raise max_tokens, or lower --reasoning-budget
Chain-of-thought returned as the answer Add --reasoning-format deepseek
429 provider concurrency limit then auto-disable Match the provider's per-model limit to --parallel
Model shows state: disabled after re-enabling Fixed in v0.6.0; before that, trust enabled and registered_models
Prefill collapses to ~50 tok/s Mixed KV cache types — use uniform q8_0

See also

Clone this wiki locally