Repository navigation
Running with the ggml-spacemit backend
With the standalone ggml-spacemit backend (-DGGML_SPACEMIT=ON), llama-cli, llama-bench
and llama-server are launched as follows:
-
Always set
-t 1. The new backend is driven by a single main thread that submits work to
the spine runtime, which then spawns its own workers.-tdoes not control the SpacemiT cores,
so keep it at 1 everywhere. -
Specify the core count only via
SPACEMIT_PERFER_CORE_ID— a comma-separated list of
physical core ids, e.g.export SPACEMIT_PERFER_CORE_ID=8,9,10,11,12,13,14,15. When unset, the
backend picks the cores itself. -
Always pass
--device SPACEMIT0so the model is actually placed on the backend.
For example:
# llama-cli
./bin/llama-cli -m model.gguf -t 1 --no-mmap -fa 1 --device SPACEMIT0 -p "Hello"
# llama-bench
./bin/llama-bench -m model.gguf -t 1 -p 128 -n 128 -fa 1 -ub 128 --device SPACEMIT0
# llama-server (multi-ASR decode with 6 cores)
SPACEMIT_PERFER_CORE_ID=10,11,12,13,14,15 \
./bin/llama-server -m model.gguf -t 1 --device SPACEMIT0 \
--media-backend smt --smt-config-dir ./model_dir --smt-multi-asr -np 4