Skip to content

v0.2.0

Latest

Choose a tag to compare

@github-actions github-actions released this 24 Sep 07:51
64316cd

Running with the ggml-spacemit backend

With the standalone ggml-spacemit backend (-DGGML_SPACEMIT=ON), llama-cli, llama-bench
and llama-server are launched as follows:

  1. Always set -t 1. The new backend is driven by a single main thread that submits work to
    the spine runtime, which then spawns its own workers. -t does not control the SpacemiT cores,
    so keep it at 1 everywhere.

  2. Specify the core count only via SPACEMIT_PERFER_CORE_ID — a comma-separated list of
    physical core ids, e.g. export SPACEMIT_PERFER_CORE_ID=8,9,10,11,12,13,14,15. When unset, the
    backend picks the cores itself.

  3. Always pass --device SPACEMIT0 so the model is actually placed on the backend.

For example:

# llama-cli
./bin/llama-cli -m model.gguf -t 1 --no-mmap -fa 1 --device SPACEMIT0 -p "Hello"

# llama-bench
./bin/llama-bench -m model.gguf -t 1 -p 128 -n 128 -fa 1 -ub 128 --device SPACEMIT0

# llama-server (multi-ASR decode with 6 cores)
SPACEMIT_PERFER_CORE_ID=10,11,12,13,14,15 \
  ./bin/llama-server -m model.gguf -t 1 --device SPACEMIT0 \
    --media-backend smt --smt-config-dir ./model_dir --smt-multi-asr -np 4