Performance of llama.cpp with FastRPC/mempool-based ggml-hexagon #83
Replies: 2 comments 1 reply
|
PP & TG Performance Comparison: dspqueue-based ggml-hexagon vs FastRPC-based ggml-hexagon Table: Eight-Model PP/TG Summary (3-round average, Qwen3.5-9B single run)
|
PP & TG Performance Comparison report on 2026-09-27Device: Qualcomm Snapdragon 8 Elite (aka 8 Gen 4) 1. Standalone testsPP&TG in dspqueue-based ggml-hexagon(aka Qualcomm's official ggml-hexagon):
PP&TG in FastRPC/mempool-based ggml-hexagon:
2. Eight-Model AB tests PP/TG Summary (3-round average, Qwen3.5-9B single run)Eight-model PP/TG Comparison Analysis (log_abtest_all_20260927-113422.txt), generated by GLM-5.2
|




Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Overview
Qualcomm Hexagon SDK exposes two distinct RPC transport mechanisms: native FastRPC and dspqueue.
Project ggml‑hexagon introduces a FastRPC/mempool‑based ggml‑hexagon backend. It coexists with the existing dspqueue‑based (aka Qualcomm's official) ggml‑hexagon implementation and demonstrates that the FastRPC-based ggml-hexagon greatly outperforms Qualcomm's official ggml-hexagon on some modern models.
This dual‑backend pattern mirrors existing patterns inside llama.cpp:
This is similar to the Apple Silicon benchmark thread, but for FastRPC-based ggml-hexagon! We'll be testing the qwen3-2b qwen3-4b spark-1b gemma4-e2b nanbeige-3b gemma4-e4b qwen3-9b model like the other thread to keep things consistent, the build script (scripts/build-run-ggmlhexagon-android.sh) automatically downloads 8 models on first run.
Attention
For developers and experts from China, please note:
The build script (
scripts/build-run-ggmlhexagon-android.sh) automatically downloads 8 models on its first run. Please ensure you can access to Hugging Face and have sufficient network bandwidth to download LLM models from Hugging Face.Instructions(template-1)
Feel free to use your preferred LLM to analyze
log_abtest_all_xxxx.txt, create the test report following the template below.PP & TG Performance Comparison: dspqueue-based ggml-hexagon vs FastRPC-based ggml-hexagon
Device: Qualcomm Snapdragon 8 Elite (aka 8 Gen 4)
Date: All AB tests complete 2026-09-15 11:49:13
llama.cpp version:
version: 0.4.1-dev (build 11594, commit 57364ab)
built with Clang 21.0.0 for Android aarch64
Table: Eight-Model PP/TG Summary (3-round average, Qwen3.5-9B single run)
Another template(template-2)
All reactions