v0.8.3-pisces.2
·
36 commits
to master
since this release
v0.8.3-pisces.2
Fork release based on upstream MnnLlmChat 0.8.3 (MNN engine synced to upstream master @ e1b8a9cb, 13 commits).
Models for Hexagon can be downloaded from https://huggingface.co/pisces312-hf/qwen3-4b-mnn-hexagon-4bit
Highlights
- Hexagon NPU (HTP) backend support: run LLMs on the Hexagon DSP — Qwen3-4B verified on device
- htp-ops skel rebuilt with upstream PWL FP16 activation optimization (
3ed5e9cb) + local Q4 matmul ops (DSP v81) - Upstream merge: Apple GPU LLM perf overhaul, 9 memory-safety fixes in operators, LLM bugfixes (LoRA-split + transformerFuseC4 garbage, LlmContext cross-thread UAF, Qwen3.5 text-only export, InternVL q/k norm export), OpenCL int32 Select/Reduction fix, prefix KV-cache file support,
llm_bench -pgreal prefill+decode benchmark
Screenshots
Hexagon backend option in model config:
Qwen3-4B running on Hexagon NPU:
Assets
MnnLlmChat-v0.8.3.2-pisces-standard-signed.apk— standard flavor, signed release (arm64-v8a, minSdk 26)- versionName
0.8.3.2/ versionCode26080101


