Skip to content

v0.8.3-pisces.2

Choose a tag to compare

@pisces312 pisces312 released this 01 Aug 12:34
· 36 commits to master since this release

v0.8.3-pisces.2

Fork release based on upstream MnnLlmChat 0.8.3 (MNN engine synced to upstream master @ e1b8a9cb, 13 commits).

Models for Hexagon can be downloaded from https://huggingface.co/pisces312-hf/qwen3-4b-mnn-hexagon-4bit

Highlights

  • Hexagon NPU (HTP) backend support: run LLMs on the Hexagon DSP — Qwen3-4B verified on device
  • htp-ops skel rebuilt with upstream PWL FP16 activation optimization (3ed5e9cb) + local Q4 matmul ops (DSP v81)
  • Upstream merge: Apple GPU LLM perf overhaul, 9 memory-safety fixes in operators, LLM bugfixes (LoRA-split + transformerFuseC4 garbage, LlmContext cross-thread UAF, Qwen3.5 text-only export, InternVL q/k norm export), OpenCL int32 Select/Reduction fix, prefix KV-cache file support, llm_bench -pg real prefill+decode benchmark

Screenshots

Hexagon backend option in model config:

Hexagon backend in model config

Qwen3-4B running on Hexagon NPU:

Qwen3-4B on Hexagon

Qwen3-4B on Hexagon

Assets

  • MnnLlmChat-v0.8.3.2-pisces-standard-signed.apk — standard flavor, signed release (arm64-v8a, minSdk 26)
  • versionName 0.8.3.2 / versionCode 26080101