Skip to content

3.6.1

Latest

Choose a tag to compare

@github-actions github-actions released this 23 Jul 02:13
d407447

MNN 3.6.1 版本发布说明

发布日期: 2026年7月

📦 安装: pip install mnn==3.6.1PyPI

📌 版本概述

MNN 3.6.1 版本聚焦于 Hexagon NPU 后端Transformer C4 Fuse 全后端性能优化稳定性修复 三大核心主题:

  • Hexagon NPU 后端:新增高通 Hexagon NPU 直接编程后端,Qwen3-0.6B(W4 对称量化,block64)在骁龙 8 Elite(Hexagon v79)上 prefill 达 2667 tok/s,为 CPU 的 7.9×。
  • Transformer C4 Fuse 全后端性能优化:新增 fused LLM 算子(RoPE / Attention / LayerNorm 等 C4 路径)并在 CPU / Metal / OpenCL / CUDA 全后端启用 Transformer C4 图;优化 C4 导出与 packed RoPE 路径;Metal 新增 LLM prefill/decode 优化 kernel 与 Fused Q4/Q8 dequant+GEMM(M_TILE=64);OpenCL 修复 NC4HW4 raster L2 cache-set 抖动并优化 int4 GEMM。
  • 稳定性修复与工具增强:新增 MNN-aware QLoRA 微调;修复 mmap 分配失败崩溃、推测解码 KVCache 索引、Arm82 LinearAttention fp16 精度等多项问题。

🚀 版本亮点

  • Hexagon 后端(高通 NPU 直接编程):Qwen3-0.6B 在骁龙 8 Elite 上 pp512 达 2667 tok/s(CPU 的 7.9×,OpenCL GPU 的 1.8×)
  • Transformer C4 图全后端启用:fused LLM 算子 + C4 后端路径覆盖 CPU / Metal / OpenCL / CUDA,包含 packed RoPE fuse 优化
  • Metal LLM Prefill 大幅优化:Fused Q4/Q8 dequant+GEMM kernel、M_TILE=64 变体(面向 M5)、prefill/decode kernel 与 shape/RoPE 基础设施重构,并新增 Metal Op Profile
  • MNN-aware QLoRA 微调:新增面向 MNN 部署的 QLoRA 微调能力

✨ 新功能

后端:Hexagon(新增)

  • Hexagon NPU 后端:新增高通 Hexagon NPU 直接编程后端(HexagonBackend / HexagonRuntime / DSP kernel),支持 Attention 等 LLM 关键算子;附带 skills/hexagon 开发文档

LLM / 推理加速

  • Fused LLM 算子与 C4 后端路径:新增 CPURoPE、SharedGather 等 fused 算子及 CPU/Metal C4 计算路径
  • C4 导出与 packed RoPE 优化:优化 C4 图导出;CPU / Metal / OpenCL 的 RoPE fuse(packed RoPE)路径优化
  • Transformer C4 图全后端启用:CPU / Metal / OpenCL / CUDA(新增 RoPEExecution.cu、LayerNorm/LinearAttention C4 适配)统一启用 C4 图
  • MNN-aware QLoRA 微调:新增 QLoRA 微调支持

GPU 后端

  • Metal LLM prefill/decode kernel 优化:prefill/decode kernel、shape/RoPE 基础设施与 MetalConvolution1x1 SIMD-group 路径重构,新增 Metal Op Profile 工具
  • Metal Fused Q4/Q8 dequant+GEMM:新增融合反量化+GEMM kernel 及 M_TILE=64 变体,优化 M5 上的 LLM prefill
  • OpenCL raster 与 int4 GEMM 优化:转置遍历顺序修复 NC4HW4 L2 cache-set 抖动;优化 int4 GEMM 性能;拆分 binary_layernorm_c4 kernel
  • Vulkan Range 算子:新增 Range 算子 image 实现
  • QNN Interp/Resize 算子:QNN 后端新增 Interp/Resize 支持

CPU 后端

  • RVV CommonOptFunction 适配:为 RISC-V RVV 架构适配 CommonOptFunction

工具

  • llm_bench -fa 选项:新增 flash attention 开关(默认 1,启用 flash attention)

🐛 缺陷修复

LLM / VLM

  • Qwen3.5 fused 推理正确性:修复 Qwen3.5 fused 推理结果错误
  • C4 fuse 保护与 Metal raster:为 C4 fuse 增加 guard 条件并修复 Metal raster 问题
  • RoPE schema 编译错误:修复 rope schema 引入的编译错误
  • MPS 大 tensor Q4 量化静默全零:修复在 MPS 上量化大 tensor 时 Q4 权重静默全零的问题
  • 通用 embedding scale 导出:修复导出时 embedding scale 缺失
  • WikiText 校准数据集加载:修复量化校准数据集加载失败
  • QNN 导出 chunk_size:修复 QNN prefill IO 生成未遵循 --chunk_size 的问题
  • decoder embedding 导出 causal mask:decoder 类 embedding 模型导出改用 causal mask

GPU 后端

  • Metal ATTENTION_C4 输出分支缺失:修复 prefill_qkv_tensor 缺少 ATTENTION_C4 输出分支
  • Metal attention kv_cache flag:修复 Metal attention 未遵循 kv_cache flag
  • OpenCL attention c4 kernel 编译:修复 attention c4 的 OpenCL kernel 编译错误
  • OpenCL c4 长 prefill attention:修复 GEMM tune 时 V-buffer 越界与 NC4HW4 输出布局缺失
  • OpenCL maskless attention setArg 溢出:修复导致 Adreno 上 SD UNet 输出错误的 setArg 溢出(issue #4653),同时修复 LinearAttention 越界写、共享 runtime gpuMode 被覆盖及 shared_gather 构建错误
  • CUDA pooling pads 与 TopKV2 largest flag:修复显式 pooling pads 与 TopKV2 largest flag 处理,并重组相关测试
  • QNN mmpose 模型:修复运行 mmpose 模型时的若干问题
  • QNN context 生成栈溢出:提高 context 生成的栈上限

CPU 后端

  • Arm82 LinearAttention fp16 精度:改用 fp32 累加器,修复 fp16 精度损失
  • 推测解码 KVCache 索引:修复 CPU KVCache value 索引错误导致 Eagle3 输出质量下降(issue #4591
  • ConvInt8 SIGSEGV:修复 Conv 缺少量化权重时 ConvInt8 崩溃
  • RVV 源文件缺失 Macro.h:补全 RISC-V RVV 源文件的头文件包含
  • RVV Int8 函数修复:修复 MNNLineDepthWiseInt8AddBiasScaleUnit 等 RVV int8 函数问题(issue #4454

框架 / 其它

  • mmap 分配失败崩溃:修复各后端 mmap 分配失败时的崩溃
  • RuntimeManager::setCache 空指针:增加 null runtime 保护
  • Android 构建缺失源文件:CMakeLists 补充 diffusion_session 与 sana_session 源文件

📚 其他改进

  • 开发流程文档:新增 development workflow guidelines
  • CI 脚本整合:整合 CI 测试脚本与文档
  • Shader 构建脚本现代化:支持复杂路径
  • NeuroPilot 代码移出公开仓库:相关实现改由内部维护
  • iOS 发布包更名:移除包名中的 CoreML 后缀
  • Copybara 同步脚本可配置化
  • PyMNN 打包更新:更新 3.6.1 发布的 PyMNN 打包配置

📊 性能基准

combined_bench

相比 3.6.0(同设备、同量化配置,模型用 3.6.1 重新导出):

  • Android 性能大幅回升:Qwen3-4B OpenCL prefill (pp512) +101.7%,Qwen3-4B CPU decode (tg128) +72.2%,Qwen3-0.6B OpenCL decode +76.1%
  • macOS Metal 全线提升:Qwen3-0.6B prefill (pp1024) +50.1%、decode +36.1%;Qwen3-4B prefill (pp512) +18.8%

测试环境: Mac mini M4 (16GB) / OnePlus PJZ110 (Snapdragon 8 Elite) | 模型: Qwen3-0.6B / Qwen3-4B | 量化: W4 (block=64) | 工具: llm_bench (rep=5, median)

Hexagon NPU 后端性能

Qwen3-0.6B(4-bit 对称量化,block64,C4)在 SM8750(骁龙 8 Elite,Hexagon v79)上三后端对比:

后端 pp512 (tok/s) tg128 (tok/s) vs CPU prefill
Hexagon NPU 2667 68.5 7.9×
OpenCL GPU 1491 71.4 4.4×
CPU 336 94.3 1.0×

Hexagon NPU 在 prefill 上优势显著(7.9×),decode 目前仍以 CPU 最优。Hexagon SDK 5.x 与 6.x 编译的 DSP skeleton 性能一致(差异 <2%,误差范围内)。


🙏 致谢

衷心感谢所有贡献者对本版本的宝贵贡献(按 GitHub login 列出):

  • @jxt1234 — Hexagon 后端 (#4633)、fused LLM 算子与 C4 后端路径 (#4554)
  • @wangzhaode — Transformer C4 图全后端启用 (#4629)、C4 导出与 packed RoPE 优化 (#4598)、QLoRA 微调 (#4560)、Qwen3.5 fused 推理修复 (#4631)、C4 fuse guard 与 Metal raster 修复 (#4607)、Vulkan Range 算子 (#4563)、Metal kv_cache flag 修复 (#4556)、embedding scale 导出修复 (#4615)
  • @v0jiuqi — Metal LLM prefill/decode kernel 与 RoPE 基础设施 (#4650)、Fused Q4/Q8 dequant+GEMM (#4651)、llm_bench -fa 选项 (#4658)、Metal ATTENTION_C4 分支修复 (#4654)、MPS Q4 全零权重修复 (#4649)、Arm82 LinearAttention fp16 修复 (#4634)、mmap 失败崩溃修复 (#4630)
  • @Qxinyu — OpenCL NC4HW4 raster / int4 GEMM 优化 (#4606)、c4 长 prefill attention 修复 (#4609)、maskless attention setArg 溢出等修复 (#4659)、attention c4 kernel 编译修复 (#4644)、QNN mmpose 修复 (#4626)、setCache 空指针保护 (#4617)
  • @Fatemanx (xly) — CUDA pooling pads 与 TopKV2 largest flag 修复 (#4627)
  • @LudovicoYIN — QNN 导出 chunk_size 修复 (#4611)、WikiText 校准数据集修复 (#4566)、QNN context 生成栈上限
  • @huluhuluu — 推测解码 KVCache 索引修复 (#4600)
  • @lanling-47 — QNN Interp/Resize 算子 (#4582)
  • @jxgxxx — RVV CommonOptFunction 适配 (#4426)
  • @Sherlockzhangjinge — RVV Int8 函数修复 (#4552)
  • @xuhaonan999 — RVV Macro.h 包含修复 (#4588)
  • @pingji2016 — Android 构建源文件修复 (#4569)
  • @RinZ27 — Shader 构建脚本现代化 (#4550)
  • @devin-lai — decoder embedding causal mask 修复 (#4548)

MNN 3.6.1 Release Notes

Release Date: July 2026

📦 Install: pip install mnn==3.6.1PyPI

📌 Overview

MNN 3.6.1 focuses on three core themes: Hexagon NPU Backend, Transformer C4 Fuse Optimization Across Backends, and Stability Fixes:

  • Hexagon NPU Backend: New Qualcomm Hexagon NPU direct-programming backend. Qwen3-0.6B (W4 symmetric, block64) reaches 2667 tok/s prefill on Snapdragon 8 Elite (Hexagon v79), 7.9× faster than CPU.
  • Transformer C4 Fuse Optimization Across Backends: New fused LLM ops (RoPE / Attention / LayerNorm C4 paths) with Transformer C4 graphs enabled across CPU / Metal / OpenCL / CUDA; optimized C4 export and packed RoPE (RoPE fuse) paths; Metal adds fused Q4/Q8 dequant+GEMM (M_TILE=64) and reworked LLM prefill/decode kernels; OpenCL fixes NC4HW4 raster L2 cache-set thrashing and optimizes int4 GEMM.
  • Stability Fixes & Tooling: New MNN-aware QLoRA finetuning; fixes for mmap allocation failure crashes, speculative-decoding KVCache indexing, Arm82 LinearAttention fp16 accuracy, and more.

🚀 Highlights

  • Hexagon Backend (Qualcomm NPU Direct Programming): Qwen3-0.6B pp512 reaches 2667 tok/s on Snapdragon 8 Elite (7.9× vs CPU, 1.8× vs OpenCL GPU)
  • Transformer C4 Graphs Across Backends: Fused LLM ops + C4 backend paths across CPU / Metal / OpenCL / CUDA, including packed RoPE fuse optimization
  • Metal LLM Prefill Optimization: Fused Q4/Q8 dequant+GEMM kernel, M_TILE=64 variant (for M5), reworked prefill/decode kernels with shape/RoPE infrastructure, plus a new Metal Op Profile tool
  • MNN-aware QLoRA Finetuning: New QLoRA finetuning targeting MNN deployment

✨ New Features

Backend: Hexagon (New)

  • Hexagon NPU Backend: New Qualcomm Hexagon NPU direct-programming backend (HexagonBackend / HexagonRuntime / DSP kernels) with LLM key ops such as Attention; ships with skills/hexagon development docs

LLM / Inference Acceleration

  • Fused LLM Ops & C4 Backend Paths: New fused ops (CPURoPE, SharedGather, etc.) and CPU/Metal C4 compute paths
  • C4 Export & Packed RoPE Optimization: Optimized C4 graph export; packed RoPE (RoPE fuse) paths across CPU / Metal / OpenCL
  • Transformer C4 Graphs Enabled Across Backends: CPU / Metal / OpenCL / CUDA (new RoPEExecution.cu, LayerNorm/LinearAttention C4 adaptation)
  • MNN-aware QLoRA Finetuning

GPU Backends

  • Metal LLM prefill/decode Kernel Optimization: Reworked prefill/decode kernels, shape/RoPE infrastructure, MetalConvolution1x1 SIMD-group paths, and a new Metal Op Profile tool
  • Metal Fused Q4/Q8 dequant+GEMM: New fused dequant+GEMM kernel with M_TILE=64 variant for LLM prefill on M5
  • OpenCL Raster & int4 GEMM Optimization: Fixed NC4HW4 L2 cache-set thrashing by transposing traversal order; optimized int4 GEMM; split binary_layernorm_c4 kernel
  • Vulkan Range Op: New image implementation
  • QNN Interp/Resize Ops: Added to QNN backend

CPU Backend

  • RVV CommonOptFunction Adaptation: Adapted CommonOptFunction for RISC-V RVV

Tools

  • llm_bench -fa Option: New flash attention toggle (default 1, flash attention enabled)

🐛 Bug Fixes

LLM / VLM

  • Qwen3.5 fused inference correctness: Fixed
  • C4 fuse guard & Metal raster: Added guard conditions for C4 fuse and fixed Metal raster
  • RoPE schema compile error: Fixed
  • Silent all-zero Q4 weights on MPS: Fixed when quantizing large tensors
  • Generic embedding scale export: Fixed missing embedding scale on export
  • WikiText calibration dataset loading: Fixed
  • QNN export chunk-size: Honor --chunk_size when generating QNN prefill IO
  • Decoder embedding export causal mask: Use causal mask for decoder embedding export

GPU Backends

  • Metal missing ATTENTION_C4 output branch: Fixed in prefill_qkv_tensor
  • Metal attention kv_cache flag: Now respected
  • OpenCL attention c4 kernel compile: Fixed
  • OpenCL c4 long-prefill attention: Fixed V-buffer OOB in GEMM tune and missing NC4HW4 output layout
  • OpenCL maskless attention setArg overflow: Fixed setArg overflow causing wrong SD UNet output on Adreno (issue #4653), plus LinearAttention OOB writes, shared runtime gpuMode clobbering, and shared_gather build error
  • CUDA pooling pads & TopKV2 largest flag: Fixed explicit pooling pads and TopKV2 largest flag handling; reorganized related tests
  • QNN mmpose model: Fixed several runtime bugs
  • QNN context generation stack limit: Raised

CPU Backend

  • Arm82 LinearAttention fp16 accuracy: Use fp32 accumulators
  • Speculative decoding KVCache indexing: Fixed CPU KVCache value index degrading Eagle3 output quality (issue #4591)
  • ConvInt8 SIGSEGV: Fixed crash when Conv lacks quantized weight
  • Missing Macro.h in RVV sources: Added
  • RVV Int8 function fixes: Fixed MNNLineDepthWiseInt8AddBiasScaleUnit and related RVV int8 functions (issue #4454)

Core / Misc

  • mmap allocation failure crashes: Fixed across all backends
  • RuntimeManager::setCache null runtime: Guarded
  • Android build missing sources: Added diffusion_session and sana_session to CMakeLists

📚 Other Improvements

  • Development workflow guidelines: Added docs
  • CI test scripts consolidation
  • Shader build scripts modernization: Handle complex paths
  • NeuroPilot implementation removed from public repository
  • iOS release package renamed: Removed CoreML suffix
  • Configurable Copybara sync script
  • PyMNN packaging update: Updated PyMNN packaging for the 3.6.1 release

📊 Performance Benchmarks

combined_bench

Compared with 3.6.0 (same devices and quantization config, models re-exported with 3.6.1):

  • Major Android recovery: Qwen3-4B OpenCL prefill (pp512) +101.7%, Qwen3-4B CPU decode (tg128) +72.2%, Qwen3-0.6B OpenCL decode +76.1%
  • macOS Metal improvements across the board: Qwen3-0.6B prefill (pp1024) +50.1%, decode +36.1%; Qwen3-4B prefill (pp512) +18.8%

Environment: Mac mini M4 (16GB) / OnePlus PJZ110 (Snapdragon 8 Elite) | Models: Qwen3-0.6B / Qwen3-4B | Quantization: W4 (block=64) | Tool: llm_bench (rep=5, median)

Hexagon NPU Backend Performance

Qwen3-0.6B (4-bit symmetric, block64, C4) on SM8750 (Snapdragon 8 Elite, Hexagon v79), three-backend comparison:

Backend pp512 (tok/s) tg128 (tok/s) vs CPU prefill
Hexagon NPU 2667 68.5 7.9×
OpenCL GPU 1491 71.4 4.4×
CPU 336 94.3 1.0×

Hexagon NPU excels at prefill (7.9×); CPU remains fastest for decode. DSP skeletons built with Hexagon SDK 5.x and 6.x perform identically (<2% difference, within noise).


🙏 Acknowledgements

We sincerely thank all contributors for their valuable contributions to this release (listed by GitHub login):


Full Changelog: 3.6.0...3.6.1