Skip to content

Releases: PocketOrca/PocketOrca-LLM

v1.4.2

Choose a tag to compare

@PocketOrca PocketOrca released this 08 Oct 06:50

PocketOrca-LLM v1.4.2 Release Notes

Fixes & Improvements

  1. Highlight: multimodal support. After the model service starts, test images and voice notes from the attachment menu on the Chat page. For images, use a vision model (e.g. Qwen3-VL / Qwen2.5-VL) — mmproj is loaded automatically when placed in the same folder as the model. JPEG/PNG are supported; convert HEIC/WebP first, and avoid very large resolutions. For voice notes, pair with a recognition model (e.g. Qwen3-ASR-0.6B+mmproj) to record and transcribe — no configuration needed. Available on NPU/GPU/MTP engines.
  2. Session persistence: save / restore / delete sessions. Restoring rehydrates the KV cache directly (no prefill) and replays the chat history for seamless continuation.
  3. Engine upgrade: llama.cpp baseline update — Q2/Q3_K quantized models now run on the NPU (Q3_K matches Q4_K_M in speed; 9B models save ~1.3 GB RAM by switching to Q3_K_M), and MTP speculative decoding ships with the official baseline.
  4. Fixed NPU startup crash on newer platforms such as the Snapdragon 8 Elite Gen 5 (upstream DMA64 mapping is incompatible with older DSP firmware; mitigated by default — thanks @qwerkilo for reporting the issue)

Install
Download PocketOrca-LLM-v1.4.2-release.apk
Requirements: Android 12 or later, arm64
Signed with the same certificate as v1.4.1 — installs directly over the existing app

v1.4.1

Choose a tag to compare

@PocketOrca PocketOrca released this 04 Oct 10:30

PocketOrca-LLM v1.4.1 Release Notes

English

Fixes & Improvements

  • Major: new MTP engine profile: tested on the Qwen3.5-4B/9B MTP family — full NPU offload + draft-model speculative decoding: Qwen3.5-4B (Samsung S25, Snapdragon 8 Elite) measured 18.9 tok/s, roughly +30% output speed, up to +60% at long context
  • KV cache quantization defaults to q8_0 on every engine profile: roughly half the KV memory
  • Sleep on idle: Off / 3 / 10 / 30 minutes — after the timeout the model and KV cache are unloaded, freeing all memory; the next request reloads automatically
  • Memory watchdog: process memory sampled every second, alert when running low
  • Thinking control: deep-thinking toggle in Chat
  • Anthropic-compatible endpoint: /v1/messages — Claude Code connects straight to the phone (ANTHROPIC_BASE_URL=http://:8080)
  • Fixed corrupted output on the Snapdragon 8 Elite (Adreno 830) GPU engine when the KV offload switch was on; the switch now ships OFF (KV stays with weights - correct output and faster). Turning it on is a RAM-saving escape hatch (passes --no-kv-offload to the engine)
  • Fixed NPU invalid device on some devices (e.g. Ace3P, Snapdragon 8 Gen 3): the system skips extracting a non-lib-prefixed vendor DSP library from the APK; the app now extracts it itself at startup

Install

  • Download PocketOrca-LLM-v1.4.1-release.apk
  • Requires Android 12 or above, arm64
  • Signed with the same certificate as v1.4.0, installs directly over it

PocketOrca-LLM v1.4.0

Choose a tag to compare

@PocketOrca PocketOrca released this 25 Sep 00:18

PocketOrca-LLM v1.4.0 Release Notes

中文 | English

Compatibility

  • Added support for older CPUs in CPU mode — tested and verified on Snapdragon 865

Fixes & Improvements

  • Highlights: six Hexagon NPU updates — NPU now supports Q4_K / Q6_K quantization: Q4_K_M models can be fully offloaded to the NPU (including the KV cache) · HMX GDN acceleration: 1.5–3× NPU prefill speedup for Qwen3.5 / 3.6 series; measured on Qwen3.5-9B Q4_K_M: 7.25 t/s generation, 54.7 t/s prefill
  • Fixed qwen3.5-series crash in CPU mode
  • Context length now supports up to 64K
  • llama.cpp baseline upgrade (58367713a)

Install

  • Download PocketOrca-LLM-v1.4.0-release.apk
  • Requirements: Android 10 or later, arm64
  • Signed with the same certificate as v1.3.8 — installs directly over the existing app

PocketOrca-LLM v1.3.8

Choose a tag to compare

@PocketOrca PocketOrca released this 16 Sep 11:23

PocketOrca-LLM v1.3.8 Release Notes

English | 中文

Compatibility

  • Minimum requirement lowered from Android 13 to Android 10: devices that previously reported "package parse error" — including Huawei HarmonyOS 4.x — now install normally (verified on Mate 70 Pro)
  • Snapdragon 8 Gen 2 GPU garbled-output fix: the GPU engine is usable again on 8 Gen 2 devices (CPU still recommended)
  • NPU: fixed support for IQ4_NL models (verified: Samsung S25 running Qwen3-8B-IQ4_NL.GGUF on the NPU engine)

New

  • Qwen3 family NPU restriction removed: the old hang was an upstream llama.cpp bug, fixed by the upgrade (verified on S25 with Qwen3-8B)
  • New notification-bar status display: idle / starting / running; the running card shows IP:port, RAM and temperature
  • Advanced NPU tuning hook: Download/npullm-htp-env.json injects environment variables into the NPU engine; takes effect after restarting the service

Fixes & Improvements

  • Android 15 long-session crash fix: foreground service type dataSync → specialUse, removing the 6-hour background quota so all-day sessions are no longer interrupted by the system
  • Chat default max generation tokens 512 → 1024
  • llama.cpp upgrade (0905 → 0915, 172 commits)

Install

  • Download PocketOrca-LLM-v1.3.8-vc72-release.apk (~45 MB, md5 d2d8f96a)
  • Requires Android 10+, arm64
  • Same signing certificate as v1.3.5 — install directly over the old version

Building from source (notes)

  • The NPU/GPU inference runtime ocl-libs.tar (~210 MB, md5 aaba86b0) is attached to this release; on Gitee it is split into 3 parts (100 MB per-file limit):
cat ocl-libs.tar.part.00 ocl-libs.tar.part.01 ocl-libs.tar.part.02 > ocl-libs.tar
md5sum ocl-libs.tar   # expect aaba86b0e99fe3f92b4383b3ba8f2377
tar xf ocl-libs.tar
  • On GitHub the attachment is a single ocl-libs.tar.

PocketOrca-LLM v1.3.5

Choose a tag to compare

@PocketOrca PocketOrca released this 14 Sep 06:39

PocketOrca-LLM v1.3.5 — First public release

English | 中文

PocketOrca-LLM turns your Snapdragon phone into an offline-capable LLM server for your local network: Hexagon NPU / Adreno GPU / CPU three-engine inference, an OpenAI-compatible endpoint, direct connection from any standard client. Free forever, no ads, no in-app purchases — data never leaves the device.

This app ships no model files — download GGUF models yourself (see the User Guide).

📦 Install

  • Download PocketOrca-LLM-v1.3.5-vc61-release.apk (md5 6421457e)
  • Requires Android 8.0+, arm64 device
  • On first launch, grant the three permissions in order: battery exemption → all-files access → notifications.

✨ Highlights

  • Three engines: Hexagon NPU (htp) / Adreno GPU (ocl) / CPU — pick per device and model quantization
  • OpenAI-compatible endpoint: http://<phone-ip>:8080/v1/chat/completions — OpenWebUI, ChatBox, SillyTavern and any standard client connect directly
  • Survives backgrounding: stays reachable through screen-off, task-swipe and unplug; auto-revives within 1.5 s if the process is killed
  • Live notification monitor: CPU / GPU / RAM / temperature / tok-s / connections
  • Sampling playground: temperature / top_k / top_p / min_p / repeat_penalty / system prompt, applied live
  • Dual-mode Chat: local streaming; or connect to any remote OpenAI-compatible API
  • API keys encrypted (Android Keystore, AES-GCM)
  • Trilingual UI (Simplified Chinese / Traditional Chinese / English) + dark/light theme
  • Privacy: loading and inference run entirely on-device; nothing is uploaded

📱 Compatibility

CPU engine — theoretically any Android phone released after 2023

Qualcomm / MediaTek / other ARMv8+ chips are all worth a try. Older chips (A53/A55 little cores) run but slowly; small quants (Q4_0 / Q4_K_M) with 4 GB+ RAM recommended. Please test and report back.

GPU engine — Adreno (OpenCL), Snapdragon 8 series

SoC Adreno Status
Snapdragon 8 Elite Gen 5 840 ✅ Verified
Snapdragon 8 Elite 830 ✅ Main test device
Snapdragon 8 Gen 5 830 ⚠️ Untested
Snapdragon 8 Gen 3 750 ✅ Verified
Snapdragon 8 Gen 2 740 ⚠️ Adreno 740 driver bug — garbled output on Q4_K quants

MediaTek (Mali) GPUs are not supported for the GPU engine — use CPU.

NPU engine — Hexagon

SoC HTP arch Status
Snapdragon 8 Elite (SM8750) v75 ✅ Verified (~23 t/s at 1B; ~10 t/s at 7B; Qwen3 excluded)
Snapdragon 8 Elite Gen 5 (SM8850) v79 ⚠️ Untested
Snapdragon 8 Gen 5 Newer Hexagon ⚠️ Untested
Snapdragon 7+ Gen 3 (SM7675) v73 ⚠️ Awaiting a test device
Snapdragon 8 Gen 2 v73 ⚠️ Incomplete upstream support — ≤4B models work, 7B waits on upstream

Upstream llama.cpp Hexagon support is early days: native NPU kernels exist only for plain 4-bit formats like Q4_0; popular formats like Q4_K_M fall back to CPU. Only Q4_0 (INT4) has been tested so far — largest verified model: Qwen2.5-7B-Q4_0.
All measured speeds were taken on a Galaxy S25 (OneUI 8.0).

Older 7-series chips (Gen 1 / Gen 2, v69) have no NPU engine — use CPU / GPU.

⚠️ Known limitations

  • Qwen3 models hang on the NPU engine (upstream llama.cpp bug); the app auto-switches to GPU
  • Snapdragon 8 Gen 2 GPU driver defect: Q4_K quants produce garbled output — use CPU instead (device driver issue; may self-heal after an OS update)
  • MediaTek devices: CPU engine only
  • Sustained inference runs hot — keep the phone plugged in and ventilated

📄 License

MIT for this app's own code. Built on llama.cpp (MIT); the Qualcomm Hexagon components in the prebuilt directories are for building and running on your own Qualcomm hardware only. See LICENSE.


PocketOrca-LLM v1.3.5 — 首个公开发布

English | 中文

PocketOrca-LLM 把你的骁龙手机变成一台可离线运行的局域网大模型服务器:Hexagon NPU / Adreno GPU / CPU 三引擎推理,OpenAI 兼容端点,任何标准客户端直连。永久免费、无广告、无内购,数据永不出设备。

本软件不包含任何模型文件,需自行下载 GGUF 格式模型(见使用手册)。

📦 安装

  • 下载 PocketOrca-LLM-v1.3.5-vc61-release.apk(md5 6421457e)
  • 系统要求:Android 8.0 及以上,arm64 设备
  • 首次启动依次授予电池豁免 / 所有文件访问 / 通知三个权限

✨ 主要功能

  • 三引擎按需选择:Hexagon NPU(htp)/ Adreno GPU(ocl)/ CPU,按机型与模型量化自由选择
  • OpenAI 兼容端点:http://<手机IP>:8080/v1/chat/completions,OpenWebUI、ChatBox、SillyTavern 等任何标准客户端直连
  • 后台稳定在线:息屏、滑卡、拔电不断连;进程被系统回收后 1.5 秒内自动复活并恢复服务
  • 通知栏实时监控:CPU / GPU / RAM / 温度 / 生成速度 / 连接数
  • 调参实验台:temperature / top_k / top_p / min_p / repeat_penalty / system prompt 即时生效
  • Chat 双模式:本地流式对话;也可直连任意远程 OpenAI 兼容 API
  • API Key 加密存储(Android Keystore,AES-GCM)
  • 三语界面(简体中文 / 繁體中文 / English)+ 明暗主题
  • 隐私:模型加载与推理全部在本机完成,无任何数据上传

📱 兼容性

CPU 引擎 — 理论上支持所有 2023 年之后上市的 Android 手机

高通 / 联发科 / 其他 ARMv8+ 均可尝试。老机型(A53/A55 小核)可跑但速度有限,建议小尺寸量化(Q4_0 / Q4_K_M)+ 4GB 以上 RAM,具体表现请自行测试并反馈。

GPU 引擎 — Adreno (OpenCL),高通 8 系

SoC Adreno 状态
骁龙 8 Elite Gen 5 840 ✅ 已验证
骁龙 8 Elite 830 ✅ 主力实测
骁龙 8 Gen 5 830 ⚠️ 待验证
骁龙 8 Gen 3 750 ✅ 已验证
骁龙 8 Gen 2 740 ⚠️ Adreno 740 驱动缺陷,Q4_K 系乱码

联发科(Mali GPU)暂不支持 GPU 引擎,请使用 CPU 引擎。

NPU 引擎 — Hexagon

SoC HTP 架构 状态
骁龙 8 Elite (SM8750) v75 ✅ 已验证(1B 约 23 t/s;7B 约 10 t/s,Qwen3 系除外)
骁龙 8 Elite Gen 5 (SM8850) v79 ⚠️ 待验证
骁龙 8 Gen 5 新代 Hexagon ⚠️ 待验证
骁龙 7+ Gen 3 (SM7675) v73 ⚠️ 待真机验证
骁龙 8 Gen 2 v73 ⚠️ 上游支持不完整,小模型(≤4B)可用,7B 等待上游修复

上游 llama.cpp 对 Hexagon 的支持尚处早期,NPU 仅对 Q4_0 等简单 4-bit 格式有原生 kernel;Q4_K_M 等常用格式会回退 CPU。现阶段实测仅覆盖 Q4_0 (INT4),作者最大验证模型为 Qwen2.5-7B-Q4_0。
实测速度均在三星 S25(OneUI 8.0)取得。

更早的 7 系(Gen 1 / Gen 2, v69)不支持 NPU 引擎,请使用 CPU / GPU 引擎。

⚠️ 已知限制

  • Qwen3 系模型在 NPU 引擎存在上游僵死问题(llama.cpp 上游 bug),App 会自动引导切换 GPU 引擎
  • 骁龙 8 Gen 2 GPU 驱动缺陷:Q4_K 系量化输出乱码,可改用 CPU(设备驱动问题,升 Android 15 可能自愈)
  • 联发科机型仅支持 CPU 引擎
  • 推理长时间高负载发热明显,建议插电并注意散热

📄 许可证

MIT(本软件自身代码)。基于 llama.cpp (MIT) 构建;预编译目录中的 Qualcomm Hexagon 组件仅限在用户自有高通设备上构建运行使用。详见 LICENSE。