Repository navigation
PocketOrca-LLM v1.3.5
PocketOrca-LLM v1.3.5 — First public release
English | 中文
PocketOrca-LLM turns your Snapdragon phone into an offline-capable LLM server for your local network: Hexagon NPU / Adreno GPU / CPU three-engine inference, an OpenAI-compatible endpoint, direct connection from any standard client. Free forever, no ads, no in-app purchases — data never leaves the device.
This app ships no model files — download GGUF models yourself (see the User Guide).
📦 Install
- Download
PocketOrca-LLM-v1.3.5-vc61-release.apk(md56421457e) - Requires Android 8.0+, arm64 device
- On first launch, grant the three permissions in order: battery exemption → all-files access → notifications.
✨ Highlights
- Three engines: Hexagon NPU (htp) / Adreno GPU (ocl) / CPU — pick per device and model quantization
- OpenAI-compatible endpoint:
http://<phone-ip>:8080/v1/chat/completions— OpenWebUI, ChatBox, SillyTavern and any standard client connect directly - Survives backgrounding: stays reachable through screen-off, task-swipe and unplug; auto-revives within 1.5 s if the process is killed
- Live notification monitor: CPU / GPU / RAM / temperature / tok-s / connections
- Sampling playground: temperature / top_k / top_p / min_p / repeat_penalty / system prompt, applied live
- Dual-mode Chat: local streaming; or connect to any remote OpenAI-compatible API
- API keys encrypted (Android Keystore, AES-GCM)
- Trilingual UI (Simplified Chinese / Traditional Chinese / English) + dark/light theme
- Privacy: loading and inference run entirely on-device; nothing is uploaded
📱 Compatibility
CPU engine — theoretically any Android phone released after 2023
Qualcomm / MediaTek / other ARMv8+ chips are all worth a try. Older chips (A53/A55 little cores) run but slowly; small quants (Q4_0 / Q4_K_M) with 4 GB+ RAM recommended. Please test and report back.
GPU engine — Adreno (OpenCL), Snapdragon 8 series
| SoC | Adreno | Status |
|---|---|---|
| Snapdragon 8 Elite Gen 5 | 840 | ✅ Verified |
| Snapdragon 8 Elite | 830 | ✅ Main test device |
| Snapdragon 8 Gen 5 | 830 | |
| Snapdragon 8 Gen 3 | 750 | ✅ Verified |
| Snapdragon 8 Gen 2 | 740 |
MediaTek (Mali) GPUs are not supported for the GPU engine — use CPU.
NPU engine — Hexagon
| SoC | HTP arch | Status |
|---|---|---|
| Snapdragon 8 Elite (SM8750) | v75 | ✅ Verified (~23 t/s at 1B; ~10 t/s at 7B; Qwen3 excluded) |
| Snapdragon 8 Elite Gen 5 (SM8850) | v79 | |
| Snapdragon 8 Gen 5 | Newer Hexagon | |
| Snapdragon 7+ Gen 3 (SM7675) | v73 | |
| Snapdragon 8 Gen 2 | v73 |
Upstream llama.cpp Hexagon support is early days: native NPU kernels exist only for plain 4-bit formats like Q4_0; popular formats like Q4_K_M fall back to CPU. Only Q4_0 (INT4) has been tested so far — largest verified model: Qwen2.5-7B-Q4_0.
All measured speeds were taken on a Galaxy S25 (OneUI 8.0).
Older 7-series chips (Gen 1 / Gen 2, v69) have no NPU engine — use CPU / GPU.
⚠️ Known limitations
- Qwen3 models hang on the NPU engine (upstream llama.cpp bug); the app auto-switches to GPU
- Snapdragon 8 Gen 2 GPU driver defect: Q4_K quants produce garbled output — use CPU instead (device driver issue; may self-heal after an OS update)
- MediaTek devices: CPU engine only
- Sustained inference runs hot — keep the phone plugged in and ventilated
📄 License
MIT for this app's own code. Built on llama.cpp (MIT); the Qualcomm Hexagon components in the prebuilt directories are for building and running on your own Qualcomm hardware only. See LICENSE.
PocketOrca-LLM v1.3.5 — 首个公开发布
English | 中文
PocketOrca-LLM 把你的骁龙手机变成一台可离线运行的局域网大模型服务器:Hexagon NPU / Adreno GPU / CPU 三引擎推理,OpenAI 兼容端点,任何标准客户端直连。永久免费、无广告、无内购,数据永不出设备。
本软件不包含任何模型文件,需自行下载 GGUF 格式模型(见使用手册)。
📦 安装
- 下载
PocketOrca-LLM-v1.3.5-vc61-release.apk(md56421457e) - 系统要求:Android 8.0 及以上,arm64 设备
- 首次启动依次授予电池豁免 / 所有文件访问 / 通知三个权限
✨ 主要功能
- 三引擎按需选择:Hexagon NPU(htp)/ Adreno GPU(ocl)/ CPU,按机型与模型量化自由选择
- OpenAI 兼容端点:
http://<手机IP>:8080/v1/chat/completions,OpenWebUI、ChatBox、SillyTavern 等任何标准客户端直连 - 后台稳定在线:息屏、滑卡、拔电不断连;进程被系统回收后 1.5 秒内自动复活并恢复服务
- 通知栏实时监控:CPU / GPU / RAM / 温度 / 生成速度 / 连接数
- 调参实验台:temperature / top_k / top_p / min_p / repeat_penalty / system prompt 即时生效
- Chat 双模式:本地流式对话;也可直连任意远程 OpenAI 兼容 API
- API Key 加密存储(Android Keystore,AES-GCM)
- 三语界面(简体中文 / 繁體中文 / English)+ 明暗主题
- 隐私:模型加载与推理全部在本机完成,无任何数据上传
📱 兼容性
CPU 引擎 — 理论上支持所有 2023 年之后上市的 Android 手机
高通 / 联发科 / 其他 ARMv8+ 均可尝试。老机型(A53/A55 小核)可跑但速度有限,建议小尺寸量化(Q4_0 / Q4_K_M)+ 4GB 以上 RAM,具体表现请自行测试并反馈。
GPU 引擎 — Adreno (OpenCL),高通 8 系
| SoC | Adreno | 状态 |
|---|---|---|
| 骁龙 8 Elite Gen 5 | 840 | ✅ 已验证 |
| 骁龙 8 Elite | 830 | ✅ 主力实测 |
| 骁龙 8 Gen 5 | 830 | |
| 骁龙 8 Gen 3 | 750 | ✅ 已验证 |
| 骁龙 8 Gen 2 | 740 |
联发科(Mali GPU)暂不支持 GPU 引擎,请使用 CPU 引擎。
NPU 引擎 — Hexagon
| SoC | HTP 架构 | 状态 |
|---|---|---|
| 骁龙 8 Elite (SM8750) | v75 | ✅ 已验证(1B 约 23 t/s;7B 约 10 t/s,Qwen3 系除外) |
| 骁龙 8 Elite Gen 5 (SM8850) | v79 | |
| 骁龙 8 Gen 5 | 新代 Hexagon | |
| 骁龙 7+ Gen 3 (SM7675) | v73 | |
| 骁龙 8 Gen 2 | v73 |
上游 llama.cpp 对 Hexagon 的支持尚处早期,NPU 仅对 Q4_0 等简单 4-bit 格式有原生 kernel;Q4_K_M 等常用格式会回退 CPU。现阶段实测仅覆盖 Q4_0 (INT4),作者最大验证模型为 Qwen2.5-7B-Q4_0。
实测速度均在三星 S25(OneUI 8.0)取得。
更早的 7 系(Gen 1 / Gen 2, v69)不支持 NPU 引擎,请使用 CPU / GPU 引擎。
⚠️ 已知限制
- Qwen3 系模型在 NPU 引擎存在上游僵死问题(llama.cpp 上游 bug),App 会自动引导切换 GPU 引擎
- 骁龙 8 Gen 2 GPU 驱动缺陷:Q4_K 系量化输出乱码,可改用 CPU(设备驱动问题,升 Android 15 可能自愈)
- 联发科机型仅支持 CPU 引擎
- 推理长时间高负载发热明显,建议插电并注意散热
📄 许可证
MIT(本软件自身代码)。基于 llama.cpp (MIT) 构建;预编译目录中的 Qualcomm Hexagon 组件仅限在用户自有高通设备上构建运行使用。详见 LICENSE。