OfflineLLM v5.1.1
Fixes the CPU performance regression introduced in 5.1.0. If 5.1.0 felt dramatically slower than 5.0.2, this is the release that fixes it — on a Pixel 6a running Qwen3.5-2B Q4_K_M, CPU generation went from under 2 tok/s back to ~10 tok/s.
Also: Android 13 support, llama.cpp 391 commits newer, and new backend diagnostics.
Still zero network permissions. Still cannot phone home.
The 5.1.0 slowdown — what actually happened
Adding the Vulkan GPU backend in 5.1.0 slowed down CPU inference, even with GPU acceleration switched off.
Registering a GPU backend is not the same as choosing to use one. With Vulkan loaded, llama.cpp still placed it in the model's device list at zero offloaded layers. Two things followed:
- The GPU's pinned host buffer took priority over the CPU's weight-repacking buffer type, so quantized weights were never repacked into the layouts ARM's dotprod/i8mm kernels need. Inference silently fell back to generic paths.
- A large Vulkan compute buffer was reserved for chats that never touched the GPU.
The fix restricts the device list to CPU when GPU offload is off, so "GPU off" now means CPU end to end. Turning the GPU toggle off in 5.1.0 did not avoid this — it wasn't something you could work around.
Two smaller wins came out of the same investigation: the logits buffer is now sized for the single token actually sampled rather than a whole batch (497 MiB → 12 MiB on a 248k-vocab model), and memory-mapped loading now defaults off, which is measurably quicker to first token on phones.
Now runs on Android 13
Minimum lowered from Android 14 to Android 13 (API 33). Nothing in the app needed API 34. targetSdk stays at 37, and nothing changes for existing users on 14+.
New: backend diagnostics
Settings → Performance → Active backend shows which ggml backend and CPU kernel variant your device actually selected, with its feature flags. Selectable text, so it pastes straight into a bug report.
DOTPROD+MATMUL_INT8→ the fast quantized-matmul kernels are running- only
NEON→ your device fell back to the armv8.0 baseline; please open an issue
The backend registry is also logged at startup now (logcat tag smollm-backends), listing every device that registered.
llama.cpp updated to b10472 — 391 upstream commits, including current ARM CPU and Vulkan backend work.
Context size is now capped by model size. Modern GGUFs advertise enormous training contexts — Qwen3.5-2B declares 262144 — and honouring that verbatim allocates a KV cache far past what a phone has. An explicit context setting in Settings still wins.
On GPU acceleration: on phones with unified memory, token generation is memory-bandwidth-bound and the GPU shares the same DRAM as the CPU, so there's no bandwidth advantage to win — measured on a Pixel 6a (Mali-G78), CPU is roughly twice as fast as full GPU offload. Adreno-class GPUs generally do better. It remains off by default; if it feels slower on your device, it probably is.
Avoid partial GPU offload in particular. Full offload produced 2 graph splits on our test model; offloading 13 of 24 layers produced 194, each one a CPU↔GPU synchronisation.
- Updated the JNI wrapper for two upstream llama.cpp API changes:
use_mmap/use_mlockbecame a singleload_modeenum, and the repeat-penalty sampler gained a leadingn_vocabargument. - Cleared the deprecated Compose/Hilt/Kotlin APIs surfaced by the dependency bump:
MenuAnchorType→ExposedDropdownMenuAnchorType,hiltViewModel()moved to its own artifact,Icons.Filled.VolumeUp→ the auto-mirrored variant, and the chat-export JSON encoder is no longer rebuilt on every call. - The version shown in About is read from
BuildConfiginstead of a hardcoded string, so it can't drift again. - SPIRV headers are passed to CMake as a cache variable rather than through
cppFlags, which applied to every C++ target and broke on paths containing spaces. .gitignoreno longer missesapp/local.properties, which carries an absolute SDK path.
Dependencies
| from → to | |
|---|---|
| AGP | 9.2.1 → 9.3.1 |
| Kotlin | 2.3.20 → 2.3.21 |
| KSP | 2.3.6 → 2.3.11 |
| Compose BOM | 2026.03.00 → 2026.08.00 |
| Lifecycle | 2.10.0 → 2.11.0 |
| Hilt | 2.59.2 → 2.60.1 |
| Hilt Navigation Compose | 1.2.0 → replaced by Hilt Lifecycle ViewModel Compose 1.4.0 |
| Room | 2.8.0 → 2.8.4 |
| Coroutines | 1.10.2 → 1.11.0 |
| Serialization | 1.10.0 → 1.11.0 |
| Navigation Compose | 2.9.7 → 2.9.8 |
| core-ktx | 1.17.0 → 1.19.0 |
| Biometric | 1.4.0-alpha06 → 1.4.0-alpha07 |
| Markdown renderer | 0.40.2 → 0.41.0 |
Kotlin stays on the 2.3 series deliberately: KSP has no 2.4.x release yet, and both Hilt and Room run through KSP. The markdown renderer is held at 0.41.0 for the same reason — 0.42.0+ is built against Kotlin 2.4.0.
📦 Install
Requires Android 13+ and arm64-v8a.
adb install OfflineLLM_v5.1.1_Signed_Release.apkOr download the APK below, allow install from your file manager, and import a GGUF model from Settings.
Release builds verify their own signing certificate at launch and refuse to run if repackaged.
Full diff: v5.1.0...v5.1.1