Repository navigation
Releases: ankit1057/sipllm
Release list
v0.5.0 — Plugin runtime + persistent context
SipLLM 0.5.0 — Plugin runtime + persistent context
Turns the bounded-memory engine into an extensible runtime with context/KV intelligence — all opt-in and default-off, so existing behavior is byte-identical (the exact fp32 path remains the numeric oracle).
Highlights
- Plugin seam —
PluginHosthosting two in-process plugins behind stable, mesh-ready contracts:- Kosh (context/token intelligence): transforms the token stream before prefill.
--koshcollapses redundant token runs, with tokens-in/out accounting. - RTK (runtime/KV intelligence): observes KV runtime state (seq-len, peak-KV bytes, steps).
--rtk.
- Kosh (context/token intelligence): transforms the token stream before prefill.
- Cross-turn context reuse (
--reuse): reuse the KV of the longest common prefix already committed and reprocess only the changed tail. A golden test proves reuse ≡ full-prefill (last-position logits within1e-4). - Persistent context (
--save-session F/--load-session F): serialize a session's committed tokens + KV cache to disk (SIPS v1) and restore it in a new process, guarded by amodel_idfingerprint (refuses a KV saved from a different model). Cross-process reuse verified end-to-end.
Notes
- Fully opt-in; without the flags, generation is byte-identical to 0.4.x.
- Verified on macOS (Apple M3): full
make -j4 all && make testgreen, including new suitestest_plugin(9) +test_reuse(3) +test_session(4); existing e2e / stress / arch / sampler suites unchanged. - Prebuilt Linux (x86_64 / aarch64) bundles are attached by CI; macOS builds from source via the installer.
Install
curl -fsSL https://raw.githubusercontent.com/ankit1057/sipllm/main/install.sh | shSee CHANGELOG.md (Waves 9–11) for full detail.
v0.4.0
SipLLM v0.4.0 — Developer Preview: streaming runtime, --ram-budget di…
sipllm v0.1.1
sipllm v0.1.1 — edge-first, CPU-first streaming GGUF inference. This release makes the engine practical on real edge devices and adds stronger, ungated models.
What's new
- Edge-safe default context. The KV cache is allocated for the full context up front, and models now advertise huge windows (Llama 3.2 = 131072 → ~8.5 GB KV cache!). The default is now capped to 4096 tokens (KV cache ~270 MB); raise it with
--ctx Nwhen you have the RAM. Verified on Llama-3.2-1B: 8590 MB → 268 MB, output unchanged. - Better bundled models (all public / ungated, all Llama-architecture, all verified running here):
llama3.2— Llama-3.2-1B-Instruct (:3balso available)smollm2— SmolLM2-1.7B-Instruct (:360mfor a tiny one)tinyllama— 1.1B (smoke tests)
- Architecture roadmap. Support for Gemma, Qwen2, Phi, Mistral, Mixtral-MoE and more is tracked in issue #7. Today the engine implements the Llama architecture only.
Install
curl -fsSL https://raw.githubusercontent.com/ankit1057/sipllm/main/install.sh | sh
sipllm run llama3.2 -p "The capital of India is"Artifacts
sipllm-0.1.1-linux-aarch64.tar.gz— prebuilt CLI + engine + tools (Android/Termux, ARM64 Linux; glibc).install.sh,checksums-linux-aarch64.txt(SHA-256).
Other platforms build from source via the installer. Prebuilt x86_64/macOS: tracked in the issues.
sipllm v0.1.0
sipllm v0.1.0 — a dependency-free streaming GGUF LLM inference engine in C++17. It sips model weights off disk one transformer layer at a time, so peak RAM stays flat (~200–400 MB) regardless of model size.
Install
One line (uses the prebuilt below on linux-aarch64; builds from source otherwise):
curl -fsSL https://raw.githubusercontent.com/ankit1057/sipllm/main/install.sh | sh
sipllm run tinyllama -p "The capital of France is"Validated against llama.cpp
TinyLlama-1.1B, prompt "The capital of France is" → both engines predict " Paris":
| Format | final logit cosine | argmax | peak RSS | result |
|---|---|---|---|---|
| F16 | 1.000000 | ✅ | 412 MB | PASS |
| Q8_0 | 0.999925 | ✅ | 269 MB | PASS |
| Q5_K_M | 0.999829 | ✅ | 223 MB | PASS |
| Q4_K_M | 0.999823 | ✅ | 215 MB | PASS |
F16 is numerically identical to llama.cpp; error grows monotonically as quantization coarsens — the signature of a faithful implementation. Reproduce with python3 golden/validate_matrix.py.
Artifacts
sipllm-0.1.0-linux-aarch64.tar.gz— prebuilt CLI + engine + tools (Android/Termux, ARM64 Linux). glibc, needs no runtime deps beyond libc/libstdc++/pthread.install.sh— the one-line installer.checksums-linux-aarch64.txt— SHA-256.
Other platforms (x86_64, macOS): the installer builds from source automatically (make + a C++17 compiler). Prebuilt x86_64/macOS binaries are on the roadmap.