Skip to content

Releases: ankit1057/sipllm

v0.5.0 — Plugin runtime + persistent context

Choose a tag to compare

@ankit1057 ankit1057 released this 14 Aug 04:52

SipLLM 0.5.0 — Plugin runtime + persistent context

Turns the bounded-memory engine into an extensible runtime with context/KV intelligence — all opt-in and default-off, so existing behavior is byte-identical (the exact fp32 path remains the numeric oracle).

Highlights

  • Plugin seam — PluginHost hosting two in-process plugins behind stable, mesh-ready contracts:
    • Kosh (context/token intelligence): transforms the token stream before prefill. --kosh collapses redundant token runs, with tokens-in/out accounting.
    • RTK (runtime/KV intelligence): observes KV runtime state (seq-len, peak-KV bytes, steps). --rtk.
  • Cross-turn context reuse (--reuse): reuse the KV of the longest common prefix already committed and reprocess only the changed tail. A golden test proves reuse ≡ full-prefill (last-position logits within 1e-4).
  • Persistent context (--save-session F / --load-session F): serialize a session's committed tokens + KV cache to disk (SIPS v1) and restore it in a new process, guarded by a model_id fingerprint (refuses a KV saved from a different model). Cross-process reuse verified end-to-end.

Notes

  • Fully opt-in; without the flags, generation is byte-identical to 0.4.x.
  • Verified on macOS (Apple M3): full make -j4 all && make test green, including new suites test_plugin (9) + test_reuse (3) + test_session (4); existing e2e / stress / arch / sampler suites unchanged.
  • Prebuilt Linux (x86_64 / aarch64) bundles are attached by CI; macOS builds from source via the installer.

Install

curl -fsSL https://raw.githubusercontent.com/ankit1057/sipllm/main/install.sh | sh

See CHANGELOG.md (Waves 9–11) for full detail.

v0.4.0

Choose a tag to compare

@github-actions github-actions released this 27 Jul 06:01
SipLLM v0.4.0 — Developer Preview: streaming runtime, --ram-budget di…

sipllm v0.1.1

Choose a tag to compare

@ankit1057 ankit1057 released this 23 Jul 04:32

sipllm v0.1.1 — edge-first, CPU-first streaming GGUF inference. This release makes the engine practical on real edge devices and adds stronger, ungated models.

What's new

  • Edge-safe default context. The KV cache is allocated for the full context up front, and models now advertise huge windows (Llama 3.2 = 131072 → ~8.5 GB KV cache!). The default is now capped to 4096 tokens (KV cache ~270 MB); raise it with --ctx N when you have the RAM. Verified on Llama-3.2-1B: 8590 MB → 268 MB, output unchanged.
  • Better bundled models (all public / ungated, all Llama-architecture, all verified running here):
    • llama3.2 — Llama-3.2-1B-Instruct (:3b also available)
    • smollm2 — SmolLM2-1.7B-Instruct (:360m for a tiny one)
    • tinyllama — 1.1B (smoke tests)
  • Architecture roadmap. Support for Gemma, Qwen2, Phi, Mistral, Mixtral-MoE and more is tracked in issue #7. Today the engine implements the Llama architecture only.

Install

curl -fsSL https://raw.githubusercontent.com/ankit1057/sipllm/main/install.sh | sh
sipllm run llama3.2 -p "The capital of India is"

Artifacts

  • sipllm-0.1.1-linux-aarch64.tar.gz — prebuilt CLI + engine + tools (Android/Termux, ARM64 Linux; glibc).
  • install.sh, checksums-linux-aarch64.txt (SHA-256).

Other platforms build from source via the installer. Prebuilt x86_64/macOS: tracked in the issues.

sipllm v0.1.0

Choose a tag to compare

@ankit1057 ankit1057 released this 23 Jul 03:12

sipllm v0.1.0 — a dependency-free streaming GGUF LLM inference engine in C++17. It sips model weights off disk one transformer layer at a time, so peak RAM stays flat (~200–400 MB) regardless of model size.

Install

One line (uses the prebuilt below on linux-aarch64; builds from source otherwise):

curl -fsSL https://raw.githubusercontent.com/ankit1057/sipllm/main/install.sh | sh
sipllm run tinyllama -p "The capital of France is"

Validated against llama.cpp

TinyLlama-1.1B, prompt "The capital of France is" → both engines predict " Paris":

Format final logit cosine argmax peak RSS result
F16 1.000000 ✅ 412 MB PASS
Q8_0 0.999925 ✅ 269 MB PASS
Q5_K_M 0.999829 ✅ 223 MB PASS
Q4_K_M 0.999823 ✅ 215 MB PASS

F16 is numerically identical to llama.cpp; error grows monotonically as quantization coarsens — the signature of a faithful implementation. Reproduce with python3 golden/validate_matrix.py.

Artifacts

  • sipllm-0.1.0-linux-aarch64.tar.gz — prebuilt CLI + engine + tools (Android/Termux, ARM64 Linux). glibc, needs no runtime deps beyond libc/libstdc++/pthread.
  • install.sh — the one-line installer.
  • checksums-linux-aarch64.txt — SHA-256.

Other platforms (x86_64, macOS): the installer builds from source automatically (make + a C++17 compiler). Prebuilt x86_64/macOS binaries are on the roadmap.