Skip to content

Releases: AvilaCarlosDev/polaris-local-ai

v0.3.0 — KV cache survives restarts

Choose a tag to compare

@AvilaCarlosDev AvilaCarlosDev released this 06 Oct 18:20

The router now saves llama-server's KV cache to disk before every restart (model swap, or freeing VRAM for SD) and restores it when the same model comes back, using llama-server's --slot-save-path.

Measured on real hardware

Same box (Radeon RX 580 8 GB), qwen2.5-7b, a 19,845-token prompt:

Scenario Before After
Swap back + 20K-token request 166.8 s 4.5 s (cached_tokens: 19,844)
Image generation → chat 167.9 s 9.5 s
Save (19,846 tokens → 577 MB file) — 0.2–1.3 s
Restore — 0.2 s
First cold request 53.4 s 53.4 s (nothing to restore)

What disappears is the huge prefill; the process restart itself (~6–10 s) still happens. Restore is same-model only (the server validates) and each warm model costs 577 MB of disk.

Hardening found while validating

  • llama-server writes a 36-byte header file even for empty slots — saves are now promoted from a temp file only when n_saved > 0, so a save against a freshly started server can no longer clobber the good 577 MB file.
  • The router's state file moved from /run (tmpfs, lost on restart) to /var/lib/llama-router/model, so a deploy or router crash no longer silently skips saves.

Full changelog: https://github.com/AvilaCarlosDev/polaris-local-ai/blob/main/CHANGELOG.md

v0.2.0 — category selector, whisper and local hero videos

Choose a tag to compare

@AvilaCarlosDev AvilaCarlosDev released this 05 Oct 23:32

Highlights

  • Category selector: every id in GET /v1/models now carries a category
    (texto, vision, multitarea, imagen, audio) and the list is sorted by
    it. clients/ia-models prints the router's models grouped by that same
    category.
  • Whisper, packaged: systemd/whisper-server.service built by setup.sh,
    advertised only when ggml-medium.bin exists — no crash loops.
  • Harmony tests: tests/test_units.py keeps the systemd units and the
    installer from drifting apart (run without root, systemd or a GPU).
  • Local hero videos: a real run of the stack (hero.mp4) and the
    FLOATING ISLAND MIRAGE voxel diorama (hero-scene-720.mp4, master in
    hero-scene.mp4), 870 frames rendered with Pillow and encoded with ffmpeg
    on the RX 580 machine. The title came from the local qwen2.5-7b.
    Measurements live in docs/evidence/.

Fixes

  • sd-server-sd15.service had no [Install] section, so enabling it silently
    did nothing after a reboot.
  • setup.sh never copied router.py / image-mcp.py into /opt/ia.
  • The 0.1.0 changelog listed units that were never part of this repository.

Assets

File What
hero.mp4 terminal demo, 1600x900, 29 s
hero-scene-720.mp4 voxel scene, 1280x720, 4.5 MB (README embed)
hero-scene.mp4 voxel scene master, 1600x900, 13.2 MB
*.webp click-through previews

v0.1.0 — first public release

Choose a tag to compare

@AvilaCarlosDev AvilaCarlosDev released this 04 Oct 22:27

First public release: a complete local AI stack on an AMD Radeon RX 580 2048SP (8 GB VRAM) with 32 GB of RAM.

Added

  • OpenAI-compatible router (router.py): one endpoint for six text/vision model ids plus image generation, with aliases (fast, general, ornith, imagen, …), Bearer auth via IA_API_KEY, SSE streaming and on-demand model swapping. The swap is serialised under a process lock so concurrent clients queue instead of restarting llama-server on top of each other.
  • Image generation: SD 3.5 Medium and SD 1.5 behind the same endpoint, with image-mcp.py as an MCP bridge for agents and clients/ia-imagen as a CLI that saves the result locally.
  • Install script (setup.sh): checks GPU, Vulkan driver, RAM and disk before touching anything, builds what is missing and installs the systemd units. It refuses to run on hardware that cannot run the stack.
  • Systemd units (systemd/): router, image bridge, Tailscale exposure and Whisper, each with its own health and ordering rules.
  • Measured docs (docs/): BENCHMARKS (cold vs warm, agent latency, all measured on real hardware), EXPECTATIONS, HARDWARE, MODELS, SETUP, TROUBLESHOOTING and HERMES (using the stack as an agent engine).
  • CI: ruff for Python lint, pytest for router contract tests (model table, aliases, image ids, unit wiring) and shellcheck for setup.sh, running on pushes and pull requests.
  • README hero capture: a real transcript — one curl chat completion and one ia-imagen generation — composited with the generated image.