Repository navigation
Releases: AvilaCarlosDev/polaris-local-ai
Releases · AvilaCarlosDev/polaris-local-ai
Release list
v0.3.0 — KV cache survives restarts
The router now saves llama-server's KV cache to disk before every restart (model swap, or freeing VRAM for SD) and restores it when the same model comes back, using llama-server's --slot-save-path.
Measured on real hardware
Same box (Radeon RX 580 8 GB), qwen2.5-7b, a 19,845-token prompt:
| Scenario | Before | After |
|---|---|---|
| Swap back + 20K-token request | 166.8 s | 4.5 s (cached_tokens: 19,844) |
| Image generation → chat | 167.9 s | 9.5 s |
| Save (19,846 tokens → 577 MB file) | — | 0.2–1.3 s |
| Restore | — | 0.2 s |
| First cold request | 53.4 s | 53.4 s (nothing to restore) |
What disappears is the huge prefill; the process restart itself (~6–10 s) still happens. Restore is same-model only (the server validates) and each warm model costs 577 MB of disk.
Hardening found while validating
- llama-server writes a 36-byte header file even for empty slots — saves are now promoted from a temp file only when
n_saved > 0, so a save against a freshly started server can no longer clobber the good 577 MB file. - The router's state file moved from
/run(tmpfs, lost on restart) to/var/lib/llama-router/model, so a deploy or router crash no longer silently skips saves.
Full changelog: https://github.com/AvilaCarlosDev/polaris-local-ai/blob/main/CHANGELOG.md
v0.2.0 — category selector, whisper and local hero videos
Highlights
- Category selector: every id in
GET /v1/modelsnow carries acategory
(texto,vision,multitarea,imagen,audio) and the list is sorted by
it.clients/ia-modelsprints the router's models grouped by that same
category. - Whisper, packaged:
systemd/whisper-server.servicebuilt bysetup.sh,
advertised only whenggml-medium.binexists — no crash loops. - Harmony tests:
tests/test_units.pykeeps the systemd units and the
installer from drifting apart (run without root, systemd or a GPU). - Local hero videos: a real run of the stack (
hero.mp4) and the
FLOATING ISLAND MIRAGE voxel diorama (hero-scene-720.mp4, master in
hero-scene.mp4), 870 frames rendered with Pillow and encoded with ffmpeg
on the RX 580 machine. The title came from the localqwen2.5-7b.
Measurements live indocs/evidence/.
Fixes
sd-server-sd15.servicehad no[Install]section, so enabling it silently
did nothing after a reboot.setup.shnever copiedrouter.py/image-mcp.pyinto/opt/ia.- The 0.1.0 changelog listed units that were never part of this repository.
Assets
| File | What |
|---|---|
hero.mp4 |
terminal demo, 1600x900, 29 s |
hero-scene-720.mp4 |
voxel scene, 1280x720, 4.5 MB (README embed) |
hero-scene.mp4 |
voxel scene master, 1600x900, 13.2 MB |
*.webp |
click-through previews |
v0.1.0 — first public release
First public release: a complete local AI stack on an AMD Radeon RX 580 2048SP (8 GB VRAM) with 32 GB of RAM.
Added
- OpenAI-compatible router (
router.py): one endpoint for six text/vision model ids plus image generation, with aliases (fast,general,ornith,imagen, …), Bearer auth viaIA_API_KEY, SSE streaming and on-demand model swapping. The swap is serialised under a process lock so concurrent clients queue instead of restartingllama-serveron top of each other. - Image generation: SD 3.5 Medium and SD 1.5 behind the same endpoint, with
image-mcp.pyas an MCP bridge for agents andclients/ia-imagenas a CLI that saves the result locally. - Install script (
setup.sh): checks GPU, Vulkan driver, RAM and disk before touching anything, builds what is missing and installs the systemd units. It refuses to run on hardware that cannot run the stack. - Systemd units (
systemd/): router, image bridge, Tailscale exposure and Whisper, each with its own health and ordering rules. - Measured docs (
docs/): BENCHMARKS (cold vs warm, agent latency, all measured on real hardware), EXPECTATIONS, HARDWARE, MODELS, SETUP, TROUBLESHOOTING and HERMES (using the stack as an agent engine). - CI:
rufffor Python lint,pytestfor router contract tests (model table, aliases, image ids, unit wiring) andshellcheckforsetup.sh, running on pushes and pull requests. - README hero capture: a real transcript — one
curlchat completion and oneia-imagengeneration — composited with the generated image.