Skip to content

Releases: psewar/skirnir

v0.13.0 — Router 0.7.0 · Agent 0.17.0

Choose a tag to compare

@psewar psewar released this 05 Oct 18:48

Router 0.7.0 · Agent 0.17.0

Game and memory pressure as facts. A game that used little VRAM and little GPU load left the node free, and a model loaded into a nearly full main memory pushed the machine into heavy paging (94 % RAM, 1.8 GB free commit, ~900 hard faults/s). The router could not see either: it inferred "busy" only from foreign VRAM and load, and did not know the main memory at all. Now the agent reports both, and the router acts on facts instead of side effects.

Agent 0.17.0

  • Heartbeat block game: {supported, running, name, via, since, folders}. Every 5 s the agent checks whether a process runs from a game folder: Steam libraries (libraryfolders.vdf) and the install folders of known game publishers from the uninstall entries (without their launchers). When a process path is not readable (anti-cheat), its name counts if a program of at least 20 MB with that name lies in a game folder (generic names such as launcher or updater excluded). The index is rebuilt every 6 h. Linux reports supported: false.
  • Heartbeat block memory: {ram_available_gib, ram_total_gib, commit_free_gib, commit_limit_gib} (Windows GlobalMemoryStatusEx; Linux only RAM, since commit is no limit there under overcommit).
  • New verb skirnir-agent spiel prints game folders, program names, match and memory once. Both blocks also appear in /health.

Router 0.7.0

  • A reported game makes the node busy with reason game after modes.busy_enter.game_sustain_s (10 s) and keeps it busy for modes.busy_exit.game_quiet_s (60 s) after it ends. gpu_busy answers name the game. Switch off with modes.busy_enter.game: false.
  • Memory pressure: when available RAM or free commit stays below modes.memory_pressure.min_free_gib (8 GiB) for modes.memory_pressure.polls heartbeats (2), nothing is loaded cold and no prewarm runs. Loaded models keep answering; nothing is unloaded. New reason code memory_pressure with Retry-After 60 s; decision log events game and memory_pressure.
  • Home Assistant sensors per node: game, free RAM, free commit.
  • All new values are settings in the UI. Agents before 0.17.0 send neither block and keep the old behaviour.

Deutsch: Der Agent meldet jetzt ein laufendes Spiel und den freien Hauptspeicher (RAM, Commit) als Fakt. Ein Spiel macht den Knoten nach 10 s busy (Grund game, 60 s Nachlauf), auch wenn es kaum VRAM belegt. Bei Speicherdruck (RAM oder Commit unter 8 GiB, zwei Heartbeats) lädt der Router nichts kalt und wärmt nicht vor; geladene Modelle antworten weiter, entladen wird nichts. Neuer Grund-Code memory_pressure, neue HA-Sensoren, neues Agent-Verb spiel.

v0.12.9 – model catalog merged field by field (router 0.6.9)

Choose a tag to compare

@psewar psewar released this 04 Oct 21:17

Router 0.6.9, agent unchanged (0.16.0). The model catalog is merged field by field.

  • What went wrong: an entry in roles.yaml (written by the UI, usually a measurement) replaced the config.yaml entry of the same model entirely. Fields that only config.yaml set did not apply: max_parallel for one model (the limit had to be set again through the UI) and a capability override for another.
  • Now: both entries are merged field by field. Size fields (weights_gib, kv_gib_per_1k, real, the measured values) still replace each other as a unit, so new weights are never combined with an old measurement. All other fields from config.yaml (max_parallel, capabilities, notes) keep applying unless roles.yaml sets them.
  • New selftest selftest_katalog_merge.py in CI.
  • UI: the full-width chart "Models actually used" collapses to its legend line while there were no requests in the last 10 minutes, instead of leaving an empty row; it comes back with the first request.
Deutsch

Router 0.6.9. Katalogeinträge aus config.yaml und roles.yaml werden Feld für Feld zusammengeführt; bisher ersetzte ein Eintrag in roles.yaml den aus config.yaml ganz. Größenangaben bleiben eine Einheit. Das Diagramm „Effektiv genutzte Modelle“ klappt ohne Anfragen in den letzten 10 Minuten auf seine Legendenzeile zusammen.

v0.12.8 – hysteresis for busy detection (router 0.6.8)

Choose a tag to compare

@psewar psewar released this 04 Oct 21:03

Router 0.6.8, agent unchanged (0.16.0). Hysteresis for busy detection.

  • What went wrong: a busy node had to keep meeting the entry condition all the time (GPU load above the threshold, 40 %, without a break for sustain_s). A single reading below it started the quiet timer, and after below_for_s the node was free again. A light game at 31–44 % load with 1–2 GiB of foreign VRAM flipped the node between busy and free every 20–60 s (29 times in under an hour). No model was loaded on the node in between, but every free window could have received a request.
  • Now: leaving busy is harder than entering it. Once busy, a node stays busy while the GPU load is at or above busy_exit.gpu_util_pct (new, default 20 %, never higher than the node's entry threshold) and foreign VRAM of at least util_requires_foreign_gib is in use. Loading screens below that are covered by below_for_s. When the game ends, its VRAM goes away and the node becomes free after the quiet period as before.
  • The end-to-end suite checks that a game with 25 % load and about 1.8 GiB of foreign VRAM keeps the node busy longer than below_for_s.
Deutsch

Router 0.6.8. Hysterese für die Busy-Erkennung: Einmal busy, bleibt ein Knoten busy, solange die GPU-Auslastung über busy_exit.gpu_util_pct (neu, Vorgabe 20 %) liegt und fremdes VRAM belegt ist. Ein leichtes Spiel mit 31–44 % Last schaltete den Knoten vorher alle 20–60 s zwischen busy und free um.

v0.12.7 – observation mode and a log line per request (router 0.6.7)

Choose a tag to compare

@psewar psewar released this 02 Oct 09:36

Router 0.6.7, agent unchanged (0.16.0). Observation mode for waiting requests and a log line per request.

  • Why: with max_parallel: 1 requests wait in the router behind agent requests that take minutes. A short everyday question could often be answered by a free smaller model in the same role; a programming task could not. Before the router acts on that, the decision log should show how often the case occurs, which classes it affects, how much waiting it would save and whether the classifications are right.
  • decision_engine.observe_blocked: true: when a request for a role has to wait, the decision engine classifies it in the background and the decision log records decision_observe: class, engine, confidence, the first free lower local tier of the same role (warm before cold, other nodes first) and would_fallback for confident classifications in observe_light (default standard, assist).
  • CPU chain only: embedding, TF-IDF and rules, 10–19 ms on the router host. The LLM stage is left out, because it would queue on the very GPU being waited for. The classification runs as its own task; the request does not wait for it.
  • decision_engine.observe_all: true: every request to /api/chat or /api/generate is classified in the background when it starts; when it ends, the router appends one line to anfragen.jsonl next to the configuration (observe_path): class, duration, waiting time, tokens, thinking, tool count, client, role, outcome. No prompt contents. A separate file so the decision log keeps its window; it rotates at 20 MB, and the response never waits for the classification. Purpose: check whether the class predicts run time before using it for Retry-After or waiting estimates.
  • Routing does not change.
Deutsch

Router 0.6.7. Beobachtungsmodus: Muss eine Anfrage an eine Rolle warten, ordnet die Decision Engine sie im Hintergrund ein (nur CPU-Kette, nie das LLM), und das Entscheidungslog hält fest, ob eine freie tiefere Stufe derselben Rolle gereicht hätte (decision_observe). Mit observe_all wird jede Anfrage eingeordnet und bekommt eine Zeile in anfragen.jsonl (Klasse, Dauer, Token, keine Inhalte). Am Routing ändert sich nichts.

v0.12.6 – interactive requests move away from a busy warm model (router 0.6.6)

Choose a tag to compare

@psewar psewar released this 02 Oct 07:25

Router 0.6.6, agent unchanged (0.16.0). Interactive requests no longer queue behind a busy warm model.

  • What went wrong: warm first chose the warm node even while it was busy. With max_parallel: 1 (since 0.6.5) a voice command could wait up to 60 s behind an agent request of several minutes, although a second node had a free tier for the same role.
  • Now: for priority interactive, if a warm model is busy at its limit, the router takes the first free local tier instead: a free warm one first, then a free cold one, and other nodes before the busy one (a fallback model on the busy node would share its GPU with the running request). Cloud tiers are never chosen on this path (cost, data class). The decision log shows reason: ausweichen. Without a busy warm model, warm first decides as before; if no tier is free, the request waits within its interactive budget (60 s).
  • Normal and batch requests keep waiting for the warm model, which is usually the better one.
  • The end-to-end suite checks both: an interactive request moves to the second node immediately, a normal one waits for the warm model.
Deutsch

Router 0.6.6. Interaktive Anfragen (Sprachbefehle) warten nicht mehr hinter einem belegten warmen Modell, wenn eine andere lokale Stufe sofort bedienen kann; im Entscheidungslog steht reason: ausweichen. Normale und Batch-Anfragen warten weiter auf das warme Modell.

v0.12.5 – models without parallel requests wait in the router (router 0.6.5)

Choose a tag to compare

@psewar psewar released this 02 Oct 06:51

Router 0.6.5, agent unchanged (0.16.0). Models without parallel requests wait in the router.

  • What went wrong: qwen3.8 (architecture qwen35) cannot run two requests at once in Ollama. When the router sent a second request while one was running, Ollama put it in its own queue and reloaded the runner as soon as the first one finished, although model and parameters were identical. Seen twice in one night, 5–7 s each, context checkpoints and prompt cache lost.
  • New catalog field models.<m>.max_parallel: at most this many concurrent requests per node for that model. A further request waits in the router's queue, with priority and aging, and the runner stays loaded. When a slot frees up, the router wakes the waiter that can actually start (a waiter for a model still at its limit stays queued). The field can also be set in the UI catalog; a measured entry keeps its measurement.
  • Waiting budget per priority class when a request sets no deadline_ms: max_wait_interactive_s 60 s, max_wait_s for normal now 600 s instead of 120 s (agent requests often take minutes), max_wait_batch_s 1800 s. The queued event in the decision log says whether a request waited for the node (max_inflight) or for the model (max_parallel).
  • New selftest selftest_max_parallel.py in CI; the end-to-end suite checks that two concurrent requests run one after the other with the limit and in parallel without it.
Deutsch

Router 0.6.5. Neues Katalogfeld max_parallel: qwen3.8 kann in Ollama keine zwei Anfragen zugleich, und Ollama lud bei einer zweiten den Runner neu. Jetzt wartet die zweite Anfrage im Router, mit Priorität, und der Runner bleibt. Wartebudget je Prioritätsklasse: interactive 60 s, normal 600 s (statt 120 s), batch 1800 s.

v0.12.4 – errors in the middle of a stream count as errors (router 0.6.4)

Choose a tag to compare

@psewar psewar released this 02 Oct 06:03

Router 0.6.4, agent unchanged (0.16.0). Errors in the middle of a stream now count as errors.

  • What went wrong: when a model writes a tool call with broken XML, Ollama ends the streamed answer with an error line (seen: {"error":"XML syntax error on line 3: unexpected EOF"}) and drops the whole answer. The router passed that line on and counted the request as ok, so these failures were invisible in metrics and logs.
  • Now: the request counts as outcome error, the router logs Ollama's message and adds a stream_error entry to the decision log. Ollama clients still get Ollama's line verbatim; Anthropic clients get an error event, OpenAI clients an error object followed by [DONE].
  • The node's circuit breaker does not count it: the node works, the model's output was broken.
Deutsch

Router 0.6.4. Bricht Ollama eine Antwort mitten im Stream mit einer Fehlerzeile ab (z. B. kaputtes Werkzeug-XML), zählt die Anfrage jetzt als Fehler, steht im Log und im Entscheidungsprotokoll (stream_error), und Anthropic- bzw. OpenAI-Clients bekommen ein Fehlerereignis in ihrer Form. Der Breaker des Knotens zählt es nicht.

v0.12.3 – TCP keepalive on the inference port (router 0.6.3)

Choose a tag to compare

@psewar psewar released this 02 Oct 00:34

Router 0.6.3, agent unchanged (0.16.0). TCP keepalive on the inference port.

  • What went wrong: a non-streaming /api/chat request ran 6.6 minutes. During that time not a single byte flowed between router and client, the connection state on the way to the client expired silently, and the finished answer never arrived. The client waited until its own timeout; the router counted the request as ok.
  • Now: the inference port sets TCP keepalive (first probe after 60 s of silence, then every 15 s, dead after 4 unanswered probes). The probes keep connection state fresh along the way, and a connection that is dead anyway now shows up as an error instead of a silent loss.
  • New selftest selftest_keepalive.py in CI checks that accepted connections inherit the options from the listening socket.
  • Clients that can stream should still do so for long requests: bytes then flow while the model thinks.
Deutsch

Router 0.6.3. TCP-Keepalive auf dem Inferenz-Port (erste Probe nach 60 s, dann alle 15 s, nach 4 ohne Antwort tot). Eine Anfrage ohne Stream lief 6,6 min ohne ein Byte, der Verbindungszustand auf dem Weg verfiel still, die fertige Antwort kam nie an, und der Router zählte sie als ok. Die Proben halten den Zustand frisch und lassen tote Verbindungen als Fehler auffallen.

v0.12.2 – router.anthropic.think (router 0.6.2)

Choose a tag to compare

@psewar psewar released this 01 Oct 11:40

Router 0.6.2, agent unchanged (0.16.0). New setting for thinking on /v1/messages.

  • router.anthropic.think (auto, on, off; in the UI settings, applies without restart). auto follows the client's thinking block as before; on and off override it. Claude Code asks for adaptive thinking on every turn, and a model that only knows on or off (such as qwen3.8) then thinks on every turn, including trivial ones like reading a file.
  • The decision log records think for every request, so a measurement can prove which setting actually reached the model.
Deutsch

Router 0.6.2. Neue Einstellung router.anthropic.think (auto folgt dem Client, on/off fest; in der UI, ohne Neustart). Claude Code verlangt bei jedem Zug adaptives Denken, qwen3.8 denkt dann bei jedem Zug. Das Entscheidungslog führt think je Anfrage.

v0.12.1 – /v1/messages accepts Claude Code's system messages (router 0.6.1)

Choose a tag to compare

@psewar psewar released this 01 Oct 07:36

Router 0.6.1, agent unchanged (0.16.0). Fix for /v1/messages with Claude Code 2.x, found in the first acceptance run against the real CLI.

  • What went wrong: Claude Code sends a message with role: system inside messages (hints to the model after the first user turn, text blocks with cache_control). The public Messages API has no such role, and router 0.6.0 rejected the request with 400 messages.1: role must be 'user' or 'assistant'.
  • Now: such a message goes on to Ollama as a system message at the same position; empty ones are dropped.
  • thinking: {type: adaptive, display: omitted}, as Claude Code sends it, gives think: true. With display: omitted the router returns thinking blocks without text, as the API does, and sends ping events while the model thinks, so the stream does not go silent for minutes.
  • Selftest and end-to-end suite cover the request shape Claude Code 2.1 sends (?beta=true, context_management, output_config).
Deutsch

Router 0.6.1. Claude Code 2.x schickt innerhalb von messages eine Nachricht mit role: system, die es in der öffentlichen API nicht gibt; 0.6.0 lehnte die Anfrage mit 400 ab. Solche Nachrichten gehen jetzt an derselben Stelle als system-Nachricht an Ollama. thinking: {type: adaptive, display: omitted} ergibt think: true und Denkblöcke ohne Text, mit ping-Ereignissen, solange das Modell denkt.