Repository navigation
Releases: psewar/skirnir
Release list
v0.13.0 — Router 0.7.0 · Agent 0.17.0
Router 0.7.0 · Agent 0.17.0
Game and memory pressure as facts. A game that used little VRAM and little GPU load left the node free, and a model loaded into a nearly full main memory pushed the machine into heavy paging (94 % RAM, 1.8 GB free commit, ~900 hard faults/s). The router could not see either: it inferred "busy" only from foreign VRAM and load, and did not know the main memory at all. Now the agent reports both, and the router acts on facts instead of side effects.
Agent 0.17.0
- Heartbeat block
game: {supported, running, name, via, since, folders}. Every 5 s the agent checks whether a process runs from a game folder: Steam libraries (libraryfolders.vdf) and the install folders of known game publishers from the uninstall entries (without their launchers). When a process path is not readable (anti-cheat), its name counts if a program of at least 20 MB with that name lies in a game folder (generic names such aslauncherorupdaterexcluded). The index is rebuilt every 6 h. Linux reportssupported: false. - Heartbeat block
memory: {ram_available_gib, ram_total_gib, commit_free_gib, commit_limit_gib}(WindowsGlobalMemoryStatusEx; Linux only RAM, since commit is no limit there under overcommit). - New verb
skirnir-agent spielprints game folders, program names, match and memory once. Both blocks also appear in/health.
Router 0.7.0
- A reported game makes the node
busywith reasongameaftermodes.busy_enter.game_sustain_s(10 s) and keeps it busy formodes.busy_exit.game_quiet_s(60 s) after it ends.gpu_busyanswers name the game. Switch off withmodes.busy_enter.game: false. - Memory pressure: when available RAM or free commit stays below
modes.memory_pressure.min_free_gib(8 GiB) formodes.memory_pressure.pollsheartbeats (2), nothing is loaded cold and no prewarm runs. Loaded models keep answering; nothing is unloaded. New reason codememory_pressurewithRetry-After60 s; decision log eventsgameandmemory_pressure. - Home Assistant sensors per node: game, free RAM, free commit.
- All new values are settings in the UI. Agents before 0.17.0 send neither block and keep the old behaviour.
Deutsch: Der Agent meldet jetzt ein laufendes Spiel und den freien Hauptspeicher (RAM, Commit) als Fakt. Ein Spiel macht den Knoten nach 10 s busy (Grund game, 60 s Nachlauf), auch wenn es kaum VRAM belegt. Bei Speicherdruck (RAM oder Commit unter 8 GiB, zwei Heartbeats) lädt der Router nichts kalt und wärmt nicht vor; geladene Modelle antworten weiter, entladen wird nichts. Neuer Grund-Code memory_pressure, neue HA-Sensoren, neues Agent-Verb spiel.
v0.12.9 – model catalog merged field by field (router 0.6.9)
Router 0.6.9, agent unchanged (0.16.0). The model catalog is merged field by field.
- What went wrong: an entry in
roles.yaml(written by the UI, usually a measurement) replaced theconfig.yamlentry of the same model entirely. Fields that onlyconfig.yamlset did not apply:max_parallelfor one model (the limit had to be set again through the UI) and a capability override for another. - Now: both entries are merged field by field. Size fields (
weights_gib,kv_gib_per_1k,real, the measured values) still replace each other as a unit, so new weights are never combined with an old measurement. All other fields fromconfig.yaml(max_parallel,capabilities, notes) keep applying unlessroles.yamlsets them. - New selftest
selftest_katalog_merge.pyin CI. - UI: the full-width chart "Models actually used" collapses to its legend line while there were no requests in the last 10 minutes, instead of leaving an empty row; it comes back with the first request.
Deutsch
Router 0.6.9. Katalogeinträge aus config.yaml und roles.yaml werden Feld für Feld zusammengeführt; bisher ersetzte ein Eintrag in roles.yaml den aus config.yaml ganz. Größenangaben bleiben eine Einheit. Das Diagramm „Effektiv genutzte Modelle“ klappt ohne Anfragen in den letzten 10 Minuten auf seine Legendenzeile zusammen.
v0.12.8 – hysteresis for busy detection (router 0.6.8)
Router 0.6.8, agent unchanged (0.16.0). Hysteresis for busy detection.
- What went wrong: a busy node had to keep meeting the entry condition all the time (GPU load above the threshold, 40 %, without a break for
sustain_s). A single reading below it started the quiet timer, and afterbelow_for_sthe node was free again. A light game at 31–44 % load with 1–2 GiB of foreign VRAM flipped the node between busy and free every 20–60 s (29 times in under an hour). No model was loaded on the node in between, but every free window could have received a request. - Now: leaving busy is harder than entering it. Once busy, a node stays busy while the GPU load is at or above
busy_exit.gpu_util_pct(new, default 20 %, never higher than the node's entry threshold) and foreign VRAM of at leastutil_requires_foreign_gibis in use. Loading screens below that are covered bybelow_for_s. When the game ends, its VRAM goes away and the node becomes free after the quiet period as before. - The end-to-end suite checks that a game with 25 % load and about 1.8 GiB of foreign VRAM keeps the node busy longer than
below_for_s.
Deutsch
Router 0.6.8. Hysterese für die Busy-Erkennung: Einmal busy, bleibt ein Knoten busy, solange die GPU-Auslastung über busy_exit.gpu_util_pct (neu, Vorgabe 20 %) liegt und fremdes VRAM belegt ist. Ein leichtes Spiel mit 31–44 % Last schaltete den Knoten vorher alle 20–60 s zwischen busy und free um.
v0.12.7 – observation mode and a log line per request (router 0.6.7)
Router 0.6.7, agent unchanged (0.16.0). Observation mode for waiting requests and a log line per request.
- Why: with
max_parallel: 1requests wait in the router behind agent requests that take minutes. A short everyday question could often be answered by a free smaller model in the same role; a programming task could not. Before the router acts on that, the decision log should show how often the case occurs, which classes it affects, how much waiting it would save and whether the classifications are right. decision_engine.observe_blocked: true: when a request for a role has to wait, the decision engine classifies it in the background and the decision log recordsdecision_observe: class, engine, confidence, the first free lower local tier of the same role (warm before cold, other nodes first) andwould_fallbackfor confident classifications inobserve_light(defaultstandard,assist).- CPU chain only: embedding, TF-IDF and rules, 10–19 ms on the router host. The LLM stage is left out, because it would queue on the very GPU being waited for. The classification runs as its own task; the request does not wait for it.
decision_engine.observe_all: true: every request to/api/chator/api/generateis classified in the background when it starts; when it ends, the router appends one line toanfragen.jsonlnext to the configuration (observe_path): class, duration, waiting time, tokens, thinking, tool count, client, role, outcome. No prompt contents. A separate file so the decision log keeps its window; it rotates at 20 MB, and the response never waits for the classification. Purpose: check whether the class predicts run time before using it forRetry-Afteror waiting estimates.- Routing does not change.
Deutsch
Router 0.6.7. Beobachtungsmodus: Muss eine Anfrage an eine Rolle warten, ordnet die Decision Engine sie im Hintergrund ein (nur CPU-Kette, nie das LLM), und das Entscheidungslog hält fest, ob eine freie tiefere Stufe derselben Rolle gereicht hätte (decision_observe). Mit observe_all wird jede Anfrage eingeordnet und bekommt eine Zeile in anfragen.jsonl (Klasse, Dauer, Token, keine Inhalte). Am Routing ändert sich nichts.
v0.12.6 – interactive requests move away from a busy warm model (router 0.6.6)
Router 0.6.6, agent unchanged (0.16.0). Interactive requests no longer queue behind a busy warm model.
- What went wrong: warm first chose the warm node even while it was busy. With
max_parallel: 1(since 0.6.5) a voice command could wait up to 60 s behind an agent request of several minutes, although a second node had a free tier for the same role. - Now: for priority
interactive, if a warm model is busy at its limit, the router takes the first free local tier instead: a free warm one first, then a free cold one, and other nodes before the busy one (a fallback model on the busy node would share its GPU with the running request). Cloud tiers are never chosen on this path (cost, data class). The decision log showsreason: ausweichen. Without a busy warm model, warm first decides as before; if no tier is free, the request waits within its interactive budget (60 s). - Normal and batch requests keep waiting for the warm model, which is usually the better one.
- The end-to-end suite checks both: an interactive request moves to the second node immediately, a normal one waits for the warm model.
Deutsch
Router 0.6.6. Interaktive Anfragen (Sprachbefehle) warten nicht mehr hinter einem belegten warmen Modell, wenn eine andere lokale Stufe sofort bedienen kann; im Entscheidungslog steht reason: ausweichen. Normale und Batch-Anfragen warten weiter auf das warme Modell.
v0.12.5 – models without parallel requests wait in the router (router 0.6.5)
Router 0.6.5, agent unchanged (0.16.0). Models without parallel requests wait in the router.
- What went wrong: qwen3.8 (architecture qwen35) cannot run two requests at once in Ollama. When the router sent a second request while one was running, Ollama put it in its own queue and reloaded the runner as soon as the first one finished, although model and parameters were identical. Seen twice in one night, 5–7 s each, context checkpoints and prompt cache lost.
- New catalog field
models.<m>.max_parallel: at most this many concurrent requests per node for that model. A further request waits in the router's queue, with priority and aging, and the runner stays loaded. When a slot frees up, the router wakes the waiter that can actually start (a waiter for a model still at its limit stays queued). The field can also be set in the UI catalog; a measured entry keeps its measurement. - Waiting budget per priority class when a request sets no
deadline_ms:max_wait_interactive_s60 s,max_wait_sfor normal now 600 s instead of 120 s (agent requests often take minutes),max_wait_batch_s1800 s. The queued event in the decision log says whether a request waited for the node (max_inflight) or for the model (max_parallel). - New selftest
selftest_max_parallel.pyin CI; the end-to-end suite checks that two concurrent requests run one after the other with the limit and in parallel without it.
Deutsch
Router 0.6.5. Neues Katalogfeld max_parallel: qwen3.8 kann in Ollama keine zwei Anfragen zugleich, und Ollama lud bei einer zweiten den Runner neu. Jetzt wartet die zweite Anfrage im Router, mit Priorität, und der Runner bleibt. Wartebudget je Prioritätsklasse: interactive 60 s, normal 600 s (statt 120 s), batch 1800 s.
v0.12.4 – errors in the middle of a stream count as errors (router 0.6.4)
Router 0.6.4, agent unchanged (0.16.0). Errors in the middle of a stream now count as errors.
- What went wrong: when a model writes a tool call with broken XML, Ollama ends the streamed answer with an error line (seen:
{"error":"XML syntax error on line 3: unexpected EOF"}) and drops the whole answer. The router passed that line on and counted the request as ok, so these failures were invisible in metrics and logs. - Now: the request counts as outcome
error, the router logs Ollama's message and adds astream_errorentry to the decision log. Ollama clients still get Ollama's line verbatim; Anthropic clients get anerrorevent, OpenAI clients an error object followed by[DONE]. - The node's circuit breaker does not count it: the node works, the model's output was broken.
Deutsch
Router 0.6.4. Bricht Ollama eine Antwort mitten im Stream mit einer Fehlerzeile ab (z. B. kaputtes Werkzeug-XML), zählt die Anfrage jetzt als Fehler, steht im Log und im Entscheidungsprotokoll (stream_error), und Anthropic- bzw. OpenAI-Clients bekommen ein Fehlerereignis in ihrer Form. Der Breaker des Knotens zählt es nicht.
v0.12.3 – TCP keepalive on the inference port (router 0.6.3)
Router 0.6.3, agent unchanged (0.16.0). TCP keepalive on the inference port.
- What went wrong: a non-streaming
/api/chatrequest ran 6.6 minutes. During that time not a single byte flowed between router and client, the connection state on the way to the client expired silently, and the finished answer never arrived. The client waited until its own timeout; the router counted the request as ok. - Now: the inference port sets TCP keepalive (first probe after 60 s of silence, then every 15 s, dead after 4 unanswered probes). The probes keep connection state fresh along the way, and a connection that is dead anyway now shows up as an error instead of a silent loss.
- New selftest
selftest_keepalive.pyin CI checks that accepted connections inherit the options from the listening socket. - Clients that can stream should still do so for long requests: bytes then flow while the model thinks.
Deutsch
Router 0.6.3. TCP-Keepalive auf dem Inferenz-Port (erste Probe nach 60 s, dann alle 15 s, nach 4 ohne Antwort tot). Eine Anfrage ohne Stream lief 6,6 min ohne ein Byte, der Verbindungszustand auf dem Weg verfiel still, die fertige Antwort kam nie an, und der Router zählte sie als ok. Die Proben halten den Zustand frisch und lassen tote Verbindungen als Fehler auffallen.
v0.12.2 – router.anthropic.think (router 0.6.2)
Router 0.6.2, agent unchanged (0.16.0). New setting for thinking on /v1/messages.
router.anthropic.think(auto,on,off; in the UI settings, applies without restart).autofollows the client'sthinkingblock as before;onandoffoverride it. Claude Code asks for adaptive thinking on every turn, and a model that only knows on or off (such as qwen3.8) then thinks on every turn, including trivial ones like reading a file.- The decision log records
thinkfor every request, so a measurement can prove which setting actually reached the model.
Deutsch
Router 0.6.2. Neue Einstellung router.anthropic.think (auto folgt dem Client, on/off fest; in der UI, ohne Neustart). Claude Code verlangt bei jedem Zug adaptives Denken, qwen3.8 denkt dann bei jedem Zug. Das Entscheidungslog führt think je Anfrage.
v0.12.1 – /v1/messages accepts Claude Code's system messages (router 0.6.1)
Router 0.6.1, agent unchanged (0.16.0). Fix for /v1/messages with Claude Code 2.x, found in the first acceptance run against the real CLI.
- What went wrong: Claude Code sends a message with
role: systeminsidemessages(hints to the model after the first user turn, text blocks withcache_control). The public Messages API has no such role, and router 0.6.0 rejected the request with 400messages.1: role must be 'user' or 'assistant'. - Now: such a message goes on to Ollama as a
systemmessage at the same position; empty ones are dropped. thinking: {type: adaptive, display: omitted}, as Claude Code sends it, givesthink: true. Withdisplay: omittedthe router returns thinking blocks without text, as the API does, and sendspingevents while the model thinks, so the stream does not go silent for minutes.- Selftest and end-to-end suite cover the request shape Claude Code 2.1 sends (
?beta=true,context_management,output_config).
Deutsch
Router 0.6.1. Claude Code 2.x schickt innerhalb von messages eine Nachricht mit role: system, die es in der öffentlichen API nicht gibt; 0.6.0 lehnte die Anfrage mit 400 ab. Solche Nachrichten gehen jetzt an derselben Stelle als system-Nachricht an Ollama. thinking: {type: adaptive, display: omitted} ergibt think: true und Denkblöcke ohne Text, mit ping-Ereignissen, solange das Modell denkt.