Skip to content

v0.12.7 – observation mode and a log line per request (router 0.6.7)

Choose a tag to compare

@psewar psewar released this 02 Oct 09:36
· 3 commits to main since this release

Router 0.6.7, agent unchanged (0.16.0). Observation mode for waiting requests and a log line per request.

  • Why: with max_parallel: 1 requests wait in the router behind agent requests that take minutes. A short everyday question could often be answered by a free smaller model in the same role; a programming task could not. Before the router acts on that, the decision log should show how often the case occurs, which classes it affects, how much waiting it would save and whether the classifications are right.
  • decision_engine.observe_blocked: true: when a request for a role has to wait, the decision engine classifies it in the background and the decision log records decision_observe: class, engine, confidence, the first free lower local tier of the same role (warm before cold, other nodes first) and would_fallback for confident classifications in observe_light (default standard, assist).
  • CPU chain only: embedding, TF-IDF and rules, 10–19 ms on the router host. The LLM stage is left out, because it would queue on the very GPU being waited for. The classification runs as its own task; the request does not wait for it.
  • decision_engine.observe_all: true: every request to /api/chat or /api/generate is classified in the background when it starts; when it ends, the router appends one line to anfragen.jsonl next to the configuration (observe_path): class, duration, waiting time, tokens, thinking, tool count, client, role, outcome. No prompt contents. A separate file so the decision log keeps its window; it rotates at 20 MB, and the response never waits for the classification. Purpose: check whether the class predicts run time before using it for Retry-After or waiting estimates.
  • Routing does not change.
Deutsch

Router 0.6.7. Beobachtungsmodus: Muss eine Anfrage an eine Rolle warten, ordnet die Decision Engine sie im Hintergrund ein (nur CPU-Kette, nie das LLM), und das Entscheidungslog hält fest, ob eine freie tiefere Stufe derselben Rolle gereicht hätte (decision_observe). Mit observe_all wird jede Anfrage eingeordnet und bekommt eine Zeile in anfragen.jsonl (Klasse, Dauer, Token, keine Inhalte). Am Routing ändert sich nichts.