Repository navigation
v0.12.6 – interactive requests move away from a busy warm model (router 0.6.6)
Router 0.6.6, agent unchanged (0.16.0). Interactive requests no longer queue behind a busy warm model.
- What went wrong: warm first chose the warm node even while it was busy. With
max_parallel: 1(since 0.6.5) a voice command could wait up to 60 s behind an agent request of several minutes, although a second node had a free tier for the same role. - Now: for priority
interactive, if a warm model is busy at its limit, the router takes the first free local tier instead: a free warm one first, then a free cold one, and other nodes before the busy one (a fallback model on the busy node would share its GPU with the running request). Cloud tiers are never chosen on this path (cost, data class). The decision log showsreason: ausweichen. Without a busy warm model, warm first decides as before; if no tier is free, the request waits within its interactive budget (60 s). - Normal and batch requests keep waiting for the warm model, which is usually the better one.
- The end-to-end suite checks both: an interactive request moves to the second node immediately, a normal one waits for the warm model.
Deutsch
Router 0.6.6. Interaktive Anfragen (Sprachbefehle) warten nicht mehr hinter einem belegten warmen Modell, wenn eine andere lokale Stufe sofort bedienen kann; im Entscheidungslog steht reason: ausweichen. Normale und Batch-Anfragen warten weiter auf das warme Modell.