Skip to content

v0.12.6 – interactive requests move away from a busy warm model (router 0.6.6)

Choose a tag to compare

@psewar psewar released this 02 Oct 07:25
· 4 commits to main since this release

Router 0.6.6, agent unchanged (0.16.0). Interactive requests no longer queue behind a busy warm model.

  • What went wrong: warm first chose the warm node even while it was busy. With max_parallel: 1 (since 0.6.5) a voice command could wait up to 60 s behind an agent request of several minutes, although a second node had a free tier for the same role.
  • Now: for priority interactive, if a warm model is busy at its limit, the router takes the first free local tier instead: a free warm one first, then a free cold one, and other nodes before the busy one (a fallback model on the busy node would share its GPU with the running request). Cloud tiers are never chosen on this path (cost, data class). The decision log shows reason: ausweichen. Without a busy warm model, warm first decides as before; if no tier is free, the request waits within its interactive budget (60 s).
  • Normal and batch requests keep waiting for the warm model, which is usually the better one.
  • The end-to-end suite checks both: an interactive request moves to the second node immediately, a normal one waits for the warm model.
Deutsch

Router 0.6.6. Interaktive Anfragen (Sprachbefehle) warten nicht mehr hinter einem belegten warmen Modell, wenn eine andere lokale Stufe sofort bedienen kann; im Entscheidungslog steht reason: ausweichen. Normale und Batch-Anfragen warten weiter auf das warme Modell.