Skip to content

v0.12.5 – models without parallel requests wait in the router (router 0.6.5)

Choose a tag to compare

@psewar psewar released this 02 Oct 06:51
· 5 commits to main since this release

Router 0.6.5, agent unchanged (0.16.0). Models without parallel requests wait in the router.

  • What went wrong: qwen3.8 (architecture qwen35) cannot run two requests at once in Ollama. When the router sent a second request while one was running, Ollama put it in its own queue and reloaded the runner as soon as the first one finished, although model and parameters were identical. Seen twice in one night, 5–7 s each, context checkpoints and prompt cache lost.
  • New catalog field models.<m>.max_parallel: at most this many concurrent requests per node for that model. A further request waits in the router's queue, with priority and aging, and the runner stays loaded. When a slot frees up, the router wakes the waiter that can actually start (a waiter for a model still at its limit stays queued). The field can also be set in the UI catalog; a measured entry keeps its measurement.
  • Waiting budget per priority class when a request sets no deadline_ms: max_wait_interactive_s 60 s, max_wait_s for normal now 600 s instead of 120 s (agent requests often take minutes), max_wait_batch_s 1800 s. The queued event in the decision log says whether a request waited for the node (max_inflight) or for the model (max_parallel).
  • New selftest selftest_max_parallel.py in CI; the end-to-end suite checks that two concurrent requests run one after the other with the limit and in parallel without it.
Deutsch

Router 0.6.5. Neues Katalogfeld max_parallel: qwen3.8 kann in Ollama keine zwei Anfragen zugleich, und Ollama lud bei einer zweiten den Runner neu. Jetzt wartet die zweite Anfrage im Router, mit Priorität, und der Runner bleibt. Wartebudget je Prioritätsklasse: interactive 60 s, normal 600 s (statt 120 s), batch 1800 s.