Repository navigation
v0.12.5 – models without parallel requests wait in the router (router 0.6.5)
Router 0.6.5, agent unchanged (0.16.0). Models without parallel requests wait in the router.
- What went wrong: qwen3.8 (architecture qwen35) cannot run two requests at once in Ollama. When the router sent a second request while one was running, Ollama put it in its own queue and reloaded the runner as soon as the first one finished, although model and parameters were identical. Seen twice in one night, 5–7 s each, context checkpoints and prompt cache lost.
- New catalog field
models.<m>.max_parallel: at most this many concurrent requests per node for that model. A further request waits in the router's queue, with priority and aging, and the runner stays loaded. When a slot frees up, the router wakes the waiter that can actually start (a waiter for a model still at its limit stays queued). The field can also be set in the UI catalog; a measured entry keeps its measurement. - Waiting budget per priority class when a request sets no
deadline_ms:max_wait_interactive_s60 s,max_wait_sfor normal now 600 s instead of 120 s (agent requests often take minutes),max_wait_batch_s1800 s. The queued event in the decision log says whether a request waited for the node (max_inflight) or for the model (max_parallel). - New selftest
selftest_max_parallel.pyin CI; the end-to-end suite checks that two concurrent requests run one after the other with the limit and in parallel without it.
Deutsch
Router 0.6.5. Neues Katalogfeld max_parallel: qwen3.8 kann in Ollama keine zwei Anfragen zugleich, und Ollama lud bei einer zweiten den Runner neu. Jetzt wartet die zweite Anfrage im Router, mit Priorität, und der Runner bleibt. Wartebudget je Prioritätsklasse: interactive 60 s, normal 600 s (statt 120 s), batch 1800 s.