Skip to content

v4.0.0

Choose a tag to compare

@SAY-5 SAY-5 released this 03 Sep 19:53
· 29 commits to main since this release

Adds micro-batching and a warm model pool. Concurrent /predict calls for the same model are queued FIFO on the event loop and run in one forward pass, bounded by a maximum batch size and wait; every forward pass is padded to a fixed row count so a request gets bit-identical output alone or in a full batch, which the swap test checks for all 2000 responses across a promote. Sparse traffic skips the wait, so single-row p50 stays at 1.7 ms. POST /admin/warm loads a candidate into a spare pool slot ahead of time and the following promote reports prewarmed with 0 s of load; role holders are never evicted. 93 tests, load test with pre-warm, canary, and mid-run swap passes with 0 drops.