Finding and solving issues with KV-cache management on Ninfer #246
co-l
started this conversation in
Show and tell
Replies: 1 comment
|
Hey! Its nice fork! Can you please make windows version? |
0 replies
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Uh oh!
There was an error while loading. Please reload this page.
Uh oh!
There was an error while loading. Please reload this page.
Hello @Neroued
Thank you for your amazing work.
I spent the last week stress-testing ninfer on my setup (5090 + 32GB RAM), and noticed issues with cache management.
I saw you worked on the same issue recently, so I'm not sure you're interested in my work.
My goal is to maximize the capacity of ninfer under a real deployment, leveraging RAM as much as possible to avoid doing re-prefill. The motto is "never re-prefill an already prefilled context EXCEPT if eviction from both VRAM and RAM is inevitable"
cache-pressure protocol
Here is the bench protocol:
All these benches are available here: https://github.com/co-l/cache-pressure ; and directly callable with
uvxas shown belowNinfer launch config:
Results with current ninfer main
All benches against the same endpoint (
--base-urldefaults tohttp://localhost:8000/v1; model auto-detected viaGET /models).# reproduce uvx --from cache-pressure cache-pressure --base-url http://localhost:8000/v1 \ --kv-size 480000 --output run.json# reproduce uvx --from cache-pressure agent-sim --base-url http://localhost:8000/v1 \ --sessions 4 --output agent-sim.json(the following turns all re-prefilled from
root, full cascade)# reproduce uvx --from cache-pressure abort-sim --base-url http://localhost:8000/v1 \ --runs 2 --salt 1789464054# reproduce (full sweep 50K -> 450K) uvx --from cache-pressure needle-test --base-url http://localhost:8000/v1Notes on the commands (matching each run's parameters):
--kv-size 480000from "capacity: 480,000"--sessions 4matches s0–s3; main/sub sizes are the defaults (150K/40K), step 5K matches the step increments.--base-urlaccordingly.My ninfer fork
Fork co-l/ninfer@legacy/validated-main-b2d27c8c, main at
b2d27c8c(= yourd4929686+ the cache-pressure fix work). Same box, same endpoint, same commands as above, run back-to-back on one server instance (dirty cache — no restart between benches).uvx --from cache-pressure agent-sim --base-url http://localhost:8000/v1 \ --sessions 4 --main-tokens 150000 --sub-tokens 40000Ground truth from the request log: all mains/finalize served via
private_endpoint, zerorootre-prefills (only the 4 cold starts + 8 cold sub windows).uvx --from cache-pressure needle-test --base-url http://localhost:8000/v1 \ --lengths 50000,100000,200000It's a large change so I prefer not to push a PR against a moving target.
But the work is there and it's making a much stable version of the KV cache management.
If you decide to use this work, no attribution needed.
Tell me what you think.
PS: on my main branch, I also added tool streaming
All reactions