v1.0.7
Adding support for multi-client runs for server-mode benchmarking.
Instead of providing a specific shape combo to the client section of runner -- e.g.,
...
--
client
--env PYTHONUNBUFFERED=1
--dataset-name random
--random-input-len 896
--random-output-len 128
--max-concurrency 1
--num-prompts 10Now we can provide a --multi parameter:
...
--
client
--env PYTHONUNBUFFERED=1
--multi 896/128/1/10,1920/128/2/20,3968/128/1/10
--multi receives a , separated list of shape combos in the format:
input size / output size / batch size / num prompts
In the example above, three combos are specified:
896/128/1/10- input size = 896, output size = 128, batch size 1, num prompts = 101920/128/2/20- input size = 1920, output size = 128, batch size 2, num prompts = 203968/128/1/10- input size = 3968, output size = 128, batch size 1, num prompts = 10
If --multi is provided, then the client script will iterate over combos and run one vllm bench serve for each combo. Each instance writes outputs to its own file, client.log.${instance}. Again, in the example above, there would be three instances -- therefore files client.log.0, client.log.1, client.log.2.
Pod / execution output:
waiting for server at PID 56 ...
done, server is ready!
...
Wed Sep 17 23:29:59 UTC 2025 -- starting 0
Wed Sep 17 23:29:59 UTC 2025 -- starting 1
Wed Sep 17 23:29:59 UTC 2025 -- clients started; waiting for completion ...
Wed Sep 17 23:29:59 UTC 2025 -- starting 2
Wed Sep 17 23:34:27 UTC 2025 -- finished 0
Wed Sep 17 23:34:49 UTC 2025 -- finished 1
Wed Sep 17 23:35:12 UTC 2025 -- finished 2
Wed Sep 17 23:35:12 UTC 2025 -- all done!
...
Each client.log.${instance} file will have the usual vllm bench output, including perf metrics.