v5.0.5
qwen3-4b with tp>1 was supportted.
RLHF:
Skipping loading weights from local files when starting up vllm engines.
LLM was replaced by AsyncLLM. So manual batching was replaced by auto-continuous batching. Which means training waitting time has be shorten as the first response time.
max_num_seqs of training was no longer needed. Adjusting max_num_seqs and num_vllm_engines of vllm to balance training speed and inferrence speed of one batch samples will benefite a highest performance.