[Bug]: Disaggregated Prefilling use different TP between prefill instance and decode instance , it will be hanged #14952

67lc · 2025-03-17T12:01:32Z

Your current environment

I changed disagg_performance_benchmark.sh as flowing

launch_disagg_prefill() {
  model="$MODEL_PATH" 
  # disagg prefill
  CUDA_VISIBLE_DEVICES=0 python3 \
    -m vllm.entrypoints.openai.api_server \
    --model $model \
    --port 8100 \
    --max-model-len 10000 \
    --tensor-parallel-size 1 \
    --dtype=half \
    --gpu-memory-utilization 0.6 \
    --kv-transfer-config \
    '{"kv_connector":"PyNcclConnector","kv_role":"kv_producer","kv_rank":0,"kv_parallel_size":2,"kv_buffer_size":5e9}' &

  CUDA_VISIBLE_DEVICES=1,2 python3 \
    -m vllm.entrypoints.openai.api_server \
    --model $model \
    --port 8200 \
    --max-model-len 10000 \
    --tensor-parallel-size 2 \
    --dtype=half \
    --gpu-memory-utilization 0.6 \
    --kv-transfer-config \
    '{"kv_connector":"PyNcclConnector","kv_role":"kv_consumer","kv_rank":1,"kv_parallel_size":2,"kv_buffer_size":5e9}' &

  wait_for_server 8100
  wait_for_server 8200
  python3 disagg_prefill_proxy_server.py &
  sleep 1
}

I have four V100s 16GB with NVlink. When I use the same tp, it's normal.

🐛 Describe the bug

Before submitting a new issue...

Make sure you already searched for relevant issues, and asked the chatbot living at the bottom right corner of the documentation page, which can answer lots of frequently asked questions.

The text was updated successfully, but these errors were encountered:

67lc added the bug label Mar 17, 2025

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[Bug]: Disaggregated Prefilling use different TP between prefill instance and decode instance , it will be hanged #14952

[Bug]: Disaggregated Prefilling use different TP between prefill instance and decode instance , it will be hanged #14952

67lc commented Mar 17, 2025 •

edited

Loading

[Bug]: Disaggregated Prefilling use different TP between prefill instance and decode instance , it will be hanged #14952

[Bug]: Disaggregated Prefilling use different TP between prefill instance and decode instance , it will be hanged #14952

Comments

67lc commented Mar 17, 2025 • edited Loading

Your current environment

🐛 Describe the bug

Before submitting a new issue...

67lc commented Mar 17, 2025 •

edited

Loading