Training speed is very slow #312

1179021477 · 2020-02-17T12:12:51Z

@RogerChern I train tridentNet_1x with resnet50 on 4 GPU (a machine with 8 GPU), and I need 2 days. Especially, when others use other left GPUs in my machines, the speed of training my models is slower. Is there any way to make training faster? Like how to construct multi-thread, etc. My machine is TITAN X (Pascal).

xchani · 2020-02-23T20:12:20Z

Our dataloader does use multi-threading to load images.
According to your description, you are sharing gpu server with others, then jobs from others may occupy cpu resource in that server, which slow down your training. Also the speed of disk(IOPS) is another major factor should be considered.

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Training speed is very slow #312

Training speed is very slow #312

1179021477 commented Feb 17, 2020

xchani commented Feb 23, 2020

Training speed is very slow #312

Training speed is very slow #312

Comments

1179021477 commented Feb 17, 2020

xchani commented Feb 23, 2020