Confusion about Deepspeed Inference #879

ZekaiGalaxy · 2024-03-25T08:26:00Z

Hi, I read the deepspeed docs and have the following confusion:

(1) What's the difference between these methods when in inferencing LLMs?

a. deepspeed.initialize and then write code to generate text

b. deepspeed.init_inference then write code to generate

c. use mii to inference

(2) Which of them are friendly for memory? For example, I want to inference 70b models, which of them support model parallelism that separates model parameters across gpus?

(3) For inference, what's the best practice now for inferencing 70b llama?

a. zero3 + cpu offload (1*a100)

b. zero3 (2*a100)

...

Thank you!

teis-e · 2024-05-17T14:19:48Z

Hello, did you find an answer?

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Confusion about Deepspeed Inference #879

Confusion about Deepspeed Inference #879

ZekaiGalaxy commented Mar 25, 2024

teis-e commented May 17, 2024

Confusion about Deepspeed Inference #879

Confusion about Deepspeed Inference #879

Comments

ZekaiGalaxy commented Mar 25, 2024

teis-e commented May 17, 2024