Support InternLM #4137

wangruohui · 2023-08-11T10:11:13Z

This PR is to support a new model named InternLM.

This model is similar to llama but with bias on qkvo matmul. So I primarily duplicate codes of llama model and add support for bias in Attention's python code.

This branch is previously checkoutted from tag v0.10.0 and I have tested locally on single GPU or multiple with TP.

wangruohui · 2023-08-11T12:49:42Z

@microsoft-github-policy-service agree

lekurile · 2023-08-25T17:18:49Z

Hi @wangruohui,

Thank you for expanding our inference capabilities and adding the container for InternLM. Can you please provide an example of a model being used with our inference test?

For example, for the BLOOM model, a command may look like this:

deepspeed --num_gpus 1 inference-test.py --model bigscience/bloom-3b --use_meta --use_kernel

It would be nice to have a similar command for InternLM so we can do some testing on our side as well.

Thanks!

wangruohui · 2023-08-30T06:53:23Z

Hello @lekurile

I am working on the test script you provided. Some note:

As InternLM is not integrated into the main branch of transformers, one still need to add trust_remote_code=True to load the model from hub. See
support trust_remote_code in inference test DeepSpeedExamples#709 for detailed modifications.
To make things easier to manage, I git clone the internlm-7b model from huggingface hub to my home directory like

git lfs install
git clone https://huggingface.co/internlm/internlm-7b

So all test commands below are based on this directory ~/internlm-7b. You may change this on your working server.

InternLM has two variants, the base model internlm-7b and the finetuned one internlm-chat-7b. Tests below use internlm-7b. Please check the model name in below test commands.
I set --greedy to make result reproducible.

Test results

HF Baseline

deepspeed --num_gpus 1 inference-test.py --model ~/internlm-7b/ --hf_baseline --greedy --trust-remote-code

generation time is 1.5033915042877197 sec

in=DeepSpeed is a machine learning framework
out=DeepSpeed is a machine learning framework for deep learning. It is a Python package that provides a set of tools for training and evaluating deep neural networks. It is designed to be easy to use and to provide a consistent interface for all of the different types of neural networks that can be trained

Use Kernels

deepspeed --num_gpus 1 inference-test.py --model ~/internlm-7b/ --use_kernel --greedy --trust_remote_code

generation time is 0.5656006336212158 sec

in=DeepSpeed is a machine learning framework
out=DeepSpeed is a machine learning framework for deep learning. It is a Python package that provides a set of tools for training and evaluating deep neural networks. It is designed to be easy to use and to provide a consistent interface for all of the different types of neural networks that can be trained

Tensor Parallel

deepspeed --num_gpus 2 inference-test.py --model ~/internlm-7b/ --use_kernel --greedy --trust_remote_code

generation time is 0.6287708282470703 sec

in=DeepSpeed is a machine learning framework
out=DeepSpeed is a machine learning framework for deep learning. It is a Python package that provides a set of tools for training and evaluating deep neural networks. It is designed to be easy to use and to provide a consistent interface for all of the different types of neural networks that can be trained

lekurile · 2023-08-30T16:51:06Z

@wangruohui Appreciate the very detailed testing and reproduction details. I've merged the latest in master and kicked off some tests. I'll also try getting the model to run on my end as well.

Thanks,
Lev

wangruohui · 2023-09-07T07:09:37Z

Hello,

Any updates?
And feel free to talk to me if you need some help at my side.

lekurile · 2023-09-11T15:47:00Z

Hello,

Any updates? And feel free to talk to me if you need some help at my side.

Hi @wangruohui,

Looks like the changes in this PR may have caused some issues with other models, specifically in the following unit tests:

unit/inference/test_inference.py::TestMPSize::test[fp32-gpt-neo] FAILED [ 80%]
unit/inference/test_inference.py::TestMPSize::test[fp16-bloom] FAILED [ 86%]
unit/inference/test_inference.py::TestMPSize::test[fp16-gpt-neo] FAILED [ 88%]

I'm suspecting this is due to changes in deepspeed/ops/transformer/inference/ds_attention.py and deepspeed/ops/transformer/inference/op_binding/qkv_gemm.py potentially breaking the behavior of GPT-Neo and BLOOM models, since all the other file changes are self-contained to InternLM.

Can you kindly test these models for compatibility with the InternLM changes on your side? I can try testing as well.

Thanks,
Lev

wangruohui · 2023-09-12T14:30:41Z

Hello @lekurile

I made some modification to make GPTNeo compatible but I cannot set up an exactly the same env to run all tests. Would you please allow the CI to run to check if the problem is solved?

deepspeed/module_inject/containers/internlm.py

* origin/master: Allow multiple inference engines in single script (microsoft#4384) adds triton flash attention2 kernel (microsoft#4337) Fix llama meta tensor loading in AutoTP and kernel injected inference (microsoft#3608) Fix min torch version (microsoft#4375) Fix multinode runner to properly append to PDSH_SSH_ARGS_APPEND (microsoft#4373) add the missing method (microsoft#4363) Openfold fix (microsoft#4368) deepspeed4science japanese blog (microsoft#4369) deepspeed4science chinese blog (microsoft#4366) Enable workflow dispatch on Torch 1.10 CI tests (microsoft#4361) Update conda env to have max pydantic version (microsoft#4362) add deepspeed4science blog link (microsoft#4364) added check to avoid undefined behavior when the input_id length is greater than max_tokens (microsoft#4349) Add the policy to run llama model from the official repo (microsoft#4313) fix deepspeed4science links (microsoft#4358) DeepSpeed4Science (microsoft#4357) Support InternLM (microsoft#4137) Pass base_dir to model files can be loaded for auto-tp/meta-tensor. (microsoft#4348)

wangruohui requested review from RezaYazdaniAminabadi, jeffra, mrwyattii, awan-10, cmikeh2 and arashb as code owners August 11, 2023 10:11

wangruohui added 7 commits August 11, 2023 20:54

correct inference with some debug codes.

b593a3d

remove prints

baf0258

update transformer import set_qkv and format

a87c97d

support some lora abstract method

8449044

fix attn_ob

5c80d83

some debug

ee1ccd3

leave orig layer set by user

bdf8a90

wangruohui force-pushed the support_internlm_0.10.0 branch from 954dac5 to bdf8a90 Compare August 11, 2023 12:55

remove debugs

76f00b6

wangruohui closed this Aug 11, 2023

wangruohui deleted the support_internlm_0.10.0 branch August 11, 2023 13:00

wangruohui restored the support_internlm_0.10.0 branch August 11, 2023 13:01

wangruohui reopened this Aug 11, 2023

molly-smith requested review from molly-smith and lekurile August 18, 2023 17:02

molly-smith self-assigned this Aug 18, 2023

lekurile self-assigned this Aug 18, 2023

Merge branch 'master' into support_internlm_0.10.0

b3bd0c6

Merge branch 'master' into support_internlm_0.10.0

835a493

move attn ob to mlp module

1ac20f0

Merge branch 'master' into support_internlm_0.10.0

d1dac2b

lekurile requested changes Sep 12, 2023

View reviewed changes

deepspeed/module_inject/containers/internlm.py Show resolved Hide resolved

wangruohui added 2 commits September 13, 2023 14:08

move import transformer

bf49fe9

init orig class only once

8796eef

molly-smith removed their assignment Sep 14, 2023

lekurile requested changes Sep 14, 2023

View reviewed changes

deepspeed/module_inject/containers/internlm.py Outdated Show resolved Hide resolved

wangruohui and others added 3 commits September 15, 2023 15:10

remove copyright

cbdc39a

Merge branch 'master' into support_internlm_0.10.0

1e9068a

Merge branch 'master' into support_internlm_0.10.0

d7a7a9a

lekurile approved these changes Sep 18, 2023

View reviewed changes

lekurile enabled auto-merge September 18, 2023 16:51

lekurile added this pull request to the merge queue Sep 18, 2023

Merged via the queue into microsoft:master with commit 367d6f9 Sep 18, 2023
16 checks passed

wangruohui mentioned this pull request Sep 1, 2023

[WIP] Support InternLM on 3rd-party inference toolboxes InternLM/lmdeploy#136

Closed

5 tasks

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Support InternLM #4137

Support InternLM #4137

wangruohui commented Aug 11, 2023

wangruohui commented Aug 11, 2023

lekurile commented Aug 25, 2023

wangruohui commented Aug 30, 2023 •

edited

Loading

lekurile commented Aug 30, 2023

wangruohui commented Sep 7, 2023

lekurile commented Sep 11, 2023 •

edited

Loading

wangruohui commented Sep 12, 2023

Support InternLM #4137

Support InternLM #4137

Conversation

wangruohui commented Aug 11, 2023

wangruohui commented Aug 11, 2023

lekurile commented Aug 25, 2023

wangruohui commented Aug 30, 2023 • edited Loading

Test results

HF Baseline

Use Kernels

Tensor Parallel

lekurile commented Aug 30, 2023

wangruohui commented Sep 7, 2023

lekurile commented Sep 11, 2023 • edited Loading

wangruohui commented Sep 12, 2023

wangruohui commented Aug 30, 2023 •

edited

Loading

lekurile commented Sep 11, 2023 •

edited

Loading