Support LLaMA training through SPMD #13

jonb377 · 2023-07-20T17:51:59Z

Code change to enable training LLaMA through SPMD.

alanwaketan

LGTM.

alanwaketan

The new change LGTM.

alanwaketan · 2023-07-24T19:09:44Z

@jonb377 As discussed in the chat, can we have a new branch called llama2-google-next-training and then land your change there?

jonb377 · 2023-07-24T19:12:45Z

Sounds good, I've changed the base to llama2-google-next-training. I'll merge this now and follow up with tensor parallelism.

alanwaketan · 2023-07-24T19:15:05Z

@jonb377 Wondering if it makes sense for you to share the command you use to run LLaMA 2 on the repo? Maybe on the "How to Run HF Transformers with SPMD" doc?

jonb377 · 2023-07-24T19:16:28Z

Sure! It's essentially the same training command, you just need to update your model config file to be llama, e.g.:

{
  "_name_or_path": "/tmp/home/config.json",
  "bos_token_id": 1,
  "eos_token_id": 2,
  "hidden_act": "silu",
  "hidden_size": 4096,
  "initializer_range": 0.02,
  "intermediate_size": 11008,
  "max_position_embeddings": 2048,
  "model_type": "llama",
  "num_attention_heads": 32,
  "num_hidden_layers": 32,
  "num_key_value_heads": 32,
  "pad_token_id": 0,
  "pretraining_tp": 1,
  "rms_norm_eps": 1e-06,
  "rope_scaling": null,
  "tie_word_embeddings": false,
  "transformers_version": "4.32.0.dev0",
  "use_cache": true,
  "vocab_size": 32000
}

alanwaketan · 2023-07-24T19:18:01Z

Thanks @jonb377! It will be great if you can share your model configs in a gs bucket just like what Alex did.

jonb377 · 2023-07-24T19:32:14Z

http://bigstore/hf-train-config, I'll also link in the doc

alanwaketan · 2023-07-24T21:44:18Z

Thanks, Jon!

* Cohere Model Release (#1) Cohere Model Release * Remove unnecessary files and code (#2) Some cleanup * Delete cohere-model directory (#3) * Make Fix (#5) * Pr fixes (#6) * fixes for pr * pr fixes for the format * pr fixes for the format * src/transformers/models/auto/tokenization_auto.py * Tokenizer test (#8) * tokenizer test * format fix * Adding Docs and other minor changes (#7) * Add modeling tests (#9) * Smol Fix (#11) * tokenization tests are fixed * format fixes * fix pr doc tests * fix pr doc tests * fix pr doc tests * fix pr style check * small changes in cohere.md * FIX: Address final comments for transformers integration (#13) * fix modeling final nits and add proper test file * for now leave empty tests * add integration test * push new test * fix modeling cohere (#14) * Update chat templates to use the new API (#15) --------- Co-authored-by: ahmetustun <ahmetustun89@gmail.com> Co-authored-by: Younes Belkada <49240599+younesbelkada@users.noreply.github.com> Co-authored-by: Matt <Rocketknight1@users.noreply.github.com>

* Initial add model additions * Test * All weights loading * Can perform full forward pass * Local and remote the same * Matching local and remote * Fixup * Idefics2Model importable; fixup docstrings * Don't skip by default * Remove deprecated use_resampler arg * Remove self.config * DecoupledLinear takes config * Tidy up * Enable eager attention and tidy up * Most tests passing * Update for batch of processed images * Add image processor * Update doc pages * Update conversion script * Remove erroneous breakpoint * Remove accidendtal spelling change * Update to reflect changes on hub - make generate work * Fix up * Image processor tests * Update tests * Add a processor * Add a processor * Update convert script * Update modeling file - remove fixmes * Bug fix * Add processing test * Use processor * Fix up * Update src/transformers/models/idefics2/modeling_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * Update src/transformers/models/idefics2/modeling_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * Fix test * Update config - PR comments and defaults align with checkpoint * Reviewer comments * Add copied froms for flahs attention * Update src/transformers/models/idefics2/modeling_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * Apply suggestions from code review Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Remove qk_layer_norm and freeze_layers functionality * Fix * Remove freeze_layer options from config * Sync with upstream main * Fix attention shapes siglip * Remove Llava-next refs - TO REBASE * Use AutoModel for text model * Add comment to explain vision embeddings * Fix issue with tie_word_embeddings * Address review comments * Fix and fix up * Chat templates for idefics * Fix copies * Fix * Add layer norms to FA2 * Fix tests * Apply suggestions from code review Co-authored-by: Victor SANH <victorsanh@gmail.com> * Fix * Review comments * Update src/transformers/models/idefics2/modeling_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * Update inputs merger * Merge weights in correct order * Update convert script * Update src/transformers/models/idefics2/processing_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * Update template * Model code examples (fix idefics too) * More review comments * Tidy up * Update processing * Fix attention mask preparation * Update inputs_merger inputs * Vectorize inputs_merger * Update src/transformers/models/idefics2/__init__.py Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com> * Update src/transformers/models/idefics2/modeling_idefics2.py * Review comments * saying bye to the `qk_layer_norms` * Simplify * Update latents * Remove erroneuous readme changes * Return images when applying chat template * Fix bug - prompt images are for a single sample * Update src/transformers/models/idefics2/modeling_idefics2.py * image splitting * fix test * some more comment * some comment * Apply suggestions from code review Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update src/transformers/models/idefics2/image_processing_idefics2.py Co-authored-by: amyeroberts <22614925+amyeroberts@users.noreply.github.com> * Update processor * Update model tests * Update src/transformers/models/idefics2/processing_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * Update src/transformers/models/idefics2/processing_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * Don't add BOS in template * Update src/transformers/models/idefics2/processing_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * Remove index in examples * Update tests to reflect #13 * Update src/transformers/models/idefics2/processing_idefics2.py Co-authored-by: Victor SANH <victorsanh@gmail.com> * PR comment - consistent typing * Update readme and model doc * Update docs * Update checkpoint references * Update examples * Fix and update tests * Small addition * Update tests - remove copied from as no ignore placement copy could be found * Update example * small fixes * Update docs/source/en/model_doc/idefics2.md Co-authored-by: Victor SANH <victorsanh@gmail.com> * Update docs/source/en/model_doc/idefics2.md Co-authored-by: Victor SANH <victorsanh@gmail.com> * Update README.md Co-authored-by: Victor SANH <victorsanh@gmail.com> * Connector model as bridge * Fix up * Fix up * Don't pass model inputs for generation kwargs update * IDEFICS-2 -> Idefics2 * Remove config archive name * IDEFICS-2 -> Idefics2 * Add back llava-next * Update readmes * Add requirements for processor tester * Use custom convert_to_rgb to avoid possible BC * Fix doc example * Fix doc example * Skip model doc tests - as model to large * More doc example - account for image splitting * Update src/transformers/image_transforms.py * Fix config doctest --------- Co-authored-by: Pablo Montalvo <39954772+molbap@users.noreply.github.com> Co-authored-by: ArthurZucker <arthur.zucker@gmail.com> Co-authored-by: Victor SANH <victorsanh@gmail.com> Co-authored-by: Arthur <48595927+ArthurZucker@users.noreply.github.com>

Support LLaMA training through SPMD

4f806cd

alanwaketan approved these changes Jul 20, 2023

View reviewed changes

Shard largest dimension

3799ac9

alanwaketan approved these changes Jul 20, 2023

View reviewed changes

Use HybridMesh

7b90ea5

jonb377 changed the base branch from master to llama2-google-next-training July 24, 2023 19:12

jonb377 merged commit e7ea6ea into llama2-google-next-training Jul 24, 2023

jonb377 deleted the jonbolin-llama-spmd branch July 24, 2023 19:13

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Support LLaMA training through SPMD #13

Support LLaMA training through SPMD #13

jonb377 commented Jul 20, 2023

alanwaketan left a comment

alanwaketan left a comment

alanwaketan commented Jul 24, 2023

jonb377 commented Jul 24, 2023

alanwaketan commented Jul 24, 2023 •

edited

Loading

jonb377 commented Jul 24, 2023

alanwaketan commented Jul 24, 2023

jonb377 commented Jul 24, 2023

alanwaketan commented Jul 24, 2023

Support LLaMA training through SPMD #13

Support LLaMA training through SPMD #13

Conversation

jonb377 commented Jul 20, 2023

alanwaketan left a comment

Choose a reason for hiding this comment

alanwaketan left a comment

Choose a reason for hiding this comment

alanwaketan commented Jul 24, 2023

jonb377 commented Jul 24, 2023

alanwaketan commented Jul 24, 2023 • edited Loading

jonb377 commented Jul 24, 2023

alanwaketan commented Jul 24, 2023

jonb377 commented Jul 24, 2023

alanwaketan commented Jul 24, 2023

alanwaketan commented Jul 24, 2023 •

edited

Loading