Automatically map deduplicated safetensors weights to their original values #501

Vinno97 · 2023-06-28T17:36:51Z

What does this PR do?

This PR automatically points tensors that were removed due to deduplication to their still existing twin.

In server.text_generation_server.utils.convert.py#convert_file, tensors that have a value equal to another tensor are removed from the list of weights. Their name, and the name of the still-existing "twin" are logged to the "metadata" dictionary. However, this dictionary was not yet used during loading. This requires explicit in-code remapping when loading the models (as mentioned in the docstring).

This PR adds some simple code to check, during loading, if a weight is one of those removed weights. It then automatically retrieves the values of its still-existing "twin" instead.

What does this fix?

We currently cannot load h2oai/h2ogpt-oig-oasst1-falcon-40b with the unmodified server, since the transformer.word_embeddings.weight weight is equal to lm_head.weight and is automatically removed. The falcon code, however, still expects this weight to exist. I could have also added some extra checks to the model itself, though that would only be a workaround.

Before submitting

This PR fixes a typo or improves the docs (you can dismiss the other checks if that's the case).
Did you read the contributor guideline,
Pull Request section?
Was this discussed/approved via a Github issue or the forum? Please add a link
to it if that's the case.
Did you make sure to update the documentation with your changes? Here are the
documentation guidelines, and
here are tips on formatting docstrings.
Did you write any new necessary tests?

Who can review?

Anyone in the community is free to review the PR once the tests have passed. Feel free to tag
members/contributors who may be interested in your PR.

This PR automatically points tensors that were removed due to deduplication to their still existing twin. In `server.text_generation_server.utils.convert.py#convert_file`, duplicated tensors are removed and logged to the "metadata" dictionary. However, this dictionary was not yet used during loading. This requires explicit remapping when loading the models (as mentioned in the docstring). What does this fix? We currently cannot load `h2oai/h2ogpt-oig-oasst1-falcon-40b` with the unmodified server, since the `transformer.word_embeddings.weight` weight is equal to `lm_head.weight` and is automatically removed.

Narsil · 2023-06-30T08:13:31Z

Hey thanks for the PR !.

Unfortunately that metadata is kept for hard debugging but it's missing crucial information, namely it doesn't recall the the tensor was a slice or not. And metadata will not necessarily be present.

I suggest a different fix:

diff --git a/server/text_generation_server/models/flash_rw.py b/server/text_generation_server/models/flash_rw.py
index 5f963bf..33079ac 100644
--- a/server/text_generation_server/models/flash_rw.py
+++ b/server/text_generation_server/models/flash_rw.py
@@ -48,7 +48,13 @@ class FlashRWSharded(FlashCausalLM):

         torch.distributed.barrier(group=self.process_group)
         filenames = weight_files(model_id, revision=revision, extension=".safetensors")
-        weights = Weights(filenames, device, dtype, process_group=self.process_group)
+        weights = Weights(
+            filenames,
+            device,
+            dtype,
+            process_group=self.process_group,
+            aliases={"transformer.word_embeddings.weight": ["lm_head.weight"]},
+        )

         config.quantize = quantize

Would that work ?

Vinno97 · 2023-07-01T22:47:29Z

Without having tested the fix (don't have access to the GPU server during the weekends), this seems like it would also fix my problem.

I initially also thought about doing this. However, it is still only a fix for this specific model. The metadata may not be a trustworthy source for tensor aliases, but it may still be a valid fallback, no? It could also be improved by adding a namespacing prefix like alias-- to the key, to prevent conflicts.

Curious to hear what you think

Narsil · 2023-07-03T09:15:18Z

However, it is still only a fix for this specific model.

Indeed, but the other one can lead to potentially catastrophic failure (loading the wrong weights) which is even worse imo.

Are you ok if I update this PR ? (If I can, otherwise I'll just create a new one with you as co-author).

Vinno97 · 2023-07-04T08:28:42Z

Alright, fair point. Silently loading the wrong weights is definitely undesired.
Sure, update this PR!

This reverts commit d6bb10f.

Vinno97 · 2023-07-06T09:39:15Z

I tested the new fix "on my machine" and it worked. A colleage used it then and it didn't work for him. The difference? In his version of model-00001-of-00018.safetensors, the transformer.word_embeddings.weight was there, but the lm_head.weight was gone.

Unless the conversion can be made to be consistent, this means we should either also alias the weights the other way around. Maybe the same as done in the current patch already, or automatically (two-way aliases by default).

I don't know why his safetensors conversion is different, I guess because the order of dictionary keys is not guaranteed to be consistent, which then affects safetensors.torch_remove_duplicate_names. Pretty sure that using its preferred_names argument to would fix this weight mapping issue, but it'd be a model-specific fix in a generic part of the code-base. Unless that can work, but it'd be nicer to keep it in flash_rw.py

- Look at `transformers` base class to check for `_key_to_ignore_on_load_missing` or `_tied_weights` which are the standard attributes to select the keys to NOT save on disk (since they are ignored) - Modified safetensors code (to be reflected in safetensors even if it's an internal function). - Will not work for trust_remote_code=True repos (like santacoder). Should help with : #555 and : #501 and #556 and #482 (comment)

- Look at `transformers` base class to check for `_key_to_ignore_on_load_missing` or `_tied_weights` which are the standard attributes to select the keys to NOT save on disk (since they are ignored) - Modified safetensors code (to be reflected in safetensors even if it's an internal function). - Will not work for trust_remote_code=True repos (like santacoder). Should help with : huggingface/text-generation-inference#555 and : huggingface/text-generation-inference#501 and huggingface/text-generation-inference#556 and huggingface/text-generation-inference#482 (comment)

lppllppl920 · 2023-07-31T01:24:04Z

Any update on this PR? I also encountered the issue of missing lm_head.weight when I try to load Falcon model with text-generation-inference.

Narsil · 2023-08-02T17:52:52Z

Hi @lppllppl920 Thanks for the ping. I'm not sure why it wasn't merged.

@OlivierDehaene

…values (#501) (#761) # What does this PR do? CI cehck for #501 This PR automatically points tensors that were removed due to deduplication to their still existing twin. In `server.text_generation_server.utils.convert.py#convert_file`, tensors that have a value equal to another tensor are removed from the list of weights. Their name, and the name of the still-existing "twin" are logged to the "metadata" dictionary. However, this dictionary was not yet used during loading. This requires explicit in-code remapping when loading the models (as mentioned in the docstring). This PR adds some simple code to check, during loading, if a weight is one of those removed weights. It then automatically retrieves the values of its still-existing "twin" instead. ## What does this fix? We currently cannot load `h2oai/h2ogpt-oig-oasst1-falcon-40b` with the unmodified server, since the `transformer.word_embeddings.weight` weight is equal to `lm_head.weight` and is automatically removed. The falcon code, however, still expects this weight to exist. I could have also added some extra checks to the model itself, though that would only be a workaround. ## Before submitting - [ ] This PR fixes a typo or improves the docs (you can dismiss the other checks if that's the case). - [ ] Did you read the [contributor guideline](https://github.com/huggingface/transformers/blob/main/CONTRIBUTING.md#start-contributing-pull-requests), Pull Request section? - [ ] Was this discussed/approved via a Github issue or the [forum](https://discuss.huggingface.co/)? Please add a link to it if that's the case. - [ ] Did you make sure to update the documentation with your changes? Here are the [documentation guidelines](https://github.com/huggingface/transformers/tree/main/docs), and [here are tips on formatting docstrings](https://github.com/huggingface/transformers/tree/main/docs#writing-source-documentation). - [ ] Did you write any new necessary tests? ## Who can review? Anyone in the community is free to review the PR once the tests have passed. Feel free to tag members/contributors who may be interested in your PR.  --------- # What does this PR do?   Fixes # (issue) ## Before submitting - [ ] This PR fixes a typo or improves the docs (you can dismiss the other checks if that's the case). - [ ] Did you read the [contributor guideline](https://github.com/huggingface/transformers/blob/main/CONTRIBUTING.md#start-contributing-pull-requests), Pull Request section? - [ ] Was this discussed/approved via a Github issue or the [forum](https://discuss.huggingface.co/)? Please add a link to it if that's the case. - [ ] Did you make sure to update the documentation with your changes? Here are the [documentation guidelines](https://github.com/huggingface/transformers/tree/main/docs), and [here are tips on formatting docstrings](https://github.com/huggingface/transformers/tree/main/docs#writing-source-documentation). - [ ] Did you write any new necessary tests? ## Who can review? Anyone in the community is free to review the PR once the tests have passed. Feel free to tag members/contributors who may be interested in your PR.  Co-authored-by: Vincent Brouwers <vincentbrouwers9@gmail.com> Co-authored-by: Vincent Brouwers <vincent.brouwers@ing.com>

lppllppl920 · 2023-08-02T20:30:38Z

Hi @lppllppl920 Thanks for the ping. I'm not sure why it wasn't merged.

Thank you!

Narsil · 2023-08-03T11:49:55Z

The fix actually doesn't work: I discovered it while testing. Fix coming soon: https://github.com/huggingface/text-generation-inference/pull/762/files#diff-2111bae5f77d998a3fe39888906b3c7be122313241ed6b69b0b0baf5abb735bbL57

- Look at `transformers` base class to check for `_key_to_ignore_on_load_missing` or `_tied_weights` which are the standard attributes to select the keys to NOT save on disk (since they are ignored) - Modified safetensors code (to be reflected in safetensors even if it's an internal function). - Will not work for trust_remote_code=True repos (like santacoder). Should help with : huggingface/text-generation-inference#555 and : huggingface/text-generation-inference#501 and huggingface/text-generation-inference#556 and huggingface/text-generation-inference#482 (comment)

Narsil added 2 commits July 4, 2023 11:30

Revert "Map deduplicated tensors via metadata"

81f234e

This reverts commit d6bb10f.

Modified fix.

742199a

Narsil requested a review from OlivierDehaene July 4, 2023 09:31

Narsil mentioned this pull request Jul 6, 2023

Attempting to harden a bit the weights choice to save on disk. #561

Merged

5 tasks

Narsil changed the base branch from main to dev August 2, 2023 17:54

Narsil approved these changes Aug 2, 2023

View reviewed changes

Narsil merged commit 9bcac46 into huggingface:dev Aug 2, 2023
2 of 5 checks passed

Narsil mentioned this pull request Aug 2, 2023

Automatically map deduplicated safetensors weights to their original values (#501) #761

Merged

10 tasks

lppllppl920 mentioned this pull request Aug 9, 2023

falcon.cpp: tensor 'lm_head.weight' is missing from model marella/ctransformers#81

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Automatically map deduplicated safetensors weights to their original values #501

Automatically map deduplicated safetensors weights to their original values #501

Vinno97 commented Jun 28, 2023

Narsil commented Jun 30, 2023

Vinno97 commented Jul 1, 2023

Narsil commented Jul 3, 2023

Vinno97 commented Jul 4, 2023

Vinno97 commented Jul 6, 2023

lppllppl920 commented Jul 31, 2023

Narsil commented Aug 2, 2023

lppllppl920 commented Aug 2, 2023

Narsil commented Aug 3, 2023

Automatically map deduplicated safetensors weights to their original values #501

Automatically map deduplicated safetensors weights to their original values #501

Conversation

Vinno97 commented Jun 28, 2023

What does this PR do?

What does this fix?

Before submitting

Who can review?

Narsil commented Jun 30, 2023

Vinno97 commented Jul 1, 2023

Narsil commented Jul 3, 2023

Vinno97 commented Jul 4, 2023

Vinno97 commented Jul 6, 2023

lppllppl920 commented Jul 31, 2023

Narsil commented Aug 2, 2023

lppllppl920 commented Aug 2, 2023

Narsil commented Aug 3, 2023