[DTensor] Made `_Partial`, `Replicate` frozen dataclasses #113919

awgu · 2023-11-17T04:39:31Z

Stack from ghstack (oldest at bottom):

This is part of the larger stack to work toward being able to cache hashes for DTensorSpec.

[ghstack-poisoned]

pytorch-bot · 2023-11-17T04:39:35Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/113919

📄 Preview Python docs built from this PR
📄 Preview C++ docs built from this PR
❓ Need help or want to give feedback on the CI? Visit the bot commands wiki or our office hours

Note: Links to docs will display an error until the docs builds have been completed.

✅ No Failures

As of commit f934854 with merge base 140c54e ():
💚 Looks good so far! There are no failures yet. 💚

This comment was automatically generated by Dr. CI and updates every 15 minutes.

wanchaol

lgtm

This is part of the larger stack to work toward being able to cache hashes for `DTensorSpec`. [ghstack-poisoned]

awgu · 2023-11-20T22:26:25Z

@pytorchbot merge

pytorchmergebot · 2023-11-20T22:28:15Z

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

Pull Request resolved: #113924 Approved by: https://github.com/wanchaol ghstack dependencies: #113919

Pull Request resolved: #114134 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924

Pull Request resolved: #113925 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134

…3930) Pull Request resolved: #113930 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134, #113925

This is a replacement for #113922. I think we can still leave the check for negative shard dimension in `compute_local_shape_and_global_offset` and replace the normalization logic with an assert. This should provide us a stack trace to see which user-facing API did not normalize the dim as expected. Pull Request resolved: #114141 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134, #113925, #113930

**Overview** Generally, I think we can try to freeze as many of these classes used in DTensor sharding propagation as possible so that we can cache hashes. This PR targets hashing `DTensorSpec`, which turns out to be relatively expensive. **Details** It looks like `tensor_meta` is only updated in `_wrap_output_spec_tensor_meta`, which only runs if the propagation was not cached: https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L137 https://github.com/pytorch/pytorch/blob/ae94c7e491e22f58d3df66571c1a568e51d70acd/torch/distributed/_tensor/sharding_prop.py#L153 In that case, I think we can cache the hash for the `DTensorSpec` and only update it when one of the hashed attributes changes, which we only really expect to happen for `tensor_meta`. To ensure correctness, we need that all hashed attributes are immutable. - `DeviceMesh` caches its hash: https://github.com/pytorch/pytorch/blob/a9134fa99a8986adf478a12db2ea5729d24554db/torch/distributed/_device_mesh.py#L181 - This PR makes each `Placement` a frozen `dataclass`, making them immutable (relying on the fact that they do not have references to any mutable objects). - `TensorMeta` is a `NamedTuple` of `torch.Size`, `Tuple[int, ...]`, and `torch.dtype`, so it is immutable: https://github.com/pytorch/pytorch/blob/9916d8a9eaaf2c05c131f2a2dbe9eabeeaa9dffc/torch/distributed/_tensor/placement_types.py#L369-L375 **Example** For some simple small GPT model: Before: 0.125 ms <img width="509" alt="Screenshot 2023-11-16 at 10 08 05 PM" src="https://github.com/pytorch/pytorch/assets/31054793/10e59401-f635-431f-80b5-1b48df3a706e"> After: 0.048 ms <img width="294" alt="Screenshot 2023-11-16 at 10 08 47 PM" src="https://github.com/pytorch/pytorch/assets/31054793/09a3b0b9-f68c-4afc-bca1-c29a4b01c2fb"> The overall Adam CPU step time decreases from 7.647 ms to 6.451 ms. Pull Request resolved: #113915 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134, #113925, #113930, #114141

This is a nit change to save one `isinstance` call for when `dim` is not `None` but the placement is not `Shard`. Pull Request resolved: #114140 Approved by: https://github.com/Skylion007, https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134, #113925, #113930, #114141, #113915

This is a forward fix for #113781. We lazily compute the hash so that we do not try to compute the hash on `SymInt`s (for the stride) during Dynamo tracing. Tested via: ``` python test/distributed/_tensor/test_dtensor_compile.py -k test_2d_fsdp_tp_ac_compile ``` Pull Request resolved: #114322 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134, #113925, #113930, #114141, #113915, #114140

ghstack-source-id: b87ed7aab72715b43f61e5313d5d485401c5a0c8 Pull Request resolved: pytorch/pytorch#113919

[DTensor] Made _Partial, Replicate frozen dataclasses

d2e70ff

[ghstack-poisoned]

awgu mentioned this pull request Nov 17, 2023

[DTensor] Renamed shard_spec -> placements in test file #113917

Closed

awgu mentioned this pull request Nov 17, 2023

[DTensor] Cached hash for DTensorSpec #113915

Closed

awgu marked this pull request as draft November 17, 2023 04:39

wanchaol approved these changes Nov 17, 2023

View reviewed changes

Update on "[DTensor] Made _Partial, Replicate frozen dataclasses"

f934854

This is part of the larger stack to work toward being able to cache hashes for `DTensorSpec`. [ghstack-poisoned]

awgu mentioned this pull request Nov 20, 2023

[DTensor] Used new placements for neg dim in from_local #114134

Closed

awgu marked this pull request as ready for review November 20, 2023 17:17

This was referenced Nov 20, 2023

[DTensor] Reduced to one isinstance call in is_shard #114140

Closed

[DTensor] Replaced neg dim normalization with assert in helper #114141

Closed

awgu added release notes: distributed (dtensor) release notes category ciflow/trunk Trigger trunk jobs on your pull request labels Nov 20, 2023

pytorchmergebot added the merging label Nov 20, 2023

pytorchmergebot added the Merged label Nov 20, 2023

pytorchmergebot closed this in 77e058f Nov 20, 2023

pytorchmergebot removed the merging label Nov 20, 2023

pytorchmergebot pushed a commit that referenced this pull request Nov 20, 2023

[DTensor] Used new placements for neg dim in redistribute (#113924)

b41ad7d

Pull Request resolved: #113924 Approved by: https://github.com/wanchaol ghstack dependencies: #113919

pytorchmergebot pushed a commit that referenced this pull request Nov 20, 2023

[DTensor] Used new placements for neg dim in from_local (#114134)

f4ffd46

Pull Request resolved: #114134 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924

pytorchmergebot pushed a commit that referenced this pull request Nov 20, 2023

[DTensor] Ensured grad_placements was tuple (#113925)

e2095a0

Pull Request resolved: #113925 Approved by: https://github.com/wanchaol ghstack dependencies: #113919, #113924, #114134

awgu mentioned this pull request Nov 22, 2023

[DTensor] Computed DTensorSpec hash lazily #114322

Closed

awgu mentioned this pull request Nov 22, 2023

[FSDP] Added DDP parity test for CPU training #114372

Closed

facebook-github-bot deleted the gh/awgu/457/head branch November 24, 2023 15:27

mgcorona pushed a commit to mgcorona/pytorch that referenced this pull request Nov 28, 2023

[DTensor] Made _Partial, Replicate frozen dataclasses

20a295f

ghstack-source-id: b87ed7aab72715b43f61e5313d5d485401c5a0c8 Pull Request resolved: pytorch/pytorch#113919

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

[DTensor] Made `_Partial`, `Replicate` frozen dataclasses #113919

[DTensor] Made `_Partial`, `Replicate` frozen dataclasses #113919

awgu commented Nov 17, 2023 •

edited

pytorch-bot bot commented Nov 17, 2023 •

edited

wanchaol left a comment

awgu commented Nov 20, 2023

pytorchmergebot commented Nov 20, 2023

[DTensor] Made _Partial, Replicate frozen dataclasses #113919

[DTensor] Made _Partial, Replicate frozen dataclasses #113919

Conversation

awgu commented Nov 17, 2023 • edited

pytorch-bot bot commented Nov 17, 2023 • edited

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/113919

✅ No Failures

wanchaol left a comment

Choose a reason for hiding this comment

awgu commented Nov 20, 2023

pytorchmergebot commented Nov 20, 2023

Merge started

[DTensor] Made `_Partial`, `Replicate` frozen dataclasses #113919

[DTensor] Made `_Partial`, `Replicate` frozen dataclasses #113919

awgu commented Nov 17, 2023 •

edited

pytorch-bot bot commented Nov 17, 2023 •

edited