Skip to content

Conversation

XilunWu
Copy link
Contributor

@XilunWu XilunWu commented May 9, 2025

Copy link

pytorch-bot bot commented May 9, 2025

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/153225

Note: Links to docs will display an error until the docs builds have been completed.

❌ 1 New Failure, 1 Cancelled Job, 1 Unrelated Failure

As of commit afa1ff7 with merge base ab829ec (image):

NEW FAILURE - The following job has failed:

CANCELLED JOB - The following job was cancelled. Please retry:

UNSTABLE - The following job is marked as unstable, possibly due to flakiness on trunk:

This comment was automatically generated by Dr. CI and updates every 15 minutes.

@pytorch-bot pytorch-bot bot added ciflow/inductor oncall: distributed Add this issue/PR to distributed oncall triage queue release notes: distributed (fsdp) release notes category release notes: distributed (checkpoint) labels May 9, 2025
@XilunWu XilunWu added better-engineering Relatively self-contained tasks for better engineering contributors topic: not user facing topic category and removed release notes: distributed (fsdp) release notes category release notes: distributed (checkpoint) labels May 9, 2025
@XilunWu XilunWu requested review from Skylion007 and cyyever May 9, 2025 01:23
@XilunWu XilunWu added the ciflow/trunk Trigger trunk jobs on your pull request label May 9, 2025
@cyyever cyyever requested a review from kwen2501 May 9, 2025 02:36
@cyyever
Copy link
Collaborator

cyyever commented May 9, 2025

I'm not an expert of distributed code. @kwen2501 can help decide.

Copy link
Contributor

@kwen2501 kwen2501 left a comment

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

…ch.distributed.tensor in test files"

cc H-Huang awgu wanchaol fegin fduwjj wz337 wconstab d4l3k

[ghstack-poisoned]
@XilunWu
Copy link
Contributor Author

XilunWu commented May 9, 2025

@pytorchbot merge

@pytorchmergebot
Copy link
Collaborator

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging
Check the merge workflow status
here

…ch.distributed.tensor in test files"

cc H-Huang awgu wanchaol fegin fduwjj wz337 wconstab d4l3k

[ghstack-poisoned]
XilunWu added a commit that referenced this pull request May 9, 2025
…ted.tensor in test files

ghstack-source-id: dcfa329
Pull Request resolved: #153225
@XilunWu
Copy link
Contributor Author

XilunWu commented May 9, 2025

@pytorchbot merge -f "unrelated ROCm state dict save/load test failure on 7 devices due to uneven sharding"

@pytorchmergebot
Copy link
Collaborator

Merge started

Your change will be merged immediately since you used the force (-f) flag, bypassing any CI checks (ETA: 1-5 minutes). Please use -f as last resort and instead consider -i/--ignore-current to continue the merge ignoring current failures. This will allow currently pending tests to finish and report signal before the merge.

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging
Check the merge workflow status
here

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

better-engineering Relatively self-contained tasks for better engineering contributors ciflow/inductor ciflow/trunk Trigger trunk jobs on your pull request Merged oncall: distributed Add this issue/PR to distributed oncall triage queue topic: not user facing topic category

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants