[CUDA] Fix truncated error messages in cudaMallocAsync Allocator #168369

galv · 2025-11-21T19:05:10Z

Previously, these error messages would get truncated when they were hit on device 0 because device is a "char" (actually, an int8_t) and therefore '0' is interpreted as the null byte to terminate a string. Essentially, it is the same issue as #123984.

There's something strange in the TORCH_CHECK_WITH macro that is causing this. I don't feel like figuring out those obscure macro details right now, though.

cc @ptrblck @msaroufim @eqy @jerryzh168 @tinglvv @nWEIdia

Previously, these error messages would get truncated when they were hit on device 0 because device is a "char" (actually, an int8_t) and therefore '0' is interpreted as the null byte to terminate a string. Essentially, it is the same issue as #123984. There's something strange in the TORCH_CHECK_WITH macro that is causing this. I don't feel like figuring out those obscure macro details right now, though.

pytorch-bot · 2025-11-21T19:05:14Z

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/168369

📄 Preview Python docs built from this PR
📄 Preview C++ docs built from this PR
❓ Need help or want to give feedback on the CI? Visit the bot commands wiki

Note: Links to docs will display an error until the docs builds have been completed.

✅ You can merge normally! (1 Unrelated Failure)

As of commit 705868c with merge base 402968e ():

FLAKY - The following job failed but was likely due to flakiness present on trunk:

trunk / linux-jammy-rocm-py3.10 / test (default, 6, 6, linux.rocm.gpu.gfx942.1) (gh) (similar failure)
Process completed with exit code 1.

This comment was automatically generated by Dr. CI and updates every 15 minutes.

eqy

I wonder how many of these are left floating around... previous fixes include e.g., 975f777

galv · 2025-11-22T21:27:48Z

@pytorchbot merge

pytorchmergebot · 2025-11-22T21:29:36Z

Merge started

Your change will be merged once all checks pass (ETA 0-4 Hours).

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

pytorchmergebot · 2025-11-22T22:00:02Z

Merge failed

Reason: 1 jobs have failed, first few of them are: trunk / linux-jammy-rocm-py3.10 / test (default, 6, 6, linux.rocm.gpu.gfx942.1)

Details for Dev Infra team

Raised by workflow job

cyyever · 2025-11-23T23:56:18Z

@pytorchbot merge -i

pytorchmergebot · 2025-11-23T23:58:08Z

Merge started

Your change will be merged while ignoring the following 1 checks: trunk / linux-jammy-rocm-py3.10 / test (default, 6, 6, linux.rocm.gpu.gfx942.1)

Learn more about merging in the wiki.

Questions? Feedback? Please reach out to the PyTorch DevX Team

Advanced Debugging

Check the merge workflow status
here

galv requested review from Aidyn-A, eqy and syed-ahmed as code owners November 21, 2025 19:05

eqy approved these changes Nov 21, 2025

View reviewed changes

eqy added module: cuda Related to torch.cuda, and CUDA support in general open source topic: bug fixes topic category topic: not user facing topic category labels Nov 21, 2025

pytorch-bot bot added the ciflow/trunk Trigger trunk jobs on your pull request label Nov 22, 2025

pytorchmergebot added the merging label Nov 22, 2025

pytorchmergebot removed the merging label Nov 22, 2025

pytorchmergebot added the merging label Nov 23, 2025

pytorchmergebot closed this in 9a38bb8 Nov 24, 2025

pytorchmergebot added Merged and removed merging labels Nov 24, 2025

lingebeng mentioned this pull request Nov 24, 2025

[CUDA][BugFix] fix truncated error messages #168942

Open

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

[CUDA] Fix truncated error messages in cudaMallocAsync Allocator #168369

[CUDA] Fix truncated error messages in cudaMallocAsync Allocator #168369

Uh oh!

galv commented Nov 21, 2025 •

edited by pytorch-bot bot

Loading

Uh oh!

pytorch-bot bot commented Nov 21, 2025 •

edited

Loading

Uh oh!

eqy left a comment

Uh oh!

galv commented Nov 22, 2025

Uh oh!

pytorchmergebot commented Nov 22, 2025

Uh oh!

pytorchmergebot commented Nov 22, 2025

Uh oh!

cyyever commented Nov 23, 2025

Uh oh!

pytorchmergebot commented Nov 23, 2025

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

[CUDA] Fix truncated error messages in cudaMallocAsync Allocator #168369

[CUDA] Fix truncated error messages in cudaMallocAsync Allocator #168369

Uh oh!

Conversation

galv commented Nov 21, 2025 • edited by pytorch-bot bot Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

pytorch-bot bot commented Nov 21, 2025 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

🔗 Helpful Links

🧪 See artifacts and rendered test results at hud.pytorch.org/pr/168369

✅ You can merge normally! (1 Unrelated Failure)

Uh oh!

eqy left a comment

Choose a reason for hiding this comment

Uh oh!

galv commented Nov 22, 2025

Uh oh!

pytorchmergebot commented Nov 22, 2025

Merge started

Uh oh!

pytorchmergebot commented Nov 22, 2025

Merge failed

Uh oh!

cyyever commented Nov 23, 2025

Uh oh!

pytorchmergebot commented Nov 23, 2025

Merge started

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

5 participants

galv commented Nov 21, 2025 •

edited by pytorch-bot bot

Loading

pytorch-bot bot commented Nov 21, 2025 •

edited

Loading