What happened:
I deployed a deployment with 64 a100 GPUs, but from the screenshot, it seems like a lot of them are stuck in Pending. The general state transition is like below, so we should not see a lot of Pods with Pending status.
Pending -> Provisioning -> Running
What you expected to happen:
How to reproduce it (as minimally and precisely as possible):
Anything else we need to know?:
Environment:
What happened:
I deployed a deployment with 64 a100 GPUs, but from the screenshot, it seems like a lot of them are stuck in Pending. The general state transition is like below, so we should not see a lot of Pods with Pending status.
What you expected to happen:
How to reproduce it (as minimally and precisely as possible):
Anything else we need to know?:
Environment: