Skip to content

GPU passthrough for micro-VM actors #762

Description

@eliranw

GPU passthrough for gVisor actors is proposed in #502. Nothing equivalent exists for micro-VM actors, and that PR gates the combination: a microvm WorkerPool requesting nvidia.com/gpu is rejected at apply time, because the pod would otherwise schedule onto a GPU node and hold a device no actor could use.

The gVisor approach does not carry over. It injects the host driver into the sandbox via CDI, which relies on the actor sharing the host kernel. A micro-VM guest has its own kernel, so the device has to be passed through — VFIO PCI passthrough for the GPU, plus getting a matching driver and user-mode libraries into the guest.

Open questions:

  • VFIO device binding and IOMMU requirements on the node, and whether that composes with the unprivileged worker posture
  • How the guest gets a driver matching the host GPU, and how it stays version-matched
  • Whether cloud-hypervisor's VFIO support covers what is needed here
  • What this means for snapshot/restore, which for micro-VM captures guest memory

Filing so the design is tracked outside the gVisor PR.

Related: #627, #502

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

Milestone

No milestone

Relationships

None yet

Development

No branches or pull requests

Issue actions