Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest
Unofficial Hacker News client; not affiliated with Y Combinator.
refibrillator · · focus · HN ↗
1) VFIO passthrough: host binds entire GPU to guest as PCI device, which only allows one VM to use the GPU, thus you sacrifice your host display too (unless you fallback to integrated graphics on cpu etc). Strongest isolation because host kernel module driver not involved.
2) virtio-gpu: guest sees paravirtual GPU and loads virgl/venus mesa driver which serializes graphics API calls and replays them on the host driver. This allows multiple VMs to use the GPU, but performance overhead can be significant, and guests can’t practically leverage lower level primitives eg NVENC without paying price of CPU readback.
3) virtio-nvgpu (this repo): guest loads standard NVIDIA user mode driver (closed source), a fake /dev/nvidia* kernel module copies ioctl bytes + handle onto queue for host kernel mode driver to execute. This also allows multiple VMs to use a GPU, but is near native speed due to low overhead. Unfortunately the tradeoff is this project has the weakest isolation, eg every guest ioctl is forwarded to the host by default, the VMM holds read/write FDs, no seccomp/caps/allowlist. With respect to There is basically no GPU related security measures here, the exposure is the same as running multiple processes using the GPU with no VM. Only caveat is these guests can’t drive a physical display, so there is some restriction of surface area but it feels incidental rather than intentional in this case.
Anyways this is a tough problem OP, I don’t want to discourage you.
Without hardware/driver support for isolation (MIG) on consumer grade NVIDIA GPUs, it won’t be possible to solve this properly for a long time.
Also a factor is that NVIDIA has no open Mesa driver to support a native context approach (guest owns GPU command buffers, host maps them) like we have for AMD/Intel.
deltoidmaximus · · focus · HN ↗
What kinds of things can a guest running undesirably applications (viruses, malware, LLM escaping a sandbox, etc) get up to with shared GPU access?
Sohcahtoa82 · · focus · HN ↗
Indeed. For me, I find the ecosystem around AI/LLMs works better on Linux than Windows, but Windows is my main OS since I'm a gamer. Being able to run GPU-accelerated AI in a VM is huge for me.
In my case though, I just use WSL which does an amazing job.
> What kinds of things can a guest running undesirably applications (viruses, malware, LLM escaping a sandbox, etc) get up to with shared GPU access?
The most obvious answer is a DoS. If my malicious VM is sharing a GPU and has full access to it, I could simply tell the GPU not to run a victim VM's workload, or manipulate it in some way. I might not be able to pivot to having a shell on their VM, but I could at least read/write their data in VRAM. If it contained secret data (custom model, or secret data being processed by AI), I could easily steal it.