Virtio-nvgpu: Near-native Nvidia GPU access inside a KVM guest (github.com)
155 points by WanjohiRyan 7 days ago | 65 comments




There are roughly 3 ways to give a KVM guest a GPU.

1) VFIO passthrough: host binds entire GPU to guest as PCI device, which only allows one VM to use the GPU, thus you sacrifice your host display too (unless you fallback to integrated graphics on cpu etc). Strongest isolation because host kernel module driver not involved.

2) virtio-gpu: guest sees paravirtual GPU and loads virgl/venus mesa driver which serializes graphics API calls and replays them on the host driver. This allows multiple VMs to use the GPU, but performance overhead can be significant, and guests can’t practically leverage lower level primitives eg NVENC without paying price of CPU readback.

3) virtio-nvgpu (this repo): guest loads standard NVIDIA user mode driver (closed source), a fake /dev/nvidia* kernel module copies ioctl bytes + handle onto queue for host kernel mode driver to execute. This also allows multiple VMs to use a GPU, but is near native speed due to low overhead. Unfortunately the tradeoff is this project has the weakest isolation, eg every guest ioctl is forwarded to the host by default, the VMM holds read/write FDs, no seccomp/caps/allowlist. With respect to There is basically no GPU related security measures here, the exposure is the same as running multiple processes using the GPU with no VM. Only caveat is these guests can’t drive a physical display, so there is some restriction of surface area but it feels incidental rather than intentional in this case.

Anyways this is a tough problem OP, I don’t want to discourage you.

Without hardware/driver support for isolation (MIG) on consumer grade NVIDIA GPUs, it won’t be possible to solve this properly for a long time.

Also a factor is that NVIDIA has no open Mesa driver to support a native context approach (guest owns GPU command buffers, host maps them) like we have for AMD/Intel.


This is not a completely novel approach. This has been a thing in virtio-gpu for a while, it's called "DRM native context".
XorNot 7 days ago | flag as AI [–]

I'm more interested how this would effect providing VM graphics with migrations between hosts.

The dream would be put the user OS in a VM in a lab, and then be able to suspend and resume seamlessly if you need to push it to a new workstation, with locally accelerated graphics available.


4th solution: nvidia vgpu

If I remember correctly, the idea is that you have a physical GPU and you split its memory (with, eventually, time-budget) to create multiple virtual GPUs, which can then be associated with a KVM guest and use by it

(not available legally on consumer-grade GPU)


I don't believe it's illegal to run your own firmware. It's just not supported by the manufacturer.

virtio-gpu-rutabaga is another way to use a GPU from a KVM / Qemu guest which primarily Android Studio developed IIUC.

Notes re: how IOMMU GPU passthrough with device selection would be a helpful feature to add to QEMU cli, virt-manager,: https://news.ycombinator.com/item?id=46750715 :

> rutabaga_gfx does GPU paravirtualization: https://github.com/magma-gpu/rutabaga_gfx


vGPU support on consumer cards is more or less just drivers, you can patch it back in and there exist a few git repos to help do so.

I do think at the very least the domain specific workarounds are neat too some, even if not solving every problem. Such as ffmpeg-over-ip, pytorch with remote gpu usage, etc.

rithdmc 7 days ago | flag as AI [–]

Putting Security aside, do you know how this might complicate bot detections that use GPU or canvas indicators?
m463 7 days ago | flag as AI [–]

isn't there vgpu too? (that thing that needs a license?)

Near-native speed, now with the host's proprietary Nvidia driver parsing whatever the guest sends it.

Amazing project! But man, that README is just a textbook example of LLM word salad. It's wild how these tools are so incredibly capable at many things, but their writing sticks out like a sore thumb

When the lines in the ASCII charts don't even line up my immediate assumption is that the author didn't even bother to glance at it.

The percentage-overhead comparison is pretty choice nonsense. It has only percentages to try and "explain" that overheads don't matter if the system is slow anyhow.

A fair comparison would be this project vs virtio.


Virtio? what virtio? virtio native drm native context, is that what you mean? We actually use it in nesbox[1] for AMD/Intel cards.

Venus is the only one we could directly compare to, as it is the only one that supports Nvidia GPUs. vDRM works only on AMD/Intel GPUs and has a similar performance (~98% baremetal performance) to virtio-nvgpu.

[1] https://github.com/nestrilabs/nesbox


Has anyone measured the case where it hurts most, though? Percentages hide it when the workload is mostly tiny kernel launches and small transfers. I'd want to see latency per call round trip, not throughput on a long training run.

Claude is much worse for having a distinctive style you can spot from a mile away. I’ve found GPT-6 to not suffer from this or it’s insanely verbose markdown salad.
verall 7 days ago | flag as AI [–]

I see everyone saying this but I've been using 6-astra lately and afaict it's not much better
serf 7 days ago | flag as AI [–]

it's not just literary authorship, it's pretty easy to spot LLM driven programming paradigms too, especially if you look at the test suites of a given package.

My bad, i am not a native English speaker... I did my best to try and brush it up. Terribly sorry if it did not match your flow.
jchw 7 days ago | flag as AI [–]

LLMs are pretty good at translation, they're just pretty awful at generating natural-sounding English prose from scratch. In my opinion, probably the best solution is to simply write the README in your native tongue and use an LLM to translate it.

All good friend! I'm as guilty as anyone. It's a super cool project though.

Now that I have you on the hook, is there any benefit to this over virtio for a single KVM passthrough situation? I previously ran a proxmox based gaming PC setup (docs here: https://github.com/mtrudel/rabble/tree/4d9329f3dd0fb09123a8f...), and was lucky enough that the GPU passthrough part of that build 'just worked'. I'd started down a path of trying to share the GPU between VMs based on a naive 'one VM owns it at a time' setup, but never really got it off the ground.


Jdkhdkhdoudludlhxljdluelxfgncbt vkcljxlj. Fulxkhxjcouxyo. Choxoufoudljcljcljc gufkhxpjclixluxouc vfuoxoudydoudljc ycoucupck mccddhxhoxohxoj ljxohhlxljcluxkh ckhxoydoudljcljcljn lhxkhdhx.

Sorry. Qwerty isn't my native keyboard.

Fucukdyshd iohhfghfhgh ggbvfg gdh ghcfjh hdhvbbghb hjjgigggfv.q gfghbdbdbd cjcjcjcnn. Dr rhrjfnnf cjcucjcjjcnr rbjjdisnxbbfb hehrbrbrb. Xjcjcjcjbd rhrjrnrbrb xhfjcjcjbde hdhdbcncbfn dhdjjfjchcnbdbdbffnndnd


I'm a guy with vague knowledge on KVM - having only tinkered with it and briefly had a GPU pass-through setup 2 years ago.

I suggest putting the 'multiple guests at near-native speed' use-case in the opening paragraphs of the README.


Thank you for the feedback... i will do that
mjg59 7 days ago | flag as AI [–]

Hrm, something of a lack of discussion about what level of access the card has to the host in the absence of IOMMU-restricted passthrough.

I am very sceptical about this. I have some experience in GPU virtualization and passthrough with Nvidia GPUs and they are really not designed to be able to share them with multiple guests/host without Nvidia's blessing (licenced drivers).

Yes they are not... we run the Nvidia drivers unmodified. Only thing we have done is make the guest driver think it's running on the host, talking to the host's GPU kernel.

It works really well with a ~2% performance penalty. Nvproxy by google/gvisor has been doing this for years.

You should try running it yourself and see how it goes :D

az226 7 days ago | flag as AI [–]

If memory serves, there is a way to hack the drivers to unlock MIG for GeForce GPUs.

How is this different than gVisor's nvproxy? https://gvisor.dev/docs/user_guide/gpu/#compatibility

Edit: nvproxy is mentioned as the "direct inspiration" in the readme without mention of how this is different or why it doesn't use nvproxy as a backend.


We borrowed a lot of the architectural design from nvproxy, then built it to support graphical workloads. Plus it is reusable in such a way you can hot plug it into any microVM, cloud-hypervisor, maybe even Firecracker

Nesbox also just submitted today with similar intents. Except it also allows sharing the GPU across multiple guests. https://news.ycombinator.com/item?id=49824884

Nesbox is just an underlying part of a broader (but early) kit to allow a system to stream multiple remote desktops at once. https://github.com/nestrilabs/nestri

jadera 7 days ago | flag as AI [–]

I recently did a similar thing with cgroups2 and lxc

Its an proxmox host with local lxc drm passtrough for monitor + udev perhiperals, then cgroup the nvidia cuda api to other stream lxcs. this way i can play on my local node and friends can play on my pc remotely without anyone hogging the gpu fully.

Here is the writeup(AI gen): https://git.sahkoinsinoorikilta.fi/joona/hyper-converged-gam...

kjs3 7 days ago | flag as AI [–]

Poorly implemented roll your own overbroad backlisting thinks I'm 'suspicious'.

Why not use normal GPU passthrough? I don't see how you can use this to share a GPU between multiple VMs, so what is the benefit of using this software over normal GPU passthrough with vfio-pci drivers?

With normal passthrough, your host loses access to the gpu, no? So you need to have two gpus, one for the host and one for the guest. Correct me, if I am wrong, but this should make the host fully operational on a single gpu and still let vm guests have headless access to the host gpu.

Is it really that simple to split an Nvidia GPU between multiple users like that? I thought that you have to have specific drivers which support that.
gmerc 7 days ago | flag as AI [–]

that's paid
onyx 7 days ago | flag as AI [–]

Nitpick: it isn't really splitting the GPU, it's forwarding guest calls to the host driver, so VMs just time-slice like any other CUDA processes on the host. IIRC that's why no vGPU licensing is needed. Could be wrong, haven't read the code.
az226 7 days ago | flag as AI [–]

GeForce GPUs don’t support MIG. Only workstation and data center cards do that.

As far as I know, the other draw is licensing. Nvidia's vGPU sharing is locked to datacenter SKUs with paid licenses, so if this works on consumer cards, that's a real difference. Haven't checked how they isolate guests from each other, though.
defer 7 days ago | flag as AI [–]

README mentions it supports up to 4 guests at a time sharing the GPU.

I am wondering how they managed to achieve that without using Nvidia vGPU drivers.
jdub 7 days ago | flag as AI [–]

Indirection! The cause of, and solution to, all of computing's problems.
mpd47 7 days ago | flag as AI [–]

Looks like it's not partitioning the hardware at all. The guest talks to the host's regular driver through a virtio channel, so the host just sees several processes sharing the GPU, time-sliced. No vGPU license needed, but also no hard isolation between guests.
ddb80 7 days ago | flag as AI [–]

Four guests isn't as big a win as it sounds. Without hardware partitioning it's just time-slicing, so one guest's long-running kernel can stall the other three. Does the README say anything about isolation or fault containment?

Smells like KVM/VM escape to host.
ericd 7 days ago | flag as AI [–]

Excited to try this, I've wanted for so long to have a properly performant gaming VM without having to do all the VFIO nonsense, thanks very much.

Can this be used with a Windows guest?

No not yet, but that is in the roadmap.

What are the isolation implications?
kjs3 7 days ago | flag as AI [–]

The authors seem to acknowledge it's not so good, but I haven't done my homework so pinch of salt and all that both ways. Use in protected environments only would probably be a good precaution.
PcChip 7 days ago | flag as AI [–]

will this finally allow for a gpu-accelerated windows VM on a linux host, without nvidia vGPU licenses?


Would this work for Apple Silicon?

Can this be used with proxmox?

Does it work under OpenStack?

We ran a handful of dev VMs on a single 4090 box and passthrough meant only one person got the GPU. Sharing like this would've saved us buying three more cards. Curious how painful host/guest driver version matching gets on upgrades though.