|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Xen Summit 2026 - "Nested Virt <-> WinPV drivers" design session notes
The following notes were taken during the design session on Thursday 17th September (https://design-sessions.xenproject.org/uid/discussion/disc_bpkf1nkpeOyiAehjw4xN/view) titled "Nested Virt <-> WinPV drivers": This is an issue identified during XenServer's nested virt development. With HyperV, your nominal Windows machine becomes an L2 guest as the root partition under HyperV (~equivalent to Xen dom0). HyperV at L1 handles things from the L2 (flowing via Xen at L0). Specifically this affects CPUID and VM(M)CALL operations - L1 handles these and doesn't give the Xen parts that the PV drivers expect - consequence is PV drivers are unable to communicate with Xen. 'Feature' of the design of nesting. A Viridian enlightenment allows L1 to provide information to L0 as to which L2 is the root partition. A workaround was implemented using this where Xen can modify the responses so the L2 still 'sees' Xen and can enable the drivers. This relies on a _currently_ invalid input to HyperV, but this may fail in future so not a safe solution and not shippable. Tangent: Bromium. Different way of solving this problem, did all hypercalls using the CPUID instruction (Intel doesn't permit you not to have an exit on CPUID, AMD permits you not to intercept but everyone does). ABI for CPUID says nothing about high halves of the registers, so Bromium puts a magic constant in high half that L0 can recognise, or if not present the nominal L1 would still answer the normal CPUID. What are people's thoughts on such an approach? Would it be restricted to root partition? Yes. HyperV doesn't seem to have good provision for nested virt on hypervisors that are not MS. Can we stick to just Viridian enlightenments in the nested case and avoid the problem? Ideal world we'd do a full VMBus implementation and abandon the PV drivers but we can't (potential legal issues with doing so. Might change if MS do an implementation in Linux which is rumoured, but no sign of this as things stand - would be complete replacement of I/O model and not practical in the timeframe of the nested virt project). Could be looked into in future, but not a short term solution. HyperV knows about PCI and enables this for the root partition, but our PV model uses hypercalls. Some questions to put to MS Never found a situation where the root partition doesn't have an identity map Ballooning will break everything? Apparently MS has an API to add/remove memory that may be a solution here. Ballooning implementation in PV drivers uses API. Concern as to whether this API is designed for HyperV and not a nested case like this? Tu would like to retain ballooning in the drivers. Andrew wants to make it an admin choice (whereas currently it must be available even in dangerous cases). Interaction with ABI enumeration questions... We have the Xen PCI platform device in HVM domains. Root partition will see this device, could we send hypercalls to it? In principal maybe. Specific BAR for hypercalls where qemu tells the hypervisor a specific range? Would mean every hypercall goes into x86 emulate! Maybe could have some sort of special flag? What would benefit be over CPUID mechanism? Don't need to check it's the root partition, can just do it by access to PCI device. Wouldn't work for PVH. Could we do something with an MMIO range? Can't have a Xen specific driver use this as HyperV comes up before Xen drivers. Could declare in ACPI and then Windows knows a driver can bind to it? We only want the root partition to be doing PV I/O to L0, not any L2 guest. In a virtio model it'd be down to HyperV/root partition which guest it passes the PCI device to.
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |