|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Re: [PATCH 0/8] Force detach PCI devices for ACPI-based and PCIe native hot-unplug
On Fri, 4 Sep 2026 07:22:10 -0400 "Michael S. Tsirkin" <mst@xxxxxxxxxx> wrote: > On Fri, Sep 04, 2026 at 01:07:09PM +0200, Igor Mammedov wrote: > > On Thu, 3 Sep 2026 16:17:20 -0400 > > "Michael S. Tsirkin" <mst@xxxxxxxxxx> wrote: > > > > > On Thu, Sep 03, 2026 at 05:28:35PM +0200, Igor Mammedov wrote: > > > > On Wed, 26 Aug 2026 09:15:47 -0700 > > > > Dongli Zhang <dongli.zhang@xxxxxxxxxx> wrote: > > > > > > > > > On Mon, Aug 24, 2026 7:44:49AM -0700, Daniel P. Berrangé wrote: > > > > > > On Sun, Aug 23, 2026 at 06:13:30PM -0700, Dongli Zhang wrote: > > > > > >> Hot-unplugging a PCI device can require cooperation from the > > > > > >> guest. For > > > > > >> ACPI PCI hotplug, QEMU notifies the guest through ACPI and the > > > > > >> guest > > > > > >> eventually writes the ACPI PCI eject register. For PCIe native > > > > > >> hotplug, > > > > > >> QEMU notifies the guest through the PCIe hotplug mechanism and > > > > > >> waits for > > > > > >> the slot unplug flow to complete. Only after that completion does > > > > > >> QEMU > > > > > >> unrealize the device and emit DEVICE_DELETED. > > > > > >> > > > > > >> This can leave a device stuck in the unplug pending state when the > > > > > >> guest > > > > > >> does not cooperate. Examples include: > > > > > >> > > > > > >> 1. The guest has panicked, or the relevant ACPI/PCI hotplug driver > > > > > >> is > > > > > >> unavailable. > > > > > >> > > > > > >> 2. The guest is stalled and cannot handle the hot-unplug event. For > > > > > >> example, stalling the Linux [irq/9-acpi] kernel thread can > > > > > >> reproduce this > > > > > >> for ACPI-based hot-unplug. > > > > > >> > > > > > >> 3. The device was attached to a slot that the guest cannot use. For > > > > > >> example, a pcie-root-port only supports slot 0. If a device is > > > > > >> added to a > > > > > >> non-zero slot below a pcie-root-port, the guest may never discover > > > > > >> the > > > > > >> device and therefore may never complete the unplug request. > > > > > > > > all of above is actually expected, no (functioning) driver => no > > > > hotplug/unplug. > > > > it's the guest problem. Once device it exposed to guest its life-cycle > > > > not longer owned by QEMU. > > > > > > > > That's what one would see in real hw as well, you press eject button > > > > but it will not do anything if OS doesn't process it. > > > > also see comment at the end. > > > > > > > > > >> > > > > > >> The non-zero slot case has also been discussed in: > > > > > >> > > > > > >> hw/pci: warn when PCIe device is plugged into non-zero slot of > > > > > >> downstream port > > > > > >> https://gitlab.com/qemu-project/qemu/-/commit/ > > > > > > ca92eb5defcf9d1c2106341744a73a03cf26e824 > > > > > >> > > > > > >> hw/pci: add comment to explain checking for available function 0 > > > > > >> in pci hotplug > > > > > >> https://gitlab.com/qemu-project/qemu/-/ > > > > > > commit/67d045a0ef5b9c5f871c3a1d87325a8a42d2b9d5 > > > > > >> > > > > > >> pci: don't skip function 0 occupancy verification for devfn auto > > > > > >> assign > > > > > >> https://gitlab.com/qemu-project/qemu/-/commit/ > > > > > > e228d62b4af29bca698ec57efdceb46f392f5444 > > > > > >> > > > > > >> For example, if root-port.1 is a pcie-root-port, the following > > > > > >> command adds > > > > > >> a vhost-scsi-pci device to an invalid slot: > > > > > >> > > > > > >> (qemu) device_add > > > > > >> vhost-scsi-pci,id=scsi01,wwpn=naa.5001405324af0985,bus=root-port.1,addr=01.0 > > > > > >> warning: PCI: slot 1 is not valid for vhost-scsi-pci, parent > > > > > >> device only allows plugging into slot 0. > > > > > > > > > > > > This rather looks like it should be a fatal error, not a mere > > > > > > warning. > > > > > > > > > > > > If I follow the commit ca92eb5def it links to > > > > > > https://bugzilla.redhat.com/show_bug.cgi?id=2128929 > > > > > > which states that this configuration is going to lead to a crash in > > > > > > QEMU on guest OS shutdown. IMHO that crash is sufficient to justify > > > > > > making this a fatal error. > > > > > > > > > > > > If we actually wanted this to remain a warning, then that shutdown > > > > > > crash would need to be fixed. > > > > > > > > > > > > > > > > Thank you very much! > > > > > > > > > > I see that the issue has been fixed. The ticket mentions the > > > > > following. > > > > > > > > > > "What I am observing is that it seems when the slot ID != 0, the > > > > > guest OS seems > > > > > to ignore this and we never seem to hit ich9_pm_device_unplug_cb()." > > > > > > > > > > Based on my experience and evaluation, ACPI-based hotplug is more > > > > > likely to > > > > > encounter an issue where the guest VM does not respond to an unplug > > > > > operation. > > > > > > > > I'm not sure it's a good idea to delete device when guest still thinks > > > > it's there > > > > (you can make guesses on QEMU side if it's in use, how useful those are > > > > is questionable). > > > > > > > > as far as I know, ACPI hotplug has no notion of surprise removal (pls > > > > educate me if it's not the case), > > > > > > why would it not? > > > > > > how do you think you can pull a laptop out of a dock? > > > I expect bus check + _STA and config space saying it is gone > > > will do exactly that. > > > > > > > > > Here's linux code: > > > static void acpiphp_check_bridge(struct acpiphp_bridge *bridge) > > > { > > > struct acpiphp_slot *slot; > > > > > > /* Bail out if the bridge is going away. */ > > > if (bridge->is_going_away) > > > return; > > > > > > if (bridge->pci_dev) > > > pm_runtime_get_sync(&bridge->pci_dev->dev); > > > > > > list_for_each_entry(slot, &bridge->slots, node) { > > > struct pci_bus *bus = slot->bus; > > > struct pci_dev *dev, *tmp; > > > > > > if (slot_no_hotplug(slot)) { > > > ; /* do nothing */ > > > } else if (device_status_valid(get_slot_status(slot))) { > > > /* remove stale devices if any */ > > > list_for_each_entry_safe_reverse(dev, tmp, > > > &bus->devices, > > > bus_list) > > > if (PCI_SLOT(dev->devfn) == slot->device) > > > trim_stale_devices(dev); > > > > > > /* configure all functions */ > > > enable_slot(slot, true); > > > } else { > > > disable_slot(slot); > > > } > > > } > > > > > > if (bridge->pci_dev) > > > pm_runtime_put(&bridge->pci_dev->dev); > > > } > > > > > > > > > so weirdly it wants bus check on a parent bus, otherwise it will > > > not trim devices? probably a bug, but easy to work around. > > > > Modern docks would use native pcie surprise removal path. > > > > As for ACPI, my old laptop, had an unlock button => _LCK > > and that relied on OS processing ACPI events, not so surprise. > > > > There might have been ACPI/hybrid docks that did surprise removal, > > but then one need to find one and model after that instead of > > just blanket force removal. (likely out come would a doc device > > support only, not an arbitrary device removal) > > > > (not the case described in this series, though. hence my request to clarify > > usecase) > > > > from what I see in spec there is _RMV method that says that device > > supports surprise removal that can be used for devices that support it. > > However I would hesitate very much to blank apply it to every PCI device. > > (it's not even realistic to ask for proving safe tear down across various > > drivers and OSes/versions) > > > > Rather than a knee jerk treatment of misconfig consequences, > > I'd rather see patches to prevent misconfig in the 1st place > > (subj to deprecation but doable). > > > > As for the cases where OS mis-behaves (apcihp thread starvation,...), > > fixing guest to follow hotplug contract is a proper place to do it. > > > > On QEMU side we have it covered as well. If unplug was not processed, > > mgmt is free to repeat action. > > > Sorry if I am unclear. I just meant that it looks like we > can support surprise removal with ACPI just by reporting > bus check events on the parent. maybe, but that ain't SPECed and might be OS specific. The way I've read the cover letter, that won't work for mentioned mis-config cases. Also what would happen on bus-check 'cleanup' would be a lottery. hence I'm for being safe here. It's better to implement native PCIE surprise removal if that's really needed. > > > > so I wouldn't do what you are proposing here at all, it's basically > > > > asking for disaster to happen. > > > > And all this is basically for dealing with abused qemu flexibility. > > > > > > > > Please (re)formulate usecase and make it more clear as what is eludes me > > > > no matter how many times i've read this cover letter. > > > > > > > > On positive note: > > > > > > > > What you can try to implement is native PCI-E support for surprise > > > > removal. > > > > How hard that would be I don't know. And I would well expect if one > > > > deviates from > > > > real hw expectations/configs (such as not 0 slot/partial func removal), > > > > one would quickly stumble upon issues as that's not what what vendors > > > > write/test > > > > drivers for. > > > > > > > > Even if it's not likely to be used in practice (guest still might not > > > > support it), > > > > it may serve as test-bed for guest drivers. > > > > > > > > > Thank you very much! > > > > > > > > > > Dongli Zhang > > > > > > > > >
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |