[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: [PATCH 2/6] x86/pass-through: no locking around pt_irq_{create,destroy}_bind()



On Mon, Sep 28, 2026 at 02:50:41PM +0200, Jan Beulich wrote:
> On 25.09.2026 16:53, Roger Pau Monné wrote:
> > On Thu, Sep 24, 2026 at 11:57:19AM +0200, Jan Beulich wrote:
> >> On 23.09.2026 12:37, Roger Pau Monné wrote:
> >>> On Tue, Sep 08, 2026 at 03:01:51PM +0200, Jan Beulich wrote:
> >>>> The questionable use of pcidevs_lock() there was discussed more than 
> >>>> once.
> >>>> It really is pointless: The functions synchronize primarily via the per-
> >>>> domain event lock. They also may already be called with the global PCI
> >>>> devices lock not held: See hvm/vmsi.c:vpci_msi_update(),
> >>>> hvm/vmsi.c:vpci_msi_arch_update(), and hvm/vmsi.c:vpci_msi_disable().
> >>>>
> >>>> Signed-off-by: Jan Beulich <jbeulich@xxxxxxxx>
> >>>>
> >>>> --- a/xen/arch/x86/domctl.c
> >>>> +++ b/xen/arch/x86/domctl.c
> >>>> @@ -636,10 +636,7 @@ long arch_do_domctl(
> >>>>              ret = -EPERM;
> >>>>          else if ( is_iommu_enabled(d) )
> >>>>          {
> >>>> -            pcidevs_lock();
> >>>>              ret = pt_irq_create_bind(d, bind);
> >>>> -            pcidevs_unlock();
> >>>
> >>> pt_irq_create_bind() might call into msixtbl_pt_register() which
> >>> requires either the pcidevs_lock() or the per-domain d->pci_lock lock
> >>> to be taken, which I think is not the case in the context here?
> >>
> >> Hmm, indeed. Not having seen the assertion there trigger kind of worries
> >> me a little. Do you agree that the change to vioapic_hwdom_map_gsi() can,
> >> otoh, be left as is?
> > 
> > Hm, I'm borderline on that one - I can't find a path where d->pci_lock
> > will be needed for legacy PCI interrupt binding, yet at the same time
> > I feel it would be better if the locking context is uniform across
> > call sites.  I guess I'm fine with the asymmetric locking context if
> > that's your preference.  Maybe worth a mention in a comment somewhere.
> 
> Maybe it's best if I get v2 out before we settle on this. The need for a
> comment may, with how v2 is done, go away. E.g. the first of the hunks
> now is
> 
> @@ -637,9 +637,13 @@ long arch_do_domctl(
>              ret = -EPERM;
>          else if ( is_iommu_enabled(d) )
>          {
> -            pcidevs_lock();
> +            if ( bind->irq_type == PT_IRQ_TYPE_MSI )
> +                read_lock(&d->pci_lock);
> +
>              ret = pt_irq_create_bind(d, bind);
> -            pcidevs_unlock();
> +
> +            if ( bind->irq_type == PT_IRQ_TYPE_MSI )
> +                read_unlock(&d->pci_lock);
>  
>              if ( ret < 0 )
>                  printk(XENLOG_G_ERR "pt_irq_create_bind failed (%ld) for 
> %pd\n",

Binding an unbinding is an expensive operation.  The legacy PCI
interrupt bindings are done only once when the device is assigned to a
domain, and afterwards all calls to XEN_DOMCTL_bind_pt_irq should be
for MSI interrupts.  I don't think taking the lock unconditionally
would be that bad, the extra penalty for the one-shot PCI legacy
binding is possibly likely fine if we can remove one conditional?

> While it's only a read-lock now, effects on parallelism aren't as bad
> anymore. Yet still I'm rather hesitant to acquire a lock when there's no
> need for doing so. In the case here we'd still impact any write-lock
> paths, i.e. first and foremost vpci_write().

It's unlikely (albeit not impossible) to have both
XEN_DOMCTL_bind_pt_irq hypercalls and vPCI against the same domain.
Either the domain uses vPCI for passthrough or it uses an external
device model IMO.

> There's possibly another somewhat related issue: vpci_read() only uses
> read_lock(), yet reads can in principle have side effects. Are we (once
> again) building upon Dom0 knowing what it's doing, and this - like many
> other aspect - being in need of auditing before DomU supported can be
> declared complete?

The point of taking d->pci_lock in write mode is to prevent accesses
to any pdevs assigned to the domain, so that the position of BARs
across any devices assigned to a domain cannot change as we have to
check for overlaps.  However for other accesses we so far have no need
to cross check like this against all devices assigned to a domain, and
hence just taking the pdev->vpci->lock (so a per-device lock) is
possibly enough, as per-device accesses are still serialized?

Thanks, Roger.



 


Rackspace

Lists.xenproject.org is hosted with RackSpace, monitoring our
servers 24x7x365 and backed by RackSpace's Fanatical Support®.