[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: [BUG] xen-swiotlb: 32-bit coherent DMA mask rejected on PV dom0 since 6.6 (CONFIG_SWIOTLB_DYNAMIC), aacraid probe fails


  • To: xen-devel@xxxxxxxxxxxxxxxxxxxx
  • From: Anton Markov <akmarkov45@xxxxxxxxx>
  • Date: Wed, 16 Sep 2026 23:35:47 +0300
  • Arc-authentication-results: i=1; mx.google.com; arc=none
  • Arc-message-signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20260327; h=to:subject:message-id:date:from:in-reply-to:references:mime-version :dkim-signature; bh=wNCZPSbPYJyFVLNjgBg2eGFHgK4UGL4TNgHPrNNY8gE=; fh=quJY5mN2l4ZorNvEoO9ngNXalhEvTdq/+W8CvHWhECs=; b=qeOYO8yvyOAdEOXVCw7EeD0F6RJ3+zbC4AGaNmwIX47Kte0RensKELhAKm6wuV7pHd p3KxnUzLzReSP0npTT91/cwVwlzHr17kRKiC1WNmDB5hUo90GoaUlIYDU7+KjXKQt9oC LtB3xNx3NVl3xK9a35D2fNze2b6vvXFdzPAEaAwD9ILbkDxtX3P0PBzyzcfH9EYs4s5i VgMBX6glnUQKw0jFJVQYqvVv0GZ04mcGQOpVomdJFbP9vN4lu4DlRJFhfgOuhABfzoZ1 JnPx7EHjY8n+steY2s0MfDtSw4mlnrHc4iNURMX+zQxopXGDpmLJoakRBGqGMlrQ8qR3 e9RQ==; darn=lists.xenproject.org
  • Arc-seal: i=1; a=rsa-sha256; t=1789590959; cv=none; d=google.com; s=arc-20260327; b=jSpGaMOONZt1Cwa9eRnC2JXc3qSoXDOgdEehjmPY7cGnLoFeIIqJZQodNLL75B1Yp/ jFbv3J7PhdvRsUmkD23pQVJO644jNJcWZDRlOX+Iea4OrlItvFF4XijgSoyJ+FLp2KwQ vfOMob/hi/h9fqTS758+dXDF7U2IoX+yaOyyeSWWTVVZJ1KGwVrQoUKUqI7QKiKRvRn3 LUAw5IxGVnZkSRLBPbW+ExA38/I7hiBTddQ+RlzYTZOzwUJZyK2TF9p/ALKz1oNq+eNe MiL6vhPWXZGruV4U+Or4bjPH8eshhp9FQn0j4lJC3MKs4eNsYFp/69oW6iIYFhM+eo7x sLsg==
  • Authentication-results: eu.smtp.expurgate.cloud; dkim=pass header.s=20251104 header.d=gmail.com header.i="@gmail.com" header.h="Content-Type:To:Subject:Message-ID:Date:From:In-Reply-To:References:MIME-Version"
  • Delivery-date: Wed, 16 Sep 2026 20:36:20 +0000
  • List-id: Xen developer discussion <xen-devel.lists.xenproject.org>

Follow-up with a second reproduction on different hardware, an exact
explanation of the failing check, and a tested patch.

My first analysis was incomplete: the problem is not that the top of
RAM lies above 4 GiB. On a PV dom0 the pseudo-physical address that
default_swiotlb_limit() returns usually has no machine frame behind it
at all, so *every* mask below 64 bits is rejected. That also explains
the July 2024 megaraid_sas report where a 63-bit mask failed:
https://lkml.iu.edu/hypermail/linux/kernel/2407.3/08254.html

Second reproduction
-------------------

  Xen:      4.21.2-pre, xen.efi via systemd-boot, PV dom0
  Hardware: ASUS N550JX, Core i7-4720HQ (Haswell), 16 GiB RAM
  dom0:     Arch Linux kernel 7.2.6-arch2-1 rebuilt with
            CONFIG_SWIOTLB_DYNAMIC=y and nothing else changed
            (the stock Arch config has it disabled and is fine)
  Memory:   host e820 top 0x42f1fffff, dom0 nr_pages 0x3c8783
            (no dom0_mem= on the Xen command line)

With the SWIOTLB_DYNAMIC kernel every driver that asks for less than a
64-bit mask fails to probe:

  i915 0000:00:02.0: [drm] *ERROR* Can't set DMA mask/consistent mask (-5)
  i915 0000:00:02.0: probe with driver i915 failed with error -5
  iwlwifi 0000:04:00.0: No suitable DMA available
  iwlwifi 0000:04:00.0: probe with driver iwlwifi failed with error -5

i915 requests a 40-bit mask on this platform, iwlwifi 36 bits with a
32-bit fallback. ahci, xhci_hcd and mei_me request 64 bits and work.
The same kernel booted natively is fine, and the stock kernel without
SWIOTLB_DYNAMIC is fine as dom0.

What the check evaluates to
---------------------------

I wrote a small out-of-tree module that calls dma_set_mask() /
dma_set_coherent_mask() on an unbound PCI device for a range of widths
and also prints what xen_swiotlb_dma_supported() is comparing against
(pfn of high_memory - 1, its p2m entry, the resulting DMA address).
Output on the SWIOTLB_DYNAMIC dom0 kernel:

  dmamask_test: xen-pv: high_memory-1 paddr=0x000000042f1fffff pfn=0x42f1ff mfn=0xffffffffffffffff (INVALID_P2M_ENTRY: unpopulated) nr_pages=0x3c8783
  dmamask_test: xen-pv: highest populated pfn below it=0x41a8b6 mfn=0x37491a; SWIOTLB_DYNAMIC limit dma addr=0xffffffffffffffff -> smallest passing mask=64 bits
  dmamask_test: 32-bit  dma_set_mask=FAIL(-5)  dma_set_coherent_mask=FAIL(-5)
  dmamask_test: 33-bit  dma_set_mask=FAIL(-5)  dma_set_coherent_mask=FAIL(-5)
  dmamask_test: 34-bit  dma_set_mask=FAIL(-5)  dma_set_coherent_mask=FAIL(-5)
  dmamask_test: 35-bit  dma_set_mask=FAIL(-5)  dma_set_coherent_mask=FAIL(-5)
  dmamask_test: 36-bit  dma_set_mask=FAIL(-5)  dma_set_coherent_mask=FAIL(-5)
  dmamask_test: 40-bit  dma_set_mask=FAIL(-5)  dma_set_coherent_mask=FAIL(-5)
  dmamask_test: 48-bit  dma_set_mask=FAIL(-5)  dma_set_coherent_mask=FAIL(-5)
  dmamask_test: 64-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)

Identical result for three different devices (00:02.0, 01:00.0,
04:00.0), so it is not device specific. On the stock kernel without
SWIOTLB_DYNAMIC all widths pass.

The chain is:

  swiotlb_init_remap(..., SWIOTLB_ANY, xen_swiotlb_fixup)
    -> io_tlb_default_mem.phys_limit = virt_to_phys(high_memory - 1)
  default_swiotlb_limit()           -> phys_limit (with SWIOTLB_DYNAMIC)
  xen_swiotlb_dma_supported()       -> xen_phys_to_dma(dev, limit) <= mask
  xen_phys_to_bus()                 -> pfn_to_bfn(0x42f1ff)
  pfn_to_mfn()                      -> INVALID_P2M_ENTRY (~0UL)
  bfn << PAGE_SHIFT | offset        -> 0xffffffffffffffff

high_memory on a PV dom0 covers the whole host e820 map, while dom0
only owns nr_pages frames of it. The top of the pseudo-physical space
is normally unpopulated (here everything above pfn 0x41a8b6 is a hole;
Xen itself and the reserved regions live above dom0's last frame). The
p2m lookup for such a pfn yields INVALID_P2M_ENTRY, which
xen_phys_to_bus() shifts into an all-ones bus address, and only a
64-bit mask can cover that.

Whether the last pfn happens to be populated depends on the host
memory map and on dom0_mem, which is why the symptom varies between
"32-bit masks fail" (my Supermicro report) and "everything below 64
bits fails" (this laptop, the 2024 megaraid_sas report).

Before v6.6 the comparison used io_tlb_default_mem.end - 1, i.e. the
end of the default pool. That pool has been made machine-contiguous
and placed below 4 GiB by xen_swiotlb_fixup(), so its machine address
is meaningful and small. On this box the pool is mapped at
pseudo-physical 0x405600000-0x409600000 (64 MiB) and its machine
frames are below 4 GiB.

Why the old value is still the right one
----------------------------------------

phys_limit only describes where *additional* pools may be allocated
when the default pool is allowed to grow. Xen always passes a remap
callback to swiotlb_init_remap(), and that clears can_grow, so under
Xen every bounce buffer comes from the default pool forever.
Translating high_memory - 1 through the p2m therefore answers a
question that never matters for Xen, and answers it with a value that
has no relation to any address a device could be handed.

The minimal change I suggested in the first mail
(io_tlb_default_mem.defpool.end - 1 in xen_swiotlb_dma_supported())
no longer compiles as is: io_tlb_default_mem has been static in
kernel/dma/swiotlb.c since the same series. The equivalent fix at the
place where the information lives is to have default_swiotlb_limit()
return the default pool end whenever the pool cannot grow. That is a
no-op for every non-Xen SWIOTLB_DYNAMIC user (they never pass remap,
so can_grow stays true) and restores the pre-6.6 behaviour for Xen.

Test result with the patch below
--------------------------------

Same laptop, same config plus the patch, booted as PV dom0:

  dmamask_test: xen-pv: high_memory-1 paddr=0x000000042f1fffff pfn=0x42f1ff mfn=0xffffffffffffffff (INVALID_P2M_ENTRY: unpopulated) nr_pages=0x3c8783
  dmamask_test: 32-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)
  dmamask_test: 33-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)
  dmamask_test: 34-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)
  dmamask_test: 35-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)
  dmamask_test: 36-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)
  dmamask_test: 40-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)
  dmamask_test: 48-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)
  dmamask_test: 64-bit  dma_set_mask=ok(0)  dma_set_coherent_mask=ok(0)

The p2m situation is unchanged (top pfn still unpopulated), but the
check now uses the default pool end. i915 and iwlwifi probe normally,
display and WiFi work, no DMA related messages in dmesg:

  software IO TLB: mapped [mem 0x0000000405600000-0x0000000409600000] (64MB)
  iwlwifi 0000:04:00.0: Detected Intel(R) Dual Band Wireless AC 7260
  iwlwifi 0000:04:00.0: loaded firmware version 17.bfb58538.0 7260-17.ucode op_mode iwlmvm

I have not yet re-tested the aacraid machine from the first mail, but
it is the same check with the same input, so I expect the same result.
Happy to confirm on that box if useful.

Patch
-----

From: Anton Markov <akmarkov45@xxxxxxxxx>
Subject: [PATCH] swiotlb: report the default pool end as limit when the pool cannot grow

Since the CONFIG_SWIOTLB_DYNAMIC work, default_swiotlb_limit() returns
io_tlb_default_mem.phys_limit, which swiotlb_init_remap() sets to
virt_to_phys(high_memory - 1) when SWIOTLB_ANY is passed.  The only
caller, xen_swiotlb_dma_supported(), compares the machine address of
that value against the mask a driver asks for.

On a Xen PV dom0 high_memory covers the whole host e820 map while dom0
owns only nr_pages of it, so the last pseudo-physical frame is normally
not populated.  pfn_to_mfn() returns INVALID_P2M_ENTRY for it and
xen_phys_to_dma() turns that into 0xffffffffffffffff.  The result is
that only a full 64-bit mask is accepted: i915 (40 bits), iwlwifi
(36 bits), aacraid and megaraid_sas (32/63 bits) all fail to probe.
Before v6.6 the end of the default pool was used and these worked.

Xen always passes a remap callback, which clears can_grow, so the
default pool is the only place bounce buffers can ever come from.
phys_limit only matters for allocating further pools.  Return the end
of the default pool whenever the pool cannot grow.

Signed-off-by: Anton Markov <akmarkov45@xxxxxxxxx>
---
--- a/kernel/dma/swiotlb.c
+++ b/kernel/dma/swiotlb.c
@@ -1680,10 +1680,18 @@
 phys_addr_t default_swiotlb_limit(void)
 {
 #ifdef CONFIG_SWIOTLB_DYNAMIC
- return io_tlb_default_mem.phys_limit;
-#else
- return io_tlb_default_mem.defpool.end - 1;
+ /*
+ * phys_limit bounds where additional pools may be allocated.  If the
+ * default pool can never grow (a remap callback was given at init,
+ * e.g. by Xen), all bounce buffers live in the default pool and its
+ * end is the real limit.  phys_limit would be high_memory - 1 in
+ * that case, which on a Xen PV domain is a pseudo-physical address
+ * that typically has no machine frame behind it at all.
+ */
+ if (io_tlb_default_mem.can_grow)
+ return io_tlb_default_mem.phys_limit;
 #endif
+ return io_tlb_default_mem.defpool.end - 1;
 }

 #ifdef CONFIG_DEBUG_FS

The test module (GPL, ~120 lines) is available on request; it takes a
bdf= parameter, probes the mask widths listed above and restores the
original masks afterwards.

Thanks,
Anton

 


Rackspace

Lists.xenproject.org is hosted with RackSpace, monitoring our
servers 24x7x365 and backed by RackSpace's Fanatical Support®.