|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Re: [BUG] xen-swiotlb: 32-bit coherent DMA mask rejected on PV dom0 since 6.6 (CONFIG_SWIOTLB_DYNAMIC), aacraid probe fails
Follow-up with a second reproduction on different hardware, an exact explanation of the failing check, and a tested patch. My first analysis was incomplete: the problem is not that the top of RAM lies above 4 GiB. On a PV dom0 the pseudo-physical address that default_swiotlb_limit() returns usually has no machine frame behind it at all, so *every* mask below 64 bits is rejected. That also explains the July 2024 megaraid_sas report where a 63-bit mask failed: https://lkml.iu.edu/hypermail/linux/kernel/2407.3/08254.html Second reproduction ------------------- Xen: 4.21.2-pre, xen.efi via systemd-boot, PV dom0 Hardware: ASUS N550JX, Core i7-4720HQ (Haswell), 16 GiB RAM dom0: Arch Linux kernel 7.2.6-arch2-1 rebuilt with CONFIG_SWIOTLB_DYNAMIC=y and nothing else changed (the stock Arch config has it disabled and is fine) Memory: host e820 top 0x42f1fffff, dom0 nr_pages 0x3c8783 (no dom0_mem= on the Xen command line) With the SWIOTLB_DYNAMIC kernel every driver that asks for less than a 64-bit mask fails to probe: i915 0000:00:02.0: [drm] *ERROR* Can't set DMA mask/consistent mask (-5) i915 0000:00:02.0: probe with driver i915 failed with error -5 iwlwifi 0000:04:00.0: No suitable DMA available iwlwifi 0000:04:00.0: probe with driver iwlwifi failed with error -5 i915 requests a 40-bit mask on this platform, iwlwifi 36 bits with a 32-bit fallback. ahci, xhci_hcd and mei_me request 64 bits and work. The same kernel booted natively is fine, and the stock kernel without SWIOTLB_DYNAMIC is fine as dom0. What the check evaluates to --------------------------- I wrote a small out-of-tree module that calls dma_set_mask() / dma_set_coherent_mask() on an unbound PCI device for a range of widths and also prints what xen_swiotlb_dma_supported() is comparing against (pfn of high_memory - 1, its p2m entry, the resulting DMA address). Output on the SWIOTLB_DYNAMIC dom0 kernel: dmamask_test: xen-pv: high_memory-1 paddr=0x000000042f1fffff pfn=0x42f1ff mfn=0xffffffffffffffff (INVALID_P2M_ENTRY: unpopulated) nr_pages=0x3c8783 dmamask_test: xen-pv: highest populated pfn below it=0x41a8b6 mfn=0x37491a; SWIOTLB_DYNAMIC limit dma addr=0xffffffffffffffff -> smallest passing mask=64 bits dmamask_test: 32-bit dma_set_mask=FAIL(-5) dma_set_coherent_mask=FAIL(-5) dmamask_test: 33-bit dma_set_mask=FAIL(-5) dma_set_coherent_mask=FAIL(-5) dmamask_test: 34-bit dma_set_mask=FAIL(-5) dma_set_coherent_mask=FAIL(-5) dmamask_test: 35-bit dma_set_mask=FAIL(-5) dma_set_coherent_mask=FAIL(-5) dmamask_test: 36-bit dma_set_mask=FAIL(-5) dma_set_coherent_mask=FAIL(-5) dmamask_test: 40-bit dma_set_mask=FAIL(-5) dma_set_coherent_mask=FAIL(-5) dmamask_test: 48-bit dma_set_mask=FAIL(-5) dma_set_coherent_mask=FAIL(-5) dmamask_test: 64-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) Identical result for three different devices (00:02.0, 01:00.0, 04:00.0), so it is not device specific. On the stock kernel without SWIOTLB_DYNAMIC all widths pass. The chain is: swiotlb_init_remap(..., SWIOTLB_ANY, xen_swiotlb_fixup) -> io_tlb_default_mem.phys_limit = virt_to_phys(high_memory - 1) default_swiotlb_limit() -> phys_limit (with SWIOTLB_DYNAMIC) xen_swiotlb_dma_supported() -> xen_phys_to_dma(dev, limit) <= mask xen_phys_to_bus() -> pfn_to_bfn(0x42f1ff) pfn_to_mfn() -> INVALID_P2M_ENTRY (~0UL) bfn << PAGE_SHIFT | offset -> 0xffffffffffffffff high_memory on a PV dom0 covers the whole host e820 map, while dom0 only owns nr_pages frames of it. The top of the pseudo-physical space is normally unpopulated (here everything above pfn 0x41a8b6 is a hole; Xen itself and the reserved regions live above dom0's last frame). The p2m lookup for such a pfn yields INVALID_P2M_ENTRY, which xen_phys_to_bus() shifts into an all-ones bus address, and only a 64-bit mask can cover that. Whether the last pfn happens to be populated depends on the host memory map and on dom0_mem, which is why the symptom varies between "32-bit masks fail" (my Supermicro report) and "everything below 64 bits fails" (this laptop, the 2024 megaraid_sas report). Before v6.6 the comparison used io_tlb_default_mem.end - 1, i.e. the end of the default pool. That pool has been made machine-contiguous and placed below 4 GiB by xen_swiotlb_fixup(), so its machine address is meaningful and small. On this box the pool is mapped at pseudo-physical 0x405600000-0x409600000 (64 MiB) and its machine frames are below 4 GiB. Why the old value is still the right one ---------------------------------------- phys_limit only describes where *additional* pools may be allocated when the default pool is allowed to grow. Xen always passes a remap callback to swiotlb_init_remap(), and that clears can_grow, so under Xen every bounce buffer comes from the default pool forever. Translating high_memory - 1 through the p2m therefore answers a question that never matters for Xen, and answers it with a value that has no relation to any address a device could be handed. The minimal change I suggested in the first mail (io_tlb_default_mem.defpool.end - 1 in xen_swiotlb_dma_supported()) no longer compiles as is: io_tlb_default_mem has been static in kernel/dma/swiotlb.c since the same series. The equivalent fix at the place where the information lives is to have default_swiotlb_limit() return the default pool end whenever the pool cannot grow. That is a no-op for every non-Xen SWIOTLB_DYNAMIC user (they never pass remap, so can_grow stays true) and restores the pre-6.6 behaviour for Xen. Test result with the patch below -------------------------------- Same laptop, same config plus the patch, booted as PV dom0: dmamask_test: xen-pv: high_memory-1 paddr=0x000000042f1fffff pfn=0x42f1ff mfn=0xffffffffffffffff (INVALID_P2M_ENTRY: unpopulated) nr_pages=0x3c8783 dmamask_test: 32-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) dmamask_test: 33-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) dmamask_test: 34-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) dmamask_test: 35-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) dmamask_test: 36-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) dmamask_test: 40-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) dmamask_test: 48-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) dmamask_test: 64-bit dma_set_mask=ok(0) dma_set_coherent_mask=ok(0) The p2m situation is unchanged (top pfn still unpopulated), but the check now uses the default pool end. i915 and iwlwifi probe normally, display and WiFi work, no DMA related messages in dmesg: software IO TLB: mapped [mem 0x0000000405600000-0x0000000409600000] (64MB) iwlwifi 0000:04:00.0: Detected Intel(R) Dual Band Wireless AC 7260 iwlwifi 0000:04:00.0: loaded firmware version 17.bfb58538.0 7260-17.ucode op_mode iwlmvm I have not yet re-tested the aacraid machine from the first mail, but it is the same check with the same input, so I expect the same result. Happy to confirm on that box if useful. Patch ----- From: Anton Markov <akmarkov45@xxxxxxxxx> Subject: [PATCH] swiotlb: report the default pool end as limit when the pool cannot grow Since the CONFIG_SWIOTLB_DYNAMIC work, default_swiotlb_limit() returns io_tlb_default_mem.phys_limit, which swiotlb_init_remap() sets to virt_to_phys(high_memory - 1) when SWIOTLB_ANY is passed. The only caller, xen_swiotlb_dma_supported(), compares the machine address of that value against the mask a driver asks for. On a Xen PV dom0 high_memory covers the whole host e820 map while dom0 owns only nr_pages of it, so the last pseudo-physical frame is normally not populated. pfn_to_mfn() returns INVALID_P2M_ENTRY for it and xen_phys_to_dma() turns that into 0xffffffffffffffff. The result is that only a full 64-bit mask is accepted: i915 (40 bits), iwlwifi (36 bits), aacraid and megaraid_sas (32/63 bits) all fail to probe. Before v6.6 the end of the default pool was used and these worked. Xen always passes a remap callback, which clears can_grow, so the default pool is the only place bounce buffers can ever come from. phys_limit only matters for allocating further pools. Return the end of the default pool whenever the pool cannot grow. Signed-off-by: Anton Markov <akmarkov45@xxxxxxxxx> --- --- a/kernel/dma/swiotlb.c +++ b/kernel/dma/swiotlb.c @@ -1680,10 +1680,18 @@ phys_addr_t default_swiotlb_limit(void) { #ifdef CONFIG_SWIOTLB_DYNAMIC - return io_tlb_default_mem.phys_limit; -#else - return io_tlb_default_mem.defpool.end - 1; + /* + * phys_limit bounds where additional pools may be allocated. If the + * default pool can never grow (a remap callback was given at init, + * e.g. by Xen), all bounce buffers live in the default pool and its + * end is the real limit. phys_limit would be high_memory - 1 in + * that case, which on a Xen PV domain is a pseudo-physical address + * that typically has no machine frame behind it at all. + */ + if (io_tlb_default_mem.can_grow) + return io_tlb_default_mem.phys_limit; #endif + return io_tlb_default_mem.defpool.end - 1; } #ifdef CONFIG_DEBUG_FS The test module (GPL, ~120 lines) is available on request; it takes a bdf= parameter, probes the mask widths listed above and restores the original masks afterwards. Thanks, Anton
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |