|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Re: [PATCH v3 13/18] x86/mm: build and maintain a sparse view of the directmap
On Wed, Oct 07, 2026 at 11:40:46AM +0100, George Dunlap wrote:
> The directmap maps all of RAM, and every root page-table carries it, so
> while any guest context runs, all guest memory is reachable from Xen's
> address space. For ASI, guest contexts are to run on a view of the
> directmap which maps only what Xen must reach through it at all times,
> leaving guest memory to transient mappings (map_domain_page()).
>
> Rather than remove the directmap for everyone, keep the full directmap
> as it is, for the idle domain and for guest contexts not using ASI, and
> build a second, sparse view: page-tables for the same virtual addresses
> (the part of the directmap PV guests can see, where everything Xen
> reaches through the directmap lives), mapping a subset of what the full
> directmap maps, each entry a copy of the full directmap's. Boot, dom0
> construction and the idle domain keep the full directmap, so nothing in
> them needs converting.
>
> The sparse view maps the xenheap, and memory never handed to the heap
> allocator (Xen's image, boot-time allocations, firmware tables). It is
> built once the boot allocator has handed its memory to the heap, from
> the full directmap minus the ranges the heap was given, plus the xenheap
> allocations made meanwhile (NUMA node heap metadata). The page-tables
> then come from the heap, and no other CPU is up. The sparse view has an
> L3 for each of its L4 slots from the start, so its L4 entries never
> change and root page-tables can take copies of them. It is then kept in
> step:
>
> - xenheap allocations are added, and taken out again before the pages
> are freed. If that fails the pages stay mapped, and are leaked with
> a warning rather than reused (the same reasoning as in "xen/page_alloc:
> Add a path for xenheap when there is no direct map" of the directmap
> removal series);
I would probably remove this last sentence, it doesn't add any
meaningful information.
> - memory handed to the heap later (boot modules, .init, memory
> hot-add) is taken out, and is not handed over if that fails;
> - a PV dom0's initrd, assigned to dom0 in place, is taken out;
I assume dom0 kernel memory is freed, and it doesn't need the same
kind of special handling the initrd needs due to it being assigned in
place.
> - page-tables from the boot allocator, which the sparse view maps with
> every other boot-time allocation, are taken out when idle_pg_table's
> hierarchy frees them, and are not freed if that fails, so that the
> sparse view never keeps a page the heap may hand out;
> - changes to the full directmap through map_pages_to_xen() and
> modify_xen_mappings() are mirrored. After boot, map_pages_to_xen()
> maps into the directmap only memory Xen keeps (stack guard and shadow
> stack pages, firmware tables) or hot-added memory about to be handed
> to the heap, which the second rule then takes out again.
>
> Common code gains four hooks for this, arch_heap_pages_init(),
> arch_xenheap_needs_scrub() and arch_xenheap_pages_{alloc,free}(), no-ops
> unless CONFIG_SPARSE_DIRECTMAP is selected. A failed
> arch_xenheap_pages_alloc() may have mapped part of the pages:
> alloc_xenheap_pages() then takes them out through
> arch_xenheap_pages_free(), as a free would, before freeing them, and
> leaks them if that fails too. xen_mapping_lookup(), built only with
> CONFIG_SPARSE_DIRECTMAP, reads an entry of either hierarchy under the
> L3 table lock.
>
> alloc_xenheap_pages() accepts pages still awaiting scrubbing
> (MEMF_no_scrub). That is harmless while the directmap maps all memory
> anyway, but the sparse view would make such a page's contents, a dying
> domain's data for instance, reachable from every context on it until
> overwritten. With a sparse view configured, arch_xenheap_needs_scrub()
> makes xenheap allocations scrub those pages first, as domain heap
> allocations do; clean pages need nothing extra.
There's another possible issue here I think: even when requesting
scrubbed pages the allocator can decide to scrub them in-place,
mapping them using map_domain_page(). Could this also cause concern,
as Xen is mapping possibly sensitive data in a context that's not
supposed to access it, even if just for the scrubbing.
> arch_mfns_in_directmap() says whether an MFN range is reachable through
> the directmap; it now answers for every context, so false once a sparse
> view is configured. Its one caller, init_node_heap(), then keeps a NUMA
> node's heap metadata in the xenheap instead of carving it out of the
> node's memory through the directmap: carved out of a range handed to the
> heap, it would be left out of the sparse view, which every heap
> allocation needs it in.
>
> The sparse view is built only when some guest type is to use it, which
> nothing can request yet: this patch changes nothing on its own.
Is it worth mentioning that modify_xen_mappings_lite() doesn't need
adjustment because it only deals with virtual addresses in the
XEN_VIRT_{START,END} range, and those are cloned in the sparse
directmap view?
>
> Assisted-by: Claude Code:claude-opus-5-5
> Signed-off-by: George Dunlap <gwd@xxxxxxxxxxxxxx>
> ---
> Changes in v3:
> - New in this version.
> ---
> xen/arch/x86/Kconfig | 1 +
> xen/arch/x86/Makefile | 1 +
> xen/arch/x86/include/asm/mm.h | 17 +-
> xen/arch/x86/include/asm/sparse-directmap.h | 63 +++
> xen/arch/x86/mm.c | 128 +++++-
> xen/arch/x86/pv/dom0_build.c | 4 +
> xen/arch/x86/setup.c | 4 +
> xen/arch/x86/sparse-directmap.c | 424 ++++++++++++++++++++
> xen/common/Kconfig | 18 +
> xen/common/page_alloc.c | 41 +-
> xen/include/xen/mm.h | 35 ++
> 11 files changed, 722 insertions(+), 14 deletions(-)
> create mode 100644 xen/arch/x86/include/asm/sparse-directmap.h
> create mode 100644 xen/arch/x86/sparse-directmap.c
I have to admit reviewing this 1K line diff is challenging. I've
tried my best, but it's likely I've missed stuff.
> diff --git a/xen/arch/x86/Kconfig b/xen/arch/x86/Kconfig
> index e5535ac484..cd72102ccc 100644
> --- a/xen/arch/x86/Kconfig
> +++ b/xen/arch/x86/Kconfig
> @@ -31,6 +31,7 @@ config X86
> select HAS_SCHED_GRANULARITY
> select HAS_SHARED_INFO
> imply HAS_SOFT_RESET
> + select HAS_SPARSE_DIRECTMAP
> select HAS_UBSAN
> select HAS_VMAP
> select HAS_VPCI if HVM
> diff --git a/xen/arch/x86/Makefile b/xen/arch/x86/Makefile
> index 8a205c3d27..22b3c8bf6b 100644
> --- a/xen/arch/x86/Makefile
> +++ b/xen/arch/x86/Makefile
> @@ -63,6 +63,7 @@ obj-y += shutdown.o
> obj-y += smp.o
> obj-y += smpboot.o
> obj-y += spec_ctrl.o
> +obj-$(CONFIG_SPARSE_DIRECTMAP) += sparse-directmap.o
> obj-y += srat.o
> obj-y += string.o
> obj-$(CONFIG_SYSCTL) += sysctl.o
> diff --git a/xen/arch/x86/include/asm/mm.h b/xen/arch/x86/include/asm/mm.h
> index 87131dd4bf..0d1e319c09 100644
> --- a/xen/arch/x86/include/asm/mm.h
> +++ b/xen/arch/x86/include/asm/mm.h
> @@ -7,6 +7,7 @@
> #include <xen/rwlock.h>
> #include <asm/io.h>
> #include <asm/page.h>
> +#include <asm/sparse-directmap.h>
> #include <asm/uaccess.h>
>
> /*
> @@ -595,6 +596,15 @@ mfn_t alloc_xen_pagetable(void);
> void free_xen_pagetable(mfn_t mfn);
> void *alloc_mapped_pagetable(mfn_t *pmfn);
>
> +int map_pages_in(l4_pgentry_t *root, unsigned long virt, mfn_t mfn,
> + unsigned long nr_mfns, pte_attr_t flags);
> +int modify_mappings_in(l4_pgentry_t *root, unsigned long s, unsigned long e,
> + pte_attr_t nf);
> +#ifdef CONFIG_SPARSE_DIRECTMAP
> +mfn_t xen_mapping_lookup(l4_pgentry_t *root, unsigned long va,
> + pte_attr_t *flags, unsigned long *nr);
> +#endif
> +
> int __sync_local_execstate(void);
>
> /* Arch-specific portion of memory_op hypercall. */
> @@ -655,12 +665,17 @@ void write_32bit_pse_identmap(uint32_t *l2);
>
> /*
> * x86 maps part of physical memory via the directmap region.
> - * Return whether the range of MFN falls in the directmap region.
> + * Return whether the range of MFN falls in the directmap region, as seen
> from
> + * every context: with a sparse directmap view configured (see
> + * asm/sparse-directmap.h), no range is, except through a xenheap allocation.
> */
> static inline bool arch_mfns_in_directmap(unsigned long mfn, unsigned long
> nr)
> {
> unsigned long eva = min(DIRECTMAP_VIRT_END, HYPERVISOR_VIRT_END);
>
> + if ( sparse_dmap_configured() )
> + return false;
> +
> return (mfn + nr) <= (virt_to_mfn(eva - 1) + 1);
> }
>
> diff --git a/xen/arch/x86/include/asm/sparse-directmap.h
> b/xen/arch/x86/include/asm/sparse-directmap.h
> new file mode 100644
> index 0000000000..48c1f3c951
> --- /dev/null
> +++ b/xen/arch/x86/include/asm/sparse-directmap.h
> @@ -0,0 +1,63 @@
> +/* SPDX-License-Identifier: GPL-2.0-only */
> +#ifndef X86_SPARSE_DIRECTMAP_H
> +#define X86_SPARSE_DIRECTMAP_H
> +
> +#include <xen/mm-frame.h>
> +#include <xen/mm-types.h>
> +#include <asm/page.h>
> +
> +#ifdef CONFIG_SPARSE_DIRECTMAP
> +
> +/*
> + * The part of the directmap the sparse view covers: the part PV guests can
> + * see, which is where everything Xen reaches through the directmap lives.
> + * SPARSE_DMAP_END_MFN is the first MFN past it.
> + */
> +#define SPARSE_DMAP_START DIRECTMAP_VIRT_START
> +#define SPARSE_DMAP_END HYPERVISOR_VIRT_END
> +#define SPARSE_DMAP_END_MFN (virt_to_mfn(SPARSE_DMAP_END - 1) + 1)
> +
> +extern bool opt_sparse_dmap_pv, opt_sparse_dmap_hvm;
> +
> +/* Root of the sparse view, once built (NULL before, or if not configured).
> */
> +extern l4_pgentry_t *sparse_dmap_root;
> +
> +static inline bool sparse_dmap_configured(void)
> +{
> + return opt_sparse_dmap_pv || opt_sparse_dmap_hvm;
> +}
> +
> +void sparse_dmap_init(void);
> +int sparse_dmap_mirror_map(unsigned long virt, mfn_t mfn, unsigned long nr,
> + pte_attr_t flags);
> +int sparse_dmap_mirror_modify(unsigned long s, unsigned long e,
> + pte_attr_t nf);
> +int sparse_dmap_remove(mfn_t mfn, unsigned long nr);
> +int sparse_dmap_drop_pagetable(mfn_t mfn);
> +
> +#else /* !CONFIG_SPARSE_DIRECTMAP */
> +
> +static inline bool sparse_dmap_configured(void) { return false; }
> +static inline void sparse_dmap_init(void) {}
> +static inline int sparse_dmap_mirror_map(unsigned long virt, mfn_t mfn,
> + unsigned long nr, pte_attr_t flags)
> +{
> + return 0;
> +}
> +static inline int sparse_dmap_mirror_modify(unsigned long s, unsigned long e,
> + pte_attr_t nf)
> +{
> + return 0;
> +}
> +static inline int sparse_dmap_remove(mfn_t mfn, unsigned long nr)
> +{
> + return 0;
> +}
> +static inline int sparse_dmap_drop_pagetable(mfn_t mfn)
> +{
> + return 0;
> +}
> +
> +#endif /* CONFIG_SPARSE_DIRECTMAP */
> +
> +#endif /* X86_SPARSE_DIRECTMAP_H */
> diff --git a/xen/arch/x86/mm.c b/xen/arch/x86/mm.c
> index 94b96c3d9b..4a1bd9d146 100644
> --- a/xen/arch/x86/mm.c
> +++ b/xen/arch/x86/mm.c
> @@ -127,6 +127,7 @@
> #include <asm/setup.h>
> #include <asm/shadow.h>
> #include <asm/shared.h>
> +#include <asm/sparse-directmap.h>
> #include <asm/trampoline.h>
> #include <asm/traps.h>
> #include <asm/x86_emulate.h>
> @@ -5394,11 +5395,104 @@ mfn_t xen_map_to_mfn(unsigned long va)
> return ret;
> }
>
> +#ifdef CONFIG_SPARSE_DIRECTMAP
> +/*
> + * Look @va up in the hierarchy rooted at @root, without allocating anything.
> + * Returns the MFN mapped at @va, or INVALID_MFN, with the mapping's flags in
> + * L1 form in *flags, and in *nr the number of pages from @va to the end of
> + * the entry that maps @va (or that finds it unmapped).
> + */
> +mfn_t xen_mapping_lookup(l4_pgentry_t *root, unsigned long va,
> + pte_attr_t *flags, unsigned long *nr)
> +{
> + bool locking = system_state > SYS_STATE_boot;
> + l4_pgentry_t l4e = root[l4_table_offset(va)];
> + const l3_pgentry_t *pl3e = NULL;
> + const l2_pgentry_t *pl2e = NULL;
> + const l1_pgentry_t *pl1e = NULL;
> + struct page_info *l3page;
> + unsigned int shift = L4_PAGETABLE_SHIFT;
> + mfn_t ret = INVALID_MFN;
> +
> + L3T_INIT(l3page);
> + *flags = 0;
> +
> + if ( !(l4e_get_flags(l4e) & _PAGE_PRESENT) )
> + goto out;
> +
> + pl3e = map_l3t_from_l4e(l4e) + l3_table_offset(va);
> + l3page = l4e_get_page(l4e);
> + L3T_LOCK(l3page);
> +
> + shift = L3_PAGETABLE_SHIFT;
> + if ( !(l3e_get_flags(*pl3e) & _PAGE_PRESENT) )
> + goto out;
> + if ( l3e_get_flags(*pl3e) & _PAGE_PSE )
> + {
> + /* The superpage PAT bit sits among the address bits. */
> + *flags = lNf_to_l1f(l3e_get_flags(*pl3e)) |
> + (l3e_get_intpte(*pl3e) & _PAGE_PSE_PAT ? _PAGE_PAT : 0);
> + ret = _mfn((l3e_get_pfn(*pl3e) & ~((1UL << (2 * PAGETABLE_ORDER)) -
> 1)) +
> + (PFN_DOWN(va) & ((1UL << (2 * PAGETABLE_ORDER)) - 1)));
> + goto out;
> + }
> +
> + pl2e = map_l2t_from_l3e(*pl3e) + l2_table_offset(va);
> + shift = L2_PAGETABLE_SHIFT;
> + if ( !(l2e_get_flags(*pl2e) & _PAGE_PRESENT) )
> + goto out;
> + if ( l2e_get_flags(*pl2e) & _PAGE_PSE )
> + {
> + *flags = lNf_to_l1f(l2e_get_flags(*pl2e)) |
> + (l2e_get_intpte(*pl2e) & _PAGE_PSE_PAT ? _PAGE_PAT : 0);
> + ret = _mfn((l2e_get_pfn(*pl2e) & ~((1UL << PAGETABLE_ORDER) - 1)) +
> + l1_table_offset(va));
> + goto out;
> + }
> +
> + pl1e = map_l1t_from_l2e(*pl2e) + l1_table_offset(va);
> + shift = PAGE_SHIFT;
> + if ( l1e_get_flags(*pl1e) & _PAGE_PRESENT )
> + {
> + *flags = l1e_get_flags(*pl1e);
> + ret = l1e_get_mfn(*pl1e);
> + }
> +
> + out:
> + L3T_UNLOCK(l3page);
> + unmap_domain_page(pl1e);
> + unmap_domain_page(pl2e);
> + unmap_domain_page(pl3e);
> + *nr = ((((va >> shift) + 1) << shift) - va) >> PAGE_SHIFT;
> + return ret;
> +}
> +#endif /* CONFIG_SPARSE_DIRECTMAP */
> +
> +/*
> + * Free a page-table page that the hierarchy rooted at @root has stopped
> + * using. idle_pg_table's hierarchy has pages from the boot allocator,
> + * which the sparse directmap view maps along with every other boot-time
> + * allocation: such a page leaves the sparse view before it reaches the
> + * heap, and is kept if it cannot. That takes the sparse view's L3 lock
> + * with an L3 lock of idle_pg_table's hierarchy held, the only order in
> + * which the two nest. The sparse view's own page-tables come from the
> + * heap, which it does not map, and are freed directly, as the caller may
> + * hold the sparse view's L3 lock; so are tables freed straight after being
> + * allocated, having lost a race to be installed.
> + */
> +static void free_pagetable_in(const l4_pgentry_t *root, mfn_t mfn)
> +{
> + if ( root == idle_pg_table && sparse_dmap_drop_pagetable(mfn) )
> + return;
> +
> + free_xen_pagetable(mfn);
> +}
> +
> /*
> * map_pages_to_xen() in the hierarchy rooted at @root: idle_pg_table's, or
> * another hierarchy sharing its virtual addresses.
> */
> -static int map_pages_in(
> +int map_pages_in(
> l4_pgentry_t *root,
> unsigned long virt,
> mfn_t mfn,
> @@ -5503,10 +5597,10 @@ static int map_pages_in(
> ol2e = l2t[i];
> if ( (l2e_get_flags(ol2e) & _PAGE_PRESENT) &&
> !(l2e_get_flags(ol2e) & _PAGE_PSE) )
> - free_xen_pagetable(l2e_get_mfn(ol2e));
> + free_pagetable_in(root, l2e_get_mfn(ol2e));
> }
> unmap_domain_page(l2t);
> - free_xen_pagetable(l3e_get_mfn(ol3e));
> + free_pagetable_in(root, l3e_get_mfn(ol3e));
> }
> }
>
> @@ -5604,7 +5698,7 @@ static int map_pages_in(
> flush_flags(l1e_get_flags(l1t[i]));
> flush_area(virt, flush_flags);
> unmap_domain_page(l1t);
> - free_xen_pagetable(l2e_get_mfn(ol2e));
> + free_pagetable_in(root, l2e_get_mfn(ol2e));
> }
> }
>
> @@ -5736,7 +5830,7 @@ static int map_pages_in(
> flush_area(virt - PAGE_SIZE,
> FLUSH_TLB_GLOBAL |
> FLUSH_ORDER(PAGETABLE_ORDER));
> - free_xen_pagetable(l2e_get_mfn(ol2e));
> + free_pagetable_in(root, l2e_get_mfn(ol2e));
> }
> else if ( locking )
> spin_unlock(&map_pgdir_lock);
> @@ -5783,7 +5877,7 @@ static int map_pages_in(
> flush_area(virt - PAGE_SIZE,
> FLUSH_TLB_GLOBAL |
> FLUSH_ORDER(2*PAGETABLE_ORDER));
> - free_xen_pagetable(l3e_get_mfn(ol3e));
> + free_pagetable_in(root, l3e_get_mfn(ol3e));
> }
> else if ( locking )
> spin_unlock(&map_pgdir_lock);
> @@ -5810,7 +5904,12 @@ int map_pages_to_xen(
> unsigned long nr_mfns,
> pte_attr_t flags)
> {
> - return map_pages_in(idle_pg_table, virt, mfn, nr_mfns, flags);
> + int rc = map_pages_in(idle_pg_table, virt, mfn, nr_mfns, flags);
> +
> + if ( !rc )
> + rc = sparse_dmap_mirror_map(virt, mfn, nr_mfns, flags);
> +
> + return rc;
> }
>
> int __init populate_pt_range(unsigned long virt, unsigned long nr_mfns)
> @@ -5833,8 +5932,8 @@ int __init populate_pt_range(unsigned long virt,
> unsigned long nr_mfns)
> * modify_mappings_in() does so in the hierarchy rooted at @root, as
> * map_pages_in() does.
> */
> -static int modify_mappings_in(l4_pgentry_t *root, unsigned long s,
> - unsigned long e, pte_attr_t nf)
> +int modify_mappings_in(l4_pgentry_t *root, unsigned long s,
> + unsigned long e, pte_attr_t nf)
> {
> bool locking = system_state > SYS_STATE_boot;
> l3_pgentry_t *pl3e = NULL;
> @@ -6044,7 +6143,7 @@ static int modify_mappings_in(l4_pgentry_t *root,
> unsigned long s,
> if ( locking )
> spin_unlock(&map_pgdir_lock);
> flush_area(NULL, FLUSH_TLB_GLOBAL); /* flush before free */
> - free_xen_pagetable(l1mfn);
> + free_pagetable_in(root, l1mfn);
> }
> else if ( locking )
> spin_unlock(&map_pgdir_lock);
> @@ -6088,7 +6187,7 @@ static int modify_mappings_in(l4_pgentry_t *root,
> unsigned long s,
> if ( locking )
> spin_unlock(&map_pgdir_lock);
> flush_area(NULL, FLUSH_TLB_GLOBAL); /* flush before free */
> - free_xen_pagetable(l2mfn);
> + free_pagetable_in(root, l2mfn);
> }
> else if ( locking )
> spin_unlock(&map_pgdir_lock);
> @@ -6114,7 +6213,12 @@ static int modify_mappings_in(l4_pgentry_t *root,
> unsigned long s,
>
> int modify_xen_mappings(unsigned long s, unsigned long e, pte_attr_t nf)
> {
> - return modify_mappings_in(idle_pg_table, s, e, nf);
> + int rc = modify_mappings_in(idle_pg_table, s, e, nf);
> +
> + if ( !rc )
> + rc = sparse_dmap_mirror_modify(s, e, nf);
> +
> + return rc;
> }
>
> int destroy_xen_mappings(unsigned long s, unsigned long e)
> diff --git a/xen/arch/x86/pv/dom0_build.c b/xen/arch/x86/pv/dom0_build.c
> index bf980eba8c..f0a1197e62 100644
> --- a/xen/arch/x86/pv/dom0_build.c
> +++ b/xen/arch/x86/pv/dom0_build.c
> @@ -22,6 +22,7 @@
> #include <asm/pv/mm.h>
> #include <asm/setup.h>
> #include <asm/shadow.h>
> +#include <asm/sparse-directmap.h>
>
> /* Allow ring-3 access in long mode as guest cannot use ring 1 ... */
> #define BASE_PROT (_PAGE_PRESENT|_PAGE_RW|_PAGE_ACCESSED|_PAGE_USER)
> @@ -651,6 +652,9 @@ static int __init dom0_construct(const struct boot_domain
> *bd)
> while ( count-- )
> if ( assign_pages(mfn_to_page(_mfn(mfn++)), 1, d, 0) )
> BUG();
> + /* The initrd is dom0's now: out of the sparse directmap. */
> + if ( sparse_dmap_remove(_mfn(initrd_mfn), PFN_UP(initrd_len)) )
> + panic("Cannot unmap dom0's initrd\n");
> /*
> * We have mapped the initrd directly into dom0, and assigned the
> * pages. Tell the boot_module handling that we've freed it, so
> the
> diff --git a/xen/arch/x86/setup.c b/xen/arch/x86/setup.c
> index a636e3205c..e52388ad28 100644
> --- a/xen/arch/x86/setup.c
> +++ b/xen/arch/x86/setup.c
> @@ -54,6 +54,7 @@
> #include <asm/spec_ctrl.h>
> #include <asm/stubs.h>
> #include <asm/tboot.h>
> +#include <asm/sparse-directmap.h>
> #include <asm/trampoline.h>
> #include <asm/traps.h>
>
> @@ -1938,6 +1939,9 @@ void asmlinkage __init noreturn __start_xen(void)
>
> system_state = SYS_STATE_boot;
>
> + /* The heap now has all memory: page-tables come from it from here on. */
> + sparse_dmap_init();
> +
> bsp_stack = cpu_alloc_stack(0); /* Needs to know IDT vs FRED */
> if ( !bsp_stack )
> panic("No memory for BSP stack\n");
> diff --git a/xen/arch/x86/sparse-directmap.c b/xen/arch/x86/sparse-directmap.c
> new file mode 100644
> index 0000000000..1eb092c3ea
> --- /dev/null
> +++ b/xen/arch/x86/sparse-directmap.c
> @@ -0,0 +1,424 @@
> +/* SPDX-License-Identifier: GPL-2.0-only */
> +/*
> + * A sparse view of the directmap.
> + *
> + * Xen maps all of RAM at DIRECTMAP_VIRT_START, and every root page-table
> + * carries those mappings. The sparse view is a second set of page-tables
> + * for the same virtual addresses (the part of the directmap PV guests can
> + * see, which is where everything Xen reaches through the directmap lives),
> + * mapping only what Xen needs to reach through the directmap at all times.
> + *
> + * Memory in the full directmap can be classified into three groups:
> + *
> + * - Memory never handed to the heap. This includes Xen's image,
> + * boot-time allocations, and firmware tables. Mapped in the sparse
> + * view.
> + * - Memory handed to the heap, and not allocated to the xenheap. Not
> + * mapped in the sparse view: guest memory will come from here.
> + * - Memory handed to the heap, then allocated to the xenheap. Mapped
> + * in the sparse view.
> + *
> + * Guest contexts selected on the command line run on root page-tables
> + * carrying the sparse view's L3s in the directmap slots.
> + *
> + * Conceptually, the sparse view starts as a copy of the full directmap
> + * (idle_pg_table's hierarchy), and then:
> + * - Removes memory when it is handed to the heap.
> + * - Adds pages back when they are allocated as xenheap pages, and removes
> + * them again when they are freed.
> + * - Mirrors changes to the full directmap (map_pages_to_xen(),
> + * modify_xen_mappings()).
> + * - Removes memory handed to a domain without passing through the heap
> + * (a PV dom0's initrd, used in place), and page-tables the boot
> + * allocator provided when idle_pg_table's hierarchy frees them.
> + *
> + * Every virtual address mapping in the sparse view should be either
> + * identical to the corresponding full view mapping, or empty.
> + *
> + * The sparse view is built once the boot allocator has handed its
> + * memory to the heap (sparse_dmap_init()). Until then, the hooks record
> + * the ranges handed to the heap and the xenheap allocations made from
> + * them, which sparse_dmap_init() leaves out and puts back.
> + */
> +
> +#include <xen/domain_page.h>
> +#include <xen/init.h>
> +#include <xen/lib.h>
> +#include <xen/mm.h>
> +#include <xen/numa.h>
> +#include <xen/sort.h>
> +
> +#include <asm/page.h>
> +#include <asm/sparse-directmap.h>
> +
> +bool __ro_after_init opt_sparse_dmap_pv;
> +bool __ro_after_init opt_sparse_dmap_hvm;
Do we need an extra opt_sparse_dmap_dom0? So that dom0 can opt-out
independently of everything else?
> +
> +/* Root of the sparse view: only ever walked and copied from, never loaded.
> */
> +static l4_pgentry_t __aligned(PAGE_SIZE) sparse_l4[L4_PAGETABLE_ENTRIES];
> +l4_pgentry_t *__ro_after_init sparse_dmap_root;
> +
> +/*
> + * Before the sparse view is built: the MFN ranges handed to the heap, which
> + * it leaves out, and the xenheap allocations made meanwhile (heap metadata
> of
> + * NUMA nodes), which it puts back.
Explicitly mentioning NUMA node metadata is likely to go out of date,
as changing that logic will most likely forget to update the comment
here.
> + */
> +struct mfn_range {
> + unsigned long s, e;
> +};
> +static struct mfn_range early_heap[256];
> +static unsigned int nr_early_heap;
> +static struct mfn_range early_xenheap[2 * MAX_NUMNODES];
> +static unsigned int nr_early_xenheap;
> +static bool early_overflow;
> +
> +/*
> + * Clip [*mfn, *mfn + *nr) to what the sparse view covers; false if nothing
> is
> + * left.
> + */
> +static bool clip_to_sparse_view(unsigned long *mfn, unsigned long *nr)
> +{
> + unsigned long end = SPARSE_DMAP_END_MFN;
> +
> + if ( *mfn >= end )
> + return false;
> + if ( *nr > end - *mfn )
> + *nr = end - *mfn;
> +
> + return *nr;
> +}
> +
> +static void early_range_add(struct mfn_range *r, unsigned int *nr,
> + unsigned int max, unsigned long s, unsigned long
> e)
> +{
> + if ( *nr && r[*nr - 1].e == s )
> + r[*nr - 1].e = e;
> + else if ( *nr < max )
> + r[(*nr)++] = (struct mfn_range){ s, e };
> + else
> + early_overflow = true;
> +}
> +
> +static void early_range_del(struct mfn_range *r, unsigned int *nr,
> + unsigned long s, unsigned long e)
> +{
> + unsigned int i;
> +
> + for ( i = 0; i < *nr; i++ )
> + if ( r[i].s == s && r[i].e == e )
> + {
> + r[i] = r[--*nr];
> + return;
> + }
> +}
> +
> +/*
> + * Copy the full directmap's mappings of MFNs [mfn, mfn + nr) into the sparse
> + * view.
> + */
> +static int add_to_sparse_view(l4_pgentry_t *root, unsigned long mfn,
> + unsigned long nr)
> +{
> + unsigned long va = (unsigned long)mfn_to_virt(mfn);
> + unsigned long end = va + (nr << PAGE_SHIFT);
> +
> + while ( va < end )
> + {
> + pte_attr_t flags;
> + unsigned long n;
> + mfn_t m = xen_mapping_lookup(idle_pg_table, va, &flags, &n);
> + int rc;
Since you are fetching the mfn from the idle (non-sparse) page-tables,
don't you want to check it matches the mfn passed as the function
parameter?
> + n = min(n, (end - va) >> PAGE_SHIFT);
> + if ( !mfn_eq(m, INVALID_MFN) &&
> + (rc = map_pages_in(root, va, m, n, flags)) )
> + return rc;
> + va += n << PAGE_SHIFT;
> + }
> +
> + return 0;
> +}
> +
> +static int remove_from_sparse_view(unsigned long mfn, unsigned long nr)
> +{
> + unsigned long va = (unsigned long)mfn_to_virt(mfn);
> +
> + return modify_mappings_in(sparse_dmap_root, va, va + (nr << PAGE_SHIFT),
> + _PAGE_NONE);
> +}
> +
> +/* Memory is about to be handed to the heap allocator. */
> +bool arch_heap_pages_init(mfn_t mfn, unsigned long nr)
> +{
> + unsigned long s = mfn_x(mfn);
> +
> + if ( !sparse_dmap_configured() || !clip_to_sparse_view(&s, &nr) )
> + return true;
> +
> + if ( !sparse_dmap_root )
> + {
> + early_range_add(early_heap, &nr_early_heap, ARRAY_SIZE(early_heap),
> + s, s + nr);
> + return true;
> + }
> +
> + if ( remove_from_sparse_view(s, nr) )
> + {
> + printk(XENLOG_ERR
> + "sparse directmap: cannot unmap MFNs %#lx-%#lx, not
> freeing\n",
> + s, s + nr - 1);
> + return false;
> + }
> +
> + return true;
> +}
> +
> +/*
> + * Xenheap pages enter the sparse view, which guest contexts run on: they
> + * must not bring along what a previous owner left in them, such as a dying
> + * domain's data still awaiting the scrubber. Pages needing scrubbing are
> + * scrubbed as they are allocated then; clean ones need nothing.
> + */
> +bool arch_xenheap_needs_scrub(void)
> +{
> + return sparse_dmap_configured();
> +}
> +
> +/* Pages have been allocated from the heap as xenheap pages. */
> +int arch_xenheap_pages_alloc(mfn_t mfn, unsigned int order)
> +{
> + unsigned long s = mfn_x(mfn), nr = 1UL << order;
> +
> + if ( !sparse_dmap_configured() )
> + return 0;
> +
> + /* Xenheap pages must be reachable through the sparse view. */
> + BUG_ON(s + nr > SPARSE_DMAP_END_MFN);
> +
> + if ( !sparse_dmap_root )
> + {
> + early_range_add(early_xenheap, &nr_early_xenheap,
> + ARRAY_SIZE(early_xenheap), s, s + nr);
> + return 0;
> + }
> +
> + return add_to_sparse_view(sparse_dmap_root, s, nr);
> +}
> +
> +/*
> + * Xenheap pages are about to be freed. A failure means the pages may still
> + * be reachable through the sparse view, and must not be reused.
> + */
> +int arch_xenheap_pages_free(mfn_t mfn, unsigned int order)
> +{
> + unsigned long s = mfn_x(mfn), nr = 1UL << order;
> +
> + if ( !sparse_dmap_configured() )
> + return 0;
> +
> + BUG_ON(s + nr > SPARSE_DMAP_END_MFN);
> +
> + if ( !sparse_dmap_root )
> + {
> + early_range_del(early_xenheap, &nr_early_xenheap, s, s + nr);
> + return 0;
> + }
> +
> + return remove_from_sparse_view(s, nr);
> +}
> +
> +/* Memory leaves Xen's hands without passing through the heap. */
> +int sparse_dmap_remove(mfn_t mfn, unsigned long nr)
> +{
> + return arch_heap_pages_init(mfn, nr) ? 0 : -ENOMEM;
> +}
> +
> +/*
> + * A page-table page of idle_pg_table's hierarchy is about to be freed. One
> + * the boot allocator provided is in the sparse view, as every boot-time
> + * allocation is: take it out before the heap may hand it on. One allocated
> + * since came from the heap and is not in the sparse view, which a lookup
> + * finds without changing, or flushing, anything. Non-zero: the page may
> + * still be mapped, and must not be freed.
> + */
> +int sparse_dmap_drop_pagetable(mfn_t mfn)
> +{
> + unsigned long s = mfn_x(mfn), nr = 1, n;
> + pte_attr_t flags;
> + int rc;
> +
> + if ( !sparse_dmap_root || !clip_to_sparse_view(&s, &nr) ||
> + mfn_eq(xen_mapping_lookup(sparse_dmap_root,
> + (unsigned long)mfn_to_virt(s), &flags,
> &n),
> + INVALID_MFN) )
> + return 0;
> +
> + rc = remove_from_sparse_view(s, 1);
> + if ( rc )
> + printk(XENLOG_WARNING
> + "Leaking page-table page %"PRI_mfn": still in the sparse
> directmap\n",
> + mfn_x(mfn));
> +
> + return rc;
> +}
> +
> +int sparse_dmap_mirror_map(unsigned long virt, mfn_t mfn, unsigned long nr,
> + pte_attr_t flags)
> +{
> + if ( !sparse_dmap_root || virt < SPARSE_DMAP_START ||
> + virt >= SPARSE_DMAP_END )
> + return 0;
> +
> + nr = min(nr, (SPARSE_DMAP_END - virt) >> PAGE_SHIFT);
> +
> + return map_pages_in(sparse_dmap_root, virt, mfn, nr, flags);
> +}
> +
> +int sparse_dmap_mirror_modify(unsigned long s, unsigned long e,
> + pte_attr_t nf)
> +{
> + unsigned long va;
> +
> + if ( !sparse_dmap_root || e <= SPARSE_DMAP_START ||
> + s >= SPARSE_DMAP_END )
> + return 0;
> +
> + s = max(s, SPARSE_DMAP_START);
> + e = min(e, SPARSE_DMAP_END);
> +
> + if ( !(nf & _PAGE_PRESENT) )
> + return modify_mappings_in(sparse_dmap_root, s, e, nf);
> +
> + /*
> + * A permission change applies where the sparse view maps anything: it
> + * never creates mappings, and the sparse view need not map all of [s,
> e).
> + */
> + for ( va = s; va < e; )
> + {
> + pte_attr_t flags;
> + unsigned long n;
> + mfn_t m = xen_mapping_lookup(sparse_dmap_root, va, &flags, &n);
> + int rc;
> +
> + n = min(n, (e - va) >> PAGE_SHIFT);
> + if ( !mfn_eq(m, INVALID_MFN) &&
> + (rc = modify_mappings_in(sparse_dmap_root, va,
> + va + (n << PAGE_SHIFT), nf)) )
> + return rc;
> + va += n << PAGE_SHIFT;
> + }
> +
> + return 0;
> +}
> +
> +static int __init cf_check cmp_mfn_range(const void *a, const void *b)
> +{
> + const struct mfn_range *l = a, *r = b;
> +
> + return l->s < r->s ? -1 : l->s > r->s;
The last arm of this ternary operator is doing an implicit conversion
from boolean to integer, I think that's not fine MISRA-wise?
> +}
> +
> +static void __init cf_check swap_mfn_range(void *a, void *b)
> +{
> + SWAP(*(struct mfn_range *)a, *(struct mfn_range *)b);
> +}
> +
> +/*
> + * Build the sparse view: the full directmap, minus the ranges handed to the
> + * heap so far, plus the xenheap allocations made from them. Called once the
> + * heap has taken over from the boot allocator, so that the page-tables come
> + * from the heap, and before any other CPU is up.
> + */
> +void __init sparse_dmap_init(void)
> +{
> + unsigned long va, mapped = 0;
> + unsigned int i, j;
> +
> + if ( !sparse_dmap_configured() )
> + return;
> +
> + if ( early_overflow )
> + {
> + printk(XENLOG_WARNING
> + "Sparse directmap: too many early heap ranges, disabled\n");
> + opt_sparse_dmap_pv = opt_sparse_dmap_hvm = false;
> + return;
> + }
> +
> + /* An L3 for every slot, so that the L4 entries never change. */
> + for ( va = SPARSE_DMAP_START; va < SPARSE_DMAP_END;
> + va += 1UL << L4_PAGETABLE_SHIFT )
> + {
> + mfn_t l3mfn;
> + void *l3t = alloc_mapped_pagetable(&l3mfn);
> +
> + if ( !l3t )
> + panic("Sparse directmap: out of memory\n");
> + unmap_domain_page(l3t);
> + sparse_l4[l4_table_offset(va)] =
> + l4e_from_mfn(l3mfn, __PAGE_HYPERVISOR);
> + }
> +
> + sort(early_heap, nr_early_heap, sizeof(*early_heap), cmp_mfn_range,
> + swap_mfn_range);
> +
> + /*
> + * The full directmap, leaf by leaf, minus the ranges added to the
> + * heap. Leaves come in address order, which is MFN order in the
> + * directmap, so the cursor into the sorted (and disjoint) heap
> + * ranges only moves forward.
> + */
> + for ( va = SPARSE_DMAP_START, j = 0; va < SPARSE_DMAP_END; )
> + {
> + pte_attr_t flags;
> + unsigned long n, s, e;
> + mfn_t m = xen_mapping_lookup(idle_pg_table, va, &flags, &n);
> +
> + n = min(n, (SPARSE_DMAP_END - va) >> PAGE_SHIFT);
> + if ( mfn_eq(m, INVALID_MFN) )
> + {
> + va += n << PAGE_SHIFT;
> + continue;
> + }
> +
> + for ( s = mfn_x(m), e = s + n; s < e; )
> + {
> + unsigned long upto = e;
> +
> + while ( j < nr_early_heap && early_heap[j].e <= s )
> + j++;
> + if ( j < nr_early_heap && early_heap[j].s <= s )
> + {
> + s = min(e, early_heap[j].e);
> + continue;
> + }
> + if ( j < nr_early_heap )
> + upto = min(e, early_heap[j].s);
> +
> + if ( map_pages_in(sparse_l4, va + ((s - mfn_x(m)) << PAGE_SHIFT),
> + _mfn(s), upto - s, flags) )
Why use map_pages_in() directly here, instead of add_to_sparse_view()?
> + panic("Sparse directmap: out of memory\n");
> + mapped += upto - s;
> + s = upto;
The above logic is complicated to follow, maybe some simple comments
will help.
> + }
> +
> + va += n << PAGE_SHIFT;
> + }
> +
> + /*
> + * Now add back in ranges from the heap which were allocated to
> + * the xenheap.
> + */
> + for ( i = 0; i < nr_early_xenheap; i++ )
> + {
> + unsigned long nr = early_xenheap[i].e - early_xenheap[i].s;
> +
> + if ( add_to_sparse_view(sparse_l4, early_xenheap[i].s, nr) )
> + panic("Sparse directmap: out of memory\n");
> + mapped += nr;
> + }
> +
> + sparse_dmap_root = sparse_l4;
This logic involves a fair amount of calls to map_pages_in(), I think
there will be no flushes in this case, because we are strictly
populating the page-tables, and hence the old PTE will always be
non-present. Worth confirming, otherwise we might need some TLB flush
avoidance, as doing TLB flushes makes no sense here, the page-tables
are not in use.
Thanks, Roger.
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |