|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] [PATCH v3 16/18] x86/domain_page: keep the maphash's slots, not their mappings
unmap_domain_page() keeps a recently unmapped slot in the vCPU's
maphash (MAPHASH_ENTRIES, 8 per vCPU) for a quick remap of the same
MFN, and leaves its PTE in place until the entry is displaced by a later
unmap or reclaimed by a full cache. Until then the page stays mapped,
although nothing refers to the mapping and the mapcache holds no
reference to the page. A page freed meanwhile and handed to another
domain stays mapped in the previous user's mapcache: in the vCPU's own
area for a domain with per-vCPU page-tables, in the domain-wide area of
a PV domain, where all of its vCPUs can reach it.
While the full directmap is loaded, that reaches nothing it does not
map anyway. On the sparse view of the directmap, Xen reaches
domain-heap memory only through mappings made for the purpose, and a
mapping kept in the maphash would leave another domain's page
reachable from a context which has nothing to do with it.
Keep the slot, not the mapping: unmapping the last reference zaps the
PTE and keeps the hash entry, and a later hit maps the page again in the
same slot. The slot has mapped nothing else in between, so that needs
no flush. Reusing the slot for another page still goes through the
garbage map, or through the replacement of a hash entry when the cache
is full, both of which flush first. A translation of the zapped PTE
left in a TLB is covered, for owned pages, by the flush the page
allocator arranges before a freed page is reused wherever a sparse view
is configured ("x86/mm: flush owned pages before reuse with the sparse
directmap view").
The cost is one PTE write per unmap into the maphash and per remap from
it. A second unmap of an entry in the maphash now finds its PTE not
present and is caught by the check unmap_domain_page() has for entries
the first unmap zapped, rather than by the reference count's ASSERT()
in debug builds only.
The change does not depend on asi= or on a sparse view of the directmap
being configured. PV vCPUs use the domain-wide mapcache on the default
command line too: in debug builds for every page, in release builds for
pages beyond the reach of the full directmap in their page-tables (5TiB,
3.5TiB with CONFIG_BIGMEM). For them the change adds the PTE writes
above. It also leaves up to MAPHASH_ENTRIES fewer pages per vCPU mapped
in the domain-wide area.
Assisted-by: Claude Code:claude-opus-5-5
Signed-off-by: George Dunlap <gwd@xxxxxxxxxxxxxx>
---
Changes in v3:
- New in this version.
---
xen/arch/x86/domain_page.c | 34 +++++++++++++++++++++++--------
xen/arch/x86/include/asm/domain.h | 5 ++++-
2 files changed, 30 insertions(+), 9 deletions(-)
diff --git a/xen/arch/x86/domain_page.c b/xen/arch/x86/domain_page.c
index 48f85b670d..eee66b848c 100644
--- a/xen/arch/x86/domain_page.c
+++ b/xen/arch/x86/domain_page.c
@@ -195,6 +195,14 @@ void *map_domain_page(mfn_t mfn)
{
idx = hashent->idx;
ASSERT(idx < cache->entries);
+ /*
+ * An entry the maphash kept holds on to its slot, not to its mapping
+ * (see unmap_domain_page()): map the page again. No flush is needed,
+ * as the slot has mapped nothing else since.
+ */
+ if ( !hashent->refcnt )
+ l1e_write(&MAPCACHE_L1ENT(idx),
+ l1e_from_mfn(mfn, __PAGE_HYPERVISOR_RW));
hashent->refcnt++;
ASSERT(hashent->refcnt);
ASSERT(mfn_eq(l1e_get_mfn(MAPCACHE_L1ENT(idx)), mfn));
@@ -227,8 +235,8 @@ void *map_domain_page(mfn_t mfn)
if ( hashent->idx != MAPHASHENT_NOTINUSE && !hashent->refcnt )
{
idx = hashent->idx;
- ASSERT(l1e_get_pfn(MAPCACHE_L1ENT(idx)) == hashent->mfn);
- l1e_write(&MAPCACHE_L1ENT(idx), l1e_empty());
+ ASSERT(!(l1e_get_flags(MAPCACHE_L1ENT(idx)) &
+ _PAGE_PRESENT));
hashent->idx = MAPHASHENT_NOTINUSE;
hashent->mfn = ~0UL;
break;
@@ -298,25 +306,35 @@ void unmap_domain_page(const void *ptr)
local_irq_save(flags);
+ /*
+ * The maphash keeps recently freed slots for a quick remap of the same
+ * page, but not their mappings: the page may change hands meanwhile, and
+ * must not stay reachable from this context then. A translation left in
+ * a TLB is covered, for owned pages, by the flush the page allocator
+ * arranges before a freed page is handed out again wherever a sparse
+ * directmap view is configured; without one, the directmap maps the page
+ * anyway.
+ */
if ( hashent->idx == idx )
{
ASSERT(hashent->mfn == mfn);
ASSERT(hashent->refcnt);
hashent->refcnt--;
+ if ( !hashent->refcnt )
+ l1e_write(&MAPCACHE_L1ENT(idx), l1e_empty());
}
else if ( !hashent->refcnt )
{
+ /* The entry replaced has no mapping left: mark its slot as garbage. */
if ( hashent->idx != MAPHASHENT_NOTINUSE )
{
- /* /First/, zap the PTE. */
- ASSERT(l1e_get_pfn(MAPCACHE_L1ENT(hashent->idx)) ==
- hashent->mfn);
- l1e_write(&MAPCACHE_L1ENT(hashent->idx), l1e_empty());
- /* /Second/, mark as garbage. */
+ ASSERT(!(l1e_get_flags(MAPCACHE_L1ENT(hashent->idx)) &
+ _PAGE_PRESENT));
set_bit(hashent->idx, cache->garbage);
}
- /* Add newly-freed mapping to the maphash. */
+ /* Add the newly-freed slot to the maphash, without its mapping. */
+ l1e_write(&MAPCACHE_L1ENT(idx), l1e_empty());
hashent->mfn = mfn;
hashent->idx = idx;
}
diff --git a/xen/arch/x86/include/asm/domain.h
b/xen/arch/x86/include/asm/domain.h
index 76a22b4040..51741f9d42 100644
--- a/xen/arch/x86/include/asm/domain.h
+++ b/xen/arch/x86/include/asm/domain.h
@@ -74,7 +74,10 @@ struct mapcache_vcpu {
/* Shadow of mapcache_domain.epoch. */
unsigned int shadow_epoch;
- /* Lock-free per-VCPU hash of recently-used mappings. */
+ /*
+ * Lock-free per-VCPU hash of recently-used mappings. An entry without
+ * references keeps its slot but not its mapping (see unmap_domain_page()).
+ */
struct vcpu_maphash_entry {
unsigned long mfn;
uint32_t idx;
--
2.55.0
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |