[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

[PATCH v3 0/6] unmap_page_range optimisation


  • To: xen-devel@xxxxxxxxxxxxxxxxxxxx
  • From: Kevin Lampis <kevin.lampis@xxxxxxxxxx>
  • Date: Thu, 1 Oct 2026 18:47:43 +0100
  • Arc-authentication-results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=citrix.com; dmarc=pass action=none header.from=citrix.com; dkim=pass header.d=citrix.com; arc=none
  • Arc-message-signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=NkdwiKtK+tR0FPw7NQxvOIvKgkeU+OyZ3OFk30VN4lo=; b=ZCCkDLlHOGrHvismthtu1t55c//IbBX4VEWOI29u3+A7IaxRs1X7TGl6Gl0e3UUyESvtrxWUeRt9uwfScUxP5OEX2U6zOprxgwtgKwQGxNZiHlN5N7O8zQIeKrvI6n+uVdx+cCWYFxQK/TZ0H5zUeorR6xCT1gFEX+yZbFFpNec1hEoju33wUgCH1SlmgDnVi8KlTTG+NiAWOJ5E5S3uVWRXh1t1E2WCwyJB4Buo6R1IUjLxlusoDmr7l/9blWwKq3OBbCVqgXGyoPrIdv/ChdGO8V4e/vQeX7AZY6/uqpaZm4K3hZ+mR/dKtrolc1NP0p/++/OAQb36D3arAmwZKQ==
  • Arc-seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=Y+Azn+hAkH4ayZAHeO3uOEKce67cW90te23kRF3QbRp1MUda0x9AHeizewMKQep/jZFDpJkyLa18WHKSOp+fkRhRAh67t9vj1f5j4Q3D+eOFcUDR0Asaspl6/EfwaZfKeCWoOlZcmWwn1ZFtcVe+pY9FHDhfd4uxVitwQWDSc7C7/GGnMcIlkgWGCySirjy/7yCcRcBuGFVpMdqH+QgJVkzr9qQEvVm9rrQDChNVnz/QO8yZQ171jgaBDD+Vqg8fuEGnsJ5NHmnj0yKCYB370f6qqqYu2+/T1f7/hssWzWZYJmm/USl7O6mC/gagAcXeddtqGU0AWLU5LxxXvPtCEw==
  • Authentication-results: eu.smtp.expurgate.cloud; dkim=pass header.s=selector1 header.d=citrix.com header.i="@citrix.com" header.h="From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck"
  • Authentication-results: mx.microsoft.com 1; dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=citrix.com;
  • Cc: jbeulich@xxxxxxxx, andrew.cooper3@xxxxxxxxxx, teddy.astie@xxxxxxxxxx, Kevin Lampis <kevin.lampis@xxxxxxxxxx>
  • Delivery-date: Thu, 01 Oct 2026 17:47:57 +0000
  • List-id: Xen developer discussion <xen-devel.lists.xenproject.org>

v1:
https://lore.kernel.org/xen-devel/20260727150615.1373200-1-kevin.lampis@xxxxxxxxxx/

Previous discussion:
RFC: unmap_page_range optimisation (avoiding emulation faults during VM 
migration)
https://lore.kernel.org/xen-devel/16133EFF-88FF-467F-B78F-E96EB148C3A5@xxxxxxxxxx/

This series adds a new MMU_PT_UPDATE_SWAP sub-op to the mmu_update hypercall
which Linux can use from a new pte_get_and_clear pv-op to significantly improve
the performance of unmap_page_range.

A microbenchmark[1] which allocates a large number of pages and then clears
them shows a performance increase from 1100ms to 640ms using the
new pte_get_and_clear pv-op.

Further profiling the microbenchmark with bpftrace[2] shows that Linux calls
unmap_page_range() 23 times with an average execution time of 48ms
dropping to 28ms when using the new operation.

Changes in v3:
- Extend swap to work with l2/l3/l4 PTEs and any value not just 0
- Other review comments

Changes in v2:
- Add a sub-op to mmu_update instead of new hypercall
- Add Linux patch
- Other review comments

[1] microbenchmark
#include <err.h>
#include <sys/mman.h>
#include <time.h>
#include <stdint.h>
#include <stdio.h>
#include <unistd.h>

static uint64_t nsec(void)
{
    struct timespec ts;
    clock_gettime(CLOCK_MONOTONIC, &ts);
    return (uint64_t)ts.tv_sec * 1e9 + ts.tv_nsec;
}

int main(int argc, char **argv)
{
    const size_t len = 1024UL * 1024 * 1024 * 4;
    const long pagesz = sysconf(_SC_PAGESIZE);

    char *p = mmap(NULL, len,
                   PROT_READ | PROT_WRITE,
                   MAP_PRIVATE | MAP_ANONYMOUS,
                   -1, 0);
    if ( p == MAP_FAILED )
        err(1, "mmap");

    /* Fault every page in */
    for (size_t i = 0; i < len; i += pagesz)
        p[i] = 1;

    uint64_t start = nsec();

    if ( madvise(p, len, MADV_DONTNEED) )
        err(1, "madvise");

    uint64_t end = nsec();

    printf("MADV_DONTNEED on %zu MB took %.3f ms\n",
           len / 1024 / 1024,
           (end - start) / 1e6);

    munmap(p, len);
    return 0;
}

[2] bpftrace
bpftrace -e '
kprobe:unmap_page_range
/comm == "a.out"/
{
    @start[tid] = nsecs;
}

kretprobe:unmap_page_range
/@start[tid] && comm == "a.out"/
{
    $delta = nsecs - @start[tid];
    @count = count();
    @total = sum($delta);
    delete(@start[tid]);
}

interval:s:10
{
    printf("avg call time = %d ns (%d calls)\n",
           (@total / @count), (uint64)@count);
    exit();
}'


Kevin Lampis (5):
  x86: Remove return value from UPDATE_ENTRY
  x86: extend update_intpte() to support atomic get-and-update
  x86: extend mod_l*_entry() to optionally return the old PTE value
  x86: extend do_mmu_update() to support returning the old PTE value
  x86: New feature flag XENFEAT_mmu_pt_update_swap.

 xen/arch/x86/mm.c               | 215 +++++++++++++++++++++++---------
 xen/arch/x86/pv/grant_table.c   |  53 ++++----
 xen/arch/x86/pv/mm.h            |  32 +++--
 xen/arch/x86/pv/ro-page-fault.c |   3 +-
 xen/common/kernel.c             |   3 +-
 xen/include/public/features.h   |   3 +
 xen/include/public/xen.h        |   1 +
 7 files changed, 208 insertions(+), 102 deletions(-)

-- 
2.52.0




 


Rackspace

Lists.xenproject.org is hosted with RackSpace, monitoring our
servers 24x7x365 and backed by RackSpace's Fanatical Support®.