|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Re: [RFC PATCH v2] x86/ACPI: optimise C3 entry on AMD CPUs
On 30.09.2026 21:30, Olivier Lambert wrote: > All Zen or newer CPU which support C3 shares cache. Its not necessary to > flush the caches in software before entering C3. This will cause drop in > performance for the cores which share some caches. ARB_DIS is not used > with current AMD C state implementation. So set related flags correctly. > > Signed-off-by: Deepak Sharma <deepak.sharma@xxxxxxx> Shouldn't this be mirrored in a From: tag? (I'll try to remember to attribute the patch that way when committing.) > Acked-by: Thomas Gleixner <tglx@xxxxxxxxxxxxx> > Signed-off-by: Rafael J. Wysocki <rafael.j.wysocki@xxxxxxxxx> > Origin: https://git.kernel.org/torvalds/p/a8fb40966f19 > [Port to Xen. Include Hygon.] > Assisted-by: Claude-Code:claude-opus-5-5 > Signed-off-by: Olivier Lambert <olivier.lambert@xxxxxxxxxx> > Reviewed-by: Jason Andryuk <jason.andryuk@xxxxxxx> Acked-by: Jan Beulich <jbeulich@xxxxxxxx> Why the RFC, btw? > --- > > Notes: > Changes in v2: > - Resend with git send-email: my mail client stripped the whitespace from > v1, so its diff could not be applied. No code change. > - Picked up Jason's Reviewed-by. > > Found on XCP-ng 8.3 (Xen 4.17.6) on a Ryzen 5 7600 (Zen 4, family 0x19) > whose _CST lists C1 (HLT) plus C2 and C3 as SYSIO states (C3 exit > latency 350us). With bm_check == 0, every C3 entry executes WBINVD. > > Measured under a VM-to-VM network load, in 25s windows: > - ~140 WBINVD/s, 65-69us each on average, 300us at worst. > - Spinning on each CPU in turn with IRQs off and reading the TSC shows > gaps of up to 222us; none with this change. The SMI count (PMC event > 0x2b, validated by counting cycles on the same counter) stayed at 0, > and NMIs didn't line up with the gaps. > - credit1's do_schedule() occasionally took up to 262us, with IRQs > off; at most 4.4us with this change. > - xenpm set-max-cstate 1 removes the stalls too, which points at C3 > entry. > > This change was booted and measured as a backport on 4.17.6 only. On > staging it is build-tested. It also covers Hygon, as Linux does today > through X86_FEATURE_ZEN; the original Linux commit was AMD-only. > > This will conflict trivially with Alejandro's pending "x86/acpi: Migrate > vendor checks to cpu_vendor()"; happy to rebase on top of it. > > Possibly a backport candidate. Perhaps, yes. > --- a/xen/arch/x86/acpi/cpu_idle.c > +++ b/xen/arch/x86/acpi/cpu_idle.c > @@ -1065,6 +1065,24 @@ static void acpi_processor_power_init_bm_check(struct > acpi_processor_flags *flag > */ > if ( c->vendor == X86_VENDOR_INTEL ) > flags->bm_control = 0; > + > + if ( (c->vendor & (X86_VENDOR_AMD | X86_VENDOR_HYGON)) && > + c->family >= 0x17 ) > + { > + /* > + * For all AMD Zen or newer CPUs that support C3, caches > + * should not be flushed by software while entering C3 > + * type state. Set bm_check to 1 so that Xen doesn't > + * need to execute cache flush operation. > + */ > + flags->bm_check = 1; > + /* > + * In current AMD C state implementation ARB_DIS is no longer > + * used. So set bm_control to zero to indicate ARB_DIS is not > + * required while entering C3 type state. > + */ > + flags->bm_control = 0; > + } > } Linux also deals with modern Centaur (VIA) and Zhaoxin (Shanghai) CPUs. We will want to mirror that too, provided we want to continue this mirroring. Alternatively we could have Linux'es xen-acpi-processor driver invoke acpi_processor_power_init_bm_check() before passing up C-state info. That would imo be better overall, with one downside: Non- Linux Dom0-s would then need to follow suit. Jan
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |