[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index]

Re: [PATCH] x86/ucode: Remove MICROCODE_UPDATE_TIMEOUT and associated panic()


  • To: Jan Beulich <jbeulich@xxxxxxxx>
  • From: Andrew Cooper <andrew.cooper3@xxxxxxxxxx>
  • Date: Wed, 19 Aug 2026 12:13:47 +0100
  • Arc-authentication-results: i=1; mx.microsoft.com 1; spf=pass smtp.mailfrom=citrix.com; dmarc=pass action=none header.from=citrix.com; dkim=pass header.d=citrix.com; arc=none
  • Arc-message-signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=7Z08gkcl2/hGSaktKo543z0fApnu6jnlkBaA+udc4/Q=; b=fDF9hbGy8lPl1uzggqir0M8dJwrwEz/wLWvQLI0NRx/1kDG7W4FV3SC54iRsR4KCrBd7YYpGXwdmRyfsgDDNtRuyL2s6dYZc7QsVG+ojkwFD5jdYpfpLQWkSScsAci/wbqZ+KNOtsUcBV7f31B1M2UYLJfhTeMCpbfRILjsO1EVSSCN+g8dFwuCVoITEJ2SXUIMWVOY0d+1Nrv7USOtHRcGX8psCJMFzQ89ggpL4TywTJO0QFZJMiHvfVCuJof1M8rSX2Hl/qRca0frIjIfn/yAyCQFZIzmJ4Diinj5TJNutuisDfD2UdHmT5lx0XNjG5JXItv3Ut7NI+JUph8g08A==
  • Arc-seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=HtIRjOuGnXbLgSCEss8tw/Ze4vzMRoAJWUMp9n360G3/S6LlK2J4/7ufEEvzBpJm0lfA+A7YYZPt/S8plA4VKYivwpT1SaDi/4DaoWv6j+yHkTi4xmAZMJS11oRuyXewFeEq95lgWCH9bqP6T4m7GXIN3nIav8cWfN6fZ5UPG5rcr7gX+AIv2GVI6E+LpuAiYynCetHtXv3/sLjrQJIk3KCKjlQlw7d9CcpfuW/ne4HTuhvNBUHdN55ogSFLO0AcEmgS/NwGtZQ1jj1IuzBpUGdHdJ56ikSTPEarRmI263Ukdi6Dm39DJIMYZfVHBrhR8tMLaRkvnZn2DJ9IAPVGSg==
  • Authentication-results: eu.smtp.expurgate.cloud; dkim=pass header.s=selector1 header.d=citrix.com header.i="@citrix.com" header.h="From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck"
  • Authentication-results: dkim=none (message not signed) header.d=none;dmarc=none action=none header.from=citrix.com;
  • Autocrypt: addr=andrew.cooper3@xxxxxxxxxx; keydata= xsFNBFLhNn8BEADVhE+Hb8i0GV6mihnnr/uiQQdPF8kUoFzCOPXkf7jQ5sLYeJa0cQi6Penp VtiFYznTairnVsN5J+ujSTIb+OlMSJUWV4opS7WVNnxHbFTPYZVQ3erv7NKc2iVizCRZ2Kxn srM1oPXWRic8BIAdYOKOloF2300SL/bIpeD+x7h3w9B/qez7nOin5NzkxgFoaUeIal12pXSR Q354FKFoy6Vh96gc4VRqte3jw8mPuJQpfws+Pb+swvSf/i1q1+1I4jsRQQh2m6OTADHIqg2E ofTYAEh7R5HfPx0EXoEDMdRjOeKn8+vvkAwhviWXTHlG3R1QkbE5M/oywnZ83udJmi+lxjJ5 YhQ5IzomvJ16H0Bq+TLyVLO/VRksp1VR9HxCzItLNCS8PdpYYz5TC204ViycobYU65WMpzWe LFAGn8jSS25XIpqv0Y9k87dLbctKKA14Ifw2kq5OIVu2FuX+3i446JOa2vpCI9GcjCzi3oHV e00bzYiHMIl0FICrNJU0Kjho8pdo0m2uxkn6SYEpogAy9pnatUlO+erL4LqFUO7GXSdBRbw5 gNt25XTLdSFuZtMxkY3tq8MFss5QnjhehCVPEpE6y9ZjI4XB8ad1G4oBHVGK5LMsvg22PfMJ ISWFSHoF/B5+lHkCKWkFxZ0gZn33ju5n6/FOdEx4B8cMJt+cWwARAQABzSlBbmRyZXcgQ29v cGVyIDxhbmRyZXcuY29vcGVyM0BjaXRyaXguY29tPsLBegQTAQgAJAIbAwULCQgHAwUVCgkI CwUWAgMBAAIeAQIXgAUCWKD95wIZAQAKCRBlw/kGpdefoHbdD/9AIoR3k6fKl+RFiFpyAhvO 59ttDFI7nIAnlYngev2XUR3acFElJATHSDO0ju+hqWqAb8kVijXLops0gOfqt3VPZq9cuHlh IMDquatGLzAadfFx2eQYIYT+FYuMoPZy/aTUazmJIDVxP7L383grjIkn+7tAv+qeDfE+txL4 SAm1UHNvmdfgL2/lcmL3xRh7sub3nJilM93RWX1Pe5LBSDXO45uzCGEdst6uSlzYR/MEr+5Z JQQ32JV64zwvf/aKaagSQSQMYNX9JFgfZ3TKWC1KJQbX5ssoX/5hNLqxMcZV3TN7kU8I3kjK mPec9+1nECOjjJSO/h4P0sBZyIUGfguwzhEeGf4sMCuSEM4xjCnwiBwftR17sr0spYcOpqET ZGcAmyYcNjy6CYadNCnfR40vhhWuCfNCBzWnUW0lFoo12wb0YnzoOLjvfD6OL3JjIUJNOmJy RCsJ5IA/Iz33RhSVRmROu+TztwuThClw63g7+hoyewv7BemKyuU6FTVhjjW+XUWmS/FzknSi dAG+insr0746cTPpSkGl3KAXeWDGJzve7/SBBfyznWCMGaf8E2P1oOdIZRxHgWj0zNr1+ooF /PzgLPiCI4OMUttTlEKChgbUTQ+5o0P080JojqfXwbPAyumbaYcQNiH1/xYbJdOFSiBv9rpt TQTBLzDKXok86M7BTQRS4TZ/ARAAkgqudHsp+hd82UVkvgnlqZjzz2vyrYfz7bkPtXaGb9H4 Rfo7mQsEQavEBdWWjbga6eMnDqtu+FC+qeTGYebToxEyp2lKDSoAsvt8w82tIlP/EbmRbDVn 7bhjBlfRcFjVYw8uVDPptT0TV47vpoCVkTwcyb6OltJrvg/QzV9f07DJswuda1JH3/qvYu0p vjPnYvCq4NsqY2XSdAJ02HrdYPFtNyPEntu1n1KK+gJrstjtw7KsZ4ygXYrsm/oCBiVW/OgU g/XIlGErkrxe4vQvJyVwg6YH653YTX5hLLUEL1NS4TCo47RP+wi6y+TnuAL36UtK/uFyEuPy wwrDVcC4cIFhYSfsO0BumEI65yu7a8aHbGfq2lW251UcoU48Z27ZUUZd2Dr6O/n8poQHbaTd 6bJJSjzGGHZVbRP9UQ3lkmkmc0+XCHmj5WhwNNYjgbbmML7y0fsJT5RgvefAIFfHBg7fTY/i kBEimoUsTEQz+N4hbKwo1hULfVxDJStE4sbPhjbsPCrlXf6W9CxSyQ0qmZ2bXsLQYRj2xqd1 bpA+1o1j2N4/au1R/uSiUFjewJdT/LX1EklKDcQwpk06Af/N7VZtSfEJeRV04unbsKVXWZAk uAJyDDKN99ziC0Wz5kcPyVD1HNf8bgaqGDzrv3TfYjwqayRFcMf7xJaL9xXedMcAEQEAAcLB XwQYAQgACQUCUuE2fwIbDAAKCRBlw/kGpdefoG4XEACD1Qf/er8EA7g23HMxYWd3FXHThrVQ HgiGdk5Yh632vjOm9L4sd/GCEACVQKjsu98e8o3ysitFlznEns5EAAXEbITrgKWXDDUWGYxd pnjj2u+GkVdsOAGk0kxczX6s+VRBhpbBI2PWnOsRJgU2n10PZ3mZD4Xu9kU2IXYmuW+e5KCA vTArRUdCrAtIa1k01sPipPPw6dfxx2e5asy21YOytzxuWFfJTGnVxZZSCyLUO83sh6OZhJkk b9rxL9wPmpN/t2IPaEKoAc0FTQZS36wAMOXkBh24PQ9gaLJvfPKpNzGD8XWR5HHF0NLIJhgg 4ZlEXQ2fVp3XrtocHqhu4UZR4koCijgB8sB7Tb0GCpwK+C4UePdFLfhKyRdSXuvY3AHJd4CP 4JzW0Bzq/WXY3XMOzUTYApGQpnUpdOmuQSfpV9MQO+/jo7r6yPbxT7CwRS5dcQPzUiuHLK9i nvjREdh84qycnx0/6dDroYhp0DFv4udxuAvt1h4wGwTPRQZerSm4xaYegEFusyhbZrI0U9tJ B8WrhBLXDiYlyJT6zOV2yZFuW47VrLsjYnHwn27hmxTC/7tvG3euCklmkn9Sl9IAKFu29RSo d5bD8kMSCYsTqtTfT6W4A3qHGvIDta3ptLYpIAOD2sY3GYq2nf3Bbzx81wZK14JdDDHUX2Rs 6+ahAA==
  • Cc: Andrew Cooper <andrew.cooper3@xxxxxxxxxx>, Roger Pau Monné <roger@xxxxxxxxxxxxxx>, Teddy Astie <teddy.astie@xxxxxxxxxx>, Xen-devel <xen-devel@xxxxxxxxxxxxxxxxxxxx>
  • Delivery-date: Wed, 19 Aug 2026 11:14:02 +0000
  • List-id: Xen developer discussion <xen-devel.lists.xenproject.org>

On 18/08/2026 3:31 pm, Jan Beulich wrote:
> On 18.08.2026 15:07, Andrew Cooper wrote:
>> Panicing in the case of a timeout turns out to be about the worst possible
>> action Xen can take.  It leaves all other APs waiting on the condition
>> variable, some in NMI context.  As a result, they fail to be shot down and
>> dump state for kexec crash analysis.
> At the same time there likely isn't much to be learned from a kexec dump, as
> the source of the issue is in the CPU, not in Xen.

The single most valuable print message I've ever added to Xen is the one
which reports which CPUs didn't respond to NMIs.  Up until now, it has
always highlighted hardware issues.

Despite my wish to remove this panic specifically, there is still
information to be gained from kexec in a similar scenario.

> Further, this code runs with the watchdog disabled. If there's truly no
> progress anymore, how would one know from the outside whether the system is
> dead altogether vs the control CPU still kicking around?

The scenario you describe can only occur if the BSP accepts the
microcode successfully, and one of the APs locks up properly.

If the BSP locks up, we never get as far as deciding to panic(), and the
system hangs already.

If we have a bad microcode, it is far more likely for the BSP to hang
than for the BSP to work one of the APs hang.

>> Microcode Loading on Granite Rapids takes about 4.5s of wallclock time, far 
>> in
>> excess of the of the arbitrary 1s Xen allows.  This time is spent in the 
>> WRMSR
>> to load the blob, and there's nothing the system can do but to sit and wait.
>> Despite the delay, the system as a whole does survive.
> For this I wonder whether the log message issues after 1 sec is adequate. If
> we know things can take this long, wouldn't we better issue the log message
> no earlier than, say, 5s * nr_sockets?
>

I have some work there not submitted yet, which would periodically print
a message.  That at least gives you slight signs of life.

But nothing involving nr_sockets.  It's far more often wrong than it is
right, and that's not how our algorithm scales.

GNR has gone from milliseconds to 4.5s.  Putting the limit at 5s is just
going to need another change in a year or two.

>> Microcode loading occures through admin operation only, so get rid of the
>> timeout completely.  It does nothing but make a bad sitaution worse.
> Worse when taking one perspective, yes, yet a silent hang also is worse than
> a panic() telling you what was wrong.
>
> This is a tough one, with - likely - no really good options. And establishing
> what's "least bad" may also be difficult.

As indicated, there's only one narrow and unlikely case where the
panic() is a true positive.

The likely true-hang case don't reach the panic(), and that only leaves
"timeout too short" which has proved to be the case on GNR.

~Andrew



 


Rackspace

Lists.xenproject.org is hosted with RackSpace, monitoring our
servers 24x7x365 and backed by RackSpace's Fanatical Support®.