From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S971118AbdDTQVy (ORCPT ); Thu, 20 Apr 2017 12:21:54 -0400 Received: from mx1.redhat.com ([209.132.183.28]:51268 "EHLO mx1.redhat.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S970993AbdDTQVo (ORCPT ); Thu, 20 Apr 2017 12:21:44 -0400 DMARC-Filter: OpenDMARC Filter v1.3.2 mx1.redhat.com 56A2A75750 Authentication-Results: ext-mx01.extmail.prod.ext.phx2.redhat.com; dmarc=none (p=none dis=none) header.from=redhat.com Authentication-Results: ext-mx01.extmail.prod.ext.phx2.redhat.com; spf=pass smtp.mailfrom=vkuznets@redhat.com DKIM-Filter: OpenDKIM Filter v2.11.0 mx1.redhat.com 56A2A75750 From: Vitaly Kuznetsov To: Boris Ostrovsky Cc: Peter Zijlstra , Prarit Bhargava , x86@kernel.org, linux-kernel@vger.kernel.org, Ingo Molnar , "H. Peter Anvin" , xen-devel@lists.xenproject.org, Thomas Gleixner , Borislav Petkov , Juergen Gross Subject: Re: [Xen-devel] [PATCH RFC] x86/smpboot: Set safer __max_logical_packages limit References: <20170420132453.19652-1-vkuznets@redhat.com> <20170420150615.ns3343rokvmc3kjt@hirez.programming.kicks-ass.net> <87fuh3xf2i.fsf@vitty.brq.redhat.com> Date: Thu, 20 Apr 2017 18:21:40 +0200 In-Reply-To: (Boris Ostrovsky's message of "Thu, 20 Apr 2017 12:01:54 -0400") Message-ID: <874lxjxd63.fsf@vitty.brq.redhat.com> User-Agent: Gnus/5.13 (Gnus v5.13) Emacs/25.1 (gnu/linux) MIME-Version: 1.0 Content-Type: text/plain X-Greylist: Sender IP whitelisted, not delayed by milter-greylist-4.5.16 (mx1.redhat.com [10.5.110.25]); Thu, 20 Apr 2017 16:21:44 +0000 (UTC) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Boris Ostrovsky writes: > On 04/20/2017 11:40 AM, Vitaly Kuznetsov wrote: >> Peter Zijlstra writes: >> >>> On Thu, Apr 20, 2017 at 03:24:53PM +0200, Vitaly Kuznetsov wrote: >>>> In this patch I suggest we set __max_logical_packages based on the >>>> max_physical_pkg_id and total_cpus, >>> So my 4 socket 144 CPU system will then get max_physical_pkg_id=144, >>> instead of 4. >>> >>> This wastes quite a bit of memory for the per-node arrays. Luckily most >>> are just pointer arrays, but still, wasting 140*8 bytes for each of >>> them. >>> >>>> this should be safe and cover all >>>> possible cases. Alternatively, we may think about eliminating the concept >>>> of __max_logical_packages completely and relying on max_physical_pkg_id/ >>>> total_cpus where we currently use topology_max_packages(). >>>> >>>> The issue could've been solved in Xen too I guess. CPUID returning >>>> x86_max_cores can be tweaked to be the lowerest(?) possible number of >>>> all logical packages of the guest. >>> This is getting ludicrous. Xen is plain broken, and instead of fixing >>> it, you propose to somehow deal with its obviously crack induced >>> behaviour :-( >> Totally agree and I don't like the solution I propose (and that's why >> this is RFC)... The problem is that there are such Xen setups in the >> wild and with the recent changes some guests will BUG() :-( >> >> Alternatively, we can just remove the BUG() and do something with CPUs >> which have their pkg >= __max_logical_packages, e.g. assign them to the >> last package. Far from ideal but will help to avoid the regression. > > Do you observe this failure on PV or HVM guest? > > We've had a number of issues with topology discovery for PV guests but > AFAIK they have been addressed (so far). I wonder though whether it > would make sense to have some sort of a callback (or an smp_ops.op) to > override native topology info, if needed. > This is HVM. I guess that CPUID handling for AMD processors in the hypervisor doesn't adjust the core information and passes it from hardware as-is while it should be tweaked. -- Vitaly