From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1757824AbdKOTSC (ORCPT ); Wed, 15 Nov 2017 14:18:02 -0500 Received: from mail-qt0-f173.google.com ([209.85.216.173]:55803 "EHLO mail-qt0-f173.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1752629AbdKOTR5 (ORCPT ); Wed, 15 Nov 2017 14:17:57 -0500 X-Google-Smtp-Source: AGs4zMZpirYtDuXLa5y61svTvKIuC/dRtovptMmsd5syQ3wtJ7i2ZMZmdJDdOAFY7Ev//quLrmamVw== Subject: Re: [PATCH] mm, meminit: Serially initialise deferred memory if trace_buf_size is specified To: Michal Hocko , Mel Gorman References: <20171115085556.fla7upm3nkydlflp@techsingularity.net> <20171115115559.rjb5hy6d6332jgjj@dhcp22.suse.cz> <20171115141329.ieoqvyoavmv6gnea@techsingularity.net> <20171115142816.zxdgkad3ch2bih6d@dhcp22.suse.cz> <20171115144314.xwdi2sbcn6m6lqdo@techsingularity.net> <20171115145716.w34jaez5ljb3fssn@dhcp22.suse.cz> Cc: Andrew Morton , linux-mm@kvack.org, linux-kernel@vger.kernel.org, koki.sanagi@us.fujitsu.com, yasu.isimatu@gmail.com From: YASUAKI ISHIMATSU Message-ID: <06a33f82-7f83-7721-50ec-87bf1370c3d4@gmail.com> Date: Wed, 15 Nov 2017 14:17:52 -0500 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:45.0) Gecko/20100101 Thunderbird/45.8.0 MIME-Version: 1.0 In-Reply-To: <20171115145716.w34jaez5ljb3fssn@dhcp22.suse.cz> Content-Type: text/plain; charset=windows-1252 Content-Transfer-Encoding: 7bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Hi Michal and Mel, To reproduce the issue, I specified the large trace buffer. The issue also occurs with trace_buf_size=12M and movable_node on 4.14.0. In my system, there are 384 CPUs and 8 nodes. So when not using movable_node boot option, kernel can use about 16GB memory for trace buffer. So Kernel boots up with trace_buf_size=12M. But when using movable_node, 6 nodes are managed as MOVABLE_ZONE in my system and kernel can use only about 4GB memory for trace buffer. So memory allocation failure of trace buffer occurs with trace_buf_size=12M and movable_node. I don't know you still think 12M is large. But the latest Fujitsu server supports 448 CPUs. The issue may occur with trace_buf_size=10M on the system. Additionally the number of CPU in a server is increasing year by year. So the issue will occurs even if we don't specify large trace buffer. Thanks, Yasuaki Ishimatsu On 11/15/2017 09:57 AM, Michal Hocko wrote: > On Wed 15-11-17 14:43:14, Mel Gorman wrote: >> On Wed, Nov 15, 2017 at 03:28:16PM +0100, Michal Hocko wrote: >>> On Wed 15-11-17 14:13:29, Mel Gorman wrote: >>> [...] >>>> I doubt anyone well. Even the original reporter appeared to pick that >>>> particular value just to trigger the OOM. >>> >>> Then why do we care at all? The trace buffer size can be configured from >>> the userspace if it is not sufficiently large IIRC. >>> >> >> I guess there is the potential that the trace buffer needs to be large >> enough early on in boot but I'm not sure why it would need to be that large >> to be honest. Bottom line, it's fairly trivial to just serialise meminit >> in the event that it's resized from command line. I'm also ok with just >> leaving this is as a "don't set the buffer that large" > > I would be reluctant to touch the code just because of insane kernel > command line option. > > That being said, I will not object or block the patch it just seems > unnecessary for most reasonable setups I can think of. If there is a > legitimate usage of such a large trace buffer then I wouldn't oppose. >