From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S932713AbdBVODR (ORCPT ); Wed, 22 Feb 2017 09:03:17 -0500 Received: from mx0b-001b2d01.pphosted.com ([148.163.158.5]:58136 "EHLO mx0a-001b2d01.pphosted.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S932360AbdBVODH (ORCPT ); Wed, 22 Feb 2017 09:03:07 -0500 Subject: Re: [PATCH] mm/cgroup: avoid panic when init with low memory To: Michal Hocko References: <1487154969-6704-1-git-send-email-ldufour@linux.vnet.ibm.com> <20170220130123.GI2431@dhcp22.suse.cz> <934d40ec-060b-4794-2fdc-35a7ea1dc9e2@linux.vnet.ibm.com> <20170220174258.GA31541@dhcp22.suse.cz> Cc: Johannes Weiner , Vladimir Davydov , cgroups@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org From: Laurent Dufour Date: Wed, 22 Feb 2017 15:02:54 +0100 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:45.0) Gecko/20100101 Thunderbird/45.7.0 MIME-Version: 1.0 In-Reply-To: <20170220174258.GA31541@dhcp22.suse.cz> Content-Type: text/plain; charset=windows-1252 Content-Transfer-Encoding: 7bit X-TM-AS-GCONF: 00 X-Content-Scanned: Fidelis XPS MAILER x-cbid: 17022214-0020-0000-0000-00000278FD11 X-IBM-AV-DETECTION: SAVI=unused REMOTE=unused XFE=unused x-cbparentid: 17022214-0021-0000-0000-00001F786B58 Message-Id: <9414873a-6c64-7b96-6251-f0ddba2b256e@linux.vnet.ibm.com> X-Proofpoint-Virus-Version: vendor=fsecure engine=2.50.10432:,, definitions=2017-02-22_08:,, signatures=0 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 suspectscore=0 malwarescore=0 phishscore=0 adultscore=0 bulkscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1612050000 definitions=main-1702220133 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 20/02/2017 18:42, Michal Hocko wrote: > On Mon 20-02-17 18:09:43, Laurent Dufour wrote: >> On 20/02/2017 14:01, Michal Hocko wrote: >>> On Wed 15-02-17 11:36:09, Laurent Dufour wrote: >>>> The system may panic when initialisation is done when almost all the >>>> memory is assigned to the huge pages using the kernel command line >>>> parameter hugepage=xxxx. Panic may occur like this: >>> >>> I am pretty sure the system might blow up in many other ways when you >>> misconfigure it and pull basically all the memory out. Anyway... >>> >>> [...] >>> >>>> This is a chicken and egg issue where the kernel try to get free >>>> memory when allocating per node data in mem_cgroup_init(), but in that >>>> path mem_cgroup_soft_limit_reclaim() is called which assumes that >>>> these data are allocated. >>>> >>>> As mem_cgroup_soft_limit_reclaim() is best effort, it should return >>>> when these data are not yet allocated. >>> >>> ... this makes some sense. Especially when there is no soft limit >>> configured. So this is a good step. I would just like to ask you to go >>> one step further. Can we make the whole soft reclaim thing uninitialized >>> until the soft limit is actually set? Soft limit is not used in cgroup >>> v2 at all and I would strongly discourage it in v1 as well. We will save >>> few bytes as a bonus. >> >> Hi Michal, and thanks for the review. >> >> I'm not familiar with that part of the kernel, so to be sure we are on >> the same line, are you suggesting to set soft_limit_tree at the first >> time mem_cgroup_write() is called to set a soft_limit field ? > > yes > >> Obviously, all callers to soft_limit_tree_node() and >> soft_limit_tree_from_page() will have to check for the return pointer to >> be NULL. > > All callers that need to access the tree unconditionally, yes. Which is > the case anyway, right? I haven't checked the check you have added is > sufficient, but we shouldn't have that many of them because some code > paths are called only when the soft limit is enabled. You're right there are not so much callers to fix. I'll send a new series containing the previous patch fixing the initial panic and another one delaying the data allocation.