From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1754781AbcEXIrm (ORCPT ); Tue, 24 May 2016 04:47:42 -0400 Received: from mail-wm0-f66.google.com ([74.125.82.66]:32961 "EHLO mail-wm0-f66.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1753093AbcEXIrk (ORCPT ); Tue, 24 May 2016 04:47:40 -0400 Date: Tue, 24 May 2016 10:47:37 +0200 From: Michal Hocko To: Vladimir Davydov Cc: Andrew Morton , Johannes Weiner , linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH] mm: memcontrol: fix possible css ref leak on oom Message-ID: <20160524084737.GC8259@dhcp22.suse.cz> References: <1464019330-7579-1-git-send-email-vdavydov@virtuozzo.com> <20160523174441.GA32715@dhcp22.suse.cz> <20160524084319.GH7917@esperanza> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20160524084319.GH7917@esperanza> User-Agent: Mutt/1.6.0 (2016-04-01) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Tue 24-05-16 11:43:19, Vladimir Davydov wrote: > On Mon, May 23, 2016 at 07:44:43PM +0200, Michal Hocko wrote: > > On Mon 23-05-16 19:02:10, Vladimir Davydov wrote: > > > mem_cgroup_oom may be invoked multiple times while a process is handling > > > a page fault, in which case current->memcg_in_oom will be overwritten > > > leaking the previously taken css reference. > > > > Have you seen this happening? I was under impression that the page fault > > paths that have oom enabled will not retry allocations. > > filemap_fault will, for readahead. I thought that the readahead is __GFP_NORETRY so we do not trigger OOM killer. > This is rather unlikely, just like the whole oom scenario, so I haven't > faced this leak in production yet, although it's pretty easy to > reproduce using a contrived test. However, even if this leak happened on > my host, I would probably not notice, because currently we have no clear > means of catching css leaks. I'm thinking about adding a file to debugfs > containing brief information about all memory cgroups, including dead > ones, so that we could at least see how many dead memory cgroups are > dangling out there. Yeah, debugfs interface would make some sense. -- Michal Hocko SUSE Labs