From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1751282AbcGLHTc (ORCPT ); Tue, 12 Jul 2016 03:19:32 -0400 Received: from mail-wm0-f49.google.com ([74.125.82.49]:38726 "EHLO mail-wm0-f49.google.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1750794AbcGLHTa (ORCPT ); Tue, 12 Jul 2016 03:19:30 -0400 Date: Tue, 12 Jul 2016 09:19:27 +0200 From: Michal Hocko To: Shayan Pooya Cc: cgroups mailinglist , LKML , linux-mm@kvack.org Subject: Re: bug in memcg oom-killer results in a hung syscall in another process in the same cgroup Message-ID: <20160712071927.GD14586@dhcp22.suse.cz> References: <20160711064150.GB5284@dhcp22.suse.cz> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: User-Agent: Mutt/1.6.0 (2016-04-01) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Mon 11-07-16 11:33:19, Shayan Pooya wrote: > >> Could you post the stack trace of the hung oom victim? Also could you > >> post the full kernel log? > > With strace, when running 500 concurrent mem-hog tasks on the same > kernel, 33 of them failed with: > > strace: ../sysdeps/nptl/fork.c:136: __libc_fork: Assertion > `THREAD_GETMEM (self, tid) != ppid' failed. > > Which is: https://sourceware.org/bugzilla/show_bug.cgi?id=15392 > And discussed before at: https://lkml.org/lkml/2015/2/6/470 but that > patch was not accepted. OK, so the problem is that the oom killed task doesn't report the futex release properly? If yes then I fail to see how that is memcg specific. Could you try to clarify what you consider a bug again, please? I am not really sure I understand this report. Thanks! -- Michal Hocko SUSE Labs