From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S933383Ab2AKP5S (ORCPT ); Wed, 11 Jan 2012 10:57:18 -0500 Received: from mx3.mail.elte.hu ([157.181.1.138]:53126 "EHLO mx3.mail.elte.hu" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S932372Ab2AKP5R (ORCPT ); Wed, 11 Jan 2012 10:57:17 -0500 Date: Wed, 11 Jan 2012 16:56:59 +0100 From: Ingo Molnar To: Peter Zijlstra Cc: David Ahern , Linus Torvalds , Eric Dumazet , Thomas Gleixner , Martin Schwidefsky , linux-kernel , Frederic Weisbecker , Suresh Siddha Subject: Re: [BUG] kernel freezes with latest tree Message-ID: <20120111155658.GB26659@elte.hu> References: <1326213442.19095.9.camel@edumazet-HP-Compaq-6005-Pro-SFF-PC> <1326214407.19095.11.camel@edumazet-HP-Compaq-6005-Pro-SFF-PC> <1326234230.2614.15.camel@edumazet-laptop> <4F0D2D9B.8030501@gmail.com> <1326272685.2442.120.camel@twins> <1326284711.2442.138.camel@twins> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <1326284711.2442.138.camel@twins> User-Agent: Mutt/1.5.21 (2010-09-15) X-ELTE-SpamScore: -2.0 X-ELTE-SpamLevel: X-ELTE-SpamCheck: no X-ELTE-SpamVersion: ELTE 2.0 X-ELTE-SpamCheck-Details: score=-2.0 required=5.9 tests=AWL,BAYES_00 autolearn=no SpamAssassin version=3.3.1 -2.0 BAYES_00 BODY: Bayes spam probability is 0 to 1% [score: 0.0000] 0.0 AWL AWL: From: address is in the auto white-list Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org * Peter Zijlstra wrote: > On Wed, 2012-01-11 at 10:04 +0100, Peter Zijlstra wrote: > > Maybe adding a few more NEED_BREAK bits > > and making it a counter and overflowing it into ABORT might be good. > > > > > > I could reproduce and confirm something like the below makes > the hang go-away. I haven't managed to fully understand why > we're stuck though because we do release the runqueue locks > and re-enable IRQs on this lock-break. Well, what happens if every CPU runs load_balance() and we keep triggering: if (loops++ > sysctl_sched_nr_migrate) { *lb_flags |= LBF_NEED_BREAK; break; } in this case load_balance() will do the retry: if (lb_flags & LBF_NEED_BREAK) { lb_flags &= ~LBF_NEED_BREAK; goto redo; } but the retry starts the loop again: list_for_each_entry_safe(p, n, &busiest_cfs_rq->tasks, se.group_node) { so nobody is able to make progress: livelock/lockup. ( This also explains why i was unable to see this in my randomized testing: my tests never extreme enough to trigger the sysctl_sched_nr_migrate threshold. ) Thanks, Ingo