From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S966091AbcBCTN3 (ORCPT ); Wed, 3 Feb 2016 14:13:29 -0500 Received: from www.linutronix.de ([62.245.132.108]:60749 "EHLO Galois.linutronix.de" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S965551AbcBCTN2 (ORCPT ); Wed, 3 Feb 2016 14:13:28 -0500 Date: Wed, 3 Feb 2016 20:12:19 +0100 (CET) From: Thomas Gleixner To: Tejun Heo cc: Mike Galbraith , LKML , Michal Hocko , Jiri Slaby , Petr Mladek , Jan Kara , Ben Hutchings , Sasha Levin , Shaohua Li , Daniel Bilik Subject: Re: [PATCH wq/for-4.5-fixes] workqueue: handle NUMA_NO_NODE for unbound pool_workqueue lookup In-Reply-To: <20160203185425.GK14091@mtj.duckdns.org> Message-ID: References: <1454424264.11183.46.camel@gmail.com> <20160203185425.GK14091@mtj.duckdns.org> User-Agent: Alpine 2.11 (DEB 23 2013-08-11) MIME-Version: 1.0 Content-Type: TEXT/PLAIN; charset=US-ASCII X-Linutronix-Spam-Score: -1.0 X-Linutronix-Spam-Level: - X-Linutronix-Spam-Status: No , -1.0 points, 5.0 required, ALL_TRUSTED=-1,SHORTCIRCUIT=-0.0001,URIBL_BLOCKED=0.001 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Wed, 3 Feb 2016, Tejun Heo wrote: > When looking up the pool_workqueue to use for an unbound workqueue, > workqueue assumes that the target CPU is always bound to a valid NUMA > node. However, currently, when a CPU goes offline, the mapping is > destroyed and cpu_to_node() returns NUMA_NO_NODE. This has always > been broken but hasn't triggered until recently. > > After 874bbfe600a6 ("workqueue: make sure delayed work run in local > cpu"), workqueue forcifully assigns the local CPU for delayed work > items without explicit target CPU to fix a different issue. This > widens the window where CPU can go offline while a delayed work item > is pending causing delayed work items dispatched with target CPU set > to an already offlined CPU. The resulting NUMA_NO_NODE mapping makes > workqueue try to queue the work item on a NULL pool_workqueue and thus > crash. > > Fix it by mapping NUMA_NO_NODE to the default pool_workqueue from > unbound_pwq_by_node(). This is a temporary workaround. The long term > solution is keeping CPU -> NODE mapping stable across CPU off/online > cycles which is in the works. > > Signed-off-by: Tejun Heo > Reported-by: Mike Galbraith > Cc: Tang Chen > Cc: Rafael J. Wysocki > Cc: Len Brown > Cc: stable@vger.kernel.org # v4.3+ 4.3+ ? Hasn't 874bbfe600a6 been backported to older stable kernels? Adding a 'Fixes: 874bbfe600a6 ...' tag is what you really want here. > diff --git a/kernel/workqueue.c b/kernel/workqueue.c > index 61a0264..f748eab 100644 > --- a/kernel/workqueue.c > +++ b/kernel/workqueue.c > @@ -570,6 +570,16 @@ static struct pool_workqueue *unbound_pwq_by_node(struct workqueue_struct *wq, > int node) > { > assert_rcu_or_wq_mutex_or_pool_mutex(wq); > + > + /* > + * XXX: @node can be NUMA_NO_NODE if CPU goes offline while a > + * delayed item is pending. The plan is to keep CPU -> NODE > + * mapping valid and stable across CPU on/offlines. Once that > + * happens, this workaround can be removed. So what happens if the complete node is offline? > + */ > + if (unlikely(node == NUMA_NO_NODE)) > + return wq->dfl_pwq; > + Thanks, tglx