From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S932542AbcFOMus (ORCPT ); Wed, 15 Jun 2016 08:50:48 -0400 Received: from mx0b-001b2d01.pphosted.com ([148.163.158.5]:4594 "EHLO mx0a-001b2d01.pphosted.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S1751903AbcFOMur (ORCPT ); Wed, 15 Jun 2016 08:50:47 -0400 X-IBM-Helo: d03dlp03.boulder.ibm.com X-IBM-MailFrom: ego@linux.vnet.ibm.com X-IBM-RcptTo: mpe@ellerman.id.au;htejun@gmail.com;peterz@infradead.org;tglx@linutronix.de;linuxppc-dev@lists.ozlabs.org;linux-kernel@vger.kernel.org Date: Wed, 15 Jun 2016 18:20:33 +0530 From: Gautham R Shenoy To: Peter Zijlstra Cc: Gautham R Shenoy , Thomas Gleixner , Tejun Heo , Michael Ellerman , Abdul Haleem , Aneesh Kumar , linuxppc-dev@lists.ozlabs.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH 2/2] workqueue:Fix affinity of an unbound worker of a node with 1 online CPU Reply-To: ego@linux.vnet.ibm.com References: <20160614112234.GF30154@twins.programming.kicks-ass.net> <20160615101936.GA31671@in.ibm.com> <20160615113249.GH30909@twins.programming.kicks-ass.net> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20160615113249.GH30909@twins.programming.kicks-ass.net> User-Agent: Mutt/1.5.23 (2014-03-12) X-TM-AS-GCONF: 00 X-Content-Scanned: Fidelis XPS MAILER x-cbid: 16061512-0008-0000-0000-000004D02F6E X-IBM-AV-DETECTION: SAVI=unused REMOTE=unused XFE=unused x-cbparentid: 16061512-0009-0000-0000-000038682B87 Message-Id: <20160615125033.GB31671@in.ibm.com> X-Proofpoint-Virus-Version: vendor=fsecure engine=2.50.10432:,, definitions=2016-06-15_07:,, signatures=0 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 suspectscore=0 malwarescore=0 phishscore=0 adultscore=0 bulkscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1604210000 definitions=main-1606150140 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Wed, Jun 15, 2016 at 01:32:49PM +0200, Peter Zijlstra wrote: > On Wed, Jun 15, 2016 at 03:49:36PM +0530, Gautham R Shenoy wrote: > > > Also, with the first patch in the series (which ensures that > > restore_unbound_workers are called *after* the new workers for the > > newly onlined CPUs are created) and without this one, you can > > reproduce this WARN_ON on both x86 and PPC by offlining all the CPUs > > of a node and bringing just one of them online. > > Ah good. > > > I am not sure about that. The workqueue creates unbound workers for a > > node via wq_update_unbound_numa() whenever the first CPU of every node > > comes online. So that seems legitimate. It then tries to affine these > > workers to the cpumask of that node. Again this seems right. As an > > optimization, it does this only when the first CPU of the node comes > > online. Since this online CPU is not yet active, and since > > nr_cpus_allowed > 1, we will hit the WARN_ON(). > > So I had another look and isn't the below a much simpler solution? > > It seems to work on my x86 with: > > for i in /sys/devices/system/cpu/cpu*/online ; do echo 0 > $i ; done > for i in /sys/devices/system/cpu/cpu*/online ; do echo 1 > $i ; done > > without complaint. Yup. This will work on PPC as well. We will no longer have the optimization in restore_unbound_workers_cpumask() but I suppose we don't lose much by resetting the affinity every time a CPU in the pool->attr->cpumask comes online. -- Thanks and Regards gautham.