From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752338AbdFLO7S (ORCPT ); Mon, 12 Jun 2017 10:59:18 -0400 Received: from Galois.linutronix.de ([146.0.238.70]:33589 "EHLO Galois.linutronix.de" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751981AbdFLO7Q (ORCPT ); Mon, 12 Jun 2017 10:59:16 -0400 Date: Mon, 12 Jun 2017 16:59:13 +0200 (CEST) From: Thomas Gleixner To: John Allen cc: LKML , Nathan Fontenot , Michael Bringmann , Peter Zijlstra Subject: Re: [RFC] cpu/hotplug: Modify lock status before making cpu hotplug callbacks In-Reply-To: Message-ID: References: User-Agent: Alpine 2.20 (DEB 67 2015-01-07) MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Wed, 7 Jun 2017, John Allen wrote: > A deadlock has been observed during a cpu hot add on powerpc machines. > > The situation is as follows: > First, in _cpu_up, we take the cpu_hotplug lock and towards the end of _cpu_up, > we make the cpu hotplug callbacks, one of which is arch_update_cpu_topology. > For most other architectures, making this call while the parent thread has the > cpu_hotplug lock seems harmless as most implementations of > arch_update_cpu_topology appear to be quite simple. However, the powerpc > implementation is significantly more complex and requires us to make a call to > stop_machine which in turn attempts to take the cpu_hotplug lock in > get_online_cpus. We then deadlock as the parent thread is waiting on the child > to complete and the child is waiting on the parent to release the lock. > > This solution attempts to resolve the issue by incrementing the cpu_hotplug > refcount and releasing the cpu_hotplug mutex in _cpu_up so that the subsequent > call to get_online_cpus can take the mutex while other invocations of _cpu_up > are still prevented from executing as the refcount is non-zero. From my testing, > this seems to resolve the deadlock, but I'm not sure if there is a reason that > get_online_cpus *should* be locked out at this point. This is horrible, really. We have code pending in -next which reworks the hotplug lock to a per cpu reader writer semaphore, so this hack is not going to work anymore. See git://git.kernel.org/pub/scm/linux/kernel/git/tip/tip.git smp/hotplug This addresses the issue of calling stomp_machine() from inside a hotplug locked region, by introducing stop_machine_cpuslocked(). That should solve your problem nicely. Thanks, tglx