From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1753383AbcBOOTr (ORCPT ); Mon, 15 Feb 2016 09:19:47 -0500 Received: from [198.137.202.9] ([198.137.202.9]:50021 "EHLO bombadil.infradead.org" rhost-flags-FAIL-FAIL-OK-OK) by vger.kernel.org with ESMTP id S1753231AbcBOOTq (ORCPT ); Mon, 15 Feb 2016 09:19:46 -0500 Date: Mon, 15 Feb 2016 15:17:55 +0100 From: Peter Zijlstra To: Joonas Lahtinen Cc: Intel graphics driver community testing & development , Linux kernel development , Ingo Molnar , David Hildenbrand , "Paul E. McKenney" , "Gautham R. Shenoy" , Chris Wilson Subject: Re: [PATCH] [RFC] kernel/cpu: Use lockref for online CPU reference counting Message-ID: <20160215141755.GG6357@twins.programming.kicks-ass.net> References: <1455539803-13913-1-git-send-email-joonas.lahtinen@linux.intel.com> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <1455539803-13913-1-git-send-email-joonas.lahtinen@linux.intel.com> User-Agent: Mutt/1.5.21 (2012-12-30) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Mon, Feb 15, 2016 at 02:36:43PM +0200, Joonas Lahtinen wrote: > Instead of implementing a custom locked reference counting, use lockref. > > Current implementation leads to a deadlock splat on Intel SKL platforms > when lockdep debugging is enabled. > > This is due to few of CPUfreq drivers (including Intel P-state) having this; > policy->rwsem is locked during driver initialization and the functions called > during init that actually apply CPU limits use get_online_cpus (because they > have other calling paths too), which will briefly lock cpu_hotplug.lock to > increase cpu_hotplug.refcount. > > On later calling path, when doing a suspend, when cpu_hotplug_begin() is called > in disable_nonboot_cpus(), callbacks to CPUfreq functions get called after, > which will lock policy->rwsem and cpu_hotplug.lock is already held by > cpu_hotplug_begin() and we do have a potential deadlock scenario reported by > our CI system (though it is a very unlikely one). See the Bugzilla link for more > details. I've been meaning to change the thing into a percpu-rwsem, I just haven't had time to look into the lockdep splat that generated.