From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S932190AbcAMJGv (ORCPT ); Wed, 13 Jan 2016 04:06:51 -0500 Received: from www.linutronix.de ([62.245.132.108]:46990 "EHLO Galois.linutronix.de" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S932067AbcAMJGr (ORCPT ); Wed, 13 Jan 2016 04:06:47 -0500 Date: Wed, 13 Jan 2016 10:05:49 +0100 (CET) From: Thomas Gleixner To: Sasha Levin cc: LKML , "Paul E. McKenney" , Peter Zijlstra Subject: Re: timers: HARDIRQ-safe -> HARDIRQ-unsafe lock order detected In-Reply-To: <56955C0F.1090005@oracle.com> Message-ID: References: <56955C0F.1090005@oracle.com> User-Agent: Alpine 2.11 (DEB 23 2013-08-11) MIME-Version: 1.0 Content-Type: TEXT/PLAIN; charset=US-ASCII X-Linutronix-Spam-Score: -1.0 X-Linutronix-Spam-Level: - X-Linutronix-Spam-Status: No , -1.0 points, 5.0 required, ALL_TRUSTED=-1,SHORTCIRCUIT=-0.0001 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Sasha, On Tue, 12 Jan 2016, Sasha Levin wrote: Cc'ing Paul, Peter > While fuzzing with trinity inside a KVM tools guest, running the latest -next > kernel, I've hit the following lockdep warning: > [ 3408.474461] Possible interrupt unsafe locking scenario: > > [ 3408.474461] > > [ 3408.475239] CPU0 CPU1 > > [ 3408.475809] ---- ---- > > [ 3408.476380] lock(&lock->wait_lock); > > [ 3408.476925] local_irq_disable(); > > [ 3408.477640] lock(&(&new_timer->it_lock)->rlock); > > [ 3408.478607] lock(&lock->wait_lock); That comes from rcu_read_unlock: rcu_read_unlock() rcu_read_unlock_special() ... rt_mutex_unlock(&rnp->boost_mtx); raw_spin_lock(&boost_mtx->wait_lock); > [ 3408.479445] > > [ 3408.479796] lock(&(&new_timer->it_lock)->rlock); So the task on CPU0 holds rnp->boost_mtx.wait_lock and then the interrupt deadlocks on the timer->it_lock. We can fix that particular issue in the posix-timer code by making the locking symetric: rcu_read_lock(); spin_lock_irq(timer->lock); ... spin_unlock_irq(timer->lock); rcu_read_unlock(); instead of: rcu_read_lock(); spin_lock_irq(timer->lock); rcu_read_unlock(); ... spin_unlock_irq(timer->lock); But the question is, whether this is the only offending code path in tree. We can avoid the hassle by making rtmutex->wait_lock irq safe. Thoughts? Thanks, tglx