From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1751916AbdJDVcA (ORCPT ); Wed, 4 Oct 2017 17:32:00 -0400 Received: from mx0b-001b2d01.pphosted.com ([148.163.158.5]:54244 "EHLO mx0a-001b2d01.pphosted.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S1751890AbdJDV3m (ORCPT ); Wed, 4 Oct 2017 17:29:42 -0400 From: "Paul E. McKenney" To: linux-kernel@vger.kernel.org Cc: mingo@kernel.org, jiangshanlai@gmail.com, dipankar@in.ibm.com, akpm@linux-foundation.org, mathieu.desnoyers@efficios.com, josh@joshtriplett.org, tglx@linutronix.de, peterz@infradead.org, rostedt@goodmis.org, dhowells@redhat.com, edumazet@google.com, fweisbec@gmail.com, oleg@redhat.com, "Paul E. McKenney" Subject: [PATCH tip/core/rcu 1/9] rcu: Provide GP ordering in face of migrations and delays Date: Wed, 4 Oct 2017 14:29:27 -0700 X-Mailer: git-send-email 2.5.2 In-Reply-To: <20171004212915.GA10089@linux.vnet.ibm.com> References: <20171004212915.GA10089@linux.vnet.ibm.com> X-TM-AS-GCONF: 00 x-cbid: 17100421-0048-0000-0000-000001EFE255 X-IBM-SpamModules-Scores: X-IBM-SpamModules-Versions: BY=3.00007843; HX=3.00000241; KW=3.00000007; PH=3.00000004; SC=3.00000233; SDB=6.00926533; UDB=6.00466100; IPR=6.00706745; BA=6.00005620; NDR=6.00000001; ZLA=6.00000005; ZF=6.00000009; ZB=6.00000000; ZP=6.00000000; ZH=6.00000000; ZU=6.00000002; MB=3.00017394; XFM=3.00000015; UTC=2017-10-04 21:29:40 X-IBM-AV-DETECTION: SAVI=unused REMOTE=unused XFE=unused x-cbparentid: 17100421-0049-0000-0000-000042C4713F Message-Id: <1507152575-11055-1-git-send-email-paulmck@linux.vnet.ibm.com> X-Proofpoint-Virus-Version: vendor=fsecure engine=2.50.10432:,, definitions=2017-10-04_09:,, signatures=0 X-Proofpoint-Spam-Details: rule=outbound_notspam policy=outbound score=0 spamscore=0 suspectscore=1 malwarescore=0 phishscore=0 adultscore=0 bulkscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.0.1-1707230000 definitions=main-1710040299 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Consider the following admittedly improbable sequence of events: o RCU is initially idle. o Task A on CPU 0 executes rcu_read_lock(). o Task B on CPU 1 executes synchronize_rcu(), which must wait on Task A: o Task B registers the callback, which starts a new grace period, awakening the grace-period kthread on CPU 3, which immediately starts a new grace period. o Task B migrates to CPU 2, which provides a quiescent state for both CPUs 1 and 2. o Both CPUs 1 and 2 take scheduling-clock interrupts, and both invoke RCU_SOFTIRQ, both thus learning of the new grace period. o Task B is delayed, perhaps by vCPU preemption on CPU 2. o CPUs 2 and 3 pass through quiescent states, which are reported to core RCU. o Task B is resumed just long enough to be migrated to CPU 3, and then is once again delayed. o Task A executes rcu_read_unlock(), exiting its RCU read-side critical section. o CPU 0 passes through a quiescent sate, which is reported to core RCU. Only CPU 1 continues to block the grace period. o CPU 1 passes through a quiescent state, which is reported to core RCU. This ends the grace period, and CPU 1 therefore invokes its callbacks, one of which awakens Task B via complete(). o Task B resumes (still on CPU 3) and starts executing wait_for_completion(), which sees that the completion has already completed, and thus does not block. It returns from the synchronize_rcu() without any ordering against the end of Task A's RCU read-side critical section. It can therefore mess up Task A's RCU read-side critical section, in theory, anyway. However, if CPU hotplug ever gets rid of stop_machine(), there will be more straightforward ways for this sort of thing to happen, so this commit adds a memory barrier in order to enforce the needed ordering. Signed-off-by: Paul E. McKenney --- kernel/rcu/update.c | 10 ++++++++++ 1 file changed, 10 insertions(+) diff --git a/kernel/rcu/update.c b/kernel/rcu/update.c index 5033b66d2753..9e599fcdd7bf 100644 --- a/kernel/rcu/update.c +++ b/kernel/rcu/update.c @@ -413,6 +413,16 @@ void __wait_rcu_gp(bool checktiny, int n, call_rcu_func_t *crcu_array, wait_for_completion(&rs_array[i].completion); destroy_rcu_head_on_stack(&rs_array[i].head); } + + /* + * If we migrated after we registered a callback, but before the + * corresponding wait_for_completion(), we might now be running + * on a CPU that has not yet noticed that the corresponding grace + * period has ended. That CPU might not yet be fully ordered + * against the completion of the grace period, so the full memory + * barrier below enforces that ordering via the completion's state. + */ + smp_mb(); /* ^^^ */ } EXPORT_SYMBOL_GPL(__wait_rcu_gp); -- 2.5.2