From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1760239AbbJ3X07 (ORCPT ); Fri, 30 Oct 2015 19:26:59 -0400 Received: from g2t4622.austin.hp.com ([15.73.212.79]:49116 "EHLO g2t4622.austin.hp.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1759723AbbJ3X05 (ORCPT ); Fri, 30 Oct 2015 19:26:57 -0400 From: Waiman Long To: Peter Zijlstra , Ingo Molnar , Thomas Gleixner , "H. Peter Anvin" Cc: x86@kernel.org, linux-kernel@vger.kernel.org, Scott J Norton , Douglas Hatch , Davidlohr Bueso , Waiman Long Subject: [PATCH tip/locking/core v9 2/6] locking/qspinlock: prefetch next node cacheline Date: Fri, 30 Oct 2015 19:26:33 -0400 Message-Id: <1446247597-61863-3-git-send-email-Waiman.Long@hpe.com> X-Mailer: git-send-email 1.7.1 In-Reply-To: <1446247597-61863-1-git-send-email-Waiman.Long@hpe.com> References: <1446247597-61863-1-git-send-email-Waiman.Long@hpe.com> Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org A queue head CPU, after acquiring the lock, will have to notify the next CPU in the wait queue that it has became the new queue head. This involves loading a new cacheline from the MCS node of the next CPU. That operation can be expensive and add to the latency of locking operation. This patch addes code to optmistically prefetch the next MCS node cacheline if the next pointer is defined and it has been spinning for the MCS lock for a while. This reduces the locking latency and improves the system throughput. Using a locking microbenchmark on a Haswell-EX system, this patch can improve throughput by about 5%. Signed-off-by: Waiman Long --- kernel/locking/qspinlock.c | 21 +++++++++++++++++++++ 1 files changed, 21 insertions(+), 0 deletions(-) diff --git a/kernel/locking/qspinlock.c b/kernel/locking/qspinlock.c index 7868418..c1c8a1a 100644 --- a/kernel/locking/qspinlock.c +++ b/kernel/locking/qspinlock.c @@ -396,6 +396,7 @@ queue: * p,*,* -> n,*,* */ old = xchg_tail(lock, tail); + next = NULL; /* * if there was a previous node; link it and wait until reaching the @@ -407,6 +408,16 @@ queue: pv_wait_node(node); arch_mcs_spin_lock_contended(&node->locked); + + /* + * While waiting for the MCS lock, the next pointer may have + * been set by another lock waiter. We optimistically load + * the next pointer & prefetch the cacheline for writing + * to reduce latency in the upcoming MCS unlock operation. + */ + next = READ_ONCE(node->next); + if (next) + prefetchw(next); } /* @@ -426,6 +437,15 @@ queue: cpu_relax(); /* + * If the next pointer is defined, we are not tail anymore. + * In this case, claim the spinlock & release the MCS lock. + */ + if (next) { + set_locked(lock); + goto mcs_unlock; + } + + /* * claim the lock: * * n,0,0 -> 0,0,1 : lock, uncontended @@ -458,6 +478,7 @@ queue: while (!(next = READ_ONCE(node->next))) cpu_relax(); +mcs_unlock: arch_mcs_spin_unlock_contended(&next->locked); pv_kick_node(lock, next); -- 1.7.1