From: Waiman Long <longman@redhat.com>
To: Alex Kogan <alex.kogan@oracle.com>, "liwei (GF)" <liwei391@huawei.com>
Cc: linux@armlinux.org.uk, Peter Zijlstra <peterz@infradead.org>,
mingo@redhat.com, will.deacon@arm.com, arnd@arndb.de,
linux-arch@vger.kernel.org, linux-arm-kernel@lists.infradead.org,
linux-kernel@vger.kernel.org,
Thomas Gleixner <tglx@linutronix.de>,
bp@alien8.de, hpa@zytor.com, x86@kernel.org,
dave.dice@oracle.com, Rahul Yadav <rahul.x.yadav@oracle.com>,
Steven Sistare <steven.sistare@oracle.com>,
Daniel Jordan <daniel.m.jordan@oracle.com>
Subject: Re: [PATCH v2 3/5] locking/qspinlock: Introduce CNA into the slow path of qspinlock
Date: Wed, 12 Jun 2019 11:05:17 -0400 [thread overview]
Message-ID: <a52a5e25-2b71-b6d9-3fa1-fb43bae1cbc1@redhat.com> (raw)
In-Reply-To: <54241445-458C-4AE2-840B-6DFCCD410399@oracle.com>
On 6/12/19 12:38 AM, Alex Kogan wrote:
> Hi, Wei.
>
>> On Jun 11, 2019, at 12:22 AM, liwei (GF) <liwei391@huawei.com> wrote:
>>
>> Hi Alex,
>>
>> On 2019/3/29 23:20, Alex Kogan wrote:
>>> In CNA, spinning threads are organized in two queues, a main queue for
>>> threads running on the same node as the current lock holder, and a
>>> secondary queue for threads running on other nodes. At the unlock time,
>>> the lock holder scans the main queue looking for a thread running on
>>> the same node. If found (call it thread T), all threads in the main queue
>>> between the current lock holder and T are moved to the end of the
>>> secondary queue, and the lock is passed to T. If such T is not found, the
>>> lock is passed to the first node in the secondary queue. Finally, if the
>>> secondary queue is empty, the lock is passed to the next thread in the
>>> main queue. For more details, see https://urldefense.proofpoint.com/v2/url?u=https-3A__arxiv.org_abs_1810.05600&d=DwICbg&c=RoP1YumCXCgaWHvlZYR8PZh8Bv7qIrMUB65eapI_JnE&r=Hvhk3F4omdCk-GE1PTOm3Kn0A7ApWOZ2aZLTuVxFK4k&m=U7mfTbYj1r2Te2BBUUNbVrRPuTa_ujlpR4GZfUsrGTM&s=Dw4O1EniF-nde4fp6RA9ISlSMOjWuqeR9OS1G0iauj0&e=.
>>>
>>> Note that this variant of CNA may introduce starvation by continuously
>>> passing the lock to threads running on the same node. This issue
>>> will be addressed later in the series.
>>>
>>> Enabling CNA is controlled via a new configuration option
>>> (NUMA_AWARE_SPINLOCKS), which is enabled by default if NUMA is enabled.
>>>
>>> Signed-off-by: Alex Kogan <alex.kogan@oracle.com>
>>> Reviewed-by: Steve Sistare <steven.sistare@oracle.com>
>>> ---
>>> arch/x86/Kconfig | 14 +++
>>> include/asm-generic/qspinlock_types.h | 13 +++
>>> kernel/locking/mcs_spinlock.h | 10 ++
>>> kernel/locking/qspinlock.c | 29 +++++-
>>> kernel/locking/qspinlock_cna.h | 173 ++++++++++++++++++++++++++++++++++
>>> 5 files changed, 236 insertions(+), 3 deletions(-)
>>> create mode 100644 kernel/locking/qspinlock_cna.h
>>>
>> (SNIP)
>>> +
>>> +static __always_inline int get_node_index(struct mcs_spinlock *node)
>>> +{
>>> + return decode_count(node->node_and_count++);
>> When nesting level is > 4, it won't return a index >= 4 here and the numa node number
>> is changed by mistake. It will go into a wrong way instead of the following branch.
>>
>>
>> /*
>> * 4 nodes are allocated based on the assumption that there will
>> * not be nested NMIs taking spinlocks. That may not be true in
>> * some architectures even though the chance of needing more than
>> * 4 nodes will still be extremely unlikely. When that happens,
>> * we fall back to spinning on the lock directly without using
>> * any MCS node. This is not the most elegant solution, but is
>> * simple enough.
>> */
>> if (unlikely(idx >= MAX_NODES)) {
>> while (!queued_spin_trylock(lock))
>> cpu_relax();
>> goto release;
>> }
> Good point.
> This patch does not handle count overflows gracefully.
> It can be easily fixed by allocating more bits for the count — we don’t really need 30 bits for #NUMA nodes.
Actually, the default setting uses 2 bits for 4-level nesting and 14
bits for cpu numbers. That means it can support up to 16k-1 cpus. It is
a limit that is likely to be exceeded in the foreseeable future.
qspinlock also supports an additional mode with 21 bits used for cpu
numbers. That can support up to 2M-1 cpus. However, this mode will be a
little bit slower. That is why we don't want to use more than 2 bits for
nesting as I have never see more than 2 level of nesting used in my
testing. So it is highly unlikely we will ever hit more than 4 levels. I
am not saying that it is impossible, though.
Cheers,
Longman
next prev parent reply other threads:[~2019-06-12 15:05 UTC|newest]
Thread overview: 38+ messages / expand[flat|nested] mbox.gz Atom feed top
2019-03-29 15:20 [PATCH v2 0/5] Add NUMA-awareness to qspinlock Alex Kogan
2019-03-29 15:20 ` [PATCH v2 1/5] locking/qspinlock: Make arch_mcs_spin_unlock_contended more generic Alex Kogan
2019-03-29 15:20 ` [PATCH v2 2/5] locking/qspinlock: Refactor the qspinlock slow path Alex Kogan
2019-03-29 15:20 ` [PATCH v2 3/5] locking/qspinlock: Introduce CNA into the slow path of qspinlock Alex Kogan
2019-04-01 9:06 ` Peter Zijlstra
2019-04-01 9:33 ` Peter Zijlstra
2019-04-03 15:53 ` Alex Kogan
2019-04-03 16:10 ` Peter Zijlstra
2019-04-01 9:21 ` Peter Zijlstra
2019-04-01 14:36 ` Waiman Long
2019-04-02 9:43 ` Peter Zijlstra
2019-04-03 15:39 ` Alex Kogan
2019-04-03 15:48 ` Waiman Long
2019-04-03 16:01 ` Peter Zijlstra
2019-04-04 5:05 ` Juergen Gross
2019-04-04 9:38 ` Peter Zijlstra
2019-04-04 18:03 ` Waiman Long
2019-06-04 23:21 ` Alex Kogan
2019-06-05 20:40 ` Peter Zijlstra
2019-06-06 15:21 ` Alex Kogan
2019-06-06 15:32 ` Waiman Long
2019-06-06 15:42 ` Waiman Long
2019-04-03 16:33 ` Waiman Long
2019-04-03 17:16 ` Peter Zijlstra
2019-04-03 17:40 ` Waiman Long
2019-04-04 2:02 ` Hanjun Guo
2019-04-04 3:14 ` Alex Kogan
2019-06-11 4:22 ` liwei (GF)
2019-06-12 4:38 ` Alex Kogan
2019-06-12 15:05 ` Waiman Long [this message]
2019-03-29 15:20 ` [PATCH v2 4/5] locking/qspinlock: Introduce starvation avoidance into CNA Alex Kogan
2019-04-02 10:37 ` Peter Zijlstra
2019-04-03 17:06 ` Alex Kogan
2019-03-29 15:20 ` [PATCH v2 5/5] locking/qspinlock: Introduce the shuffle reduction optimization " Alex Kogan
2019-04-01 9:09 ` [PATCH v2 0/5] Add NUMA-awareness to qspinlock Peter Zijlstra
2019-04-03 17:13 ` Alex Kogan
2019-07-03 11:58 ` Jan Glauber
2019-07-12 8:12 ` Hanjun Guo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=a52a5e25-2b71-b6d9-3fa1-fb43bae1cbc1@redhat.com \
--to=longman@redhat.com \
--cc=alex.kogan@oracle.com \
--cc=arnd@arndb.de \
--cc=bp@alien8.de \
--cc=daniel.m.jordan@oracle.com \
--cc=dave.dice@oracle.com \
--cc=hpa@zytor.com \
--cc=linux-arch@vger.kernel.org \
--cc=linux-arm-kernel@lists.infradead.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux@armlinux.org.uk \
--cc=liwei391@huawei.com \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=rahul.x.yadav@oracle.com \
--cc=steven.sistare@oracle.com \
--cc=tglx@linutronix.de \
--cc=will.deacon@arm.com \
--cc=x86@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®