From: David Laight <David.Laight@ACULAB.COM>
To: "linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"peterz@infradead.org" <peterz@infradead.org>,
"longman@redhat.com" <longman@redhat.com>
Cc: "mingo@redhat.com" <mingo@redhat.com>,
"will@kernel.org" <will@kernel.org>,
"boqun.feng@gmail.com" <boqun.feng@gmail.com>,
"'Linus Torvalds'" <torvalds@linux-foundation.org>,
"'xinhui.pan@linux.vnet.ibm.com'" <xinhui.pan@linux.vnet.ibm.com>,
"'virtualization@lists.linux-foundation.org'"
<virtualization@lists.linux-foundation.org>,
'Zeng Heng' <zengheng4@huawei.com>
Subject: [PATCH next 0/5] locking/osq_lock: Optimisations to osq_lock code
Date: Fri, 29 Dec 2023 20:51:46 +0000 [thread overview]
Message-ID: <73a4b31c9c874081baabad9e5f2e5204@AcuMS.aculab.com> (raw)
Zeng Heng noted that heavy use of the osq (optimistic spin queue) code
used rather more cpu than might be expected. See:
https://lore.kernel.org/lkml/202312210155.Wc2HUK8C-lkp@intel.com/T/#mcc46eedd1ef22a0d668828b1d088508c9b1875b8
Part of the problem is there is a pretty much guaranteed cache line reload
reading node->prev->cpu for the vcpu_is_preempted() check (even on bare metal)
in the wakeup path which slows it down.
(On bare metal the hypervisor call is patched out, but the argument is still read.)
Careful analysis shows that it isn't necessary to dirty the per-cpu data
on the fast-path osq_lock() path. This may be slightly beneficial.
The code also uses this_cpu_ptr() to get the address of the per-cpu data.
On x86-64 (at least) this is implemented as:
&per_cpu_data[smp_processor_id()]->member
ie there is a real function call, an array index and an add.
However if raw_cpu_read() can used then (which is typically just an offset
from register - %gs for x86-64) the code will be faster.
Putting the address of the per-cpu node into itself means that only one
cache line need be loaded.
I can't see a list of per-cpu data initialisation functions, so the fields
are initialised on the first osq_lock() call.
The last patch avoids the cache line reload calling vcpu_is_preempted()
by simply saving node->prev->cpu as node->prev_cpu and updating it when
node->prev changes.
This is simpler than the patch proposed by Waimon.
David Laight (5):
Move the definition of optimistic_spin_node into osf_lock.c
Avoid dirtying the local cpu's 'node' in the osq_lock() fast path.
Clarify osq_wait_next()
Optimise per-cpu data accesses.
Optimise vcpu_is_preempted() check.
include/linux/osq_lock.h | 5 ----
kernel/locking/osq_lock.c | 61 +++++++++++++++++++++------------------
2 files changed, 33 insertions(+), 33 deletions(-)
--
2.17.1
-
Registered Address Lakeside, Bramley Road, Mount Farm, Milton Keynes, MK1 1PT, UK
Registration No: 1397386 (Wales)
next reply other threads:[~2023-12-29 20:52 UTC|newest]
Thread overview: 31+ messages / expand[flat|nested] mbox.gz Atom feed top
2023-12-29 20:51 David Laight [this message]
2023-12-29 20:53 ` [PATCH next 1/5] locking/osq_lock: Move the definition of optimistic_spin_node into osf_lock.c David Laight
2023-12-30 1:59 ` Waiman Long
2023-12-29 20:54 ` [PATCH next 2/5] locking/osq_lock: Avoid dirtying the local cpu's 'node' in the osq_lock() fast path David Laight
2023-12-29 20:56 ` [PATCH next 3/5] locking/osq_lock: Clarify osq_wait_next() David Laight
2023-12-29 22:54 ` Linus Torvalds
2023-12-30 2:54 ` Waiman Long
2023-12-29 20:57 ` [PATCH next 4/5] locking/osq_lock: Optimise per-cpu data accesses David Laight
2023-12-30 3:08 ` Waiman Long
2023-12-30 11:09 ` Ingo Molnar
2023-12-30 11:35 ` David Laight
2023-12-31 3:04 ` Waiman Long
2023-12-31 10:36 ` David Laight
2023-12-30 20:37 ` Ingo Molnar
2023-12-30 22:47 ` David Laight
2023-12-30 20:41 ` Linus Torvalds
2023-12-30 20:59 ` Linus Torvalds
2023-12-31 11:56 ` David Laight
2023-12-31 11:41 ` David Laight
2023-12-29 20:58 ` [PATCH next 5/5] locking/osq_lock: Optimise vcpu_is_preempted() check David Laight
2023-12-30 3:13 ` Waiman Long
2023-12-30 15:57 ` Waiman Long
2023-12-30 22:37 ` David Laight
2023-12-29 22:11 ` [PATCH next 2/5] locking/osq_lock: Avoid dirtying the local cpu's 'node' in the osq_lock() fast path David Laight
2023-12-30 3:20 ` Waiman Long
2023-12-30 15:49 ` David Laight
2024-01-02 18:53 ` Boqun Feng
2024-01-02 23:32 ` David Laight
2023-12-30 19:40 ` [PATCH next 0/5] locking/osq_lock: Optimisations to osq_lock code Linus Torvalds
2023-12-30 22:39 ` David Laight
2023-12-31 2:14 ` Waiman Long
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=73a4b31c9c874081baabad9e5f2e5204@AcuMS.aculab.com \
--to=david.laight@aculab.com \
--cc=boqun.feng@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=longman@redhat.com \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=torvalds@linux-foundation.org \
--cc=virtualization@lists.linux-foundation.org \
--cc=will@kernel.org \
--cc=xinhui.pan@linux.vnet.ibm.com \
--cc=zengheng4@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®