From: David Laight <David.Laight@ACULAB.COM>
To: 'Linus Torvalds' <torvalds@linux-foundation.org>
Cc: "linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>,
"peterz@infradead.org" <peterz@infradead.org>,
"longman@redhat.com" <longman@redhat.com>,
"mingo@redhat.com" <mingo@redhat.com>,
"will@kernel.org" <will@kernel.org>,
"boqun.feng@gmail.com" <boqun.feng@gmail.com>,
"xinhui.pan@linux.vnet.ibm.com" <xinhui.pan@linux.vnet.ibm.com>,
"virtualization@lists.linux-foundation.org"
<virtualization@lists.linux-foundation.org>,
Zeng Heng <zengheng4@huawei.com>
Subject: RE: [PATCH next 4/5] locking/osq_lock: Optimise per-cpu data accesses.
Date: Sun, 31 Dec 2023 11:56:05 +0000 [thread overview]
Message-ID: <5cff1ac6581142228886f78d54d836cf@AcuMS.aculab.com> (raw)
In-Reply-To: <CAHk-=wjsO=ODvUcwi=SPSzvsxW7Gj+3OU8q4CfHa+zMcivF6Bw@mail.gmail.com>
From: Linus Torvalds
> Sent: 30 December 2023 20:59
>
> On Sat, 30 Dec 2023 at 12:41, Linus Torvalds
> <torvalds@linux-foundation.org> wrote:
> >
> > UNTESTED patch to just do the "this_cpu_write()" parts attached.
> > Again, note how we do end up doing that this_cpu_ptr conversion later
> > anyway, but at least it's off the critical path.
>
> Also note that while 'this_cpu_ptr()' doesn't exactly generate lovely
> code, it really is still better than caching a value in memory.
>
> At least the memory location that 'this_cpu_ptr()' accesses is
> slightly more likely to be hot (and is right next to the cpu number,
> iirc).
I was only going to access the 'self' field in code that required
the 'node' cache line be present.
>
> That said, I think we should fix this_cpu_ptr() to not ever generate
> that disgusting cltq just because the cpu pointer has the wrong
> signedness. I don't quite know how to do it, but this:
>
> -#define per_cpu_offset(x) (__per_cpu_offset[x])
> +#define per_cpu_offset(x) (__per_cpu_offset[(unsigned)(x)])
>
> at least helps a *bit*. It gets rid of the cltq, at least, but if
> somebody actually passes in an 'unsigned long' cpuid, it would cause
> an unnecessary truncation.
Doing the conversion using arithmetic might help, so:
__per_cpu_offset[(x) + 0u]
> And gcc still generates
>
> subl $1, %eax #, cpu_nr
> addq __per_cpu_offset(,%rax,8), %rcx
>
> instead of just doing
>
> addq __per_cpu_offset-8(,%rax,8), %rcx
>
> because it still needs to clear the upper 32 bits and doesn't know
> that the 'xchg()' already did that.
Not only that, you need to do the 'subl' after converting to 64 bits.
Otherwise the wrong location is read were cpu_nr to be zero.
I've tried that - but it still failed.
> Oh well. I guess even without the -1/+1 games by the OSQ code, we
> would still end up with a "movl" just to do that upper bits clearing
> that the compiler doesn't know is unnecessary.
>
> I don't think we have any reasonable way to tell the compiler that the
> register output of our xchg() inline asm has the upper 32 bits clear.
It could be done for a 32bit unsigned xchg() - just make the return
type unsigned 64bit.
But that won't work for the signed exchange - and 'atomic_t' is signed.
OTOH I'd guess this code could use 'unsigned int' instead of atomic_t?
David
-
Registered Address Lakeside, Bramley Road, Mount Farm, Milton Keynes, MK1 1PT, UK
Registration No: 1397386 (Wales)
next prev parent reply other threads:[~2023-12-31 11:56 UTC|newest]
Thread overview: 31+ messages / expand[flat|nested] mbox.gz Atom feed top
2023-12-29 20:51 [PATCH next 0/5] locking/osq_lock: Optimisations to osq_lock code David Laight
2023-12-29 20:53 ` [PATCH next 1/5] locking/osq_lock: Move the definition of optimistic_spin_node into osf_lock.c David Laight
2023-12-30 1:59 ` Waiman Long
2023-12-29 20:54 ` [PATCH next 2/5] locking/osq_lock: Avoid dirtying the local cpu's 'node' in the osq_lock() fast path David Laight
2023-12-29 20:56 ` [PATCH next 3/5] locking/osq_lock: Clarify osq_wait_next() David Laight
2023-12-29 22:54 ` Linus Torvalds
2023-12-30 2:54 ` Waiman Long
2023-12-29 20:57 ` [PATCH next 4/5] locking/osq_lock: Optimise per-cpu data accesses David Laight
2023-12-30 3:08 ` Waiman Long
2023-12-30 11:09 ` Ingo Molnar
2023-12-30 11:35 ` David Laight
2023-12-31 3:04 ` Waiman Long
2023-12-31 10:36 ` David Laight
2023-12-30 20:37 ` Ingo Molnar
2023-12-30 22:47 ` David Laight
2023-12-30 20:41 ` Linus Torvalds
2023-12-30 20:59 ` Linus Torvalds
2023-12-31 11:56 ` David Laight [this message]
2023-12-31 11:41 ` David Laight
2023-12-29 20:58 ` [PATCH next 5/5] locking/osq_lock: Optimise vcpu_is_preempted() check David Laight
2023-12-30 3:13 ` Waiman Long
2023-12-30 15:57 ` Waiman Long
2023-12-30 22:37 ` David Laight
2023-12-29 22:11 ` [PATCH next 2/5] locking/osq_lock: Avoid dirtying the local cpu's 'node' in the osq_lock() fast path David Laight
2023-12-30 3:20 ` Waiman Long
2023-12-30 15:49 ` David Laight
2024-01-02 18:53 ` Boqun Feng
2024-01-02 23:32 ` David Laight
2023-12-30 19:40 ` [PATCH next 0/5] locking/osq_lock: Optimisations to osq_lock code Linus Torvalds
2023-12-30 22:39 ` David Laight
2023-12-31 2:14 ` Waiman Long
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=5cff1ac6581142228886f78d54d836cf@AcuMS.aculab.com \
--to=david.laight@aculab.com \
--cc=boqun.feng@gmail.com \
--cc=linux-kernel@vger.kernel.org \
--cc=longman@redhat.com \
--cc=mingo@redhat.com \
--cc=peterz@infradead.org \
--cc=torvalds@linux-foundation.org \
--cc=virtualization@lists.linux-foundation.org \
--cc=will@kernel.org \
--cc=xinhui.pan@linux.vnet.ibm.com \
--cc=zengheng4@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®