From: David Laight <david.laight.linux@gmail.com>
To: Kees Cook <kees@kernel.org>
Cc: jim.cromie@gmail.com, Andrew Morton <akpm@linux-foundation.org>,
Petr Mladek <pmladek@suse.com>,
Zhen Lei <thunder.leizhen@huawei.com>,
Luis Chamberlain <mcgrof@kernel.org>,
Andrey Grodzovsky <andrey.grodzovsky@crowdstrike.com>,
Steven Rostedt <rostedt@goodmis.org>,
Lorenzo Stoakes <ljs@kernel.org>,
Masahiro Yamada <masahiroy@kernel.org>,
Jiri Olsa <olsajiri@gmail.com>,
linux-kernel@vger.kernel.org, linux-kbuild@vger.kernel.org,
bpf@vger.kernel.org
Subject: Re: [PATCH v7 3/3] kallsyms: Unroll 24-bit sequence reconstruction in get_symbol_seq()
Date: Wed, 30 Sep 2026 10:13:07 +0100 [thread overview]
Message-ID: <20260930101307.4477e264@pumpkin> (raw)
In-Reply-To: <202609292329.87EE58BCF8@keescook>
On Tue, 29 Sep 2026 23:56:12 -0700
Kees Cook <kees@kernel.org> wrote:
> On Tue, Sep 29, 2026 at 12:07:32PM -0600, Jim Cromie via B4 Relay wrote:
> > Mark get_symbol_seq() as static inline and unroll the 3-byte extraction
> > into direct byte shifts: (p[0] << 16) | (p[1] << 8) | p[2]. This
> > eliminates loop induction variable maintenance and allows the compiler
> > to generate direct loads and constant shifts.
>
> So... did you actually see any difference in the output binary? This is
> such a small loop I'd expect every compiler to both inline and unroll
> it already. *time passes* I've checked; I was mostly right. :)
>
> It's already inlined and unrolled by the compilers, but GCC weirdly
> didn't see through the array math:
>
> At -O2 GCC (and Clang, mostly the same) on x86_64:
>
> before after
> lea (%rax,%rax,2),%eax lea (%rax,%rax,2),%edx
> mov %rax,%rdx movslq %edx,%rdx
> movzbl seqs(%rax),%eax movzbl seqs(%rdx),%eax
> lea 0x1(%rdx),%ecx movzbl seqs+2(%rdx),%ecx
> add $0x2,%edx movzbl seqs+1(%rdx),%edx
> movzbl seqs(%rcx),%ecx shl $0x10,%eax
> shl $0x8,%eax shl $0x8,%edx
> movzbl seqs(%rdx),%edx or %ecx,%eax
> or %ecx,%eax or %edx,%eax
> shl $0x8,%eax
> or %edx,%eax
>
> So it's actually the loss of the "+ i" part that does it.
A quick 'register dependency chain count' gives the old as 7 and the
new as 6.
The annoying 'movslq' might be removable by making the parameter
either 'unsigned int' or 'long' (and might go away if the function
in inlined).
>
> > -static unsigned int get_symbol_seq(int index)
> > +static inline unsigned int get_symbol_seq(int index)
> > {
> > - unsigned int i, seq = 0;
> > + const u8 *p = &kallsyms_seqs_of_names[3 * index];
> >
> > - for (i = 0; i < 3; i++)
> > - seq = (seq << 8) | kallsyms_seqs_of_names[3 * index + i];
> > -
> > - return seq;
> > + return (p[0] << 16) | (p[1] << 8) | p[2];
> > }
>
> I'm on the fence about readability, but I think it's improved.
Ditto.
There might be a get_unaligned_be24() but that doesn't help readability.
The code would be better if the value were host-endian.
gcc 16+ then uses a 16bit read for two of the bytes.
That also lets the table be defined as:
extern struct { unsigned int v:24; } kallsyms_seqs_of_names[];
so that accesses can just be kallsyms_seqs_of_names[n].v
(Although even gcc 16 seems to do three reads in that case.)
If the cpu supports misaligned accesses you can do a 32bit read
and an 8bit shift.
David
>
> Reviewed-by: Kees Cook <kees@kernel.org>
>
next prev parent reply other threads:[~2026-09-30 9:13 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-29 18:07 [PATCH v7 0/3] kallsyms: Accelerate symbol name lookups by ~7x Jim Cromie via B4 Relay
2026-09-29 18:07 ` [PATCH v7 1/3] kallsyms: Match compressed tokens on the fly during binary search Jim Cromie via B4 Relay
2026-09-30 1:17 ` bot+bpf-ci
2026-09-30 6:26 ` Kees Cook
2026-09-29 18:07 ` [PATCH v7 2/3] kallsyms: Increase marker density to 16:1 to accelerate lookups Jim Cromie via B4 Relay
2026-09-30 6:28 ` Kees Cook
2026-09-29 18:07 ` [PATCH v7 3/3] kallsyms: Unroll 24-bit sequence reconstruction in get_symbol_seq() Jim Cromie via B4 Relay
2026-09-30 6:56 ` Kees Cook
2026-09-30 9:13 ` David Laight [this message]
2026-09-29 18:45 ` [PATCH v7 0/3] kallsyms: Accelerate symbol name lookups by ~7x Andrew Morton
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20260930101307.4477e264@pumpkin \
--to=david.laight.linux@gmail.com \
--cc=akpm@linux-foundation.org \
--cc=andrey.grodzovsky@crowdstrike.com \
--cc=bpf@vger.kernel.org \
--cc=jim.cromie@gmail.com \
--cc=kees@kernel.org \
--cc=linux-kbuild@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=ljs@kernel.org \
--cc=masahiroy@kernel.org \
--cc=mcgrof@kernel.org \
--cc=olsajiri@gmail.com \
--cc=pmladek@suse.com \
--cc=rostedt@goodmis.org \
--cc=thunder.leizhen@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®