* [PATCH v5 0/4] Switch get/put unaligned to use memcpy
@ 2025-10-16 20:51 Ian Rogers
2025-10-16 20:51 ` [PATCH v5 1/4] parisc: Inline a type punning version of get_unaligned_le32 Ian Rogers
` (3 more replies)
0 siblings, 4 replies; 23+ messages in thread
From: Ian Rogers @ 2025-10-16 20:51 UTC (permalink / raw)
To: James E.J. Bottomley, Helge Deller, Andy Lutomirski,
Thomas Gleixner, Vincenzo Frascino, Ian Rogers,
Arnaldo Carvalho de Melo, linux-parisc, linux-kernel,
Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld
The existing type punning approach with packed structs requires
-fno-strict-aliasing to be passed to the compiler for
correctness. This is true in the kernel tree but was not true in the
tools directory until this patch from Eric Biggers <ebiggers@google.com>:
https://lore.kernel.org/lkml/20250625202311.23244-2-ebiggers@kernel.org/
Requiring -fno-strict-aliasing seems unfortunate and so this patch
makes the unaligned code work via memcpy rather than type punning with
the packed attribute.
v5: add a patch to make parisc still use a punned version of
get_unaligned_le32 for an unusual boot case they have. This is
untested but suggested as necessary by:
https://lore.kernel.org/lkml/202509051042.7KOze0fZ-lkp@intel.com/
I wasn't clear if this work was picked up, but I don't see it in
v6.18-rc1 and so I'm resending rebased as v5.
v4: switch the type/expression variable __get_unaligned_ctrl_type that
is used by _Generic to be a pointer to avoid 0 vs NULL usage
warnings - always use NULL and dereference the type. This should
also hopefully address analysis bots complaints.
v3: switch to __unqual_scalar_typeof, reducing the code, and use an
uninitialized variable rather than a cast of 0 to try to avoid a
sparse warning about not using NULL. The code is trying to
navigate a minefield of uninitialized and casting warnings,
hopefully the best balance has been struck, but the code will fail
for cases like:
const void *val = get_unaligned((const void * const *)ptr);
due to __unqual_scalar_typeof leaving the 2nd const of the cast in
place. Thankfully no code does this - tested with an
allyesconfig. Support would be achievable by using void* as a
default case in __unqual_scalar_typeof, it just doesn't seem worth
it for a fairly unusual const case.
v2: switch memcpy to __builtin_memcpy to avoid potential/disallowed
memcpy calls in vdso caused by -fno-builtin. Reported by
Christophe Leroy <christophe.leroy@csgroup.eu>:
https://lore.kernel.org/lkml/c57de5bf-d55c-48c5-9dfa-e2fb844dafe9@csgroup.eu/
Ian Rogers (4):
parisc: Inline a type punning version of get_unaligned_le32
vdso: Switch get/put unaligned from packed struct to memcpy
tools headers: Update the linux/unaligned.h copy with the kernel
sources
tools headers: Remove unneeded ignoring of warnings in unaligned.h
arch/parisc/boot/compressed/misc.c | 15 +++++++++-
include/vdso/unaligned.h | 41 ++++++++++++++++++++++++----
tools/include/linux/compiler_types.h | 22 +++++++++++++++
tools/include/linux/unaligned.h | 4 ---
tools/include/vdso/unaligned.h | 41 ++++++++++++++++++++++++----
5 files changed, 106 insertions(+), 17 deletions(-)
--
2.51.0.858.gf9c4a03a3a-goog
^ permalink raw reply [flat|nested] 23+ messages in thread* [PATCH v5 1/4] parisc: Inline a type punning version of get_unaligned_le32 2025-10-16 20:51 [PATCH v5 0/4] Switch get/put unaligned to use memcpy Ian Rogers @ 2025-10-16 20:51 ` Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] parisc: Inline a type punning version of get_unaligned_le32() tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 2025-10-16 20:51 ` [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Ian Rogers ` (2 subsequent siblings) 3 siblings, 2 replies; 23+ messages in thread From: Ian Rogers @ 2025-10-16 20:51 UTC (permalink / raw) To: James E.J. Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Ian Rogers, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld Reading the byte/char output_len with get_unaligned_le32 can trigger compiler warnings due to the size read. Avoid these warnings by using type punning. This avoids issues when switching get_unaligned_t to __builtin_memcpy. Signed-off-by: Ian Rogers <irogers@google.com> --- arch/parisc/boot/compressed/misc.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-) diff --git a/arch/parisc/boot/compressed/misc.c b/arch/parisc/boot/compressed/misc.c index 9c83bd06ef15..111f267230a1 100644 --- a/arch/parisc/boot/compressed/misc.c +++ b/arch/parisc/boot/compressed/misc.c @@ -278,6 +278,19 @@ static void parse_elf(void *output) free(phdrs); } +/* + * The regular get_unaligned_le32 uses __builtin_memcpy which can trigger + * warnings when reading a byte/char output_len as an integer, as the size of a + * char is less than that of an integer. Use type punning and the packed + * attribute, which requires -fno-strict-aliasing, to work around the problem. + */ +static u32 punned_get_unaligned_le32(const void *p) +{ + const struct { __le32 x; } __packed * __get_pptr = p; + + return le32_to_cpu(__get_pptr->x); +} + asmlinkage unsigned long __visible decompress_kernel(unsigned int started_wide, unsigned int command_line, const unsigned int rd_start, @@ -309,7 +322,7 @@ asmlinkage unsigned long __visible decompress_kernel(unsigned int started_wide, * leave 2 MB for the stack. */ vmlinux_addr = (unsigned long) &_ebss + 2*1024*1024; - vmlinux_len = get_unaligned_le32(&output_len); + vmlinux_len = punned_get_unaligned_le32(&output_len); output = (char *) vmlinux_addr; /* -- 2.51.0.858.gf9c4a03a3a-goog ^ permalink raw reply [flat|nested] 23+ messages in thread
* [tip: timers/vdso] parisc: Inline a type punning version of get_unaligned_le32() 2025-10-16 20:51 ` [PATCH v5 1/4] parisc: Inline a type punning version of get_unaligned_le32 Ian Rogers @ 2026-01-13 13:47 ` tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 1 sibling, 0 replies; 23+ messages in thread From: tip-bot2 for Ian Rogers @ 2026-01-13 13:47 UTC (permalink / raw) To: linux-tip-commits; +Cc: Ian Rogers, Thomas Gleixner, x86, linux-kernel The following commit has been merged into the timers/vdso branch of tip: Commit-ID: feffe3a9d4b0d1672db4c16457f029f1c22b35da Gitweb: https://git.kernel.org/tip/feffe3a9d4b0d1672db4c16457f029f1c22b35da Author: Ian Rogers <irogers@google.com> AuthorDate: Thu, 16 Oct 2025 13:51:23 -07:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Tue, 13 Jan 2026 14:46:00 +01:00 parisc: Inline a type punning version of get_unaligned_le32() Reading the byte/char output_len with get_unaligned_le32() can trigger compiler warnings due to the size read. Avoid these warnings by using type punning. This avoids issues when switching get_unaligned_t() to __builtin_memcpy(). Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Link: https://patch.msgid.link/20251016205126.2882625-2-irogers@google.com --- arch/parisc/boot/compressed/misc.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-) diff --git a/arch/parisc/boot/compressed/misc.c b/arch/parisc/boot/compressed/misc.c index 9c83bd0..111f267 100644 --- a/arch/parisc/boot/compressed/misc.c +++ b/arch/parisc/boot/compressed/misc.c @@ -278,6 +278,19 @@ static void parse_elf(void *output) free(phdrs); } +/* + * The regular get_unaligned_le32 uses __builtin_memcpy which can trigger + * warnings when reading a byte/char output_len as an integer, as the size of a + * char is less than that of an integer. Use type punning and the packed + * attribute, which requires -fno-strict-aliasing, to work around the problem. + */ +static u32 punned_get_unaligned_le32(const void *p) +{ + const struct { __le32 x; } __packed * __get_pptr = p; + + return le32_to_cpu(__get_pptr->x); +} + asmlinkage unsigned long __visible decompress_kernel(unsigned int started_wide, unsigned int command_line, const unsigned int rd_start, @@ -309,7 +322,7 @@ asmlinkage unsigned long __visible decompress_kernel(unsigned int started_wide, * leave 2 MB for the stack. */ vmlinux_addr = (unsigned long) &_ebss + 2*1024*1024; - vmlinux_len = get_unaligned_le32(&output_len); + vmlinux_len = punned_get_unaligned_le32(&output_len); output = (char *) vmlinux_addr; /* ^ permalink raw reply [flat|nested] 23+ messages in thread
* [tip: timers/vdso] parisc: Inline a type punning version of get_unaligned_le32() 2025-10-16 20:51 ` [PATCH v5 1/4] parisc: Inline a type punning version of get_unaligned_le32 Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] parisc: Inline a type punning version of get_unaligned_le32() tip-bot2 for Ian Rogers @ 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 1 sibling, 0 replies; 23+ messages in thread From: tip-bot2 for Ian Rogers @ 2026-01-14 8:01 UTC (permalink / raw) To: linux-tip-commits; +Cc: Ian Rogers, Thomas Gleixner, x86, linux-kernel The following commit has been merged into the timers/vdso branch of tip: Commit-ID: df0f9a664be55a8529362a1ada847a19a91e4807 Gitweb: https://git.kernel.org/tip/df0f9a664be55a8529362a1ada847a19a91e4807 Author: Ian Rogers <irogers@google.com> AuthorDate: Thu, 16 Oct 2025 13:51:23 -07:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Wed, 14 Jan 2026 08:56:41 +01:00 parisc: Inline a type punning version of get_unaligned_le32() Reading the byte/char output_len with get_unaligned_le32() can trigger compiler warnings due to the size read. Avoid these warnings by using type punning. This avoids issues when switching get_unaligned_t() to __builtin_memcpy(). Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Link: https://patch.msgid.link/20251016205126.2882625-2-irogers@google.com --- arch/parisc/boot/compressed/misc.c | 15 ++++++++++++++- 1 file changed, 14 insertions(+), 1 deletion(-) diff --git a/arch/parisc/boot/compressed/misc.c b/arch/parisc/boot/compressed/misc.c index 9c83bd0..111f267 100644 --- a/arch/parisc/boot/compressed/misc.c +++ b/arch/parisc/boot/compressed/misc.c @@ -278,6 +278,19 @@ static void parse_elf(void *output) free(phdrs); } +/* + * The regular get_unaligned_le32 uses __builtin_memcpy which can trigger + * warnings when reading a byte/char output_len as an integer, as the size of a + * char is less than that of an integer. Use type punning and the packed + * attribute, which requires -fno-strict-aliasing, to work around the problem. + */ +static u32 punned_get_unaligned_le32(const void *p) +{ + const struct { __le32 x; } __packed * __get_pptr = p; + + return le32_to_cpu(__get_pptr->x); +} + asmlinkage unsigned long __visible decompress_kernel(unsigned int started_wide, unsigned int command_line, const unsigned int rd_start, @@ -309,7 +322,7 @@ asmlinkage unsigned long __visible decompress_kernel(unsigned int started_wide, * leave 2 MB for the stack. */ vmlinux_addr = (unsigned long) &_ebss + 2*1024*1024; - vmlinux_len = get_unaligned_le32(&output_len); + vmlinux_len = punned_get_unaligned_le32(&output_len); output = (char *) vmlinux_addr; /* ^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2025-10-16 20:51 [PATCH v5 0/4] Switch get/put unaligned to use memcpy Ian Rogers 2025-10-16 20:51 ` [PATCH v5 1/4] parisc: Inline a type punning version of get_unaligned_le32 Ian Rogers @ 2025-10-16 20:51 ` Ian Rogers 2025-10-19 17:24 ` David Laight ` (3 more replies) 2025-10-16 20:51 ` [PATCH v5 3/4] tools headers: Update the linux/unaligned.h copy with the kernel sources Ian Rogers 2025-10-16 20:51 ` [PATCH v5 4/4] tools headers: Remove unneeded ignoring of warnings in unaligned.h Ian Rogers 3 siblings, 4 replies; 23+ messages in thread From: Ian Rogers @ 2025-10-16 20:51 UTC (permalink / raw) To: James E.J. Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Ian Rogers, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld Type punning is necessary for get/put unaligned but the use of a packed struct violates strict aliasing rules, requiring -fno-strict-aliasing to be passed to the C compiler. Switch to using memcpy so that -fno-strict-aliasing isn't necessary. Signed-off-by: Ian Rogers <irogers@google.com> --- include/vdso/unaligned.h | 41 ++++++++++++++++++++++++++++++++++------ 1 file changed, 35 insertions(+), 6 deletions(-) diff --git a/include/vdso/unaligned.h b/include/vdso/unaligned.h index ff0c06b6513e..9076483c9fbb 100644 --- a/include/vdso/unaligned.h +++ b/include/vdso/unaligned.h @@ -2,14 +2,43 @@ #ifndef __VDSO_UNALIGNED_H #define __VDSO_UNALIGNED_H -#define __get_unaligned_t(type, ptr) ({ \ - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ - __get_pptr->x; \ +#include <linux/compiler_types.h> + +/** + * __get_unaligned_t - read an unaligned value from memory. + * @type: the type to load from the pointer. + * @ptr: the pointer to load from. + * + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior + * from approaches like type punning that require -fno-strict-aliasing in order + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a + * non-const type - you can't memcpy into a const type. The + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required + * expression rather than type, a pointer is used to avoid warnings about mixing + * the use of 0 and NULL. The void* cast silences ubsan warnings. + */ +#define __get_unaligned_t(type, ptr) ({ \ + type *__get_unaligned_ctrl_type __always_unused = NULL; \ + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ + sizeof(__get_unaligned_val)); \ + __get_unaligned_val; \ }) -#define __put_unaligned_t(type, val, ptr) do { \ - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ - __put_pptr->x = (val); \ +/** + * __put_unaligned_t - write an unaligned value to memory. + * @type: the type of the value to store. + * @val: the value to store. + * @ptr: the pointer to store to. + * + * Use memcpy to affect an unaligned type sized store avoiding undefined + * behavior from approaches like type punning that require -fno-strict-aliasing + * in order to be correct. The void* cast silences ubsan warnings. + */ +#define __put_unaligned_t(type, val, ptr) do { \ + type __put_unaligned_val = (val); \ + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ + sizeof(__put_unaligned_val)); \ } while (0) #endif /* __VDSO_UNALIGNED_H */ -- 2.51.0.858.gf9c4a03a3a-goog ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2025-10-16 20:51 ` [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Ian Rogers @ 2025-10-19 17:24 ` David Laight 2026-01-13 13:47 ` [tip: timers/vdso] vdso: Switch get/put_unaligned() from packed struct to memcpy() tip-bot2 for Ian Rogers ` (2 subsequent siblings) 3 siblings, 0 replies; 23+ messages in thread From: David Laight @ 2025-10-19 17:24 UTC (permalink / raw) To: Ian Rogers Cc: James E.J. Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld On Thu, 16 Oct 2025 13:51:24 -0700 Ian Rogers <irogers@google.com> wrote: > Type punning is necessary for get/put unaligned but the use of a > packed struct violates strict aliasing rules, requiring > -fno-strict-aliasing to be passed to the C compiler. Switch to using > memcpy so that -fno-strict-aliasing isn't necessary. Does the compiler always manage to optimise everything away? You really do need it to generate the code for a misaligned memory access. You might be better off removing the 'strict-aliasing' warning by 'laundering' the pointer through an integer type (probably long). David > > Signed-off-by: Ian Rogers <irogers@google.com> > --- > include/vdso/unaligned.h | 41 ++++++++++++++++++++++++++++++++++------ > 1 file changed, 35 insertions(+), 6 deletions(-) > > diff --git a/include/vdso/unaligned.h b/include/vdso/unaligned.h > index ff0c06b6513e..9076483c9fbb 100644 > --- a/include/vdso/unaligned.h > +++ b/include/vdso/unaligned.h > @@ -2,14 +2,43 @@ > #ifndef __VDSO_UNALIGNED_H > #define __VDSO_UNALIGNED_H > > -#define __get_unaligned_t(type, ptr) ({ \ > - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ > - __get_pptr->x; \ > +#include <linux/compiler_types.h> > + > +/** > + * __get_unaligned_t - read an unaligned value from memory. > + * @type: the type to load from the pointer. > + * @ptr: the pointer to load from. > + * > + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior > + * from approaches like type punning that require -fno-strict-aliasing in order > + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a > + * non-const type - you can't memcpy into a const type. The > + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required > + * expression rather than type, a pointer is used to avoid warnings about mixing > + * the use of 0 and NULL. The void* cast silences ubsan warnings. > + */ > +#define __get_unaligned_t(type, ptr) ({ \ > + type *__get_unaligned_ctrl_type __always_unused = NULL; \ > + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ > + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ > + sizeof(__get_unaligned_val)); \ > + __get_unaligned_val; \ > }) > > -#define __put_unaligned_t(type, val, ptr) do { \ > - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ > - __put_pptr->x = (val); \ > +/** > + * __put_unaligned_t - write an unaligned value to memory. > + * @type: the type of the value to store. > + * @val: the value to store. > + * @ptr: the pointer to store to. > + * > + * Use memcpy to affect an unaligned type sized store avoiding undefined > + * behavior from approaches like type punning that require -fno-strict-aliasing > + * in order to be correct. The void* cast silences ubsan warnings. > + */ > +#define __put_unaligned_t(type, val, ptr) do { \ > + type __put_unaligned_val = (val); \ > + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ > + sizeof(__put_unaligned_val)); \ > } while (0) > > #endif /* __VDSO_UNALIGNED_H */ ^ permalink raw reply [flat|nested] 23+ messages in thread
* [tip: timers/vdso] vdso: Switch get/put_unaligned() from packed struct to memcpy() 2025-10-16 20:51 ` [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Ian Rogers 2025-10-19 17:24 ` David Laight @ 2026-01-13 13:47 ` tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 2026-09-28 15:39 ` [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Stefan Kerkmann 3 siblings, 0 replies; 23+ messages in thread From: tip-bot2 for Ian Rogers @ 2026-01-13 13:47 UTC (permalink / raw) To: linux-tip-commits; +Cc: Ian Rogers, Thomas Gleixner, x86, linux-kernel The following commit has been merged into the timers/vdso branch of tip: Commit-ID: e04a494143bab7ea804fe1ebe286701ee8288e4a Gitweb: https://git.kernel.org/tip/e04a494143bab7ea804fe1ebe286701ee8288e4a Author: Ian Rogers <irogers@google.com> AuthorDate: Thu, 16 Oct 2025 13:51:24 -07:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Tue, 13 Jan 2026 14:46:00 +01:00 vdso: Switch get/put_unaligned() from packed struct to memcpy() Type punning is necessary for get/put_unaligned() but the use of a packed struct violates strict aliasing rules, requiring -fno-strict-aliasing to be passed to the C compiler. Switch to using memcpy() so that -fno-strict-aliasing isn't necessary. Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Link: https://patch.msgid.link/20251016205126.2882625-3-irogers@google.com --- include/vdso/unaligned.h | 41 +++++++++++++++++++++++++++++++++------ 1 file changed, 35 insertions(+), 6 deletions(-) diff --git a/include/vdso/unaligned.h b/include/vdso/unaligned.h index ff0c06b..9076483 100644 --- a/include/vdso/unaligned.h +++ b/include/vdso/unaligned.h @@ -2,14 +2,43 @@ #ifndef __VDSO_UNALIGNED_H #define __VDSO_UNALIGNED_H -#define __get_unaligned_t(type, ptr) ({ \ - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ - __get_pptr->x; \ +#include <linux/compiler_types.h> + +/** + * __get_unaligned_t - read an unaligned value from memory. + * @type: the type to load from the pointer. + * @ptr: the pointer to load from. + * + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior + * from approaches like type punning that require -fno-strict-aliasing in order + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a + * non-const type - you can't memcpy into a const type. The + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required + * expression rather than type, a pointer is used to avoid warnings about mixing + * the use of 0 and NULL. The void* cast silences ubsan warnings. + */ +#define __get_unaligned_t(type, ptr) ({ \ + type *__get_unaligned_ctrl_type __always_unused = NULL; \ + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ + sizeof(__get_unaligned_val)); \ + __get_unaligned_val; \ }) -#define __put_unaligned_t(type, val, ptr) do { \ - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ - __put_pptr->x = (val); \ +/** + * __put_unaligned_t - write an unaligned value to memory. + * @type: the type of the value to store. + * @val: the value to store. + * @ptr: the pointer to store to. + * + * Use memcpy to affect an unaligned type sized store avoiding undefined + * behavior from approaches like type punning that require -fno-strict-aliasing + * in order to be correct. The void* cast silences ubsan warnings. + */ +#define __put_unaligned_t(type, val, ptr) do { \ + type __put_unaligned_val = (val); \ + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ + sizeof(__put_unaligned_val)); \ } while (0) #endif /* __VDSO_UNALIGNED_H */ ^ permalink raw reply [flat|nested] 23+ messages in thread
* [tip: timers/vdso] vdso: Switch get/put_unaligned() from packed struct to memcpy() 2025-10-16 20:51 ` [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Ian Rogers 2025-10-19 17:24 ` David Laight 2026-01-13 13:47 ` [tip: timers/vdso] vdso: Switch get/put_unaligned() from packed struct to memcpy() tip-bot2 for Ian Rogers @ 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 2026-09-28 15:39 ` [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Stefan Kerkmann 3 siblings, 0 replies; 23+ messages in thread From: tip-bot2 for Ian Rogers @ 2026-01-14 8:01 UTC (permalink / raw) To: linux-tip-commits; +Cc: Ian Rogers, Thomas Gleixner, x86, linux-kernel The following commit has been merged into the timers/vdso branch of tip: Commit-ID: a339671db64b12bb02492557d2b0658811286277 Gitweb: https://git.kernel.org/tip/a339671db64b12bb02492557d2b0658811286277 Author: Ian Rogers <irogers@google.com> AuthorDate: Thu, 16 Oct 2025 13:51:24 -07:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Wed, 14 Jan 2026 08:56:41 +01:00 vdso: Switch get/put_unaligned() from packed struct to memcpy() Type punning is necessary for get/put_unaligned() but the use of a packed struct violates strict aliasing rules, requiring -fno-strict-aliasing to be passed to the C compiler. Switch to using memcpy() so that -fno-strict-aliasing isn't necessary. Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Link: https://patch.msgid.link/20251016205126.2882625-3-irogers@google.com --- include/vdso/unaligned.h | 41 +++++++++++++++++++++++++++++++++------ 1 file changed, 35 insertions(+), 6 deletions(-) diff --git a/include/vdso/unaligned.h b/include/vdso/unaligned.h index ff0c06b..9076483 100644 --- a/include/vdso/unaligned.h +++ b/include/vdso/unaligned.h @@ -2,14 +2,43 @@ #ifndef __VDSO_UNALIGNED_H #define __VDSO_UNALIGNED_H -#define __get_unaligned_t(type, ptr) ({ \ - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ - __get_pptr->x; \ +#include <linux/compiler_types.h> + +/** + * __get_unaligned_t - read an unaligned value from memory. + * @type: the type to load from the pointer. + * @ptr: the pointer to load from. + * + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior + * from approaches like type punning that require -fno-strict-aliasing in order + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a + * non-const type - you can't memcpy into a const type. The + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required + * expression rather than type, a pointer is used to avoid warnings about mixing + * the use of 0 and NULL. The void* cast silences ubsan warnings. + */ +#define __get_unaligned_t(type, ptr) ({ \ + type *__get_unaligned_ctrl_type __always_unused = NULL; \ + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ + sizeof(__get_unaligned_val)); \ + __get_unaligned_val; \ }) -#define __put_unaligned_t(type, val, ptr) do { \ - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ - __put_pptr->x = (val); \ +/** + * __put_unaligned_t - write an unaligned value to memory. + * @type: the type of the value to store. + * @val: the value to store. + * @ptr: the pointer to store to. + * + * Use memcpy to affect an unaligned type sized store avoiding undefined + * behavior from approaches like type punning that require -fno-strict-aliasing + * in order to be correct. The void* cast silences ubsan warnings. + */ +#define __put_unaligned_t(type, val, ptr) do { \ + type __put_unaligned_val = (val); \ + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ + sizeof(__put_unaligned_val)); \ } while (0) #endif /* __VDSO_UNALIGNED_H */ ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2025-10-16 20:51 ` [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Ian Rogers ` (2 preceding siblings ...) 2026-01-14 8:01 ` tip-bot2 for Ian Rogers @ 2026-09-28 15:39 ` Stefan Kerkmann 2026-09-28 15:58 ` Ian Rogers 2026-10-06 12:31 ` Marc Kleine-Budde 3 siblings, 2 replies; 23+ messages in thread From: Stefan Kerkmann @ 2026-09-28 15:39 UTC (permalink / raw) To: Ian Rogers, James E.J. Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld Hi Ian, On 10/16/25 22:51, Ian Rogers wrote: > Type punning is necessary for get/put unaligned but the use of a > packed struct violates strict aliasing rules, requiring > -fno-strict-aliasing to be passed to the C compiler. Switch to using > memcpy so that -fno-strict-aliasing isn't necessary. > > Signed-off-by: Ian Rogers <irogers@google.com> > --- > include/vdso/unaligned.h | 41 ++++++++++++++++++++++++++++++++++------ > 1 file changed, 35 insertions(+), 6 deletions(-) > > diff --git a/include/vdso/unaligned.h b/include/vdso/unaligned.h > index ff0c06b6513e..9076483c9fbb 100644 > --- a/include/vdso/unaligned.h > +++ b/include/vdso/unaligned.h > @@ -2,14 +2,43 @@ > #ifndef __VDSO_UNALIGNED_H > #define __VDSO_UNALIGNED_H > > -#define __get_unaligned_t(type, ptr) ({ \ > - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ > - __get_pptr->x; \ > +#include <linux/compiler_types.h> > + > +/** > + * __get_unaligned_t - read an unaligned value from memory. > + * @type: the type to load from the pointer. > + * @ptr: the pointer to load from. > + * > + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior > + * from approaches like type punning that require -fno-strict-aliasing in order > + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a > + * non-const type - you can't memcpy into a const type. The > + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required > + * expression rather than type, a pointer is used to avoid warnings about mixing > + * the use of 0 and NULL. The void* cast silences ubsan warnings. > + */ > +#define __get_unaligned_t(type, ptr) ({ \ > + type *__get_unaligned_ctrl_type __always_unused = NULL; \ > + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ > + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ > + sizeof(__get_unaligned_val)); \ > + __get_unaligned_val; \ > }) > > -#define __put_unaligned_t(type, val, ptr) do { \ > - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ > - __put_pptr->x = (val); \ > +/** > + * __put_unaligned_t - write an unaligned value to memory. > + * @type: the type of the value to store. > + * @val: the value to store. > + * @ptr: the pointer to store to. > + * > + * Use memcpy to affect an unaligned type sized store avoiding undefined > + * behavior from approaches like type punning that require -fno-strict-aliasing > + * in order to be correct. The void* cast silences ubsan warnings. > + */ > +#define __put_unaligned_t(type, val, ptr) do { \ > + type __put_unaligned_val = (val); \ > + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ > + sizeof(__put_unaligned_val)); \ > } while (0) > > #endif /* __VDSO_UNALIGNED_H */ commit a339671db64b ("vdso: Switch get/put_unaligned() from packed struct to memcpy()"), which landed in 7.0, causes a performance regression on an NXP i.MX25 (ARMv5TE) SoC. I found it while updating a client's board from 6.12 to 7.0. A fio 4k randwrite benchmark on a NAND storage with UBI and UBIFS filesystem was the only workload that showed a clear regression between those two versions, so I bisected with it: perf stat -e irq:irq_handler_entry --filter 'irq == 49' -a \ -- \ fio --name=rw \ --filename=/var/stat/testfile \ --size=8M \ --rw=randwrite \ --bs=4k \ --direct=0 \ --fsync=1 \ --numjobs=4 \ --group_reporting | kernel | irq_handler_entry | fio bw | | ------ | ----------------- | -------- | | 6.12 | 155783 | 401KiB/s | | 6.13 | 162317 | 395KiB/s | | 6.14 | 168741 | 392KiB/s | | 6.15 | 168556 | 401KiB/s | | 6.16 | 166090 | 403KiB/s | | 6.17 | 162541 | 385KiB/s | | 6.18 | 157527 | 386KiB/s | | 6.19 | 183675 | 381KiB/s | | 7.0 | 190372 | 297KiB/s | The bisect targeted the large drop between 6.19 and 7.0; the smaller 6.17 regression predates this commit and is unrelated. Reverting a339671db64b restores throughput to the 6.17 level (~385 KiB/s). The commit is still present in 7.3-rc5, and the same codegen problem reproduces there. Digging deeper, I built 7.3-rc5 with my config and GCC 16.2, with and without the commit, and compared the object files: 114 of them differ. As <vdso/unaligned.h> is included by <linux/unaligned.h>, every get/put_unaligned() call site depends on it transitively. GCC did not inline __builtin_memcpy() and turned it into a function call, e.g. in crypto/crc32c.c (__chksum_finup(), inlined into chksum_digest()): Without the commit: <chksum_digest>: str lr, [sp, #-0x4]! sub sp, sp, #12 str lr, [sp, #-0x4]! bl 0xc0 <chksum_digest+0xc> @ imm = #-0x8 R_ARM_CALL __gnu_mcount_nc ldr r0, [r0] str r3, [sp, #0x4] ldr r0, [r0, #0x20] bl 0xd0 <chksum_digest+0x1c> @ imm = #-0x8 R_ARM_CALL crc32c mvn r2, r0 mov r0, #0 ldr r3, [sp, #0x4] lsr r12, r2, #8 lsr r1, r2, #16 strb r2, [r3] lsr r2, r2, #24 strb r12, [r3, #0x1] strb r1, [r3, #0x2] strb r2, [r3, #0x3] add sp, sp, #12 ldr pc, [sp], #4 With the commit: <chksum_digest>: push {r4, lr} sub sp, sp, #8 str lr, [sp, #-0x4]! bl 0x124 <chksum_digest+0xc> @ imm = #-0x8 R_ARM_CALL __gnu_mcount_nc ldr r0, [r0] ldr r12, [pc, #0x54] @ 0x188 <chksum_digest+0x70> ldr r0, [r0, #0x20] mov r4, r3 ldr r12, [r12] str r12, [sp, #0x4] mov r12, #0 bl 0x144 <chksum_digest+0x2c> @ imm = #-0x8 R_ARM_CALL crc32c mvn r3, r0 mov r2, #4 mov r0, r4 mov r1, sp str r3, [sp] bl 0x15c <chksum_digest+0x44> @ imm = #-0x8 R_ARM_CALL memcpy ldr r3, [pc, #0x20] @ 0x188 <chksum_digest+0x70> ldr r2, [r3] ldr r3, [sp, #0x4] eors r2, r3, r2 mov r3, #0 bne 0x184 <chksum_digest+0x6c> @ imm = #0x8 mov r0, #0 add sp, sp, #8 pop {r4, pc} bl 0x184 <chksum_digest+0x6c> @ imm = #-0x8 R_ARM_CALL __stack_chk_fail 188: 00 00 00 00 .word 0x00000000 R_ARM_ABS32 __stack_chk_guard Is this an accepted trade-off? My understanding is that the kernel is always built with -fno-strict-aliasing, so the packed-struct type punning was well defined there, and the __packed annotation is what lets GCC generate valid code for the unaligned access. Best regards, Stefan -- Pengutronix e.K. | Stefan Kerkmann | Steuerwalder Str. 21 | https://www.pengutronix.de/ | 31137 Hildesheim, Germany | Phone: +49-5121-206917-128 | Amtsgericht Hildesheim, HRA 2686 | Fax: +49-5121-206917-9 | ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2026-09-28 15:39 ` [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Stefan Kerkmann @ 2026-09-28 15:58 ` Ian Rogers 2026-10-06 12:31 ` Marc Kleine-Budde 1 sibling, 0 replies; 23+ messages in thread From: Ian Rogers @ 2026-09-28 15:58 UTC (permalink / raw) To: Stefan Kerkmann Cc: James E.J. Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld On Mon, Sep 28, 2026 at 8:39 AM Stefan Kerkmann <s.kerkmann@pengutronix.de> wrote: > > Hi Ian, > > On 10/16/25 22:51, Ian Rogers wrote: > > Type punning is necessary for get/put unaligned but the use of a > > packed struct violates strict aliasing rules, requiring > > -fno-strict-aliasing to be passed to the C compiler. Switch to using > > memcpy so that -fno-strict-aliasing isn't necessary. > > > > Signed-off-by: Ian Rogers <irogers@google.com> > > --- > > include/vdso/unaligned.h | 41 ++++++++++++++++++++++++++++++++++------ > > 1 file changed, 35 insertions(+), 6 deletions(-) > > > > diff --git a/include/vdso/unaligned.h b/include/vdso/unaligned.h > > index ff0c06b6513e..9076483c9fbb 100644 > > --- a/include/vdso/unaligned.h > > +++ b/include/vdso/unaligned.h > > @@ -2,14 +2,43 @@ > > #ifndef __VDSO_UNALIGNED_H > > #define __VDSO_UNALIGNED_H > > > > -#define __get_unaligned_t(type, ptr) ({ \ > > - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ > > - __get_pptr->x; \ > > +#include <linux/compiler_types.h> > > + > > +/** > > + * __get_unaligned_t - read an unaligned value from memory. > > + * @type: the type to load from the pointer. > > + * @ptr: the pointer to load from. > > + * > > + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior > > + * from approaches like type punning that require -fno-strict-aliasing in order > > + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a > > + * non-const type - you can't memcpy into a const type. The > > + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required > > + * expression rather than type, a pointer is used to avoid warnings about mixing > > + * the use of 0 and NULL. The void* cast silences ubsan warnings. > > + */ > > +#define __get_unaligned_t(type, ptr) ({ \ > > + type *__get_unaligned_ctrl_type __always_unused = NULL; \ > > + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ > > + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ > > + sizeof(__get_unaligned_val)); \ > > + __get_unaligned_val; \ > > }) > > > > -#define __put_unaligned_t(type, val, ptr) do { \ > > - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ > > - __put_pptr->x = (val); \ > > +/** > > + * __put_unaligned_t - write an unaligned value to memory. > > + * @type: the type of the value to store. > > + * @val: the value to store. > > + * @ptr: the pointer to store to. > > + * > > + * Use memcpy to affect an unaligned type sized store avoiding undefined > > + * behavior from approaches like type punning that require -fno-strict-aliasing > > + * in order to be correct. The void* cast silences ubsan warnings. > > + */ > > +#define __put_unaligned_t(type, val, ptr) do { \ > > + type __put_unaligned_val = (val); \ > > + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ > > + sizeof(__put_unaligned_val)); \ > > } while (0) > > > > #endif /* __VDSO_UNALIGNED_H */ > > commit a339671db64b ("vdso: Switch get/put_unaligned() from packed struct to > memcpy()"), which landed in 7.0, causes a performance regression on an NXP > i.MX25 (ARMv5TE) SoC. > > I found it while updating a client's board from 6.12 to 7.0. A fio 4k randwrite > benchmark on a NAND storage with UBI and UBIFS filesystem was the only workload > that showed a clear regression between those two versions, so I bisected with > it: > > perf stat -e irq:irq_handler_entry --filter 'irq == 49' -a \ > -- \ > fio --name=rw \ > --filename=/var/stat/testfile \ > --size=8M \ > --rw=randwrite \ > --bs=4k \ > --direct=0 \ > --fsync=1 \ > --numjobs=4 \ > --group_reporting > > | kernel | irq_handler_entry | fio bw | > | ------ | ----------------- | -------- | > | 6.12 | 155783 | 401KiB/s | > | 6.13 | 162317 | 395KiB/s | > | 6.14 | 168741 | 392KiB/s | > | 6.15 | 168556 | 401KiB/s | > | 6.16 | 166090 | 403KiB/s | > | 6.17 | 162541 | 385KiB/s | > | 6.18 | 157527 | 386KiB/s | > | 6.19 | 183675 | 381KiB/s | > | 7.0 | 190372 | 297KiB/s | > > The bisect targeted the large drop between 6.19 and 7.0; the smaller 6.17 > regression predates this commit and is unrelated. Reverting a339671db64b > restores throughput to the 6.17 level (~385 KiB/s). The commit is still > present in 7.3-rc5, and the same codegen problem reproduces there. > > Digging deeper, I built 7.3-rc5 with my config and GCC 16.2, with and without > the commit, and compared the object files: 114 of them differ. As > <vdso/unaligned.h> is included by <linux/unaligned.h>, every > get/put_unaligned() call site depends on it transitively. GCC did not inline > __builtin_memcpy() and turned it into a function call, e.g. in crypto/crc32c.c > (__chksum_finup(), inlined into chksum_digest()): > > Without the commit: > > <chksum_digest>: > str lr, [sp, #-0x4]! > sub sp, sp, #12 > str lr, [sp, #-0x4]! > bl 0xc0 <chksum_digest+0xc> @ imm = #-0x8 > R_ARM_CALL __gnu_mcount_nc > ldr r0, [r0] > str r3, [sp, #0x4] > ldr r0, [r0, #0x20] > bl 0xd0 <chksum_digest+0x1c> @ imm = #-0x8 > R_ARM_CALL crc32c > mvn r2, r0 > mov r0, #0 > ldr r3, [sp, #0x4] > lsr r12, r2, #8 > lsr r1, r2, #16 > strb r2, [r3] > lsr r2, r2, #24 > strb r12, [r3, #0x1] > strb r1, [r3, #0x2] > strb r2, [r3, #0x3] > add sp, sp, #12 > ldr pc, [sp], #4 > > With the commit: > > <chksum_digest>: > push {r4, lr} > sub sp, sp, #8 > str lr, [sp, #-0x4]! > bl 0x124 <chksum_digest+0xc> @ imm = #-0x8 > R_ARM_CALL __gnu_mcount_nc > ldr r0, [r0] > ldr r12, [pc, #0x54] @ 0x188 <chksum_digest+0x70> > ldr r0, [r0, #0x20] > mov r4, r3 > ldr r12, [r12] > str r12, [sp, #0x4] > mov r12, #0 > bl 0x144 <chksum_digest+0x2c> @ imm = #-0x8 > R_ARM_CALL crc32c > mvn r3, r0 > mov r2, #4 > mov r0, r4 > mov r1, sp > str r3, [sp] > bl 0x15c <chksum_digest+0x44> @ imm = #-0x8 > R_ARM_CALL memcpy > ldr r3, [pc, #0x20] @ 0x188 <chksum_digest+0x70> > ldr r2, [r3] > ldr r3, [sp, #0x4] > eors r2, r3, r2 > mov r3, #0 > bne 0x184 <chksum_digest+0x6c> @ imm = #0x8 > mov r0, #0 > add sp, sp, #8 > pop {r4, pc} > bl 0x184 <chksum_digest+0x6c> @ imm = #-0x8 > R_ARM_CALL __stack_chk_fail > 188: 00 00 00 00 .word 0x00000000 > R_ARM_ABS32 __stack_chk_guard > > Is this an accepted trade-off? My understanding is that the kernel is always > built with -fno-strict-aliasing, so the packed-struct type punning was well > defined there, and the __packed annotation is what lets GCC generate valid > code for the unaligned access. Hi Stefan, and sorry for the performance regression! I largely work on the perf tool which is in the tools/ directory but uses kernel header files, such as for unaligned accesses. My work was motivated by trying to avoid using -fno-strict-aliasing in the perf tool. As the compiler is given __builtin_memcpy of a fixed memory size then the lowering should match that of using the packed struct. Are there compiler flags you are using that disable compiler optimizations? This feels like a compiler bug, a workaround is to use the packed struct helpers in linux/unaligned/packed_struct.h Thanks, Ian > Best regards, > Stefan > > -- > Pengutronix e.K. | Stefan Kerkmann | > Steuerwalder Str. 21 | https://www.pengutronix.de/ | > 31137 Hildesheim, Germany | Phone: +49-5121-206917-128 | > Amtsgericht Hildesheim, HRA 2686 | Fax: +49-5121-206917-9 | > ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2026-09-28 15:39 ` [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Stefan Kerkmann 2026-09-28 15:58 ` Ian Rogers @ 2026-10-06 12:31 ` Marc Kleine-Budde 2026-10-07 10:59 ` Arnd Bergmann 1 sibling, 1 reply; 23+ messages in thread From: Marc Kleine-Budde @ 2026-10-06 12:31 UTC (permalink / raw) To: Stefan Kerkmann, arnd Cc: Ian Rogers, James E.J. Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld [-- Attachment #1: Type: text/plain, Size: 8561 bytes --] Cc+ Arnd On 28.09.2026 17:39:15, Stefan Kerkmann wrote: > Hi Ian, > > On 10/16/25 22:51, Ian Rogers wrote: > > Type punning is necessary for get/put unaligned but the use of a > > packed struct violates strict aliasing rules, requiring > > -fno-strict-aliasing to be passed to the C compiler. Switch to using > > memcpy so that -fno-strict-aliasing isn't necessary. > > > > Signed-off-by: Ian Rogers <irogers@google.com> > > --- > > include/vdso/unaligned.h | 41 ++++++++++++++++++++++++++++++++++------ > > 1 file changed, 35 insertions(+), 6 deletions(-) > > > > diff --git a/include/vdso/unaligned.h b/include/vdso/unaligned.h > > index ff0c06b6513e..9076483c9fbb 100644 > > --- a/include/vdso/unaligned.h > > +++ b/include/vdso/unaligned.h > > @@ -2,14 +2,43 @@ > > #ifndef __VDSO_UNALIGNED_H > > #define __VDSO_UNALIGNED_H > > > > -#define __get_unaligned_t(type, ptr) ({ \ > > - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ > > - __get_pptr->x; \ > > +#include <linux/compiler_types.h> > > + > > +/** > > + * __get_unaligned_t - read an unaligned value from memory. > > + * @type: the type to load from the pointer. > > + * @ptr: the pointer to load from. > > + * > > + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior > > + * from approaches like type punning that require -fno-strict-aliasing in order > > + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a > > + * non-const type - you can't memcpy into a const type. The > > + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required > > + * expression rather than type, a pointer is used to avoid warnings about mixing > > + * the use of 0 and NULL. The void* cast silences ubsan warnings. > > + */ > > +#define __get_unaligned_t(type, ptr) ({ \ > > + type *__get_unaligned_ctrl_type __always_unused = NULL; \ > > + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ > > + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ > > + sizeof(__get_unaligned_val)); \ > > + __get_unaligned_val; \ > > }) > > > > -#define __put_unaligned_t(type, val, ptr) do { \ > > - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ > > - __put_pptr->x = (val); \ > > +/** > > + * __put_unaligned_t - write an unaligned value to memory. > > + * @type: the type of the value to store. > > + * @val: the value to store. > > + * @ptr: the pointer to store to. > > + * > > + * Use memcpy to affect an unaligned type sized store avoiding undefined > > + * behavior from approaches like type punning that require -fno-strict-aliasing > > + * in order to be correct. The void* cast silences ubsan warnings. > > + */ > > +#define __put_unaligned_t(type, val, ptr) do { \ > > + type __put_unaligned_val = (val); \ > > + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ > > + sizeof(__put_unaligned_val)); \ > > } while (0) > > > > #endif /* __VDSO_UNALIGNED_H */ > > commit a339671db64b ("vdso: Switch get/put_unaligned() from packed struct to > memcpy()"), which landed in 7.0, causes a performance regression on an NXP > i.MX25 (ARMv5TE) SoC. > > I found it while updating a client's board from 6.12 to 7.0. A fio 4k randwrite > benchmark on a NAND storage with UBI and UBIFS filesystem was the only workload > that showed a clear regression between those two versions, so I bisected with > it: > > perf stat -e irq:irq_handler_entry --filter 'irq == 49' -a \ > -- \ > fio --name=rw \ > --filename=/var/stat/testfile \ > --size=8M \ > --rw=randwrite \ > --bs=4k \ > --direct=0 \ > --fsync=1 \ > --numjobs=4 \ > --group_reporting > > | kernel | irq_handler_entry | fio bw | > | ------ | ----------------- | -------- | > | 6.12 | 155783 | 401KiB/s | > | 6.13 | 162317 | 395KiB/s | > | 6.14 | 168741 | 392KiB/s | > | 6.15 | 168556 | 401KiB/s | > | 6.16 | 166090 | 403KiB/s | > | 6.17 | 162541 | 385KiB/s | > | 6.18 | 157527 | 386KiB/s | > | 6.19 | 183675 | 381KiB/s | > | 7.0 | 190372 | 297KiB/s | > > The bisect targeted the large drop between 6.19 and 7.0; the smaller 6.17 > regression predates this commit and is unrelated. Reverting a339671db64b > restores throughput to the 6.17 level (~385 KiB/s). The commit is still > present in 7.3-rc5, and the same codegen problem reproduces there. > > Digging deeper, I built 7.3-rc5 with my config and GCC 16.2, with and without > the commit, and compared the object files: 114 of them differ. As > <vdso/unaligned.h> is included by <linux/unaligned.h>, every > get/put_unaligned() call site depends on it transitively. GCC did not inline > __builtin_memcpy() and turned it into a function call, e.g. in crypto/crc32c.c > (__chksum_finup(), inlined into chksum_digest()): > > Without the commit: > > <chksum_digest>: > str lr, [sp, #-0x4]! > sub sp, sp, #12 > str lr, [sp, #-0x4]! > bl 0xc0 <chksum_digest+0xc> @ imm = #-0x8 > R_ARM_CALL __gnu_mcount_nc > ldr r0, [r0] > str r3, [sp, #0x4] > ldr r0, [r0, #0x20] > bl 0xd0 <chksum_digest+0x1c> @ imm = #-0x8 > R_ARM_CALL crc32c > mvn r2, r0 > mov r0, #0 > ldr r3, [sp, #0x4] > lsr r12, r2, #8 > lsr r1, r2, #16 > strb r2, [r3] > lsr r2, r2, #24 > strb r12, [r3, #0x1] > strb r1, [r3, #0x2] > strb r2, [r3, #0x3] > add sp, sp, #12 > ldr pc, [sp], #4 > > With the commit: > > <chksum_digest>: > push {r4, lr} > sub sp, sp, #8 > str lr, [sp, #-0x4]! > bl 0x124 <chksum_digest+0xc> @ imm = #-0x8 > R_ARM_CALL __gnu_mcount_nc > ldr r0, [r0] > ldr r12, [pc, #0x54] @ 0x188 <chksum_digest+0x70> > ldr r0, [r0, #0x20] > mov r4, r3 > ldr r12, [r12] > str r12, [sp, #0x4] > mov r12, #0 > bl 0x144 <chksum_digest+0x2c> @ imm = #-0x8 > R_ARM_CALL crc32c > mvn r3, r0 > mov r2, #4 > mov r0, r4 > mov r1, sp > str r3, [sp] > bl 0x15c <chksum_digest+0x44> @ imm = #-0x8 > R_ARM_CALL memcpy > ldr r3, [pc, #0x20] @ 0x188 <chksum_digest+0x70> > ldr r2, [r3] > ldr r3, [sp, #0x4] > eors r2, r3, r2 > mov r3, #0 > bne 0x184 <chksum_digest+0x6c> @ imm = #0x8 > mov r0, #0 > add sp, sp, #8 > pop {r4, pc} > bl 0x184 <chksum_digest+0x6c> @ imm = #-0x8 > R_ARM_CALL __stack_chk_fail > 188: 00 00 00 00 .word 0x00000000 > R_ARM_ABS32 __stack_chk_guard > > Is this an accepted trade-off? My understanding is that the kernel is always > built with -fno-strict-aliasing, so the packed-struct type punning was well > defined there, and the __packed annotation is what lets GCC generate valid > code for the unaligned access. > > Best regards, > Stefan > > -- > Pengutronix e.K. | Stefan Kerkmann | > Steuerwalder Str. 21 | https://www.pengutronix.de/ | > 31137 Hildesheim, Germany | Phone: +49-5121-206917-128 | > Amtsgericht Hildesheim, HRA 2686 | Fax: +49-5121-206917-9 | > > -- Pengutronix e.K. | Marc Kleine-Budde | Embedded Linux | https://www.pengutronix.de | Vertretung Nürnberg | Phone: +49-5121-206917-129 | Amtsgericht Hildesheim, HRA 2686 | Fax: +49-5121-206917-9 | [-- Attachment #2: signature.asc --] [-- Type: application/pgp-signature, Size: 228 bytes --] ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2026-10-06 12:31 ` Marc Kleine-Budde @ 2026-10-07 10:59 ` Arnd Bergmann 2026-10-09 21:28 ` Ian Rogers 0 siblings, 1 reply; 23+ messages in thread From: Arnd Bergmann @ 2026-10-07 10:59 UTC (permalink / raw) To: Marc Kleine-Budde, Stefan Kerkmann Cc: Ian Rogers, James E . J . Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Alexander Viro, Christophe Leroy, Jason A . Donenfeld On Tue, Oct 6, 2026, at 14:31, Marc Kleine-Budde wrote: > Cc+ Arnd Thanks for the Cc! > On 28.09.2026 17:39:15, Stefan Kerkmann wrote: >> Hi Ian, >> >> On 10/16/25 22:51, Ian Rogers wrote: >> > Type punning is necessary for get/put unaligned but the use of a >> > packed struct violates strict aliasing rules, requiring >> > -fno-strict-aliasing to be passed to the C compiler. Switch to using >> > memcpy so that -fno-strict-aliasing isn't necessary. >> > >> > Signed-off-by: Ian Rogers <irogers@google.com> >> > --- > >> Is this an accepted trade-off? My understanding is that the kernel is always >> built with -fno-strict-aliasing, so the packed-struct type punning was well >> defined there, and the __packed annotation is what lets GCC generate valid >> code for the unaligned access. The previous upstream version was the result of endless discussions, and it looks like changing it to the memcpy version was premature. At the time we unified all architectures to use a common implentation, this was the only one that resulted in correct and fast code on all architectures, so I don't understand why this was just applied without including everyone who was involved in coming up with the version that was replaced. My feeling is that we should just revert this. I'm not sure about the motivation for the change. It sounds like this was meant to be used in userland code, and that clashed with assumptions we make in the kernel, but I don't think that is sufficient reason for regressing kernel code. Arnd ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2026-10-07 10:59 ` Arnd Bergmann @ 2026-10-09 21:28 ` Ian Rogers 2026-10-10 10:26 ` Christophe Leroy (CS GROUP) 2026-10-10 14:06 ` Arnd Bergmann 0 siblings, 2 replies; 23+ messages in thread From: Ian Rogers @ 2026-10-09 21:28 UTC (permalink / raw) To: Arnd Bergmann, Thomas Gleixner Cc: Marc Kleine-Budde, Stefan Kerkmann, James E . J . Bottomley, Helge Deller, Andy Lutomirski, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Alexander Viro, Christophe Leroy, Jason A . Donenfeld On Wed, Oct 7, 2026 at 4:00 AM Arnd Bergmann <arnd@arndb.de> wrote: > > On Tue, Oct 6, 2026, at 14:31, Marc Kleine-Budde wrote: > > Cc+ Arnd > > Thanks for the Cc! > > > On 28.09.2026 17:39:15, Stefan Kerkmann wrote: > >> Hi Ian, > >> > >> On 10/16/25 22:51, Ian Rogers wrote: > >> > Type punning is necessary for get/put unaligned but the use of a > >> > packed struct violates strict aliasing rules, requiring > >> > -fno-strict-aliasing to be passed to the C compiler. Switch to using > >> > memcpy so that -fno-strict-aliasing isn't necessary. > >> > > >> > Signed-off-by: Ian Rogers <irogers@google.com> > >> > --- > > > >> Is this an accepted trade-off? My understanding is that the kernel is always > >> built with -fno-strict-aliasing, so the packed-struct type punning was well > >> defined there, and the __packed annotation is what lets GCC generate valid > >> code for the unaligned access. > > The previous upstream version was the result of endless discussions, and it > looks like changing it to the memcpy version was premature. At the time we > unified all architectures to use a common implentation, this was the only one > that resulted in correct and fast code on all architectures, so I don't > understand why this was just applied without including everyone who was > involved in coming up with the version that was replaced. > > My feeling is that we should just revert this. I'm not sure about the > motivation for the change. It sounds like this was meant to be > used in userland code, and that clashed with assumptions we make > in the kernel, but I don't think that is sufficient reason for > regressing kernel code. So the original motivation for the change was that the perf tool had an OpenSSL dependency for the sake of doing a hash when copying jitted code into a fake ELF file for the purpose of disassembly and symbolization. The kernel contained the same hash function, and using the kernel function allowed the perf tool to avoid an OpenSSL dependency. Linus asked the perf tool to minimize its dependencies, so we pursued this change. The kernel hash function used the get_unaligned code for unaligned memory accesses, meaning the change required the perf tool to add -fno-strict-aliasing to its build flags until we could have a strict aliasing safe get_unaligned. The strict aliasing safe code is what we're discussing here and when originally written it sat on the mailing list not really doing anything. To my surprise Thomas Gleixner picked it up 6 months later, and I believe something other than the perf tool motivated this. There was an issue with the original series on Power IIRC, they had a char global variable that they knew was an int through linker tricks. The change caused a correct compiler error of a 4 byte copy to a byte sized address, and I forget if we used pragmas or a local packed struct get_unaligned implementation to work around this. I didn't have a way to replicate the problem locally, I was glad others were helping with the series! Am i going to be able to convince people builtin_memcpy is a better unaligned choice than a packed struct? Well memcpy is the expected C way to solve this problem, and C compilers like clang implicitly emit memcpy intrinsics when copying things like aggregate values. The packed struct requires a cast to a pointer type violating strict aliasing rules, but has been sound for many years because of -fno-strict-aliasing. We'd like the perf tool to be clean for things like undefined behavior sanitizers, but separate userland code could support this. I was trying to be as useful as possible and I don't think the patches were rushed at any point. What the bug report shows is that there is a lowering problem for memcpy with GCC on ARMv5, presumably as ARMv5 lacks unaligned memory operations. It seems the packed code shows how we can teach the faster unaligned lowering to GCC for ARMv5. Fixing the lowering issue in GCC will likely win performance improvements elsewhere on code compiled for ARMv5, as optimal memcpy codegen is an expectation for things like copying an aggregate value. Fixing GCC would be best but a workaround is to use the packed unaligned functions: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/linux/unaligned/packed_struct.h I guess this could be a pain if there are regressions elsewhere in the kernel that also need fixing. I believe I spoke to Stefan at LPC on Tuesday and explained all of this, suggesting the packed unaligned functions as a workaround. It would be interesting to hear of other motivations for memcpy that Thomas may know. Perhaps some config value defaulted to memcpy and switching to packed structs works for everyone. Maybe separating the user and kernel code makes most sense. Thanks, Ian > Arnd ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2026-10-09 21:28 ` Ian Rogers @ 2026-10-10 10:26 ` Christophe Leroy (CS GROUP) 2026-10-10 12:20 ` Ian Rogers 2026-10-10 14:06 ` Arnd Bergmann 1 sibling, 1 reply; 23+ messages in thread From: Christophe Leroy (CS GROUP) @ 2026-10-10 10:26 UTC (permalink / raw) To: Ian Rogers, Arnd Bergmann, Thomas Gleixner Cc: Marc Kleine-Budde, Stefan Kerkmann, James E . J . Bottomley, Helge Deller, Andy Lutomirski, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Alexander Viro, Jason A . Donenfeld Le 09/10/2026 à 23:28, Ian Rogers a écrit : > On Wed, Oct 7, 2026 at 4:00 AM Arnd Bergmann <arnd@arndb.de> wrote: >> >> On Tue, Oct 6, 2026, at 14:31, Marc Kleine-Budde wrote: >>> Cc+ Arnd >> >> Thanks for the Cc! >> >>> On 28.09.2026 17:39:15, Stefan Kerkmann wrote: >>>> Hi Ian, >>>> >>>> On 10/16/25 22:51, Ian Rogers wrote: >>>>> Type punning is necessary for get/put unaligned but the use of a >>>>> packed struct violates strict aliasing rules, requiring >>>>> -fno-strict-aliasing to be passed to the C compiler. Switch to using >>>>> memcpy so that -fno-strict-aliasing isn't necessary. >>>>> >>>>> Signed-off-by: Ian Rogers <irogers@google.com> >>>>> --- >>> >>>> Is this an accepted trade-off? My understanding is that the kernel is always >>>> built with -fno-strict-aliasing, so the packed-struct type punning was well >>>> defined there, and the __packed annotation is what lets GCC generate valid >>>> code for the unaligned access. >> >> The previous upstream version was the result of endless discussions, and it >> looks like changing it to the memcpy version was premature. At the time we >> unified all architectures to use a common implentation, this was the only one >> that resulted in correct and fast code on all architectures, so I don't >> understand why this was just applied without including everyone who was >> involved in coming up with the version that was replaced. >> >> My feeling is that we should just revert this. I'm not sure about the >> motivation for the change. It sounds like this was meant to be >> used in userland code, and that clashed with assumptions we make >> in the kernel, but I don't think that is sufficient reason for >> regressing kernel code. > > So the original motivation for the change was that the perf tool had > an OpenSSL dependency for the sake of doing a hash when copying jitted > code into a fake ELF file for the purpose of disassembly and > symbolization. The kernel contained the same hash function, and using > the kernel function allowed the perf tool to avoid an OpenSSL > dependency. Linus asked the perf tool to minimize its dependencies, so > we pursued this change. The kernel hash function used the > get_unaligned code for unaligned memory accesses, meaning the change > required the perf tool to add -fno-strict-aliasing to its build flags > until we could have a strict aliasing safe get_unaligned. The strict > aliasing safe code is what we're discussing here and when originally > written it sat on the mailing list not really doing anything. To my > surprise Thomas Gleixner picked it up 6 months later, and I believe > something other than the perf tool motivated this. > > There was an issue with the original series on Power IIRC, they had a > char global variable that they knew was an int through linker tricks. > The change caused a correct compiler error of a 4 byte copy to a byte > sized address, and I forget if we used pragmas or a local packed > struct get_unaligned implementation to work around this. I didn't have > a way to replicate the problem locally, I was glad others were helping > with the series! > I can't remember any special issue with powerpc, I've looked into the history and couldn't find anything either. As far as I can see the only problem we had was with your version v1 where you were using memcpy() instead of builtin_memcpy(), leading to a VDSO link failure due to missing memcpy() function. But if you think about other problems with powerpc let me know and I'll look at it. > VDSO build fails with this patch: > > VDSO32L arch/powerpc/kernel/vdso/vdso32.so.dbg > arch/powerpc/kernel/vdso/vdso32.so.dbg: dynamic relocations are not > supported > make[2]: *** [arch/powerpc/kernel/vdso/Makefile:79: arch/powerpc/kernel/ > vdso/vdso32.so.dbg] Error 1 > > Behind the relocation issue, calling memcpy() for a single 4-bytes word > kills performance. > > 170: 7f e4 fb 78 mr r4,r31 > 174: 38 a0 00 04 li r5,4 > 178: 38 61 00 10 addi r3,r1,16 > 17c: 93 81 00 10 stw r28,16(r1) > 180: 48 00 00 01 bl 180 <__c_kernel_getrandom+0x180> > 180: R_PPC_REL24 memcpy > 184: 38 81 00 10 addi r4,r1,16 Christophe ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2026-10-10 10:26 ` Christophe Leroy (CS GROUP) @ 2026-10-10 12:20 ` Ian Rogers 0 siblings, 0 replies; 23+ messages in thread From: Ian Rogers @ 2026-10-10 12:20 UTC (permalink / raw) To: Christophe Leroy (CS GROUP) Cc: Arnd Bergmann, Thomas Gleixner, Marc Kleine-Budde, Stefan Kerkmann, James E . J . Bottomley, Helge Deller, Andy Lutomirski, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Alexander Viro, Jason A . Donenfeld On Sat, Oct 10, 2026 at 3:26 AM Christophe Leroy (CS GROUP) <chleroy@kernel.org> wrote: > > Le 09/10/2026 à 23:28, Ian Rogers a écrit : > > On Wed, Oct 7, 2026 at 4:00 AM Arnd Bergmann <arnd@arndb.de> wrote: > >> > >> On Tue, Oct 6, 2026, at 14:31, Marc Kleine-Budde wrote: > >>> Cc+ Arnd > >> > >> Thanks for the Cc! > >> > >>> On 28.09.2026 17:39:15, Stefan Kerkmann wrote: > >>>> Hi Ian, > >>>> > >>>> On 10/16/25 22:51, Ian Rogers wrote: > >>>>> Type punning is necessary for get/put unaligned but the use of a > >>>>> packed struct violates strict aliasing rules, requiring > >>>>> -fno-strict-aliasing to be passed to the C compiler. Switch to using > >>>>> memcpy so that -fno-strict-aliasing isn't necessary. > >>>>> > >>>>> Signed-off-by: Ian Rogers <irogers@google.com> > >>>>> --- > >>> > >>>> Is this an accepted trade-off? My understanding is that the kernel is always > >>>> built with -fno-strict-aliasing, so the packed-struct type punning was well > >>>> defined there, and the __packed annotation is what lets GCC generate valid > >>>> code for the unaligned access. > >> > >> The previous upstream version was the result of endless discussions, and it > >> looks like changing it to the memcpy version was premature. At the time we > >> unified all architectures to use a common implentation, this was the only one > >> that resulted in correct and fast code on all architectures, so I don't > >> understand why this was just applied without including everyone who was > >> involved in coming up with the version that was replaced. > >> > >> My feeling is that we should just revert this. I'm not sure about the > >> motivation for the change. It sounds like this was meant to be > >> used in userland code, and that clashed with assumptions we make > >> in the kernel, but I don't think that is sufficient reason for > >> regressing kernel code. > > > > So the original motivation for the change was that the perf tool had > > an OpenSSL dependency for the sake of doing a hash when copying jitted > > code into a fake ELF file for the purpose of disassembly and > > symbolization. The kernel contained the same hash function, and using > > the kernel function allowed the perf tool to avoid an OpenSSL > > dependency. Linus asked the perf tool to minimize its dependencies, so > > we pursued this change. The kernel hash function used the > > get_unaligned code for unaligned memory accesses, meaning the change > > required the perf tool to add -fno-strict-aliasing to its build flags > > until we could have a strict aliasing safe get_unaligned. The strict > > aliasing safe code is what we're discussing here and when originally > > written it sat on the mailing list not really doing anything. To my > > surprise Thomas Gleixner picked it up 6 months later, and I believe > > something other than the perf tool motivated this. > > > > There was an issue with the original series on Power IIRC, they had a > > char global variable that they knew was an int through linker tricks. > > The change caused a correct compiler error of a 4 byte copy to a byte > > sized address, and I forget if we used pragmas or a local packed > > struct get_unaligned implementation to work around this. I didn't have > > a way to replicate the problem locally, I was glad others were helping > > with the series! > > > > I can't remember any special issue with powerpc, I've looked into the > history and couldn't find anything either. > > As far as I can see the only problem we had was with your version v1 > where you were using memcpy() instead of builtin_memcpy(), leading to a > VDSO link failure due to missing memcpy() function. > > But if you think about other problems with powerpc let me know and I'll > look at it. Thanks Christophe! The issues weren't performance regressions, as here with ARMv5, but build bots that failed due to warnings about the memcpy source or destination being invalid. You are right that PowerPC wasn't the architecture that failed; it was PA-RISC. Here is what I could find in the LKML history: PA-RISC build error: https://lore.kernel.org/lkml/CAP-5=fVEp8UPS-B3X=66AwYbbbs8559AHEObpQZSVnSnVYSxhA@mail.gmail.com/ PA-RISC workaround: https://lore.kernel.org/lkml/20251016205126.2882625-2-irogers@google.com/ MIPS failure with suggested workaround: https://lore.kernel.org/lkml/CAP-5=fVBZPBz6J1omfvSS4JLkueRGqCdouLji6gsFLTeu9nJTw@mail.gmail.com/ Thanks, Ian > > VDSO build fails with this patch: > > > > VDSO32L arch/powerpc/kernel/vdso/vdso32.so.dbg > > arch/powerpc/kernel/vdso/vdso32.so.dbg: dynamic relocations are not > > supported > > make[2]: *** [arch/powerpc/kernel/vdso/Makefile:79: arch/powerpc/kernel/ > > vdso/vdso32.so.dbg] Error 1 > > > > Behind the relocation issue, calling memcpy() for a single 4-bytes word > > kills performance. > > > > 170: 7f e4 fb 78 mr r4,r31 > > 174: 38 a0 00 04 li r5,4 > > 178: 38 61 00 10 addi r3,r1,16 > > 17c: 93 81 00 10 stw r28,16(r1) > > 180: 48 00 00 01 bl 180 <__c_kernel_getrandom+0x180> > > 180: R_PPC_REL24 memcpy > > 184: 38 81 00 10 addi r4,r1,16 > > Christophe ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2026-10-09 21:28 ` Ian Rogers 2026-10-10 10:26 ` Christophe Leroy (CS GROUP) @ 2026-10-10 14:06 ` Arnd Bergmann 2026-10-10 15:04 ` Ian Rogers 1 sibling, 1 reply; 23+ messages in thread From: Arnd Bergmann @ 2026-10-10 14:06 UTC (permalink / raw) To: Ian Rogers, Thomas Gleixner Cc: Marc Kleine-Budde, Stefan Kerkmann, James E . J . Bottomley, Helge Deller, Andy Lutomirski, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Alexander Viro, Christophe Leroy, Jason A . Donenfeld On Fri, Oct 9, 2026, at 23:28, Ian Rogers wrote: > On Wed, Oct 7, 2026 at 4:00 AM Arnd Bergmann <arnd@arndb.de> wrote: >> On Tue, Oct 6, 2026, at 14:31, Marc Kleine-Budde wrote: >> > On 28.09.2026 17:39:15, Stefan Kerkmann wrote: >> The previous upstream version was the result of endless discussions, and it >> looks like changing it to the memcpy version was premature. At the time we >> unified all architectures to use a common implentation, this was the only one >> that resulted in correct and fast code on all architectures, so I don't >> understand why this was just applied without including everyone who was >> involved in coming up with the version that was replaced. >> >> My feeling is that we should just revert this. I'm not sure about the >> motivation for the change. It sounds like this was meant to be >> used in userland code, and that clashed with assumptions we make >> in the kernel, but I don't think that is sufficient reason for >> regressing kernel code. > > So the original motivation for the change was that the perf tool had > an OpenSSL dependency for the sake of doing a hash when copying jitted > code into a fake ELF file for the purpose of disassembly and > symbolization. The kernel contained the same hash function, and using > the kernel function allowed the perf tool to avoid an OpenSSL > dependency. Linus asked the perf tool to minimize its dependencies, so > we pursued this change. The kernel hash function used the > get_unaligned code for unaligned memory accesses, meaning the change > required the perf tool to add -fno-strict-aliasing to its build flags > until we could have a strict aliasing safe get_unaligned. The strict > aliasing safe code is what we're discussing here and when originally > written it sat on the mailing list not really doing anything. To my > surprise Thomas Gleixner picked it up 6 months later, and I believe > something other than the perf tool motivated this. > > There was an issue with the original series on Power IIRC, they had a > char global variable that they knew was an int through linker tricks. > The change caused a correct compiler error of a 4 byte copy to a byte > sized address, and I forget if we used pragmas or a local packed > struct get_unaligned implementation to work around this. I didn't have > a way to replicate the problem locally, I was glad others were helping > with the series! Ok, thanks for explaining the background. > Am i going to be able to convince people builtin_memcpy is a better > unaligned choice than a packed struct? Well memcpy is the expected C > way to solve this problem, and C compilers like clang implicitly emit > memcpy intrinsics when copying things like aggregate values. The > packed struct requires a cast to a pointer type violating strict > aliasing rules, but has been sound for many years because of > -fno-strict-aliasing. We'd like the perf tool to be clean for things > like undefined behavior sanitizers, but separate userland code could > support this. I was trying to be as useful as possible and I don't > think the patches were rushed at any point. Aside from the performance regression, I see more issues with using __builtin_memcpy() here: - since the compiler is allowed to always turn __builtin_memcpy() into an extern memcpy() call, it looks invalid to use this in the vdso, which is not allowed to call any functions. - the original version used memmove() instead of memcpy(). While this was removed in bf067edf5d2f ("openrisc: always use unaligned-struct header"), this was surely intentional at the time. - the use of __unqual_scalar_typeof() turns a relatively simple expression into a much larger amount of preprocessed code, which tends to confuse the inlining choices and compile speed, especially when this is mixed with other macros that expand the arguments multiple times. > What the bug report shows is that there is a lowering problem for > memcpy with GCC on ARMv5, presumably as ARMv5 lacks unaligned memory > operations. It seems the packed code shows how we can teach the faster > unaligned lowering to GCC for ARMv5. Fixing the lowering issue in GCC > will likely win performance improvements elsewhere on code compiled > for ARMv5, as optimal memcpy codegen is an expectation for things like > copying an aggregate value. As far as I can tell, this only happens when building with -Os, and I don't even think the decision to call the external memcpy() is necessarily wrong here. The same thing happens on mips32, mips64, openrisc, riscv32, riscv64, sh4, sparc32, and sparc64. Some of these also use an out-of-line bswap32 in put_unaligned_be32() when building with -Os. > Fixing GCC would be best but a workaround is to use the packed > unaligned functions: > https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/linux/unaligned/packed_struct.h > I guess this could be a pain if there are regressions elsewhere in the > kernel that also need fixing. I hadn't realized that we still have the extra copy for those, my guess is that we had planned to remove that but somehow never converted the last remaining user in tools/include/linux/jhash.h after commit 50b4233a22b1 ("include/linux/jhash.h: replace __get_unaligned_cpu32 in jhash function") did the second-to-last. > I believe I spoke to Stefan at LPC on Tuesday and explained all of > this, suggesting the packed unaligned functions as a workaround. It > would be interesting to hear of other motivations for memcpy that > Thomas may know. Perhaps some config value defaulted to memcpy and > switching to packed structs works for everyone. Maybe separating the > user and kernel code makes most sense. Right, that seems easy enough to do: the include/vdso/ headers are not meant for user consumption in the first place, so it would make sense to use the struct variant there, and the include/linux/unaligned/packed_struct.h version could provide the memcpy() based code for tools/ and get deleted from the kernel internal version. It's just complicated a bit by the fact that the current code does the exact oppposite ;-) Arnd ^ permalink raw reply [flat|nested] 23+ messages in thread
* Re: [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy 2026-10-10 14:06 ` Arnd Bergmann @ 2026-10-10 15:04 ` Ian Rogers 0 siblings, 0 replies; 23+ messages in thread From: Ian Rogers @ 2026-10-10 15:04 UTC (permalink / raw) To: Arnd Bergmann Cc: Thomas Gleixner, Marc Kleine-Budde, Stefan Kerkmann, James E . J . Bottomley, Helge Deller, Andy Lutomirski, Vincenzo Frascino, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Alexander Viro, Christophe Leroy, Jason A . Donenfeld On Sat, Oct 10, 2026 at 7:06 AM Arnd Bergmann <arnd@arndb.de> wrote: > > On Fri, Oct 9, 2026, at 23:28, Ian Rogers wrote: > > On Wed, Oct 7, 2026 at 4:00 AM Arnd Bergmann <arnd@arndb.de> wrote: > >> On Tue, Oct 6, 2026, at 14:31, Marc Kleine-Budde wrote: > >> > On 28.09.2026 17:39:15, Stefan Kerkmann wrote: > >> The previous upstream version was the result of endless discussions, and it > >> looks like changing it to the memcpy version was premature. At the time we > >> unified all architectures to use a common implentation, this was the only one > >> that resulted in correct and fast code on all architectures, so I don't > >> understand why this was just applied without including everyone who was > >> involved in coming up with the version that was replaced. > >> > >> My feeling is that we should just revert this. I'm not sure about the > >> motivation for the change. It sounds like this was meant to be > >> used in userland code, and that clashed with assumptions we make > >> in the kernel, but I don't think that is sufficient reason for > >> regressing kernel code. > > > > So the original motivation for the change was that the perf tool had > > an OpenSSL dependency for the sake of doing a hash when copying jitted > > code into a fake ELF file for the purpose of disassembly and > > symbolization. The kernel contained the same hash function, and using > > the kernel function allowed the perf tool to avoid an OpenSSL > > dependency. Linus asked the perf tool to minimize its dependencies, so > > we pursued this change. The kernel hash function used the > > get_unaligned code for unaligned memory accesses, meaning the change > > required the perf tool to add -fno-strict-aliasing to its build flags > > until we could have a strict aliasing safe get_unaligned. The strict > > aliasing safe code is what we're discussing here and when originally > > written it sat on the mailing list not really doing anything. To my > > surprise Thomas Gleixner picked it up 6 months later, and I believe > > something other than the perf tool motivated this. > > > > There was an issue with the original series on Power IIRC, they had a > > char global variable that they knew was an int through linker tricks. > > The change caused a correct compiler error of a 4 byte copy to a byte > > sized address, and I forget if we used pragmas or a local packed > > struct get_unaligned implementation to work around this. I didn't have > > a way to replicate the problem locally, I was glad others were helping > > with the series! > > Ok, thanks for explaining the background. > > > Am i going to be able to convince people builtin_memcpy is a better > > unaligned choice than a packed struct? Well memcpy is the expected C > > way to solve this problem, and C compilers like clang implicitly emit > > memcpy intrinsics when copying things like aggregate values. The > > packed struct requires a cast to a pointer type violating strict > > aliasing rules, but has been sound for many years because of > > -fno-strict-aliasing. We'd like the perf tool to be clean for things > > like undefined behavior sanitizers, but separate userland code could > > support this. I was trying to be as useful as possible and I don't > > think the patches were rushed at any point. > > Aside from the performance regression, I see more issues with > using __builtin_memcpy() here: > > - since the compiler is allowed to always turn __builtin_memcpy() > into an extern memcpy() call, it looks invalid to use this > in the vdso, which is not allowed to call any functions. > > - the original version used memmove() instead of memcpy(). While > this was removed in bf067edf5d2f ("openrisc: always use > unaligned-struct header"), this was surely intentional at the time. > > - the use of __unqual_scalar_typeof() turns a relatively simple > expression into a much larger amount of preprocessed code, which > tends to confuse the inlining choices and compile speed, especially > when this is mixed with other macros that expand the arguments > multiple times. > > > What the bug report shows is that there is a lowering problem for > > memcpy with GCC on ARMv5, presumably as ARMv5 lacks unaligned memory > > operations. It seems the packed code shows how we can teach the faster > > unaligned lowering to GCC for ARMv5. Fixing the lowering issue in GCC > > will likely win performance improvements elsewhere on code compiled > > for ARMv5, as optimal memcpy codegen is an expectation for things like > > copying an aggregate value. > > As far as I can tell, this only happens when building with -Os, > and I don't even think the decision to call the external memcpy() > is necessarily wrong here. The same thing happens on mips32, mips64, > openrisc, riscv32, riscv64, sh4, sparc32, and sparc64. Some of these > also use an out-of-line bswap32 in put_unaligned_be32() when building > with -Os. > > > Fixing GCC would be best but a workaround is to use the packed > > unaligned functions: > > https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/include/linux/unaligned/packed_struct.h > > I guess this could be a pain if there are regressions elsewhere in the > > kernel that also need fixing. > > I hadn't realized that we still have the extra copy for those, > my guess is that we had planned to remove that but somehow never > converted the last remaining user in tools/include/linux/jhash.h > after commit 50b4233a22b1 ("include/linux/jhash.h: replace > __get_unaligned_cpu32 in jhash function") did the second-to-last. > > > I believe I spoke to Stefan at LPC on Tuesday and explained all of > > this, suggesting the packed unaligned functions as a workaround. It > > would be interesting to hear of other motivations for memcpy that > > Thomas may know. Perhaps some config value defaulted to memcpy and > > switching to packed structs works for everyone. Maybe separating the > > user and kernel code makes most sense. > > Right, that seems easy enough to do: the include/vdso/ headers > are not meant for user consumption in the first place, so it would > make sense to use the struct variant there, and the > include/linux/unaligned/packed_struct.h version could provide the > memcpy() based code for tools/ and get deleted from the kernel > internal version. > > It's just complicated a bit by the fact that the current code does > the exact oppposite ;-) So completely fwiw, I speak to the compiler folks within Google and there has always been unanimous support for code using memcpy over type punning. I'm not sure where the inlining decisions would be impacted, as to my understanding those are made later in the compiler pipeline and things like __unqual_scalar_typeof are almost at the stage of the preprocessor. If -Os is not inlining small fixed size memcpys, this sounds like another compiler bug. It would also be completely valid for the compiler to lower packed struct accesses to memcpys. I was trying to reach an implementation that avoided undefined behavior and benefited both user and kernel code. Performance regressions aren't a benefit and fixing compilers can be tricky, so work remains. I'd prefer to avoid undefined behavior at least in user land. I'd also prefer to avoid config values, so using packed structs for __KERNEL__ code is a possibility. It seems a little sad that we've lived with a strict-aliasing-compliant implementation for a year, and we could regress that for the sake of ARMv5. I don't feel strongly enough to argue for kernel land; hopefully, Thomas can say what motivated moving the patch series forward as I'm pretty certain it wouldn't be undefined behavior in the perf tool :-) Thanks, Ian > Arnd ^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v5 3/4] tools headers: Update the linux/unaligned.h copy with the kernel sources 2025-10-16 20:51 [PATCH v5 0/4] Switch get/put unaligned to use memcpy Ian Rogers 2025-10-16 20:51 ` [PATCH v5 1/4] parisc: Inline a type punning version of get_unaligned_le32 Ian Rogers 2025-10-16 20:51 ` [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Ian Rogers @ 2025-10-16 20:51 ` Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] " tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 2025-10-16 20:51 ` [PATCH v5 4/4] tools headers: Remove unneeded ignoring of warnings in unaligned.h Ian Rogers 3 siblings, 2 replies; 23+ messages in thread From: Ian Rogers @ 2025-10-16 20:51 UTC (permalink / raw) To: James E.J. Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Ian Rogers, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld To pick up the changes in: vdso: Switch get/put unaligned from packed struct to memcpy As the code is dependent on __unqual_scalar_typeof, update the tools version of compiler_types.h to include this. Signed-off-by: Ian Rogers <irogers@google.com> --- tools/include/linux/compiler_types.h | 22 +++++++++++++++ tools/include/vdso/unaligned.h | 41 ++++++++++++++++++++++++---- 2 files changed, 57 insertions(+), 6 deletions(-) diff --git a/tools/include/linux/compiler_types.h b/tools/include/linux/compiler_types.h index d09f9dc172a4..890982283a5e 100644 --- a/tools/include/linux/compiler_types.h +++ b/tools/include/linux/compiler_types.h @@ -40,4 +40,26 @@ #define asm_goto_output(x...) asm goto(x) #endif +/* + * __unqual_scalar_typeof(x) - Declare an unqualified scalar type, leaving + * non-scalar types unchanged. + */ +/* + * Prefer C11 _Generic for better compile-times and simpler code. Note: 'char' + * is not type-compatible with 'signed char', and we define a separate case. + */ +#define __scalar_type_to_expr_cases(type) \ + unsigned type: (unsigned type)0, \ + signed type: (signed type)0 + +#define __unqual_scalar_typeof(x) typeof( \ + _Generic((x), \ + char: (char)0, \ + __scalar_type_to_expr_cases(char), \ + __scalar_type_to_expr_cases(short), \ + __scalar_type_to_expr_cases(int), \ + __scalar_type_to_expr_cases(long), \ + __scalar_type_to_expr_cases(long long), \ + default: (x))) + #endif /* __LINUX_COMPILER_TYPES_H */ diff --git a/tools/include/vdso/unaligned.h b/tools/include/vdso/unaligned.h index ff0c06b6513e..9076483c9fbb 100644 --- a/tools/include/vdso/unaligned.h +++ b/tools/include/vdso/unaligned.h @@ -2,14 +2,43 @@ #ifndef __VDSO_UNALIGNED_H #define __VDSO_UNALIGNED_H -#define __get_unaligned_t(type, ptr) ({ \ - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ - __get_pptr->x; \ +#include <linux/compiler_types.h> + +/** + * __get_unaligned_t - read an unaligned value from memory. + * @type: the type to load from the pointer. + * @ptr: the pointer to load from. + * + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior + * from approaches like type punning that require -fno-strict-aliasing in order + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a + * non-const type - you can't memcpy into a const type. The + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required + * expression rather than type, a pointer is used to avoid warnings about mixing + * the use of 0 and NULL. The void* cast silences ubsan warnings. + */ +#define __get_unaligned_t(type, ptr) ({ \ + type *__get_unaligned_ctrl_type __always_unused = NULL; \ + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ + sizeof(__get_unaligned_val)); \ + __get_unaligned_val; \ }) -#define __put_unaligned_t(type, val, ptr) do { \ - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ - __put_pptr->x = (val); \ +/** + * __put_unaligned_t - write an unaligned value to memory. + * @type: the type of the value to store. + * @val: the value to store. + * @ptr: the pointer to store to. + * + * Use memcpy to affect an unaligned type sized store avoiding undefined + * behavior from approaches like type punning that require -fno-strict-aliasing + * in order to be correct. The void* cast silences ubsan warnings. + */ +#define __put_unaligned_t(type, val, ptr) do { \ + type __put_unaligned_val = (val); \ + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ + sizeof(__put_unaligned_val)); \ } while (0) #endif /* __VDSO_UNALIGNED_H */ -- 2.51.0.858.gf9c4a03a3a-goog ^ permalink raw reply [flat|nested] 23+ messages in thread
* [tip: timers/vdso] tools headers: Update the linux/unaligned.h copy with the kernel sources 2025-10-16 20:51 ` [PATCH v5 3/4] tools headers: Update the linux/unaligned.h copy with the kernel sources Ian Rogers @ 2026-01-13 13:47 ` tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 1 sibling, 0 replies; 23+ messages in thread From: tip-bot2 for Ian Rogers @ 2026-01-13 13:47 UTC (permalink / raw) To: linux-tip-commits; +Cc: Ian Rogers, Thomas Gleixner, x86, linux-kernel The following commit has been merged into the timers/vdso branch of tip: Commit-ID: 029a9504d871ce72adf9a13e5f1ce66f1db48be3 Gitweb: https://git.kernel.org/tip/029a9504d871ce72adf9a13e5f1ce66f1db48be3 Author: Ian Rogers <irogers@google.com> AuthorDate: Thu, 16 Oct 2025 13:51:25 -07:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Tue, 13 Jan 2026 14:46:00 +01:00 tools headers: Update the linux/unaligned.h copy with the kernel sources To pick up the changes in: vdso: Switch get/put_unaligned() from packed struct to memcpy As the code is dependent on __unqual_scalar_typeof, update also the tools version of compiler_types.h to include this. Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Link: https://patch.msgid.link/20251016205126.2882625-4-irogers@google.com --- tools/include/linux/compiler_types.h | 22 ++++++++++++++- tools/include/vdso/unaligned.h | 41 +++++++++++++++++++++++---- 2 files changed, 57 insertions(+), 6 deletions(-) diff --git a/tools/include/linux/compiler_types.h b/tools/include/linux/compiler_types.h index d09f9dc..8909822 100644 --- a/tools/include/linux/compiler_types.h +++ b/tools/include/linux/compiler_types.h @@ -40,4 +40,26 @@ #define asm_goto_output(x...) asm goto(x) #endif +/* + * __unqual_scalar_typeof(x) - Declare an unqualified scalar type, leaving + * non-scalar types unchanged. + */ +/* + * Prefer C11 _Generic for better compile-times and simpler code. Note: 'char' + * is not type-compatible with 'signed char', and we define a separate case. + */ +#define __scalar_type_to_expr_cases(type) \ + unsigned type: (unsigned type)0, \ + signed type: (signed type)0 + +#define __unqual_scalar_typeof(x) typeof( \ + _Generic((x), \ + char: (char)0, \ + __scalar_type_to_expr_cases(char), \ + __scalar_type_to_expr_cases(short), \ + __scalar_type_to_expr_cases(int), \ + __scalar_type_to_expr_cases(long), \ + __scalar_type_to_expr_cases(long long), \ + default: (x))) + #endif /* __LINUX_COMPILER_TYPES_H */ diff --git a/tools/include/vdso/unaligned.h b/tools/include/vdso/unaligned.h index ff0c06b..9076483 100644 --- a/tools/include/vdso/unaligned.h +++ b/tools/include/vdso/unaligned.h @@ -2,14 +2,43 @@ #ifndef __VDSO_UNALIGNED_H #define __VDSO_UNALIGNED_H -#define __get_unaligned_t(type, ptr) ({ \ - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ - __get_pptr->x; \ +#include <linux/compiler_types.h> + +/** + * __get_unaligned_t - read an unaligned value from memory. + * @type: the type to load from the pointer. + * @ptr: the pointer to load from. + * + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior + * from approaches like type punning that require -fno-strict-aliasing in order + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a + * non-const type - you can't memcpy into a const type. The + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required + * expression rather than type, a pointer is used to avoid warnings about mixing + * the use of 0 and NULL. The void* cast silences ubsan warnings. + */ +#define __get_unaligned_t(type, ptr) ({ \ + type *__get_unaligned_ctrl_type __always_unused = NULL; \ + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ + sizeof(__get_unaligned_val)); \ + __get_unaligned_val; \ }) -#define __put_unaligned_t(type, val, ptr) do { \ - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ - __put_pptr->x = (val); \ +/** + * __put_unaligned_t - write an unaligned value to memory. + * @type: the type of the value to store. + * @val: the value to store. + * @ptr: the pointer to store to. + * + * Use memcpy to affect an unaligned type sized store avoiding undefined + * behavior from approaches like type punning that require -fno-strict-aliasing + * in order to be correct. The void* cast silences ubsan warnings. + */ +#define __put_unaligned_t(type, val, ptr) do { \ + type __put_unaligned_val = (val); \ + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ + sizeof(__put_unaligned_val)); \ } while (0) #endif /* __VDSO_UNALIGNED_H */ ^ permalink raw reply [flat|nested] 23+ messages in thread
* [tip: timers/vdso] tools headers: Update the linux/unaligned.h copy with the kernel sources 2025-10-16 20:51 ` [PATCH v5 3/4] tools headers: Update the linux/unaligned.h copy with the kernel sources Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] " tip-bot2 for Ian Rogers @ 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 1 sibling, 0 replies; 23+ messages in thread From: tip-bot2 for Ian Rogers @ 2026-01-14 8:01 UTC (permalink / raw) To: linux-tip-commits; +Cc: Ian Rogers, Thomas Gleixner, x86, linux-kernel The following commit has been merged into the timers/vdso branch of tip: Commit-ID: 1d7cf255eefbb479d0eea9aa3b6372a1e52f8c62 Gitweb: https://git.kernel.org/tip/1d7cf255eefbb479d0eea9aa3b6372a1e52f8c62 Author: Ian Rogers <irogers@google.com> AuthorDate: Thu, 16 Oct 2025 13:51:25 -07:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Wed, 14 Jan 2026 08:56:41 +01:00 tools headers: Update the linux/unaligned.h copy with the kernel sources To pick up the changes in: vdso: Switch get/put_unaligned() from packed struct to memcpy As the code is dependent on __unqual_scalar_typeof, update also the tools version of compiler_types.h to include this. Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Link: https://patch.msgid.link/20251016205126.2882625-4-irogers@google.com --- tools/include/linux/compiler_types.h | 22 ++++++++++++++- tools/include/vdso/unaligned.h | 41 +++++++++++++++++++++++---- 2 files changed, 57 insertions(+), 6 deletions(-) diff --git a/tools/include/linux/compiler_types.h b/tools/include/linux/compiler_types.h index d09f9dc..8909822 100644 --- a/tools/include/linux/compiler_types.h +++ b/tools/include/linux/compiler_types.h @@ -40,4 +40,26 @@ #define asm_goto_output(x...) asm goto(x) #endif +/* + * __unqual_scalar_typeof(x) - Declare an unqualified scalar type, leaving + * non-scalar types unchanged. + */ +/* + * Prefer C11 _Generic for better compile-times and simpler code. Note: 'char' + * is not type-compatible with 'signed char', and we define a separate case. + */ +#define __scalar_type_to_expr_cases(type) \ + unsigned type: (unsigned type)0, \ + signed type: (signed type)0 + +#define __unqual_scalar_typeof(x) typeof( \ + _Generic((x), \ + char: (char)0, \ + __scalar_type_to_expr_cases(char), \ + __scalar_type_to_expr_cases(short), \ + __scalar_type_to_expr_cases(int), \ + __scalar_type_to_expr_cases(long), \ + __scalar_type_to_expr_cases(long long), \ + default: (x))) + #endif /* __LINUX_COMPILER_TYPES_H */ diff --git a/tools/include/vdso/unaligned.h b/tools/include/vdso/unaligned.h index ff0c06b..9076483 100644 --- a/tools/include/vdso/unaligned.h +++ b/tools/include/vdso/unaligned.h @@ -2,14 +2,43 @@ #ifndef __VDSO_UNALIGNED_H #define __VDSO_UNALIGNED_H -#define __get_unaligned_t(type, ptr) ({ \ - const struct { type x; } __packed * __get_pptr = (typeof(__get_pptr))(ptr); \ - __get_pptr->x; \ +#include <linux/compiler_types.h> + +/** + * __get_unaligned_t - read an unaligned value from memory. + * @type: the type to load from the pointer. + * @ptr: the pointer to load from. + * + * Use memcpy to affect an unaligned type sized load avoiding undefined behavior + * from approaches like type punning that require -fno-strict-aliasing in order + * to be correct. As type may be const, use __unqual_scalar_typeof to map to a + * non-const type - you can't memcpy into a const type. The + * __get_unaligned_ctrl_type gives __unqual_scalar_typeof its required + * expression rather than type, a pointer is used to avoid warnings about mixing + * the use of 0 and NULL. The void* cast silences ubsan warnings. + */ +#define __get_unaligned_t(type, ptr) ({ \ + type *__get_unaligned_ctrl_type __always_unused = NULL; \ + __unqual_scalar_typeof(*__get_unaligned_ctrl_type) __get_unaligned_val; \ + __builtin_memcpy(&__get_unaligned_val, (void *)(ptr), \ + sizeof(__get_unaligned_val)); \ + __get_unaligned_val; \ }) -#define __put_unaligned_t(type, val, ptr) do { \ - struct { type x; } __packed * __put_pptr = (typeof(__put_pptr))(ptr); \ - __put_pptr->x = (val); \ +/** + * __put_unaligned_t - write an unaligned value to memory. + * @type: the type of the value to store. + * @val: the value to store. + * @ptr: the pointer to store to. + * + * Use memcpy to affect an unaligned type sized store avoiding undefined + * behavior from approaches like type punning that require -fno-strict-aliasing + * in order to be correct. The void* cast silences ubsan warnings. + */ +#define __put_unaligned_t(type, val, ptr) do { \ + type __put_unaligned_val = (val); \ + __builtin_memcpy((void *)(ptr), &__put_unaligned_val, \ + sizeof(__put_unaligned_val)); \ } while (0) #endif /* __VDSO_UNALIGNED_H */ ^ permalink raw reply [flat|nested] 23+ messages in thread
* [PATCH v5 4/4] tools headers: Remove unneeded ignoring of warnings in unaligned.h 2025-10-16 20:51 [PATCH v5 0/4] Switch get/put unaligned to use memcpy Ian Rogers ` (2 preceding siblings ...) 2025-10-16 20:51 ` [PATCH v5 3/4] tools headers: Update the linux/unaligned.h copy with the kernel sources Ian Rogers @ 2025-10-16 20:51 ` Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] " tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 3 siblings, 2 replies; 23+ messages in thread From: Ian Rogers @ 2025-10-16 20:51 UTC (permalink / raw) To: James E.J. Bottomley, Helge Deller, Andy Lutomirski, Thomas Gleixner, Vincenzo Frascino, Ian Rogers, Arnaldo Carvalho de Melo, linux-parisc, linux-kernel, Eric Biggers, Al Viro, Christophe Leroy, Jason A. Donenfeld Now the get/put unaligned use memcpy the -Wpacked and -Wattributes warnings don't need disabling. Signed-off-by: Ian Rogers <irogers@google.com> --- tools/include/linux/unaligned.h | 4 ---- 1 file changed, 4 deletions(-) diff --git a/tools/include/linux/unaligned.h b/tools/include/linux/unaligned.h index 395a4464fe73..d51ddafed138 100644 --- a/tools/include/linux/unaligned.h +++ b/tools/include/linux/unaligned.h @@ -6,9 +6,6 @@ * This is the most generic implementation of unaligned accesses * and should work almost anywhere. */ -#pragma GCC diagnostic push -#pragma GCC diagnostic ignored "-Wpacked" -#pragma GCC diagnostic ignored "-Wattributes" #include <vdso/unaligned.h> #define get_unaligned(ptr) __get_unaligned_t(typeof(*(ptr)), (ptr)) @@ -143,6 +140,5 @@ static inline u64 get_unaligned_be48(const void *p) { return __get_unaligned_be48(p); } -#pragma GCC diagnostic pop #endif /* __LINUX_UNALIGNED_H */ -- 2.51.0.858.gf9c4a03a3a-goog ^ permalink raw reply [flat|nested] 23+ messages in thread
* [tip: timers/vdso] tools headers: Remove unneeded ignoring of warnings in unaligned.h 2025-10-16 20:51 ` [PATCH v5 4/4] tools headers: Remove unneeded ignoring of warnings in unaligned.h Ian Rogers @ 2026-01-13 13:47 ` tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 1 sibling, 0 replies; 23+ messages in thread From: tip-bot2 for Ian Rogers @ 2026-01-13 13:47 UTC (permalink / raw) To: linux-tip-commits; +Cc: Ian Rogers, Thomas Gleixner, x86, linux-kernel The following commit has been merged into the timers/vdso branch of tip: Commit-ID: 576d8a7a985dde4f51a4561dddc0d8734494d3de Gitweb: https://git.kernel.org/tip/576d8a7a985dde4f51a4561dddc0d8734494d3de Author: Ian Rogers <irogers@google.com> AuthorDate: Thu, 16 Oct 2025 13:51:26 -07:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Tue, 13 Jan 2026 14:46:01 +01:00 tools headers: Remove unneeded ignoring of warnings in unaligned.h Now that get/put_unaligned() use memcpy() the -Wpacked and -Wattributes warnings don't need disabling anymore. Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Link: https://patch.msgid.link/20251016205126.2882625-5-irogers@google.com --- tools/include/linux/unaligned.h | 4 ---- 1 file changed, 4 deletions(-) diff --git a/tools/include/linux/unaligned.h b/tools/include/linux/unaligned.h index 395a446..d51ddaf 100644 --- a/tools/include/linux/unaligned.h +++ b/tools/include/linux/unaligned.h @@ -6,9 +6,6 @@ * This is the most generic implementation of unaligned accesses * and should work almost anywhere. */ -#pragma GCC diagnostic push -#pragma GCC diagnostic ignored "-Wpacked" -#pragma GCC diagnostic ignored "-Wattributes" #include <vdso/unaligned.h> #define get_unaligned(ptr) __get_unaligned_t(typeof(*(ptr)), (ptr)) @@ -143,6 +140,5 @@ static inline u64 get_unaligned_be48(const void *p) { return __get_unaligned_be48(p); } -#pragma GCC diagnostic pop #endif /* __LINUX_UNALIGNED_H */ ^ permalink raw reply [flat|nested] 23+ messages in thread
* [tip: timers/vdso] tools headers: Remove unneeded ignoring of warnings in unaligned.h 2025-10-16 20:51 ` [PATCH v5 4/4] tools headers: Remove unneeded ignoring of warnings in unaligned.h Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] " tip-bot2 for Ian Rogers @ 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 1 sibling, 0 replies; 23+ messages in thread From: tip-bot2 for Ian Rogers @ 2026-01-14 8:01 UTC (permalink / raw) To: linux-tip-commits; +Cc: Ian Rogers, Thomas Gleixner, x86, linux-kernel The following commit has been merged into the timers/vdso branch of tip: Commit-ID: 10a62a0611f5544d209446acfde5beb7b27773c7 Gitweb: https://git.kernel.org/tip/10a62a0611f5544d209446acfde5beb7b27773c7 Author: Ian Rogers <irogers@google.com> AuthorDate: Thu, 16 Oct 2025 13:51:26 -07:00 Committer: Thomas Gleixner <tglx@kernel.org> CommitterDate: Wed, 14 Jan 2026 08:56:41 +01:00 tools headers: Remove unneeded ignoring of warnings in unaligned.h Now that get/put_unaligned() use memcpy() the -Wpacked and -Wattributes warnings don't need disabling anymore. Signed-off-by: Ian Rogers <irogers@google.com> Signed-off-by: Thomas Gleixner <tglx@kernel.org> Link: https://patch.msgid.link/20251016205126.2882625-5-irogers@google.com --- tools/include/linux/unaligned.h | 4 ---- 1 file changed, 4 deletions(-) diff --git a/tools/include/linux/unaligned.h b/tools/include/linux/unaligned.h index 395a446..d51ddaf 100644 --- a/tools/include/linux/unaligned.h +++ b/tools/include/linux/unaligned.h @@ -6,9 +6,6 @@ * This is the most generic implementation of unaligned accesses * and should work almost anywhere. */ -#pragma GCC diagnostic push -#pragma GCC diagnostic ignored "-Wpacked" -#pragma GCC diagnostic ignored "-Wattributes" #include <vdso/unaligned.h> #define get_unaligned(ptr) __get_unaligned_t(typeof(*(ptr)), (ptr)) @@ -143,6 +140,5 @@ static inline u64 get_unaligned_be48(const void *p) { return __get_unaligned_be48(p); } -#pragma GCC diagnostic pop #endif /* __LINUX_UNALIGNED_H */ ^ permalink raw reply [flat|nested] 23+ messages in thread
end of thread, other threads:[~2026-10-10 15:04 UTC | newest] Thread overview: 23+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2025-10-16 20:51 [PATCH v5 0/4] Switch get/put unaligned to use memcpy Ian Rogers 2025-10-16 20:51 ` [PATCH v5 1/4] parisc: Inline a type punning version of get_unaligned_le32 Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] parisc: Inline a type punning version of get_unaligned_le32() tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 2025-10-16 20:51 ` [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Ian Rogers 2025-10-19 17:24 ` David Laight 2026-01-13 13:47 ` [tip: timers/vdso] vdso: Switch get/put_unaligned() from packed struct to memcpy() tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 2026-09-28 15:39 ` [Regression] [PATCH v5 2/4] vdso: Switch get/put unaligned from packed struct to memcpy Stefan Kerkmann 2026-09-28 15:58 ` Ian Rogers 2026-10-06 12:31 ` Marc Kleine-Budde 2026-10-07 10:59 ` Arnd Bergmann 2026-10-09 21:28 ` Ian Rogers 2026-10-10 10:26 ` Christophe Leroy (CS GROUP) 2026-10-10 12:20 ` Ian Rogers 2026-10-10 14:06 ` Arnd Bergmann 2026-10-10 15:04 ` Ian Rogers 2025-10-16 20:51 ` [PATCH v5 3/4] tools headers: Update the linux/unaligned.h copy with the kernel sources Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] " tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers 2025-10-16 20:51 ` [PATCH v5 4/4] tools headers: Remove unneeded ignoring of warnings in unaligned.h Ian Rogers 2026-01-13 13:47 ` [tip: timers/vdso] " tip-bot2 for Ian Rogers 2026-01-14 8:01 ` tip-bot2 for Ian Rogers
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®