mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH v3 0/1] arch: arm64: Implement unaligned atomic emulation
@ 2026-09-30  2:01 André Almeida
  2026-09-30  2:01 ` [RFC PATCH v3 1/1] " André Almeida
  0 siblings, 1 reply; 5+ messages in thread
From: André Almeida @ 2026-09-30  2:01 UTC (permalink / raw)
  To: Catalin Marinas, Will Deacon, Billy Laws, Mark Rutland,
	Mark Brown, Ryan Houdek
  Cc: linux-arm-kernel, linux-kernel, kernel-dev, André Almeida

This patch proposes adding kernel-side emulation for unaligned atomic
instructions on ARM64. This is intended for x86 emulators (like FEX)
that struggle to effectively handle such operations in userspace alone.
Such handling is required as x86 permits such unaligned accesses (albeit
sometimes with a performance penalty as in the case of split-locks[1])
but ARM64 does not and will raise a bus error. Due to the weaker memory
model, ARM64 requires a lot of sync instructions to keep the consistency.

The patch is a reduced version of the real effort, to make easier to review
the proposed approach here. Some optimizations and instructions were left
for future revisions of this patchset. The full picture takes advantage of
all the LRCPC1/2/3/4 instructions + FEAT_LSE, according to the support of a
given processor.

User applications that wish to enable support for this can use the new
pctrl() flag `PR_ARM64_UNALIGN_ATOMIC_EMULATE`.

 * Why should we add the ability to deal with x86 shenanigans in ARM64?

To increase ARM64 support for more use cases, such as gaming. The vast
majority of games are proprietary apps compiled to Windows/x86, and there's
a huge effort on emulating this stack so games can run properly on ARM64.
Apple did an extra step on this and even added hardware extensions just for
this use case. Why don't just use x86? ARM64 is a much better platform for
embedded use cases such as the Steam Frame, so this isn't really an option
here.

 * Precedent:

Both XNU and NT kernels support unaligned atomic emulation for their
respective x86 emulators. Given that "Apple Silicon" has the advantage of
having the full TSO model enabled, they need to translate much less memory
access to atomic/sync instructions. Windows additionally supports 'volatile
metadata', which is emitted by newer versions of MSVC to inform
emulators which specific load/store accesses require atomic handling
[3]. FEX supports this together with an extension mechanism [4] which
can be manually populated to avoid e.g. the aforementioned Assassin's
Creed slowdown.

Emulators like FEX attempt to emulate this in userspace, but with
caveats in two areas:

 * Performance

It should first be noted that due to x86's TSO (total store order) memory
model, ARM64 synchronization instructions (such as LL/SC and atomics) must
be used for all memory accesses. This results in unaligned
loads/stores being much more common than one would expect and the
overhead of emulating them significantly impacting performance.  For
this common case of unaligned loads/stores, code backpatching is used in
FEX to avoid repeated overhead from handling the same faulting access.
This replaces faulting unaligned sync ARM64 instructions with regular
load/stores and memory barriers.  This comes at a cost of introducing
significant performance problems if a function like memcpy ends up being
patched because it very infrequently happens to be used with unaligned
memory. This is severe enough to make games like Mirror's Edge and
Assassin's Creed: Origin unplayable without application-specific
configuration. LSE2 helps a lot here, but it's limited to a 16B
granularity, so it doesn't cover all cases.

Microbenchmarks[2] measure more than 4x decrease in overhead with
kernel-side handling compared to userspace, and this figure is currently
even larger when FEX is ran under Wine. Such a dramatic decrease would
make it reasonable for FEX to default to the no-backpatching path and
provide consistent performance.

 * Correctness:

x86 atomic accesses can cross 16-byte (LSE2) granules, but there is no
ARM64 instruction that would allow for direct atomic emulation of this.
As such, a lock must be taken for correctness. While this is easy to
emulate in userspace within a process, correct atomic cross-granule
operations on shared memory mapped into multiple processes would require
a lock shared between all FEX instances which cannot be implemented
safely in userspace as is (robust futexes do not work here as they are
under the control of the emulated x86 program). Note, this is a less
coarse restriction than split locks on x86, which are only concerned
with accesses crossing a 64 byte cacheline size.

This implementation is a RFC so we can learn more about how to make this
code upstream and what the maintainers think of such feature being
merged here. The code is a simplified version of the original work done
by Billy Laws, where we accept just a subset of 64bit atomic
instructions that are enough to be used with a benchmark tool[2], and
this is the proposed interface being used by FEX: [5].

If you want to read even further, FEX developers came up with a good post about
all the details of emulating the memory model: https://fex-emu.com/Scourge-of-emulation/

Thanks!
	André

[1] https://lwn.net/Articles/911219/
[2] https://gitlab.freedesktop.org/freedesktop/snippets/-/snippets/7875
[3] https://learn.microsoft.com/en-us/cpp/build/reference/volatile?view=msvc-170
[4] https://github.com/FEX-Emu/FEX/pull/4773
[5] https://github.com/FEX-Emu/FEX/pull/4985

---
Changelog:
 - Refactored asm code
 - Add Kcofig option
 - Used more common reg code
 v2: https://lore.kernel.org/lkml/20251117160841.334224-1-andrealmeid@igalia.com/

 - Added a check for LSE Atomic instruction for the prctl()
 v1: https://lore.kernel.org/lkml/20251106160735.2638485-1-andrealmeid@igalia.com/

André Almeida (1):
  arch: arm64: Implement unaligned atomic emulation

 arch/arm64/Kconfig                   |   6 +
 arch/arm64/include/asm/exception.h   |   1 +
 arch/arm64/include/asm/processor.h   |   5 +
 arch/arm64/include/asm/rwonce.h      |  14 +-
 arch/arm64/include/asm/thread_info.h |   1 +
 arch/arm64/kernel/Makefile           |   3 +-
 arch/arm64/kernel/process.c          |  15 +
 arch/arm64/kernel/unaligned_atomic.c | 521 +++++++++++++++++++++++++++
 arch/arm64/mm/fault.c                |  10 +
 include/uapi/linux/prctl.h           |   5 +
 kernel/sys.c                         |   7 +-
 11 files changed, 579 insertions(+), 9 deletions(-)
 create mode 100644 arch/arm64/kernel/unaligned_atomic.c

-- 
2.55.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* [RFC PATCH v3 1/1] arch: arm64: Implement unaligned atomic emulation
  2026-09-30  2:01 [RFC PATCH v3 0/1] arch: arm64: Implement unaligned atomic emulation André Almeida
@ 2026-09-30  2:01 ` André Almeida
  2026-09-30  7:48   ` Will Deacon
  0 siblings, 1 reply; 5+ messages in thread
From: André Almeida @ 2026-09-30  2:01 UTC (permalink / raw)
  To: Catalin Marinas, Will Deacon, Billy Laws, Mark Rutland,
	Mark Brown, Ryan Houdek
  Cc: linux-arm-kernel, linux-kernel, kernel-dev, André Almeida

Implement support for emulating unaligned atomic operations on arm64.
User applications that wish to enable support for this should use the
pctrl() flag `PR_ARM64_UNALIGN_ATOMIC_EMULATE`.

Signed-off-by: André Almeida <andrealmeid@igalia.com>
---
 arch/arm64/Kconfig                   |   6 +
 arch/arm64/include/asm/exception.h   |   1 +
 arch/arm64/include/asm/processor.h   |   5 +
 arch/arm64/include/asm/rwonce.h      |  14 +-
 arch/arm64/include/asm/thread_info.h |   1 +
 arch/arm64/kernel/Makefile           |   3 +-
 arch/arm64/kernel/process.c          |  15 +
 arch/arm64/kernel/unaligned_atomic.c | 521 +++++++++++++++++++++++++++
 arch/arm64/mm/fault.c                |  10 +
 include/uapi/linux/prctl.h           |   5 +
 kernel/sys.c                         |   7 +-
 11 files changed, 579 insertions(+), 9 deletions(-)
 create mode 100644 arch/arm64/kernel/unaligned_atomic.c

diff --git a/arch/arm64/Kconfig b/arch/arm64/Kconfig
index b3afe0688919..b8991da50a16 100644
--- a/arch/arm64/Kconfig
+++ b/arch/arm64/Kconfig
@@ -2417,6 +2417,12 @@ config ARM64_CONTPTE
 	  bit, for any mappings that meet the size and alignment requirements.
 	  This reduces TLB pressure and improves performance.
 
+config ARM64_UNALIGNED_ATOMIC_EMULATION
+	bool "Enable unaligned atomic emulation"
+	help
+	  Enable support for handling userspace unaligned atomic instructions
+	  with kernel side emulation, instead of raising a SIGBUS.
+
 endmenu # "Kernel Features"
 
 menu "Boot options"
diff --git a/arch/arm64/include/asm/exception.h b/arch/arm64/include/asm/exception.h
index a2da3cb21c24..f6dd1b9afe69 100644
--- a/arch/arm64/include/asm/exception.h
+++ b/arch/arm64/include/asm/exception.h
@@ -82,6 +82,7 @@ void do_sp_pc_abort(unsigned long addr, unsigned long esr, struct pt_regs *regs)
 void bad_el0_sync(struct pt_regs *regs, int reason, unsigned long esr);
 void do_el0_cp15(unsigned long esr, struct pt_regs *regs);
 int do_compat_alignment_fixup(unsigned long addr, struct pt_regs *regs);
+int do_unaligned_atomic_fixup(struct pt_regs *regs, u64 *fault_addr);
 void do_el0_svc(struct pt_regs *regs);
 void do_el0_svc_compat(struct pt_regs *regs);
 void do_el0_fpac(struct pt_regs *regs, unsigned long esr);
diff --git a/arch/arm64/include/asm/processor.h b/arch/arm64/include/asm/processor.h
index c2a627f39314..da748012a0ed 100644
--- a/arch/arm64/include/asm/processor.h
+++ b/arch/arm64/include/asm/processor.h
@@ -445,5 +445,10 @@ int set_tsc_mode(unsigned int val);
 #define GET_TSC_CTL(adr)        get_tsc_mode((adr))
 #define SET_TSC_CTL(val)        set_tsc_mode((val))
 
+#ifdef CONFIG_ARM64_UNALIGNED_ATOMIC_EMULATION
+int set_unalign_atomic_ctl(unsigned int val);
+#define ARM64_SET_UNALIGN_ATOMIC_CTL(val)	set_unalign_atomic_ctl((val))
+#endif
+
 #endif /* __ASSEMBLER__ */
 #endif /* __ASM_PROCESSOR_H */
diff --git a/arch/arm64/include/asm/rwonce.h b/arch/arm64/include/asm/rwonce.h
index 0f3a01d30f66..3c62a1b52e9d 100644
--- a/arch/arm64/include/asm/rwonce.h
+++ b/arch/arm64/include/asm/rwonce.h
@@ -5,13 +5,6 @@
 #ifndef __ASM_RWONCE_H
 #define __ASM_RWONCE_H
 
-#if defined(CONFIG_LTO) && !defined(__ASSEMBLER__)
-
-#include <linux/compiler_types.h>
-#include <asm/alternative-macros.h>
-
-#ifndef BUILD_VDSO
-
 #define __LOAD_RCPC(sfx, regs...)					\
 	ALTERNATIVE(							\
 		"ldar"	#sfx "\t" #regs,				\
@@ -19,6 +12,13 @@
 		"ldapr"	#sfx "\t" #regs,				\
 	ARM64_HAS_LDAPR)
 
+#if defined(CONFIG_LTO) && !defined(__ASSEMBLER__)
+
+#include <linux/compiler_types.h>
+#include <asm/alternative-macros.h>
+
+#ifndef BUILD_VDSO
+
 /*
  * Replace this with typeof_unqual() when minimum compiler versions are
  * increased to GCC 14 and Clang 19. For the time being, we need this
diff --git a/arch/arm64/include/asm/thread_info.h b/arch/arm64/include/asm/thread_info.h
index 5d7fe3e153c8..c280264f64dd 100644
--- a/arch/arm64/include/asm/thread_info.h
+++ b/arch/arm64/include/asm/thread_info.h
@@ -88,6 +88,7 @@ void arch_setup_new_exec(void);
 #define TIF_KERNEL_FPSTATE	29	/* Task is in a kernel mode FPSIMD section */
 #define TIF_TSC_SIGSEGV		30	/* SIGSEGV on counter-timer access */
 #define TIF_LAZY_MMU_PENDING	31	/* Ops pending for lazy mmu mode exit */
+#define TIF_UNALIGN_ATOMIC_EMULATE	32
 
 #define _TIF_SIGPENDING		(1 << TIF_SIGPENDING)
 #define _TIF_NEED_RESCHED	(1 << TIF_NEED_RESCHED)
diff --git a/arch/arm64/kernel/Makefile b/arch/arm64/kernel/Makefile
index d2690c3ec528..d1dfa7c7cf47 100644
--- a/arch/arm64/kernel/Makefile
+++ b/arch/arm64/kernel/Makefile
@@ -34,7 +34,7 @@ obj-y			:= debug-monitors.o entry.o irq.o fpsimd.o		\
 			   cpufeature.o alternative.o cacheinfo.o		\
 			   smp.o smp_spin_table.o topology.o smccc-call.o	\
 			   syscall.o proton-pack.o idle.o patching.o pi/	\
-			   rsi.o jump_label.o
+			   rsi.o jump_label.o unaligned_atomic.o
 
 obj-$(CONFIG_COMPAT)			+= sys32.o signal32.o			\
 					   sys_compat.o
@@ -72,6 +72,7 @@ obj-$(CONFIG_ARM64_MPAM)		+= mpam.o
 obj-$(CONFIG_ARM64_MTE)			+= mte.o
 obj-y					+= vdso-wrap.o
 obj-$(CONFIG_COMPAT_VDSO)		+= vdso32-wrap.o
+obj-$(CONFIG_ARM64_UNALIGNED_ATOMIC_EMULATION) += unaligned_atomic.o
 
 # Force dependency (vdso*-wrap.S includes vdso.so through incbin)
 $(obj)/vdso-wrap.o: $(obj)/vdso/vdso.so
diff --git a/arch/arm64/kernel/process.c b/arch/arm64/kernel/process.c
index 581f80e9b9b7..3a2cf093c0d2 100644
--- a/arch/arm64/kernel/process.c
+++ b/arch/arm64/kernel/process.c
@@ -1003,3 +1003,18 @@ int set_tsc_mode(unsigned int val)
 
 	return do_set_tsc_mode(val);
 }
+
+int set_unalign_atomic_ctl(unsigned int val)
+{
+	unsigned long valid_mask = PR_ARM64_UNALIGN_ATOMIC_EMULATE;
+
+	if (val & ~valid_mask)
+		return -EINVAL;
+
+	if (!cpus_have_final_cap(ARM64_HAS_LSE_ATOMICS))
+		return -EINVAL;
+
+	update_thread_flag(TIF_UNALIGN_ATOMIC_EMULATE, val & PR_ARM64_UNALIGN_ATOMIC_EMULATE);
+
+	return 0;
+}
diff --git a/arch/arm64/kernel/unaligned_atomic.c b/arch/arm64/kernel/unaligned_atomic.c
new file mode 100644
index 000000000000..5f176226b733
--- /dev/null
+++ b/arch/arm64/kernel/unaligned_atomic.c
@@ -0,0 +1,521 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/*
+ * Unaligned atomic emulation by André Almeida <andrealmeid@igalia.com>
+ * Derived from original work by Billy Laws <blaws05@gmail.com>
+ */
+
+#include <linux/cacheflush.h>
+#include <linux/compiler.h>
+#include <linux/errno.h>
+#include <linux/init.h>
+#include <linux/kernel.h>
+#include <linux/semaphore.h>
+#include <linux/uaccess.h>
+
+#include <asm-generic/rwonce.h>
+#include <asm/asm-extable.h>
+#include <asm/atomic_ll_sc.h>
+#include <asm/exception.h>
+#include <asm/ptrace.h>
+#include <asm/traps.h>
+
+struct fault_info {
+	int error;
+	u64 addr;
+	u64 size;
+};
+
+#define ATOMIC_MEM_MASK 0x3b200c00
+#define ATOMIC_MEM_INST 0x38200000
+
+#define RCPC2_MASK  0x3fe00c00
+#define LDAPUR_INST 0x19400000
+#define STLUR_INST  0x19000000
+
+#define LDAXR_MASK 0x3ffffc00
+
+#define LDAXP_MASK 0xbfff8000
+#define LDAXP_INST 0x887f8000
+
+#define LDAR_INST  0x08dffc00
+#define LDAPR_INST 0x38bfc000
+#define STLR_INST  0x089ffc00
+
+#define ATOMIC_ADD_OP	0x0
+#define ATOMIC_CLR_OP	0x1
+#define ATOMIC_EOR_OP	0x2
+#define ATOMIC_SET_OP	0x3
+#define ATOMIC_SWAP_OP	0x8
+
+#define get_addr_reg(instr) ((instr >> 5) & 0b11111)
+#define get_size(instr) (1 << (instr >> 30))
+
+static DEFINE_MUTEX(buslock);
+
+static struct fault_info do_load_acquire_128(u64 addr, u128 *result)
+{
+	u64 lower, upper, orig_addr = addr;
+	int ret;
+
+	addr = (u64)__uaccess_mask_ptr((void __user *)addr);
+
+	if (!access_ok((void __user *)orig_addr, 16))
+		return (struct fault_info) {.error = -EFAULT, .addr = orig_addr, .size = 16};
+
+	uaccess_enable_privileged();
+	asm volatile(
+		"1: ldaxp 	%[lower], %[upper], [%[addr]]\n"
+		"   clrex\n"
+		"2:\n"
+		_ASM_EXTABLE_UACCESS_ERR(1b, 2b, %w[ret])
+		: [lower] "=r" (lower),
+		  [upper] "=r" (upper),
+		  [ret] "+r" (ret)
+		: [addr] "r" (addr)
+		: "memory");
+	uaccess_disable_privileged();
+
+	if (ret)
+		return (struct fault_info) {.error = ret, .addr = orig_addr, .size = 16};
+
+	*result = (u128)upper << 64 | lower;
+
+	return (struct fault_info) {0};
+}
+
+/*
+ * Do a load acquire at address `addr` and save it at `result`
+ */
+static struct fault_info do_load_acquire_64(u64 addr, u64 *result)
+{
+	u64 orig_addr = addr;
+	int ret = 0;
+
+	addr = (u64)__uaccess_mask_ptr((void __user *)addr);
+
+	if (!access_ok((void __user *)orig_addr, 8))
+		return (struct fault_info) {.error = -EFAULT, .addr = orig_addr, .size = 8};
+
+	uaccess_enable_privileged();
+	asm volatile(
+		"1: " __LOAD_RCPC(, %[result], [%[addr]]) "\n"
+		"2:\n"
+		_ASM_EXTABLE_UACCESS_ERR(1b, 2b, %w[ret])
+		: [result] "=r" (*result),
+		  [ret] "+r" (ret)
+		: [addr] "r" (addr)
+		: "memory");
+	uaccess_disable_privileged();
+
+	return (struct fault_info) {.error = ret, .addr = orig_addr, .size = 8};
+}
+
+/*
+ * Do a cmpxchg taking care of handling user access error
+ */
+static struct fault_info do_store_cas_64(u64 *expected, u64 val, u64 addr)
+{
+	u64 orig_addr = addr, tmp, oldval = 0;
+	int ret = 0;
+
+	addr = (u64)__uaccess_mask_ptr((void __user *)addr);
+
+	if (!access_ok((void __user *)orig_addr, 8))
+		return (struct fault_info) {.error = -EFAULT, .addr = orig_addr, .size = 8};
+
+	uaccess_enable_privileged();
+	asm volatile(
+		"prfm		pstl1strm, [%[addr]]		\n"
+		"1: ldxr	%[oldval], [%[addr]]		\n"
+		"   cmp		%[oldval], %[expected]		\n"
+		"   b.ne 	3f				\n"
+		"2: stlxr 	%w[tmp], %[val], [%[addr]]	\n"
+		"   cbnz 	%w[tmp], 1b			\n"
+		"3:						\n"
+		_ASM_EXTABLE_UACCESS_ERR(1b, 3b, %w[ret])
+		_ASM_EXTABLE_UACCESS_ERR(2b, 3b, %w[ret])
+		: [tmp] "=&r" (tmp),
+		  [oldval] "=&r" (oldval),
+		  [ret] "+r" (ret)
+		: [addr] "r" (addr),
+		  [expected] "r" (*expected),
+		  [val] "r" (val)
+		: "memory", "cc");
+	uaccess_disable_privileged();
+
+	if (ret)
+		return (struct fault_info) {.error = ret, .addr = orig_addr, .size = 8};
+
+	if (oldval != *expected) {
+		*expected = oldval;
+		return (struct fault_info) {.error = -EAGAIN};
+	}
+
+	return (struct fault_info) {0};
+}
+
+/*
+ * If possible, do one 128 bit load. Otherwise, do two 64 bit loads and combine
+ * the results.
+ */
+static struct fault_info do_load_64(u64 addr, u64 *result)
+{
+	u64 alignment_mask = 0b1111, align_offset = addr & alignment_mask;
+	struct fault_info fi = {0};
+	u128 tmp;
+
+	/* The address crosses a 16 byte boundary and needs two loads */
+	if (align_offset > 8) {
+		u64 alignment = addr & 0b111, upper_val, lower_val;
+
+		addr &= ~0b111ULL;
+
+		fi = do_load_acquire_64(addr + 8, &upper_val);
+		if (fi.error)
+			return fi;
+
+		fi = do_load_acquire_64(addr, &lower_val);
+		if (fi.error)
+			return fi;
+
+		tmp = ((u128) upper_val << 64) | lower_val;
+		*result = tmp >> (alignment * 8);
+	} else {
+		addr &= ~alignment_mask;
+
+		fi = do_load_acquire_128(addr, &tmp);
+		if (fi.error)
+			return fi;
+
+		*result = tmp >> (align_offset * 8);
+	}
+
+	return fi;
+}
+
+static u64 do_atomic_mem_op(u8 op, u64 src_val, u64 desired)
+{
+	switch (op) {
+	case ATOMIC_ADD_OP:
+		return src_val + desired;
+	case ATOMIC_CLR_OP:
+		return src_val & ~desired;
+	case ATOMIC_EOR_OP:
+		return src_val & desired;
+	case ATOMIC_SET_OP:
+		return src_val ^ desired;
+	case ATOMIC_SWAP_OP:
+		return desired;
+	default:
+		BUG();
+	}
+
+	/* Unreachable */
+	return 0;
+}
+
+static struct fault_info load_cas(u64 desired_src, u64 addr, u8 op, u64 alignment,
+				     u128 *aux_desired, u128 *aux_expected,
+				     u128 *aux_actual)
+{
+	u128 tmp_expected, tmp_desired, tmp_actual, mask = ~0ULL, desired, expected, neg_mask;
+	u64 addr_upper, load_order_upper, load_order_lower;
+	struct fault_info fi;
+
+	mask <<= alignment * 8;
+	addr_upper = addr + 8;
+	neg_mask = ~mask;
+
+	fi = do_load_acquire_64(addr_upper, &load_order_upper);
+	if (fi.error)
+		return fi;
+	fi = do_load_acquire_64(addr, &load_order_lower);
+	if (fi.error)
+		return fi;
+
+	tmp_actual = (u128)load_order_upper << 64 | load_order_lower;
+
+	desired = do_atomic_mem_op(op, tmp_actual >> (alignment * 8), desired_src);
+	expected = (tmp_actual >> (alignment * 8));
+
+	tmp_expected = tmp_actual;
+	tmp_expected &= neg_mask;
+	tmp_expected |= expected << (alignment * 8);
+
+	tmp_desired = tmp_expected;
+	tmp_desired &= neg_mask;
+	tmp_desired |= desired << (alignment * 8);
+
+	*aux_desired = tmp_desired;
+	*aux_expected = tmp_expected;
+	*aux_actual = tmp_actual;
+
+	return (struct fault_info) {0};
+}
+
+/*
+ * After CAS failed, check if we need to try again or if we should return error.
+ * Returns true if needs to retry.
+ */
+static bool handle_fail(u128 tmp_expected, u128 tmp_desired, u128 mask,
+			u64 *result, bool retry, bool tear,
+			u64 alignment)
+{
+	u128 neg_mask = ~mask,
+	     failed_result_our_bits = tmp_expected & mask,
+	     failed_result_not_our_bits = tmp_expected & neg_mask,
+	     failed_desired_not_our_bits = tmp_desired & neg_mask;
+	u64 failed_result = failed_result_our_bits >> (alignment * 8);
+
+	/*
+	 * If the bits changed weren't part of our regular CAS, then we retry.
+	 */
+	if ((failed_result_not_our_bits ^ failed_desired_not_our_bits) != 0)
+		return true;
+
+	/*
+	 * This happens in the case that between load and CAS that something has
+	 * store our desired in to the memory location. This means our CAS fails
+	 * because what we wanted to store was already stored.
+	 */
+	if (retry) {
+		/* If we are retrying and tearing then we can't do anything */
+		if (tear) {
+			*result = failed_result;
+			return false;
+		} else {
+			return true;
+		}
+	} else {
+		/*
+		 * With we were called without retry, then we have failed
+		 * regardless of tear. CAS failed but handled successfully
+		 */
+		*result = failed_result;
+	}
+
+	return false;
+}
+
+/*
+ * This instruction reads a 64-bit doubleword from memory, and compares it
+ * against the value held in a first register. If the comparison is equal, the
+ * value in a second register is written to memory. If the comparison is not
+ * equal, the architecture permits writing the value read from the location to
+ * memory.
+ *
+ * To handle an unaligned CAS, the code first loads two 64-bit words. Then, if
+ * what is found in the word is the same as the expected value, the code tries
+ * to do two 64-bit writes in the address. If both stores works then return
+ * success. If only the first store works, then the code is in a "tear" state
+ * and returns with the word found in the address.
+ *
+ * TODO: here we threat every operation as it's a cross cache address,
+ * meaning that we need two 64 bit ops to make it work. That works for
+ * all cases, but it's slower and unnecessary for the cases that doesn't
+ * cross it and can do a single 128 bit operation.
+ */
+static struct fault_info do_cas_64(bool retry, u64 desired_src, u64 expected_src,
+				   u64 addr, u8 op, u64 *result)
+{
+	u128 tmp_expected, tmp_desired, tmp_actual, mask = ~0ULL;
+	u64 alignment = addr & 0b111, addr_upper;
+	struct fault_info fi;
+	bool tear = false;
+
+	mask <<= alignment * 8;
+	addr &= ~0b111ULL;
+	addr_upper = addr + 8;
+
+	/*
+	 * TODO: take this lock only if we need to emulate a split lock
+	 */
+	guard(mutex)(&buslock);
+
+retry:
+	fi = load_cas(desired_src, addr, op, alignment, &tmp_desired, &tmp_expected, &tmp_actual);
+	if (fi.error)
+		return fi;
+
+	if (tmp_expected == tmp_actual) {
+		u128 expected = (tmp_actual >> (alignment * 8));
+		u64 tmp_expected_lower = tmp_expected,
+		    tmp_expected_upper = tmp_expected >> 64,
+		    tmp_desired_lower = tmp_desired,
+		    tmp_desired_upper = tmp_desired >> 64;
+
+		fi = do_store_cas_64(&tmp_expected_upper, tmp_desired_upper, addr_upper);
+		if (fi.error && fi.error != -EAGAIN)
+			return fi;
+
+		if (fi.error == 0) {
+			fi = do_store_cas_64(&tmp_expected_lower, tmp_desired_lower, addr);
+			if (fi.error && fi.error != -EAGAIN)
+				return fi;
+
+			/* Both store() worked, CAS succeeded */
+			if (fi.error == 0) {
+				*result = expected;
+				return fi;
+			/* A partial store() happened, tear state */
+			} else {
+				tear = true;
+			}
+		}
+
+		tmp_expected = tmp_expected_upper;
+		tmp_expected <<= 64;
+		tmp_expected |= tmp_expected_lower;
+	} else {
+		tmp_expected = tmp_actual;
+	}
+
+	if (handle_fail(tmp_expected, tmp_desired, mask, result, retry, tear, alignment))
+		goto retry;
+
+	return (struct fault_info) {0};
+}
+
+/*
+ * For a giving atomic memory operation, parse it and call the desired op
+ */
+static struct fault_info handle_atomic_mem_op(u32 instr, struct pt_regs *regs)
+{
+	u32 size = get_size(instr), result_reg = instr & 0b11111,
+		   source_reg = (instr >> 16) & 0b11111,
+		   addr_reg = get_addr_reg(instr);
+	u64 addr = pt_regs_read_reg(regs, addr_reg);
+	u8 op = (instr >> 12) & 0xf;
+	struct fault_info fi = {0};
+
+	if (size == 8) {
+		u64 res = 0;
+
+		switch (op) {
+		case ATOMIC_ADD_OP:
+		case ATOMIC_CLR_OP:
+		case ATOMIC_EOR_OP:
+		case ATOMIC_SET_OP:
+		case ATOMIC_SWAP_OP:
+			break;
+		default:
+			pr_warn("Unhandled atomic mem op 0x%02x\n", op);
+			return (struct fault_info) {.error = -EINVAL};
+		}
+
+		fi = do_cas_64(true, pt_regs_read_reg(regs, source_reg), 0, addr, op, &res);
+
+		/*
+		 * If the operation succeeded and our dest reg is not zero, we
+		 * need to update the result reg with what was in memory
+		 */
+		if (!fi.error)
+			pt_regs_write_reg(regs, result_reg, res);
+	} else {
+		fi.error = -EINVAL;
+	}
+
+	return fi;
+}
+
+static struct fault_info handle_atomic_load(u32 instr, struct pt_regs *regs, s64 offset)
+{
+	u32 size = get_size(instr), result_reg = instr & 0b11111,
+		   addr_reg = get_addr_reg(instr);
+	u64 addr = pt_regs_read_reg(regs, addr_reg) + offset, res;
+
+	struct fault_info fi = {0};
+
+	if (size == 8) {
+		fi = do_load_64(addr, &res);
+		if (!fi.error)
+			pt_regs_write_reg(regs, result_reg, res);
+	} else {
+		fi.error = -EINVAL;
+	}
+
+	return fi;
+}
+
+static struct fault_info handle_atomic_store(u32 instr, struct pt_regs *regs, s64 offset)
+{
+	u32 size = get_size(instr), data_reg = instr & 0x1F, addr_reg = get_addr_reg(instr);
+	u64 addr = pt_regs_read_reg(regs, addr_reg) + offset;
+	struct fault_info fi = {0};
+
+	if (size == 8) {
+		u64 res;
+
+		fi = do_cas_64(false, pt_regs_read_reg(regs, data_reg), 0, addr, ATOMIC_SWAP_OP, &res);
+	} else {
+		fi.error = -EINVAL;
+	}
+
+	return fi;
+}
+
+static struct fault_info decode_instruction(u32 instr, struct pt_regs *regs)
+{
+	struct fault_info fi = {0};
+	s32 offset = (s32)(instr) << 11 >> 23, size = get_size(instr);
+
+	/* TODO: We only support 64-bit instructions for now */
+	if (size != 8)
+		goto exit;
+
+	if ((instr & LDAXR_MASK) == LDAR_INST || (instr & LDAXR_MASK) == LDAPR_INST)
+		return handle_atomic_load(instr, regs, 0);
+
+	if ((instr & RCPC2_MASK) == LDAPUR_INST)
+		return handle_atomic_load(instr, regs, offset);
+
+	if ((instr & LDAXR_MASK) == STLR_INST)
+		return  handle_atomic_store(instr, regs, 0);
+
+	if ((instr & RCPC2_MASK) == STLUR_INST)
+		return handle_atomic_store(instr, regs, offset);
+
+	if ((instr & ATOMIC_MEM_MASK) == ATOMIC_MEM_INST)
+		return handle_atomic_mem_op(instr, regs);
+
+	/* TODO: Handle CASAL and CASPAL as well */
+
+exit:
+	fi.error = -EINVAL;
+	return fi;
+}
+
+/*
+ * After a fault access happened, move the pointer to the end of the bad section
+ */
+static void increment_fault_address(u64 **fault_addr, int size)
+{
+	u64 *addr = *fault_addr;
+	u8 tmp;
+	int i;
+
+	for (i = 0; i < size - 1 && !get_user(tmp, (u8 __user *)(addr + i)); i++)
+		(*fault_addr)++;
+}
+
+int do_unaligned_atomic_fixup(struct pt_regs *regs, u64 *fault_addr)
+{
+	u32 *pc = (u32 *)regs->pc, instr;
+	struct fault_info fi;
+
+	if (get_user(instr, (u32 __user *)(pc)))
+		return -EFAULT;
+
+	fi = decode_instruction(instr, regs);
+
+	if (fi.error) {
+		*fault_addr = fi.addr;
+		if (fi.error == -EFAULT)
+			increment_fault_address(&fault_addr, fi.size);
+
+		return fi.error;
+	}
+
+	arm64_skip_faulting_instruction(regs, 4);
+	return 0;
+}
diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c
index 0b52557652be..4aba725c1fd1 100644
--- a/arch/arm64/mm/fault.c
+++ b/arch/arm64/mm/fault.c
@@ -854,6 +854,16 @@ static int do_alignment_fault(unsigned long far, unsigned long esr,
 	if (IS_ENABLED(CONFIG_COMPAT_ALIGNMENT_FIXUPS) &&
 	    compat_user_mode(regs))
 		return do_compat_alignment_fixup(far, regs);
+
+	if (user_mode(regs) && test_thread_flag(TIF_UNALIGN_ATOMIC_EMULATE)) {
+		u64 page_fault_address;
+		int ret = do_unaligned_atomic_fixup(regs, &page_fault_address);
+
+		if (!ret)
+			return 0;
+		else if (ret == -EFAULT)
+			return do_translation_fault(page_fault_address, ESR_ELx_FSC_FAULT | ESR_ELx_WNR, regs);
+	}
 	do_bad_area(far, esr, regs);
 	return 0;
 }
diff --git a/include/uapi/linux/prctl.h b/include/uapi/linux/prctl.h
index 55ca31b522fc..21171af687cf 100644
--- a/include/uapi/linux/prctl.h
+++ b/include/uapi/linux/prctl.h
@@ -421,4 +421,9 @@ struct prctl_mm_map {
 # define PR_SET_COMPAT_INPUT_DISABLE  0
 # define PR_SET_COMPAT_INPUT_ENABLE   1
 
+#define PR_ARM64_SET_UNALIGN_ATOMIC 0x46455849
+# define PR_ARM64_UNALIGN_ATOMIC_EMULATE	(1UL << 0)
+# define PR_ARM64_UNALIGN_ATOMIC_BACKPATCH	(1UL << 1)
+# define PR_ARM64_UNALIGN_ATOMIC_STRICT_SPLIT_LOCKS	(1UL << 2)
+
 #endif /* _LINUX_PRCTL_H */
diff --git a/kernel/sys.c b/kernel/sys.c
index 0b8bb0d4803c..60e64512a11d 100644
--- a/kernel/sys.c
+++ b/kernel/sys.c
@@ -158,7 +158,9 @@
 #ifndef PPC_SET_DEXCR_ASPECT
 # define PPC_SET_DEXCR_ASPECT(a, b, c)	(-EINVAL)
 #endif
-
+#ifndef ARM64_SET_UNALIGN_ATOMIC_CTL
+# define ARM64_SET_UNALIGN_ATOMIC_CTL(a)		(-EINVAL)
+#endif
 /*
  * this is where the system-wide overflow UID and GID are defined, for
  * architectures that now have 32-bit UID/GID but didn't in the past
@@ -2922,6 +2924,9 @@ SYSCALL_DEFINE5(prctl, int, option, unsigned long, arg2, unsigned long, arg3,
 		else
 			return -EINVAL;
 		break;		
+	case PR_ARM64_SET_UNALIGN_ATOMIC:
+		error = ARM64_SET_UNALIGN_ATOMIC_CTL(arg2);
+		break;
 	default:
 		trace_task_prctl_unknown(option, arg2, arg3, arg4, arg5);
 		error = -EINVAL;
-- 
2.55.0


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [RFC PATCH v3 1/1] arch: arm64: Implement unaligned atomic emulation
  2026-09-30  2:01 ` [RFC PATCH v3 1/1] " André Almeida
@ 2026-09-30  7:48   ` Will Deacon
  2026-09-30 10:29     ` André Almeida
  0 siblings, 1 reply; 5+ messages in thread
From: Will Deacon @ 2026-09-30  7:48 UTC (permalink / raw)
  To: André Almeida
  Cc: Catalin Marinas, Billy Laws, Mark Rutland, Mark Brown,
	Ryan Houdek, linux-arm-kernel, linux-kernel, kernel-dev

On Tue, Sep 29, 2026 at 11:01:38PM -0300, André Almeida wrote:
> Implement support for emulating unaligned atomic operations on arm64.
> User applications that wish to enable support for this should use the
> pctrl() flag `PR_ARM64_UNALIGN_ATOMIC_EMULATE`.
> 
> Signed-off-by: André Almeida <andrealmeid@igalia.com>
> ---
>  arch/arm64/Kconfig                   |   6 +
>  arch/arm64/include/asm/exception.h   |   1 +
>  arch/arm64/include/asm/processor.h   |   5 +
>  arch/arm64/include/asm/rwonce.h      |  14 +-
>  arch/arm64/include/asm/thread_info.h |   1 +
>  arch/arm64/kernel/Makefile           |   3 +-
>  arch/arm64/kernel/process.c          |  15 +
>  arch/arm64/kernel/unaligned_atomic.c | 521 +++++++++++++++++++++++++++
>  arch/arm64/mm/fault.c                |  10 +
>  include/uapi/linux/prctl.h           |   5 +
>  kernel/sys.c                         |   7 +-
>  11 files changed, 579 insertions(+), 9 deletions(-)
>  create mode 100644 arch/arm64/kernel/unaligned_atomic.c

No.

I already explained to you why this doesn't work:

https://lore.kernel.org/r/aV1YnOetDHhKe4hz@willie-the-truck

Will

^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [RFC PATCH v3 1/1] arch: arm64: Implement unaligned atomic emulation
  2026-09-30  7:48   ` Will Deacon
@ 2026-09-30 10:29     ` André Almeida
  2026-09-30 13:29       ` Will Deacon
  0 siblings, 1 reply; 5+ messages in thread
From: André Almeida @ 2026-09-30 10:29 UTC (permalink / raw)
  To: Will Deacon
  Cc: Catalin Marinas, Billy Laws, Mark Rutland, Mark Brown,
	Ryan Houdek, linux-arm-kernel, linux-kernel, kernel-dev

Hi Will, thanks for the quickly reply,

Em 30/09/2026 04:48, Will Deacon escreveu:
> On Tue, Sep 29, 2026 at 11:01:38PM -0300, André Almeida wrote:
>> Implement support for emulating unaligned atomic operations on arm64.
>> User applications that wish to enable support for this should use the
>> pctrl() flag `PR_ARM64_UNALIGN_ATOMIC_EMULATE`.
>>
>> Signed-off-by: André Almeida <andrealmeid@igalia.com>
>> ---
>>   arch/arm64/Kconfig                   |   6 +
>>   arch/arm64/include/asm/exception.h   |   1 +
>>   arch/arm64/include/asm/processor.h   |   5 +
>>   arch/arm64/include/asm/rwonce.h      |  14 +-
>>   arch/arm64/include/asm/thread_info.h |   1 +
>>   arch/arm64/kernel/Makefile           |   3 +-
>>   arch/arm64/kernel/process.c          |  15 +
>>   arch/arm64/kernel/unaligned_atomic.c | 521 +++++++++++++++++++++++++++
>>   arch/arm64/mm/fault.c                |  10 +
>>   include/uapi/linux/prctl.h           |   5 +
>>   kernel/sys.c                         |   7 +-
>>   11 files changed, 579 insertions(+), 9 deletions(-)
>>   create mode 100644 arch/arm64/kernel/unaligned_atomic.c
> 
> No.
> 
> I already explained to you why this doesn't work:
> 
> https://lore.kernel.org/r/aV1YnOetDHhKe4hz@willie-the-truck

Indeed, last time you raised some points, and then Ryan replied them. Is 
there any specific point that doesn't work? I couldn't find a reply for 
Ryan's answers:

https://lore.kernel.org/all/CABnRqDf5EQUoXu=pJ6mj4-JfwAzEfcAE2cYrNzJANFycx7cMUA@mail.gmail.com/

> 
> Will


^ permalink raw reply	[flat|nested] 5+ messages in thread

* Re: [RFC PATCH v3 1/1] arch: arm64: Implement unaligned atomic emulation
  2026-09-30 10:29     ` André Almeida
@ 2026-09-30 13:29       ` Will Deacon
  0 siblings, 0 replies; 5+ messages in thread
From: Will Deacon @ 2026-09-30 13:29 UTC (permalink / raw)
  To: André Almeida
  Cc: Catalin Marinas, Billy Laws, Mark Rutland, Mark Brown,
	Ryan Houdek, linux-arm-kernel, linux-kernel, kernel-dev

On Wed, Sep 30, 2026 at 07:29:22AM -0300, André Almeida wrote:
> Em 30/09/2026 04:48, Will Deacon escreveu:
> > On Tue, Sep 29, 2026 at 11:01:38PM -0300, André Almeida wrote:
> > > Implement support for emulating unaligned atomic operations on arm64.
> > > User applications that wish to enable support for this should use the
> > > pctrl() flag `PR_ARM64_UNALIGN_ATOMIC_EMULATE`.
> > > 
> > > Signed-off-by: André Almeida <andrealmeid@igalia.com>
> > > ---
> > >   arch/arm64/Kconfig                   |   6 +
> > >   arch/arm64/include/asm/exception.h   |   1 +
> > >   arch/arm64/include/asm/processor.h   |   5 +
> > >   arch/arm64/include/asm/rwonce.h      |  14 +-
> > >   arch/arm64/include/asm/thread_info.h |   1 +
> > >   arch/arm64/kernel/Makefile           |   3 +-
> > >   arch/arm64/kernel/process.c          |  15 +
> > >   arch/arm64/kernel/unaligned_atomic.c | 521 +++++++++++++++++++++++++++
> > >   arch/arm64/mm/fault.c                |  10 +
> > >   include/uapi/linux/prctl.h           |   5 +
> > >   kernel/sys.c                         |   7 +-
> > >   11 files changed, 579 insertions(+), 9 deletions(-)
> > >   create mode 100644 arch/arm64/kernel/unaligned_atomic.c
> > 
> > No.
> > 
> > I already explained to you why this doesn't work:
> > 
> > https://lore.kernel.org/r/aV1YnOetDHhKe4hz@willie-the-truck
> 
> Indeed, last time you raised some points, and then Ryan replied them. Is
> there any specific point that doesn't work? I couldn't find a reply for
> Ryan's answers:
> 
> https://lore.kernel.org/all/CABnRqDf5EQUoXu=pJ6mj4-JfwAzEfcAE2cYrNzJANFycx7cMUA@mail.gmail.com/

So rather than get involved in the discussion, you did nothing for almost
a year and then resent the exact same patch? Why?

I don't think the implementation is correct and I don't think we should
be emulating this either. I hope I made that clear last year. Ryan
thinks it's "fine" due to the locking, but I don't see how that helps
with the example I gave.

You are apparently the author of this code, so you need to get involved
in the discussion if you want to move the needle.

Will

^ permalink raw reply	[flat|nested] 5+ messages in thread

end of thread, other threads:[~2026-09-30 13:29 UTC | newest]

Thread overview: 5+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30  2:01 [RFC PATCH v3 0/1] arch: arm64: Implement unaligned atomic emulation André Almeida
2026-09-30  2:01 ` [RFC PATCH v3 1/1] " André Almeida
2026-09-30  7:48   ` Will Deacon
2026-09-30 10:29     ` André Almeida
2026-09-30 13:29       ` Will Deacon

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®