From: Jinjie Ruan <ruanjinjie@huawei.com>
To: Petr Mladek <pmladek@suse.com>
Cc: <bcrl@kvack.org>, <viro@zeniv.linux.org.uk>, <brauner@kernel.org>,
<jack@suse.cz>, <tytso@mit.edu>, <adilger.kernel@dilger.ca>,
<libaokun@linux.alibaba.com>, <ojaswin@linux.ibm.com>,
<ritesh.list@gmail.com>, <yi.zhang@huawei.com>,
<sforshee@kernel.org>, <akpm@linux-foundation.org>,
<rostedt@goodmis.org>, <andriy.shevchenko@linux.intel.com>,
<linux@rasmusvillemoes.dk>, <senozhatsky@chromium.org>,
<kees@kernel.org>, <tglx@kernel.org>,
<linux-fsdevel@vger.kernel.org>, <linux-aio@kvack.org>,
<linux-kernel@vger.kernel.org>, <linux-ext4@vger.kernel.org>,
<linux-riscv@lists.infradead.org>
Subject: Re: [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance
Date: Mon, 7 Sep 2026 19:29:06 +0800 [thread overview]
Message-ID: <75aeca50-74bf-4888-9dc8-7ac3c5af32ca@huawei.com> (raw)
In-Reply-To: <apq5iLOF8lAQ_ZVU@pathway.suse.cz>
在 2026/9/4 20:28, Petr Mladek 写道:
> Adding Risc-V list into Cc.
>
> On Wed 2026-09-02 15:47:57, Jinjie Ruan wrote:
>> Hi,
>>
>> This series converts some existing smp_wmb()/smp_rmb() barrier pairs to
>> smp_store_release()/smp_load_acquire() across various subsystems.
>>
>> Background
>> ==========
>>
>> Many architectures support load acquire and store release instructions
>> which can replace explicit memory barriers and save cycles. As noted
>> in the ARM architecture reference [1]:
>>
>> "Weaker ordering requirements that are imposed by Load-Acquire and
>> Store-Release instructions allow for micro-architectural
>> optimizations, which could reduce some of the performance impacts
>> that are otherwise imposed by an explicit memory barrier.
>>
>> If the ordering requirement is satisfied using either a Load-Acquire
>> or Store-Release, then it would be preferable to use these
>> instructions instead of a DMB."
>>
>> On arm64, a typical seqcount [2] read loop requires 13 cycles with DMB
>> barriers. Replacing the read barrier with smp_load_acquire() reduces
>> this to 8 cycles on an Ampere Altra.
>
> I wonder if this is true on all other architectures:
>
> + It seems that Arm gets the gain because the instruction
> does both load/store + barrier. It helps even when
> the barrier is full.
>
> + Some other architectures need two instructions. One for the
> load/store and the other for the barrier. But the barrier
> is weaker, it synchronizes just reads or just writes.
>
> For example, I see the following in riscv/include/asm/barrier.h:
>
> <paste riscv/include/asm/barrier.h>
> #define smp_mb() RISCV_FENCE(rw, rw)
> #define smp_rmb() RISCV_FENCE(r, r)
> #define smp_wmb() RISCV_FENCE(w, w)
>
> #define smp_store_release(p, v) \
> do { \
> RISCV_FENCE(rw, w); \
> WRITE_ONCE(*p, v); \
> } while (0)
>
> #define smp_load_acquire(p) \
> ({ \
> typeof(*p) ___p1 = READ_ONCE(*p); \
> RISCV_FENCE(r, rw); \
> ___p1; \
> })
> </paste riscv/include/asm/barrier.h>
>
> I wonder whether:
>
> + RISCV_FENCE(r, r) is faster than RISCV_FENCE(r, rw)
> + RISCV_FENCE(w, w) is faster than RISCV_FENCE(rw, w)
>
> so it might cause performance regression there...
Hi Petr,
Thanks for the detailed analysis. You're right that on some
architectures the acquire/release variants use a slightly heavier
fence than the plain smp_wmb()/smp_rmb() pair. The full picture by
architecture:
arm64: Improvement (DMB ISHST/ISHLD + STR/LDR → STLR/LDAR)
x86: Neutral (both are compiler barriers)
s390: Neutral (both are compiler barriers)
ppc64: Neutral (both use lwsync)
loongarch: Slightly heavier fence (DBAR(o_w_w) → DBAR(orw_w),
DBAR(or_r_) → DBAR(or_rw))
riscv: Slightly heavier fence (fence w,w → fence rw,w,
fence r,r → fence r,rw)
So on Loongarch and Riscv, there may be a slight performance regression.
Regards,
Jinjie
>
> Best Regards,
> Petr
>
>> We also observed significant barrier overhead while profiling Unxibench
>> syscall test on arm64: a single getuid() call is ~8ns slower than on
>> a comparable x86 system, with the dominant cost in map_id_up()'s smp_rmb(),
>> which is a DMB ISHLD on arm64. Converting it to smp_load_acquire() allows
>> the use of LDAR, eliminating the measurable overhead.
>>
>> This motivated a broader search for existing barrier pairs that can
>> be converted to the lighter acquire/release semantics.
>>
>> Changes
>> =======
>>
>> Each patch in this series targets a specific barrier pair where the
>> publish/subscribe pattern is already present:
>>
>> - Writers populate data, then publish a flag/count/pointer via
>> smp_store_release()
>>
>> - Readers load the flag/count/pointer via smp_load_acquire(), then
>> consume the data
>>
>> This preserves the existing memory ordering guarantees while allowing
>> architectures with native acquire/release instructions (e.g. arm64's
>> STLR/LDAR) to avoid the cost of full one-way barriers (DMB ISHST/ISHLD).
>> On architectures without native support, the generated code is
>> generally no worse than the explicit barrier pair.
>>
>> The conversions are mechanical and no functional change is intended.
>>
>> Testing (Kunpeng HIP09 arm64 server)
>> ====================================
>>
>> 1. UNIXBENCH syscall
>> Baseline: 715.27
>> Patched: 718.83
>> Improvement: +0.50%
>>
>> 2. fs/aio (fio + null_blk, 4 jobs):
>> Baseline: 1441k IOPS, 86.46us
>> Patched: 1452k IOPS, 85.80us
>> Improvement: ~0.8%
>>
>> Both improvements are consistent across runs and align with the
>> expected savings from replacing DMB with LDAR/STLR on arm64.
>>
>> [1]: https://support.arm.com/documentation/102336/0100/Load-Acquire-and-Store-Release-instructions
>> [2]: https://github.com/torvalds/linux/commit/d0dd066a0fa26d55c19ace9e89dedd9504c5bcba
>>
>> Changes in v3:
>> - Add Reviewed-by.
>> - Split out network patch set as Kuniyuki suggested.
>> - Link to v2: https://lore.kernel.org/all/20260901024234.135119-1-ruanjinjie@huawei.com/
>>
>> Changes in v2:
>> - Fix pre-existing issue for ext4 and 8021q [3].
>> - Fix missing copy_mnt_idmap() udapte [3].
>> - Drop nacked isotp patch.
>> - Add test data.
>> - Add Reviewed-by and update fs patch as Jan suggested.
>>
>> [3]: https://sashiko.dev/#/patchset/20260825095422.3166067-1-ruanjinjie%40huawei.com
>>
>> Jinjie Ruan (8):
>> user_namespace: Use acquire/release for nr_extents synchronization
>> lib/vsprintf: Use acquire/release for ptr_key publication
>> fs: aio: Use acquire/release for ring->tail publication
>> fs: Use acquire/release for fdtable resize synchronization
>> pidfs: Use test_bit_acquire() for attr flag tests
>> super: Use acquire for SB_BORN check in super_cache_count()
>> ext4: Fix out-of-bounds read in ext4_get_group_info()
>> ext4: Convert group-count barrier protocol to acquire/release
>>
>> fs/aio.c | 10 ++++------
>> fs/ext4/balloc.c | 2 +-
>> fs/ext4/ext4.h | 10 +++-------
>> fs/ext4/mballoc.c | 6 ++----
>> fs/ext4/resize.c | 19 +++++++++++--------
>> fs/file.c | 10 ++++------
>> fs/mnt_idmapping.c | 5 ++---
>> fs/pidfs.c | 6 ++----
>> fs/super.c | 6 ++----
>> kernel/user_namespace.c | 24 +++++++++++++-----------
>> lib/vsprintf.c | 11 ++++-------
>> 11 files changed, 48 insertions(+), 61 deletions(-)
>>
>> --
>> 2.34.1
next prev parent reply other threads:[~2026-09-07 11:29 UTC|newest]
Thread overview: 22+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-09-02 7:47 Jinjie Ruan
2026-09-02 7:47 ` [PATCH v3 1/8] user_namespace: Use acquire/release for nr_extents synchronization Jinjie Ruan
2026-09-02 10:09 ` Bradley Morgan
2026-09-02 13:03 ` Andy Shevchenko
2026-09-02 13:04 ` Andy Shevchenko
2026-09-03 1:42 ` Jinjie Ruan
2026-09-03 2:16 ` Jinjie Ruan
2026-09-02 7:47 ` [PATCH v3 2/8] lib/vsprintf: Use acquire/release for ptr_key publication Jinjie Ruan
2026-09-04 13:13 ` Petr Mladek
2026-09-02 7:48 ` [PATCH v3 3/8] fs: aio: Use acquire/release for ring->tail publication Jinjie Ruan
2026-09-02 7:48 ` [PATCH v3 4/8] fs: Use acquire/release for fdtable resize synchronization Jinjie Ruan
2026-09-02 7:48 ` [PATCH v3 5/8] pidfs: Use test_bit_acquire() for attr flag tests Jinjie Ruan
2026-09-02 7:48 ` [PATCH v3 6/8] super: Use acquire for SB_BORN check in super_cache_count() Jinjie Ruan
2026-09-02 7:48 ` [PATCH v3 7/8] ext4: Fix out-of-bounds read in ext4_get_group_info() Jinjie Ruan
2026-09-02 7:48 ` [PATCH v3 8/8] ext4: Convert group-count barrier protocol to acquire/release Jinjie Ruan
2026-09-02 14:19 ` [PATCH v3 0/8] Convert barrier pairs to acquire/release for better performance Theodore Tso
2026-09-02 14:56 ` Jan Kara
2026-09-04 7:20 ` Jinjie Ruan
2026-09-04 12:28 ` Petr Mladek
2026-09-07 11:29 ` Jinjie Ruan [this message]
2026-09-07 12:59 ` David Laight
2026-09-09 7:11 ` Jinjie Ruan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=75aeca50-74bf-4888-9dc8-7ac3c5af32ca@huawei.com \
--to=ruanjinjie@huawei.com \
--cc=adilger.kernel@dilger.ca \
--cc=akpm@linux-foundation.org \
--cc=andriy.shevchenko@linux.intel.com \
--cc=bcrl@kvack.org \
--cc=brauner@kernel.org \
--cc=jack@suse.cz \
--cc=kees@kernel.org \
--cc=libaokun@linux.alibaba.com \
--cc=linux-aio@kvack.org \
--cc=linux-ext4@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-riscv@lists.infradead.org \
--cc=linux@rasmusvillemoes.dk \
--cc=ojaswin@linux.ibm.com \
--cc=pmladek@suse.com \
--cc=ritesh.list@gmail.com \
--cc=rostedt@goodmis.org \
--cc=senozhatsky@chromium.org \
--cc=sforshee@kernel.org \
--cc=tglx@kernel.org \
--cc=tytso@mit.edu \
--cc=viro@zeniv.linux.org.uk \
--cc=yi.zhang@huawei.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®