* [PATCH] locking/rwsem: Remove reader optimistic lock stealing
@ 2026-05-21 9:59 Peng Wang
2026-05-22 2:08 ` Waiman Long
2026-05-31 7:27 ` [PATCH] locking/rwsem: Remove reader optimistic lock stealing kernel test robot
0 siblings, 2 replies; 8+ messages in thread
From: Peng Wang @ 2026-05-21 9:59 UTC (permalink / raw)
To: peterz, mingo, will, boqun, longman, dbueso; +Cc: linux-kernel, Peng Wang
Reader optimistic lock stealing, introduced by commit 1a728dff855a
("locking/rwsem: Enable reader optimistic lock stealing") and made more
aggressive by commit 617f3ef95177 ("locking/rwsem: Remove reader
optimistic spinning"), allows a reader entering the slowpath to bypass
the wait queue and acquire the lock directly when WRITER_LOCKED and
HANDOFF bits are not set.
This causes severe writer starvation in workloads where readers hold
the lock for extended periods, such as Direct I/O operations which
hold inode->i_rwsem for the entire duration of iomap_dio_rw(). A
common example is log-structured storage where one thread appends via
DIO writes while another thread tails the log via DIO reads -- a
pattern seen in database redo-log replay and shared-storage
replication.
The problem occurs because:
1. A reader entering the slowpath (due to RWSEM_FLAG_WAITERS being set)
can still steal the lock as the steal condition only checks
WRITER_LOCKED and HANDOFF, not WAITERS. This is inconsistent with
the fast path which already blocks readers when WAITERS is set (via
RWSEM_READ_FAILED_MASK).
2. In the window between the last reader releasing the lock and the
waiting writer being scheduled (~10-100us), a new reader can steal
the lock in ~50-100ns. This race is structurally inevitable due to
the 1000x speed difference between atomic operations vs context
switching.
3. Each stolen read lock is held for the full DIO duration (potentially
milliseconds), and the pattern repeats until the 4ms HANDOFF timeout.
This effectively taxes every write operation with a ~4ms penalty.
Performance impact measured with a DIO mixed read/write workload
(1 writer + 1 reader, O_DIRECT, ext4):
NVMe SSD: Before After
Write-only baseline: 397 MB/s 397 MB/s (no change)
Mixed write throughput: 11 MB/s 350 MB/s (+31x)
Mixed write latency: 880 us 23 us (-38x)
Mixed read throughput: 95 MB/s 95 MB/s (no change)
Fixes: 617f3ef95177 ("locking/rwsem: Remove reader optimistic spinning")
Signed-off-by: Peng Wang <peng_wang@linux.alibaba.com>
---
kernel/locking/lock_events_list.h | 1 -
kernel/locking/rwsem.c | 33 ---------------------------------
2 files changed, 34 deletions(-)
diff --git a/kernel/locking/lock_events_list.h b/kernel/locking/lock_events_list.h
index 97fb6f3f840a..35b45576bee4 100644
--- a/kernel/locking/lock_events_list.h
+++ b/kernel/locking/lock_events_list.h
@@ -65,7 +65,6 @@ LOCK_EVENT(rwsem_opt_lock) /* # of opt-acquired write locks */
LOCK_EVENT(rwsem_opt_fail) /* # of failed optspins */
LOCK_EVENT(rwsem_opt_nospin) /* # of disabled optspins */
LOCK_EVENT(rwsem_rlock) /* # of read locks acquired */
-LOCK_EVENT(rwsem_rlock_steal) /* # of read locks by lock stealing */
LOCK_EVENT(rwsem_rlock_fast) /* # of fast read locks acquired */
LOCK_EVENT(rwsem_rlock_fail) /* # of failed read lock acquisitions */
LOCK_EVENT(rwsem_rlock_handoff) /* # of read lock handoffs */
diff --git a/kernel/locking/rwsem.c b/kernel/locking/rwsem.c
index bda5577339c0..40b141c5765f 100644
--- a/kernel/locking/rwsem.c
+++ b/kernel/locking/rwsem.c
@@ -1017,42 +1017,9 @@ static struct rw_semaphore __sched *
rwsem_down_read_slowpath(struct rw_semaphore *sem, long count, unsigned int state)
{
long adjustment = -RWSEM_READER_BIAS;
- long rcnt = (count >> RWSEM_READER_SHIFT);
struct rwsem_waiter waiter, *first;
DEFINE_WAKE_Q(wake_q);
- /*
- * To prevent a constant stream of readers from starving a sleeping
- * writer, don't attempt optimistic lock stealing if the lock is
- * very likely owned by readers.
- */
- if ((atomic_long_read(&sem->owner) & RWSEM_READER_OWNED) &&
- (rcnt > 1) && !(count & RWSEM_WRITER_LOCKED))
- goto queue;
-
- /*
- * Reader optimistic lock stealing.
- */
- if (!(count & (RWSEM_WRITER_LOCKED | RWSEM_FLAG_HANDOFF))) {
- rwsem_set_reader_owned(sem);
- lockevent_inc(rwsem_rlock_steal);
-
- /*
- * Wake up other readers in the wait queue if it is
- * the first reader.
- */
- if ((rcnt == 1) && (count & RWSEM_FLAG_WAITERS)) {
- raw_spin_lock_irq(&sem->wait_lock);
- if (sem->first_waiter)
- rwsem_mark_wake(sem, RWSEM_WAKE_READ_OWNED,
- &wake_q);
- raw_spin_unlock_irq(&sem->wait_lock);
- wake_up_q(&wake_q);
- }
- return sem;
- }
-
-queue:
waiter.task = current;
waiter.type = RWSEM_WAITING_FOR_READ;
waiter.timeout = jiffies + RWSEM_WAIT_TIMEOUT;
--
2.39.3
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] locking/rwsem: Remove reader optimistic lock stealing
2026-05-21 9:59 [PATCH] locking/rwsem: Remove reader optimistic lock stealing Peng Wang
@ 2026-05-22 2:08 ` Waiman Long
2026-05-22 3:33 ` Peng Wang
2026-05-22 8:55 ` Peter Zijlstra
2026-05-31 7:27 ` [PATCH] locking/rwsem: Remove reader optimistic lock stealing kernel test robot
1 sibling, 2 replies; 8+ messages in thread
From: Waiman Long @ 2026-05-22 2:08 UTC (permalink / raw)
To: Peng Wang, peterz, mingo, will, boqun, dbueso; +Cc: linux-kernel
On 5/21/26 5:59 AM, Peng Wang wrote:
> Reader optimistic lock stealing, introduced by commit 1a728dff855a
> ("locking/rwsem: Enable reader optimistic lock stealing") and made more
> aggressive by commit 617f3ef95177 ("locking/rwsem: Remove reader
> optimistic spinning"), allows a reader entering the slowpath to bypass
> the wait queue and acquire the lock directly when WRITER_LOCKED and
> HANDOFF bits are not set.
>
> This causes severe writer starvation in workloads where readers hold
> the lock for extended periods, such as Direct I/O operations which
> hold inode->i_rwsem for the entire duration of iomap_dio_rw(). A
> common example is log-structured storage where one thread appends via
> DIO writes while another thread tails the log via DIO reads -- a
> pattern seen in database redo-log replay and shared-storage
> replication.
It is generally assume that reader lock critical section is shorter than
that of writer. In this particular case, does the reader critical
section run longer than the writer's one?
Reader lock stealing should only happen if the previous lock owner is a
writer. So readers and writer should at most alternately own the lock if
there are many readers waiting. Of course, if a reader own the lock, it
will wake up the remaining readers in the wait queue.
>
> The problem occurs because:
>
> 1. A reader entering the slowpath (due to RWSEM_FLAG_WAITERS being set)
> can still steal the lock as the steal condition only checks
> WRITER_LOCKED and HANDOFF, not WAITERS. This is inconsistent with
> the fast path which already blocks readers when WAITERS is set (via
> RWSEM_READ_FAILED_MASK).
>
> 2. In the window between the last reader releasing the lock and the
> waiting writer being scheduled (~10-100us), a new reader can steal
> the lock in ~50-100ns. This race is structurally inevitable due to
> the 1000x speed difference between atomic operations vs context
> switching.
>
> 3. Each stolen read lock is held for the full DIO duration (potentially
> milliseconds), and the pattern repeats until the 4ms HANDOFF timeout.
> This effectively taxes every write operation with a ~4ms penalty.
>
> Performance impact measured with a DIO mixed read/write workload
> (1 writer + 1 reader, O_DIRECT, ext4):
>
> NVMe SSD: Before After
> Write-only baseline: 397 MB/s 397 MB/s (no change)
> Mixed write throughput: 11 MB/s 350 MB/s (+31x)
> Mixed write latency: 880 us 23 us (-38x)
> Mixed read throughput: 95 MB/s 95 MB/s (no change)
Does the reader wait for something else and taking other locks after
acquiring the read lock? Can you point me of the reader critical section
for this particular workload?
Cheers,
Longman
>
> Fixes: 617f3ef95177 ("locking/rwsem: Remove reader optimistic spinning")
> Signed-off-by: Peng Wang <peng_wang@linux.alibaba.com>
> ---
> kernel/locking/lock_events_list.h | 1 -
> kernel/locking/rwsem.c | 33 ---------------------------------
> 2 files changed, 34 deletions(-)
>
> diff --git a/kernel/locking/lock_events_list.h b/kernel/locking/lock_events_list.h
> index 97fb6f3f840a..35b45576bee4 100644
> --- a/kernel/locking/lock_events_list.h
> +++ b/kernel/locking/lock_events_list.h
> @@ -65,7 +65,6 @@ LOCK_EVENT(rwsem_opt_lock) /* # of opt-acquired write locks */
> LOCK_EVENT(rwsem_opt_fail) /* # of failed optspins */
> LOCK_EVENT(rwsem_opt_nospin) /* # of disabled optspins */
> LOCK_EVENT(rwsem_rlock) /* # of read locks acquired */
> -LOCK_EVENT(rwsem_rlock_steal) /* # of read locks by lock stealing */
> LOCK_EVENT(rwsem_rlock_fast) /* # of fast read locks acquired */
> LOCK_EVENT(rwsem_rlock_fail) /* # of failed read lock acquisitions */
> LOCK_EVENT(rwsem_rlock_handoff) /* # of read lock handoffs */
> diff --git a/kernel/locking/rwsem.c b/kernel/locking/rwsem.c
> index bda5577339c0..40b141c5765f 100644
> --- a/kernel/locking/rwsem.c
> +++ b/kernel/locking/rwsem.c
> @@ -1017,42 +1017,9 @@ static struct rw_semaphore __sched *
> rwsem_down_read_slowpath(struct rw_semaphore *sem, long count, unsigned int state)
> {
> long adjustment = -RWSEM_READER_BIAS;
> - long rcnt = (count >> RWSEM_READER_SHIFT);
> struct rwsem_waiter waiter, *first;
> DEFINE_WAKE_Q(wake_q);
>
> - /*
> - * To prevent a constant stream of readers from starving a sleeping
> - * writer, don't attempt optimistic lock stealing if the lock is
> - * very likely owned by readers.
> - */
> - if ((atomic_long_read(&sem->owner) & RWSEM_READER_OWNED) &&
> - (rcnt > 1) && !(count & RWSEM_WRITER_LOCKED))
> - goto queue;
> -
> - /*
> - * Reader optimistic lock stealing.
> - */
> - if (!(count & (RWSEM_WRITER_LOCKED | RWSEM_FLAG_HANDOFF))) {
> - rwsem_set_reader_owned(sem);
> - lockevent_inc(rwsem_rlock_steal);
> -
> - /*
> - * Wake up other readers in the wait queue if it is
> - * the first reader.
> - */
> - if ((rcnt == 1) && (count & RWSEM_FLAG_WAITERS)) {
> - raw_spin_lock_irq(&sem->wait_lock);
> - if (sem->first_waiter)
> - rwsem_mark_wake(sem, RWSEM_WAKE_READ_OWNED,
> - &wake_q);
> - raw_spin_unlock_irq(&sem->wait_lock);
> - wake_up_q(&wake_q);
> - }
> - return sem;
> - }
> -
> -queue:
> waiter.task = current;
> waiter.type = RWSEM_WAITING_FOR_READ;
> waiter.timeout = jiffies + RWSEM_WAIT_TIMEOUT;
> --
> 2.39.3
>
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] locking/rwsem: Remove reader optimistic lock stealing
2026-05-22 2:08 ` Waiman Long
@ 2026-05-22 3:33 ` Peng Wang
2026-05-22 8:55 ` Peter Zijlstra
1 sibling, 0 replies; 8+ messages in thread
From: Peng Wang @ 2026-05-22 3:33 UTC (permalink / raw)
To: Waiman Long; +Cc: peterz, mingo, will, boqun, dbueso, linux-kernel
On Thu, May 21, 2026 at 10:08:58PM -0400, Waiman Long wrote:
> On 5/21/26 5:59 AM, Peng Wang wrote:
> > Reader optimistic lock stealing, introduced by commit 1a728dff855a
> > ("locking/rwsem: Enable reader optimistic lock stealing") and made more
> > aggressive by commit 617f3ef95177 ("locking/rwsem: Remove reader
> > optimistic spinning"), allows a reader entering the slowpath to bypass
> > the wait queue and acquire the lock directly when WRITER_LOCKED and
> > HANDOFF bits are not set.
> >
> > This causes severe writer starvation in workloads where readers hold
> > the lock for extended periods, such as Direct I/O operations which
> > hold inode->i_rwsem for the entire duration of iomap_dio_rw(). A
> > common example is log-structured storage where one thread appends via
> > DIO writes while another thread tails the log via DIO reads -- a
> > pattern seen in database redo-log replay and shared-storage
> > replication.
>
> It is generally assume that reader lock critical section is shorter than
> that of writer. In this particular case, does the reader critical section
> run longer than the writer's one?
Hi Longman,
Thanks for the quick reply.
In this workload both reader and writer hold i_rwsem across iomap_dio_rw(),
so the critical sections are not short, and they cover the full disk I/O latency
about hundreds of microseconds on NVMe
>
> Reader lock stealing should only happen if the previous lock owner is a
> writer. So readers and writer should at most alternately own the lock if
> there are many readers waiting. Of course, if a reader own the lock, it will
> wake up the remaining readers in the wait queue.
Because the steal (an atomic operation, ~50ns) is orders of magnitude faster than
the writer wakeup path (context switch, ~10-100us), the new reader wins this race
every time the lock becomes free.
This repeats with each subsequent reader arrival until the HANDOFF timeout.
>
> >
> > The problem occurs because:
> >
> > 1. A reader entering the slowpath (due to RWSEM_FLAG_WAITERS being set)
> > can still steal the lock as the steal condition only checks
> > WRITER_LOCKED and HANDOFF, not WAITERS. This is inconsistent with
> > the fast path which already blocks readers when WAITERS is set (via
> > RWSEM_READ_FAILED_MASK).
> >
> > 2. In the window between the last reader releasing the lock and the
> > waiting writer being scheduled (~10-100us), a new reader can steal
> > the lock in ~50-100ns. This race is structurally inevitable due to
> > the 1000x speed difference between atomic operations vs context
> > switching.
> >
> > 3. Each stolen read lock is held for the full DIO duration (potentially
> > milliseconds), and the pattern repeats until the 4ms HANDOFF timeout.
> > This effectively taxes every write operation with a ~4ms penalty.
> >
> > Performance impact measured with a DIO mixed read/write workload
> > (1 writer + 1 reader, O_DIRECT, ext4):
> >
> > NVMe SSD: Before After
> > Write-only baseline: 397 MB/s 397 MB/s (no change)
> > Mixed write throughput: 11 MB/s 350 MB/s (+31x)
> > Mixed write latency: 880 us 23 us (-38x)
> > Mixed read throughput: 95 MB/s 95 MB/s (no change)
>
> Does the reader wait for something else and taking other locks after
> acquiring the read lock? Can you point me of the reader critical section for
> this particular workload?
The reader does not wait on anything else or take other locks after acquiring i_rwsem.
The critical section is as below:
ext4_dio_read_iter() (fs/ext4/file.c:70)
inode_lock_shared(inode);
ret = iomap_dio_rw(iocb, to, &ext4_iomap_ops, NULL, 0, NULL, 0);
inode_unlock_shared(inode);
The lock is held across iomap_dio_rw() which submits the bio and waits for I/O completion synchronously.
The hold time is essentially the device latency for a single DIO read.
>
> Cheers,
> Longman
>
> >
> > Fixes: 617f3ef95177 ("locking/rwsem: Remove reader optimistic spinning")
> > Signed-off-by: Peng Wang <peng_wang@linux.alibaba.com>
> > ---
> > kernel/locking/lock_events_list.h | 1 -
> > kernel/locking/rwsem.c | 33 ---------------------------------
> > 2 files changed, 34 deletions(-)
> >
> > diff --git a/kernel/locking/lock_events_list.h b/kernel/locking/lock_events_list.h
> > index 97fb6f3f840a..35b45576bee4 100644
> > --- a/kernel/locking/lock_events_list.h
> > +++ b/kernel/locking/lock_events_list.h
> > @@ -65,7 +65,6 @@ LOCK_EVENT(rwsem_opt_lock) /* # of opt-acquired write locks */
> > LOCK_EVENT(rwsem_opt_fail) /* # of failed optspins */
> > LOCK_EVENT(rwsem_opt_nospin) /* # of disabled optspins */
> > LOCK_EVENT(rwsem_rlock) /* # of read locks acquired */
> > -LOCK_EVENT(rwsem_rlock_steal) /* # of read locks by lock stealing */
> > LOCK_EVENT(rwsem_rlock_fast) /* # of fast read locks acquired */
> > LOCK_EVENT(rwsem_rlock_fail) /* # of failed read lock acquisitions */
> > LOCK_EVENT(rwsem_rlock_handoff) /* # of read lock handoffs */
> > diff --git a/kernel/locking/rwsem.c b/kernel/locking/rwsem.c
> > index bda5577339c0..40b141c5765f 100644
> > --- a/kernel/locking/rwsem.c
> > +++ b/kernel/locking/rwsem.c
> > @@ -1017,42 +1017,9 @@ static struct rw_semaphore __sched *
> > rwsem_down_read_slowpath(struct rw_semaphore *sem, long count, unsigned int state)
> > {
> > long adjustment = -RWSEM_READER_BIAS;
> > - long rcnt = (count >> RWSEM_READER_SHIFT);
> > struct rwsem_waiter waiter, *first;
> > DEFINE_WAKE_Q(wake_q);
> >
> > - /*
> > - * To prevent a constant stream of readers from starving a sleeping
> > - * writer, don't attempt optimistic lock stealing if the lock is
> > - * very likely owned by readers.
> > - */
> > - if ((atomic_long_read(&sem->owner) & RWSEM_READER_OWNED) &&
> > - (rcnt > 1) && !(count & RWSEM_WRITER_LOCKED))
> > - goto queue;
> > -
> > - /*
> > - * Reader optimistic lock stealing.
> > - */
> > - if (!(count & (RWSEM_WRITER_LOCKED | RWSEM_FLAG_HANDOFF))) {
> > - rwsem_set_reader_owned(sem);
> > - lockevent_inc(rwsem_rlock_steal);
> > -
> > - /*
> > - * Wake up other readers in the wait queue if it is
> > - * the first reader.
> > - */
> > - if ((rcnt == 1) && (count & RWSEM_FLAG_WAITERS)) {
> > - raw_spin_lock_irq(&sem->wait_lock);
> > - if (sem->first_waiter)
> > - rwsem_mark_wake(sem, RWSEM_WAKE_READ_OWNED,
> > - &wake_q);
> > - raw_spin_unlock_irq(&sem->wait_lock);
> > - wake_up_q(&wake_q);
> > - }
> > - return sem;
> > - }
> > -
> > -queue:
> > waiter.task = current;
> > waiter.type = RWSEM_WAITING_FOR_READ;
> > waiter.timeout = jiffies + RWSEM_WAIT_TIMEOUT;
> > --
> > 2.39.3
> >
>
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] locking/rwsem: Remove reader optimistic lock stealing
2026-05-22 2:08 ` Waiman Long
2026-05-22 3:33 ` Peng Wang
@ 2026-05-22 8:55 ` Peter Zijlstra
2026-05-22 9:14 ` Peter Zijlstra
` (2 more replies)
1 sibling, 3 replies; 8+ messages in thread
From: Peter Zijlstra @ 2026-05-22 8:55 UTC (permalink / raw)
To: Waiman Long; +Cc: Peng Wang, mingo, will, boqun, dbueso, linux-kernel
On Thu, May 21, 2026 at 10:08:58PM -0400, Waiman Long wrote:
> On 5/21/26 5:59 AM, Peng Wang wrote:
> > Reader optimistic lock stealing, introduced by commit 1a728dff855a
> > ("locking/rwsem: Enable reader optimistic lock stealing") and made more
> > aggressive by commit 617f3ef95177 ("locking/rwsem: Remove reader
> > optimistic spinning"), allows a reader entering the slowpath to bypass
> > the wait queue and acquire the lock directly when WRITER_LOCKED and
> > HANDOFF bits are not set.
> >
> > This causes severe writer starvation in workloads where readers hold
> > the lock for extended periods, such as Direct I/O operations which
> > hold inode->i_rwsem for the entire duration of iomap_dio_rw(). A
> > common example is log-structured storage where one thread appends via
> > DIO writes while another thread tails the log via DIO reads -- a
> > pattern seen in database redo-log replay and shared-storage
> > replication.
>
> It is generally assume that reader lock critical section is shorter than
> that of writer. In this particular case, does the reader critical section
> run longer than the writer's one?
Well, that and writers are assumed to be rare. Reader-writer setups
where writers are common or even dominant make little sense. And that
seems to be exactly this. Then again, it isn't unreasonable to expect it
to not perform significantly worse than an exclusive lock.
> Reader lock stealing should only happen if the previous lock owner is a
> writer. So readers and writer should at most alternately own the lock if
> there are many readers waiting. Of course, if a reader own the lock, it will
> wake up the remaining readers in the wait queue.
Anyway, IIRC I've mentioned phase change locks many times before. And
what we have here is an asymmetric phase change. The timeout causes a
change to writers, but any one writer completing then switches back to
reader dominance.
Perhaps look at evening out the phase change. Retain the 'no-steal'
phase for an equal duration.
Also, 4ms is an eternity, that might need tweaking too.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] locking/rwsem: Remove reader optimistic lock stealing
2026-05-22 8:55 ` Peter Zijlstra
@ 2026-05-22 9:14 ` Peter Zijlstra
2026-05-22 10:06 ` Peng Wang
2026-06-04 14:14 ` [RFC PATCH v2] locking/rwsem: Block reader stealing during writer phase Peng Wang
2 siblings, 0 replies; 8+ messages in thread
From: Peter Zijlstra @ 2026-05-22 9:14 UTC (permalink / raw)
To: Waiman Long; +Cc: Peng Wang, mingo, will, boqun, dbueso, linux-kernel
On Fri, May 22, 2026 at 10:55:13AM +0200, Peter Zijlstra wrote:
> On Thu, May 21, 2026 at 10:08:58PM -0400, Waiman Long wrote:
> > On 5/21/26 5:59 AM, Peng Wang wrote:
> > > Reader optimistic lock stealing, introduced by commit 1a728dff855a
> > > ("locking/rwsem: Enable reader optimistic lock stealing") and made more
> > > aggressive by commit 617f3ef95177 ("locking/rwsem: Remove reader
> > > optimistic spinning"), allows a reader entering the slowpath to bypass
> > > the wait queue and acquire the lock directly when WRITER_LOCKED and
> > > HANDOFF bits are not set.
> > >
> > > This causes severe writer starvation in workloads where readers hold
> > > the lock for extended periods, such as Direct I/O operations which
> > > hold inode->i_rwsem for the entire duration of iomap_dio_rw(). A
> > > common example is log-structured storage where one thread appends via
> > > DIO writes while another thread tails the log via DIO reads -- a
> > > pattern seen in database redo-log replay and shared-storage
> > > replication.
> >
> > It is generally assume that reader lock critical section is shorter than
> > that of writer. In this particular case, does the reader critical section
> > run longer than the writer's one?
>
> Well, that and writers are assumed to be rare. Reader-writer setups
> where writers are common or even dominant make little sense. And that
> seems to be exactly this. Then again, it isn't unreasonable to expect it
> to not perform significantly worse than an exclusive lock.
>
> > Reader lock stealing should only happen if the previous lock owner is a
> > writer. So readers and writer should at most alternately own the lock if
> > there are many readers waiting. Of course, if a reader own the lock, it will
> > wake up the remaining readers in the wait queue.
>
> Anyway, IIRC I've mentioned phase change locks many times before. And
> what we have here is an asymmetric phase change. The timeout causes a
> change to writers, but any one writer completing then switches back to
> reader dominance.
>
> Perhaps look at evening out the phase change. Retain the 'no-steal'
> phase for an equal duration.
>
> Also, 4ms is an eternity, that might need tweaking too.
Also, perhaps reduce MAX_READERS_WAKEUP when in the 'writer' phase of
things.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] locking/rwsem: Remove reader optimistic lock stealing
2026-05-22 8:55 ` Peter Zijlstra
2026-05-22 9:14 ` Peter Zijlstra
@ 2026-05-22 10:06 ` Peng Wang
2026-06-04 14:14 ` [RFC PATCH v2] locking/rwsem: Block reader stealing during writer phase Peng Wang
2 siblings, 0 replies; 8+ messages in thread
From: Peng Wang @ 2026-05-22 10:06 UTC (permalink / raw)
To: Peter Zijlstra; +Cc: Waiman Long, mingo, will, boqun, dbueso, linux-kernel
On Fri, May 22, 2026 at 10:55:13AM +0200, Peter Zijlstra wrote:
> On Thu, May 21, 2026 at 10:08:58PM -0400, Waiman Long wrote:
> > On 5/21/26 5:59 AM, Peng Wang wrote:
> > > Reader optimistic lock stealing, introduced by commit 1a728dff855a
> > > ("locking/rwsem: Enable reader optimistic lock stealing") and made more
> > > aggressive by commit 617f3ef95177 ("locking/rwsem: Remove reader
> > > optimistic spinning"), allows a reader entering the slowpath to bypass
> > > the wait queue and acquire the lock directly when WRITER_LOCKED and
> > > HANDOFF bits are not set.
> > >
> > > This causes severe writer starvation in workloads where readers hold
> > > the lock for extended periods, such as Direct I/O operations which
> > > hold inode->i_rwsem for the entire duration of iomap_dio_rw(). A
> > > common example is log-structured storage where one thread appends via
> > > DIO writes while another thread tails the log via DIO reads -- a
> > > pattern seen in database redo-log replay and shared-storage
> > > replication.
> >
> > It is generally assume that reader lock critical section is shorter than
> > that of writer. In this particular case, does the reader critical section
> > run longer than the writer's one?
>
> Well, that and writers are assumed to be rare. Reader-writer setups
> where writers are common or even dominant make little sense. And that
> seems to be exactly this. Then again, it isn't unreasonable to expect it
> to not perform significantly worse than an exclusive lock.
>
> > Reader lock stealing should only happen if the previous lock owner is a
> > writer. So readers and writer should at most alternately own the lock if
> > there are many readers waiting. Of course, if a reader own the lock, it will
> > wake up the remaining readers in the wait queue.
>
> Anyway, IIRC I've mentioned phase change locks many times before. And
> what we have here is an asymmetric phase change. The timeout causes a
> change to writers, but any one writer completing then switches back to
> reader dominance.
>
> Perhaps look at evening out the phase change. Retain the 'no-steal'
> phase for an equal duration.
>
> Also, 4ms is an eternity, that might need tweaking too.
Hi Peter,
I agree this is an asymmetric phase change. Looking at the code, the asymmetry
appears to exist only through the steal path, which is only reachable when
RWSEM_FLAG_WAITERS triggers slowpath entry (since READ_FAILED_MASK includes WAITERS,
and the steal condition requires WRITER_LOCKED=0 and HANDOFF=0).
With steal removed, the wait queue with phase-fair batch wakeup seems to provide naturally
symmetric transitions without explicit phase management:
- Writer at queue head -> writer runs
- Reader at queue head -> all readers batch-woken in parallel
- Alternation governed by arrival order
A read-heavy mmap benchmark (16 threads, short critical sections) showed no measurable difference
with steal removed (308k vs 306k ops/sec), as readers succeed via fast path when there is no contention.
Regarding the timeout: agreed 4ms feels too long for fast storage.
With steal removed, handoff would only matter as a backstop when a writer enters the queue while readers are already active.
Would the simple removal be acceptable, or would you prefer a more structured phase-change approach?
Best regards,
Peng
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH] locking/rwsem: Remove reader optimistic lock stealing
2026-05-21 9:59 [PATCH] locking/rwsem: Remove reader optimistic lock stealing Peng Wang
2026-05-22 2:08 ` Waiman Long
@ 2026-05-31 7:27 ` kernel test robot
1 sibling, 0 replies; 8+ messages in thread
From: kernel test robot @ 2026-05-31 7:27 UTC (permalink / raw)
To: Peng Wang
Cc: oe-lkp, lkp, linux-kernel, peterz, mingo, will, boqun, longman,
dbueso, Peng Wang, oliver.sang
Hello,
kernel test robot noticed a 19.3% regression of fio.write_iops on:
commit: b3bfaf6505df7705d1d7aa21e09c19a98b10dc84 ("[PATCH] locking/rwsem: Remove reader optimistic lock stealing")
url: https://github.com/intel-lab-lkp/linux/commits/Peng-Wang/locking-rwsem-Remove-reader-optimistic-lock-stealing/20260521-183741
base: https://git.kernel.org/cgit/linux/kernel/git/tip/tip.git 43a037d4fa6d7e7f9437e069f857a5e06eb2dff7
patch link: https://lore.kernel.org/all/20260521095926.29363-1-peng_wang@linux.alibaba.com/
patch subject: [PATCH] locking/rwsem: Remove reader optimistic lock stealing
testcase: fio-basic
config: x86_64-rhel-9.4
compiler: gcc-14
test machine: 64 threads 2 sockets Intel(R) Xeon(R) Gold 6346 CPU @ 3.10GHz (Ice Lake) with 256G memory
parameters:
runtime: 300s
disk: 1HDD
fs: btrfs
nr_task: 100%
test_size: 128G
rw: randwrite
bs: 4k
ioengine: ftruncate
cpufreq_governor: performance
In addition to that, the commit also has significant impact on the following tests:
+------------------+-------------------------------------------------------------+
| testcase: change | stress-ng: stress-ng.shm-sysv.ops_per_sec 14.3% improvement |
| test parameters | cpufreq_governor=performance |
| | nr_threads=100% |
| | test=shm-sysv |
| | testtime=60s |
+------------------+-------------------------------------------------------------+
If you fix the issue in a separate patch/commit (i.e. not just a new version of
the same patch/commit), kindly add following tags
| Reported-by: kernel test robot <oliver.sang@intel.com>
| Closes: https://lore.kernel.org/oe-lkp/202605291757.b455c22c-lkp@intel.com
Details are as below:
-------------------------------------------------------------------------------------------------->
The kernel config and materials to reproduce are available at:
https://download.01.org/0day-ci/archive/20260529/202605291757.b455c22c-lkp@intel.com
=========================================================================================
bs/compiler/cpufreq_governor/disk/fs/ioengine/kconfig/nr_task/rootfs/runtime/rw/tbox_group/test_size/testcase:
4k/gcc-14/performance/1HDD/btrfs/ftruncate/x86_64-rhel-9.4/100%/debian-13-x86_64-20250902.cgz/300s/randwrite/lkp-icl-2sp9/128G/fio-basic
commit:
43a037d4fa ("seqlock: Allow UBSAN_ALIGNMENT to fail optimizing")
b3bfaf6505 ("[PATCH] locking/rwsem: Remove reader optimistic lock stealing")
43a037d4fa6d7e7f b3bfaf6505df7705d1d7aa21e09
---------------- ---------------------------
%stddev %change %stddev
\ | \
0.04 ± 36% +0.1 0.10 ± 2% fio.latency_1000us%
9.89 +7.2 17.10 fio.latency_100us%
17.86 ± 2% -17.3 0.52 ± 20% fio.latency_10us%
4.63 ± 6% -4.2 0.44 ± 2% fio.latency_20us%
39.20 +19.2 58.39 fio.latency_250us%
0.13 ± 45% -0.1 0.06 ± 6% fio.latency_2us%
5.17 ± 8% -4.9 0.24 ± 46% fio.latency_4us%
16.12 +4.5 20.57 ± 2% fio.latency_500us%
4.52 ± 9% -2.8 1.73 ± 3% fio.latency_50us%
1.86 ± 4% -1.6 0.26 ± 3% fio.latency_750us%
80.70 +23.6% 99.76 fio.time.elapsed_time
80.70 +23.6% 99.76 fio.time.elapsed_time.max
3398 ± 24% -50.1% 1696 ± 21% fio.time.involuntary_context_switches
1275 ± 3% +30.2% 1660 fio.time.system_time
31138255 +37.1% 42697561 fio.time.voluntary_context_switches
1633 -19.3% 1318 fio.write_bw_MBps
449877 -8.3% 412330 fio.write_clat_95%_ns
561152 -12.2% 492885 fio.write_clat_99%_ns
150967 +24.5% 187975 fio.write_clat_mean_ns
159157 -19.1% 128723 fio.write_clat_stddev
418163 -19.3% 337623 fio.write_iops
3.889e+09 +20.6% 4.689e+09 cpuidle..time
33940620 +37.1% 46542688 cpuidle..usage
25.27 ± 3% +5.8% 26.74 ± 2% iostat.cpu.system
1.05 +16.2% 1.22 iostat.cpu.user
128.58 +14.9% 147.79 uptime.boot
6572 +12.0% 7362 uptime.idle
0.75 ± 2% +0.1 0.81 ± 2% mpstat.cpu.all.irq%
0.13 ± 2% -0.0 0.12 ± 2% mpstat.cpu.all.soft%
1.03 +0.2 1.20 mpstat.cpu.all.usr%
183168 ± 5% +7.7% 197182 ± 2% numa-numastat.node0.numa_hit
707493 ± 4% +9.4% 774114 ± 2% numa-numastat.node1.local_node
750282 +9.4% 821047 numa-numastat.node1.numa_hit
375447 +9.7% 411736 sched_debug.cpu.nr_switches.avg
386408 +9.5% 422976 sched_debug.cpu.nr_switches.max
340142 ± 2% +9.4% 372192 sched_debug.cpu.nr_switches.min
4394759 +10.1% 4838493 vmstat.memory.cache
742204 +11.4% 826992 vmstat.system.cs
427361 +10.7% 472956 vmstat.system.in
380208 ± 67% +111.5% 804091 ± 22% numa-meminfo.node1.Active
380195 ± 67% +111.5% 804078 ± 22% numa-meminfo.node1.Active(anon)
20479 ±143% +928.0% 210520 ± 16% numa-meminfo.node1.Mapped
134972 ± 4% +334.2% 586041 ± 3% numa-meminfo.node1.Shmem
1016 ± 2% +6.2% 1078 ± 2% turbostat.Avg_MHz
28.30 ± 2% +1.7 30.04 ± 2% turbostat.Busy%
35943585 +36.1% 48919629 turbostat.IRQ
0.06 +0.0 0.07 turbostat.POLL%
2.05 -10.7% 1.83 perf-sched.total_wait_and_delay.average.ms
635441 +8.6% 690012 perf-sched.total_wait_and_delay.count.ms
2.05 -10.8% 1.83 perf-sched.total_wait_time.average.ms
2.05 -10.7% 1.83 perf-sched.wait_and_delay.avg.ms.[unknown].[unknown].[unknown].[unknown].[unknown]
635441 +8.6% 690012 perf-sched.wait_and_delay.count.[unknown].[unknown].[unknown].[unknown].[unknown]
2.05 -10.8% 1.83 perf-sched.wait_time.avg.ms.[unknown].[unknown].[unknown].[unknown].[unknown]
183285 ± 5% +7.7% 197373 ± 2% numa-vmstat.node0.numa_hit
95133 ± 67% +111.4% 201074 ± 22% numa-vmstat.node1.nr_active_anon
5148 ±143% +923.4% 52686 ± 16% numa-vmstat.node1.nr_mapped
33827 ± 4% +333.3% 146565 ± 3% numa-vmstat.node1.nr_shmem
95133 ± 67% +111.4% 201074 ± 22% numa-vmstat.node1.nr_zone_active_anon
750372 +9.4% 821016 numa-vmstat.node1.numa_hit
707583 ± 4% +9.4% 774083 ± 2% numa-vmstat.node1.numa_local
841564 +54.2% 1297873 meminfo.Active
841532 +54.2% 1297841 meminfo.Active(anon)
85390 ± 2% +17.1% 99975 meminfo.AnonHugePages
4263847 +10.7% 4718785 meminfo.Cached
3400733 +14.1% 3880266 meminfo.Committed_AS
101799 +182.4% 287453 ± 5% meminfo.Mapped
6510661 +6.6% 6942224 meminfo.Memused
148510 +306.2% 603279 ± 2% meminfo.Shmem
768695 +11.0% 853548 perf-stat.i.context-switches
5070 ± 7% +35.3% 6858 perf-stat.i.cpu-migrations
2575 ± 2% +9.1% 2809 ± 4% perf-stat.i.minor-faults
2575 ± 2% +9.1% 2809 ± 4% perf-stat.i.page-faults
759278 +11.3% 845041 perf-stat.ps.context-switches
5014 ± 7% +35.5% 6795 perf-stat.ps.cpu-migrations
2544 ± 2% +9.3% 2781 ± 4% perf-stat.ps.minor-faults
2544 ± 2% +9.3% 2781 ± 4% perf-stat.ps.page-faults
210458 +54.2% 324493 proc-vmstat.nr_active_anon
1066202 +10.7% 1179895 proc-vmstat.nr_file_pages
25450 +182.5% 71894 ± 5% proc-vmstat.nr_mapped
37204 +305.5% 150855 ± 2% proc-vmstat.nr_shmem
210458 +54.2% 324493 proc-vmstat.nr_zone_active_anon
2110 ± 50% +202.1% 6377 ± 22% proc-vmstat.numa_hint_faults_local
934523 +9.2% 1020627 proc-vmstat.numa_hit
868232 +9.9% 954306 proc-vmstat.numa_local
1003773 +9.0% 1093947 proc-vmstat.pgalloc_normal
295121 +26.3% 372800 ± 3% proc-vmstat.pgfault
301174 ± 3% +13.5% 341864 ± 3% proc-vmstat.pgfree
16096 ± 5% +18.3% 19037 ± 6% proc-vmstat.pgreuse
12.71 ± 80% -10.8 1.92 ±142% perf-profile.calltrace.cycles-pp.__ordered_events__flush.perf_session__process_user_event.perf_session__process_events.record__finish_output.cmd_record
12.71 ± 80% -10.8 1.92 ±142% perf-profile.calltrace.cycles-pp.cmd_record
12.71 ± 80% -10.8 1.92 ±142% perf-profile.calltrace.cycles-pp.ordered_events__deliver_event.__ordered_events__flush.perf_session__process_user_event.perf_session__process_events.record__finish_output
12.71 ± 80% -10.8 1.92 ±142% perf-profile.calltrace.cycles-pp.perf_session__deliver_event.ordered_events__deliver_event.__ordered_events__flush.perf_session__process_user_event.perf_session__process_events
12.71 ± 80% -10.8 1.92 ±142% perf-profile.calltrace.cycles-pp.perf_session__process_events.record__finish_output.cmd_record
12.71 ± 80% -10.8 1.92 ±142% perf-profile.calltrace.cycles-pp.perf_session__process_user_event.perf_session__process_events.record__finish_output.cmd_record
12.71 ± 80% -10.8 1.92 ±142% perf-profile.calltrace.cycles-pp.record__finish_output.cmd_record
12.70 ± 80% -10.8 1.92 ±142% perf-profile.calltrace.cycles-pp.build_id__mark_dso_hit.perf_session__deliver_event.ordered_events__deliver_event.__ordered_events__flush.perf_session__process_user_event
12.30 ± 78% -10.4 1.92 ±142% perf-profile.calltrace.cycles-pp.sample__for_each_callchain_node.build_id__mark_dso_hit.perf_session__deliver_event.ordered_events__deliver_event.__ordered_events__flush
12.15 ± 78% -10.2 1.92 ±142% perf-profile.calltrace.cycles-pp.__thread__resolve_callchain.sample__for_each_callchain_node.build_id__mark_dso_hit.perf_session__deliver_event.ordered_events__deliver_event
10.31 ± 85% -8.4 1.92 ±142% perf-profile.calltrace.cycles-pp.thread__resolve_callchain_sample.__thread__resolve_callchain.sample__for_each_callchain_node.build_id__mark_dso_hit.perf_session__deliver_event
10.15 ± 84% -8.2 1.92 ±142% perf-profile.calltrace.cycles-pp.add_callchain_ip.thread__resolve_callchain_sample.__thread__resolve_callchain.sample__for_each_callchain_node.build_id__mark_dso_hit
5.59 ±107% -4.7 0.88 ±223% perf-profile.calltrace.cycles-pp.cmd_record.perf_c2c__record.handle_internal_command.main
5.59 ±107% -4.7 0.88 ±223% perf-profile.calltrace.cycles-pp.perf_c2c__record.handle_internal_command.main
7.33 ± 84% -4.5 2.80 ±148% perf-profile.calltrace.cycles-pp.handle_internal_command.main
7.33 ± 84% -4.5 2.80 ±148% perf-profile.calltrace.cycles-pp.main
5.92 ±126% -4.2 1.76 ±223% perf-profile.calltrace.cycles-pp.handle_mm_fault.do_user_addr_fault.exc_page_fault.asm_exc_page_fault
3.95 ± 72% -4.0 0.00 perf-profile.calltrace.cycles-pp.acpi_idle_do_entry.acpi_idle_enter.cpuidle_enter_state.cpuidle_enter.cpuidle_idle_call
3.95 ± 72% -4.0 0.00 perf-profile.calltrace.cycles-pp.acpi_idle_enter.cpuidle_enter_state.cpuidle_enter.cpuidle_idle_call.do_idle
3.95 ± 72% -4.0 0.00 perf-profile.calltrace.cycles-pp.acpi_safe_halt.acpi_idle_do_entry.acpi_idle_enter.cpuidle_enter_state.cpuidle_enter
3.95 ± 72% -4.0 0.00 perf-profile.calltrace.cycles-pp.cpuidle_enter.cpuidle_idle_call.do_idle.cpu_startup_entry.start_secondary
3.95 ± 72% -4.0 0.00 perf-profile.calltrace.cycles-pp.cpuidle_enter_state.cpuidle_enter.cpuidle_idle_call.do_idle.cpu_startup_entry
3.95 ± 72% -4.0 0.00 perf-profile.calltrace.cycles-pp.cpuidle_idle_call.do_idle.cpu_startup_entry.start_secondary.common_startup_64
7.31 ± 94% -3.6 3.67 ±160% perf-profile.calltrace.cycles-pp.asm_exc_page_fault
6.27 ±118% -3.5 2.80 ±148% perf-profile.calltrace.cycles-pp.do_user_addr_fault.exc_page_fault.asm_exc_page_fault
6.27 ±118% -3.5 2.80 ±148% perf-profile.calltrace.cycles-pp.exc_page_fault.asm_exc_page_fault
5.92 ±126% -3.3 2.63 ±223% perf-profile.calltrace.cycles-pp.__handle_mm_fault.handle_mm_fault.do_user_addr_fault.exc_page_fault.asm_exc_page_fault
4.10 ± 72% -3.2 0.88 ±223% perf-profile.calltrace.cycles-pp.common_startup_64
4.10 ± 72% -3.2 0.88 ±223% perf-profile.calltrace.cycles-pp.cpu_startup_entry.start_secondary.common_startup_64
4.10 ± 72% -3.2 0.88 ±223% perf-profile.calltrace.cycles-pp.do_idle.cpu_startup_entry.start_secondary.common_startup_64
4.10 ± 72% -3.2 0.88 ±223% perf-profile.calltrace.cycles-pp.start_secondary.common_startup_64
3.42 ±100% +0.7 4.17 ±223% perf-profile.calltrace.cycles-pp.__free_event.perf_event_release_kernel.perf_release.__fput.task_work_run
3.42 ±100% +0.7 4.17 ±223% perf-profile.calltrace.cycles-pp.sw_perf_event_destroy.__free_event.perf_event_release_kernel.perf_release.__fput
20.04 ± 76% -16.4 3.67 ±160% perf-profile.children.cycles-pp.cmd_record
12.71 ± 80% -10.8 1.92 ±142% perf-profile.children.cycles-pp.__ordered_events__flush
12.71 ± 80% -10.8 1.92 ±142% perf-profile.children.cycles-pp.ordered_events__deliver_event
12.71 ± 80% -10.8 1.92 ±142% perf-profile.children.cycles-pp.perf_session__deliver_event
12.71 ± 80% -10.8 1.92 ±142% perf-profile.children.cycles-pp.perf_session__process_events
12.71 ± 80% -10.8 1.92 ±142% perf-profile.children.cycles-pp.perf_session__process_user_event
12.71 ± 80% -10.8 1.92 ±142% perf-profile.children.cycles-pp.record__finish_output
12.50 ± 79% -10.6 1.92 ±142% perf-profile.children.cycles-pp.build_id__mark_dso_hit
12.50 ± 79% -10.6 1.92 ±142% perf-profile.children.cycles-pp.sample__for_each_callchain_node
12.15 ± 78% -10.2 1.92 ±142% perf-profile.children.cycles-pp.__thread__resolve_callchain
10.31 ± 85% -8.4 1.92 ±142% perf-profile.children.cycles-pp.add_callchain_ip
10.31 ± 85% -8.4 1.92 ±142% perf-profile.children.cycles-pp.thread__resolve_callchain_sample
7.94 ±132% -7.9 0.00 perf-profile.children.cycles-pp.ksys_write
7.94 ±132% -7.9 0.00 perf-profile.children.cycles-pp.vfs_write
5.59 ±107% -4.7 0.88 ±223% perf-profile.children.cycles-pp.perf_c2c__record
7.33 ± 84% -4.5 2.80 ±148% perf-profile.children.cycles-pp.handle_internal_command
7.33 ± 84% -4.5 2.80 ±148% perf-profile.children.cycles-pp.main
4.44 ± 59% -4.4 0.00 perf-profile.children.cycles-pp.__mmap_region
4.44 ± 59% -4.4 0.00 perf-profile.children.cycles-pp.do_mmap
4.44 ± 59% -4.4 0.00 perf-profile.children.cycles-pp.vm_mmap_pgoff
3.95 ± 72% -4.0 0.00 perf-profile.children.cycles-pp.acpi_idle_do_entry
3.95 ± 72% -4.0 0.00 perf-profile.children.cycles-pp.acpi_idle_enter
3.95 ± 72% -4.0 0.00 perf-profile.children.cycles-pp.acpi_safe_halt
3.95 ± 72% -4.0 0.00 perf-profile.children.cycles-pp.cpuidle_enter
3.95 ± 72% -4.0 0.00 perf-profile.children.cycles-pp.cpuidle_enter_state
3.95 ± 72% -4.0 0.00 perf-profile.children.cycles-pp.cpuidle_idle_call
3.95 ± 72% -4.0 0.00 perf-profile.children.cycles-pp.pv_native_safe_halt
3.92 ±107% -3.9 0.00 perf-profile.children.cycles-pp.map_symbol__exit
6.27 ±118% -3.6 2.63 ±223% perf-profile.children.cycles-pp.handle_mm_fault
7.87 ± 86% -3.3 4.55 ±169% perf-profile.children.cycles-pp.asm_exc_page_fault
5.92 ±126% -3.3 2.63 ±223% perf-profile.children.cycles-pp.__handle_mm_fault
4.10 ± 72% -3.2 0.88 ±223% perf-profile.children.cycles-pp.common_startup_64
4.10 ± 72% -3.2 0.88 ±223% perf-profile.children.cycles-pp.cpu_startup_entry
4.10 ± 72% -3.2 0.88 ±223% perf-profile.children.cycles-pp.do_idle
4.10 ± 72% -3.2 0.88 ±223% perf-profile.children.cycles-pp.start_secondary
6.47 ±115% -2.8 3.67 ±160% perf-profile.children.cycles-pp.do_user_addr_fault
6.47 ±115% -2.8 3.67 ±160% perf-profile.children.cycles-pp.exc_page_fault
3.98 ± 74% -1.6 2.38 ±223% perf-profile.children.cycles-pp.its_return_thunk
3.42 ±100% +0.7 4.17 ±223% perf-profile.children.cycles-pp.__free_event
3.42 ±100% +0.7 4.17 ±223% perf-profile.children.cycles-pp.sw_perf_event_destroy
***************************************************************************************************
=========================================================================================
compiler/cpufreq_governor/kconfig/nr_threads/rootfs/tbox_group/test/testcase/testtime:
gcc-14/performance/x86_64-rhel-9.4/100%/debian-13-x86_64-20250902.cgz/lkp-gnr-2sp4/shm-sysv/stress-ng/60s
commit:
43a037d4fa ("seqlock: Allow UBSAN_ALIGNMENT to fail optimizing")
b3bfaf6505 ("[PATCH] locking/rwsem: Remove reader optimistic lock stealing")
43a037d4fa6d7e7f b3bfaf6505df7705d1d7aa21e09
---------------- ---------------------------
%stddev %change %stddev
\ | \
1251446 -75.9% 301654 stress-ng.shm-sysv.nanosecs_per_shmat_call
3450142 -68.9% 1074649 ± 2% stress-ng.shm-sysv.nanosecs_per_shmdt_call
322286 -22.9% 248566 stress-ng.shm-sysv.nanosecs_per_shmget_call
219186 +14.3% 250564 stress-ng.shm-sysv.ops
3658 +14.3% 4181 stress-ng.shm-sysv.ops_per_sec
7507 +45.8% 10944 stress-ng.time.involuntary_context_switches
14229412 +16.9% 16632333 ± 2% stress-ng.time.minor_page_faults
1785 +49.8% 2675 stress-ng.time.percent_of_cpu_this_job_got
1036 +51.4% 1569 stress-ng.time.system_time
1605102 +24.8% 2003722 stress-ng.time.voluntary_context_switches
12183165 +9.5% 13338260 cpuidle..usage
6385 ± 3% +25.9% 8037 ± 9% perf-c2c.HITM.local
154924 ± 3% +19.3% 184900 ± 3% meminfo.Mapped
56527 ± 3% +9.7% 61983 ± 2% meminfo.PageTables
18.15 ± 11% +47.5% 26.76 ± 6% vmstat.procs.r
58533 +22.9% 71928 vmstat.system.cs
239900 +15.2% 276246 vmstat.system.in
0.71 +0.1 0.78 mpstat.cpu.all.irq%
0.14 +0.0 0.17 mpstat.cpu.all.nice%
6.89 +3.5 10.36 mpstat.cpu.all.sys%
11.83 ± 4% +36.7% 16.17 ± 3% mpstat.max_utilization_pct
5045408 ± 3% +18.5% 5977715 ± 2% numa-numastat.node0.local_node
5226760 ± 2% +16.6% 6096126 numa-numastat.node0.numa_hit
4935973 ± 3% +14.2% 5637088 numa-numastat.node1.local_node
5018567 ± 3% +15.2% 5782749 numa-numastat.node1.numa_hit
5226696 ± 2% +16.6% 6096100 numa-vmstat.node0.numa_hit
5045344 ± 3% +18.5% 5977690 ± 2% numa-vmstat.node0.numa_local
5018578 ± 3% +15.2% 5782937 numa-vmstat.node1.numa_hit
4935984 ± 3% +14.2% 5637277 numa-vmstat.node1.numa_local
44.00 -25.4% 32.83 perf-sched.total_wait_and_delay.average.ms
136083 +37.6% 187199 perf-sched.total_wait_and_delay.count.ms
43.82 -25.5% 32.65 perf-sched.total_wait_time.average.ms
44.00 -25.4% 32.83 perf-sched.wait_and_delay.avg.ms.[unknown].[unknown].[unknown].[unknown].[unknown]
136083 +37.6% 187199 perf-sched.wait_and_delay.count.[unknown].[unknown].[unknown].[unknown].[unknown]
43.82 -25.5% 32.65 perf-sched.wait_time.avg.ms.[unknown].[unknown].[unknown].[unknown].[unknown]
60660 +23.5% 74920 perf-stat.i.context-switches
3628 +51.4% 5491 perf-stat.i.cpu-migrations
240822 +16.3% 280126 ± 2% perf-stat.i.minor-faults
247803 +16.3% 288104 ± 2% perf-stat.i.page-faults
59664 +23.3% 73595 perf-stat.ps.context-switches
3578 +50.8% 5396 perf-stat.ps.cpu-migrations
236274 +16.3% 274869 ± 2% perf-stat.ps.minor-faults
243169 +16.3% 282723 ± 2% perf-stat.ps.page-faults
38758 ± 3% +19.3% 46256 ± 3% proc-vmstat.nr_mapped
14176 ± 3% +9.3% 15495 ± 2% proc-vmstat.nr_page_table_pages
34663 ± 17% +34.6% 46646 ± 19% proc-vmstat.numa_hint_faults
27962 ± 23% +51.1% 42243 ± 16% proc-vmstat.numa_hint_faults_local
10246657 ± 2% +15.9% 11879228 proc-vmstat.numa_hit
9982711 ± 2% +16.4% 11615157 proc-vmstat.numa_local
43773 ± 20% +37.0% 59949 ± 15% proc-vmstat.numa_pte_updates
10940550 ± 2% +15.5% 12640941 proc-vmstat.pgalloc_normal
14977157 +15.9% 17360370 ± 2% proc-vmstat.pgfault
10590608 ± 2% +16.2% 12302205 proc-vmstat.pgfree
1242845 +14.3% 1420964 proc-vmstat.unevictable_pgs_scanned
339.83 +40.1% 476.17 turbostat.Avg_MHz
9.14 +3.6 12.78 turbostat.Busy%
4.17 +1.6 5.73 turbostat.C1E%
59.56 +3.6 63.17 turbostat.C6%
27.42 -8.7 18.68 turbostat.C6P%
7.09 +27.2% 9.02 turbostat.CPU%c1
76.31 -9.7% 68.91 turbostat.CPU%c6
15351025 +15.5% 17723005 turbostat.IRQ
1748282 +26.1% 2204598 turbostat.NMI
158.43 +9.1% 172.81 turbostat.PkgWatt
5.20 +2.0% 5.30 turbostat.RAMWatt
15202 ± 21% +45.4% 22111 ± 9% sched_debug.cfs_rq:/.avg_vruntime.avg
44395 ± 8% +24.7% 55374 ± 9% sched_debug.cfs_rq:/.avg_vruntime.max
7949 ± 26% +69.1% 13440 ± 20% sched_debug.cfs_rq:/.avg_vruntime.min
0.05 ± 9% +50.7% 0.07 ± 18% sched_debug.cfs_rq:/.h_nr_queued.avg
0.21 ± 8% +15.2% 0.24 ± 8% sched_debug.cfs_rq:/.h_nr_queued.stddev
0.05 ± 10% +51.8% 0.07 ± 18% sched_debug.cfs_rq:/.h_nr_runnable.avg
0.21 ± 8% +16.3% 0.24 ± 8% sched_debug.cfs_rq:/.h_nr_runnable.stddev
15419 ± 7% +45.6% 22447 ± 18% sched_debug.cfs_rq:/.load.avg
0.05 ± 9% +50.7% 0.07 ± 18% sched_debug.cfs_rq:/.nr_queued.avg
0.21 ± 8% +15.2% 0.24 ± 8% sched_debug.cfs_rq:/.nr_queued.stddev
88.44 ± 7% +20.7% 106.78 ± 7% sched_debug.cfs_rq:/.runnable_avg.avg
2088 ± 57% -72.9% 564.91 ±109% sched_debug.cfs_rq:/.sum_weight.avg
435307 ± 48% -73.3% 116029 ± 98% sched_debug.cfs_rq:/.sum_weight.max
29171 ± 52% -73.3% 7775 ±101% sched_debug.cfs_rq:/.sum_weight.stddev
87.80 ± 7% +20.8% 106.05 ± 7% sched_debug.cfs_rq:/.util_avg.avg
15202 ± 21% +45.4% 22111 ± 9% sched_debug.cfs_rq:/.zero_vruntime.avg
44395 ± 8% +24.7% 55374 ± 9% sched_debug.cfs_rq:/.zero_vruntime.max
7949 ± 26% +69.1% 13440 ± 20% sched_debug.cfs_rq:/.zero_vruntime.min
1181 -5.1% 1120 sched_debug.cpu.clock_task.stddev
118497 +13.1% 134053 sched_debug.cpu.curr->pid.max
0.00 ± 5% +10.8% 0.00 ± 8% sched_debug.cpu.next_balance.stddev
0.04 ± 8% +53.3% 0.07 ± 18% sched_debug.cpu.nr_running.avg
0.21 ± 7% +16.0% 0.24 ± 8% sched_debug.cpu.nr_running.stddev
8607 +19.6% 10297 sched_debug.cpu.nr_switches.avg
6329 ± 2% +21.8% 7706 ± 9% sched_debug.cpu.nr_switches.min
0.67 -10.3% 0.60 ± 2% sched_debug.cpu.nr_uninterruptible.avg
11.11 ± 8% +35.4% 15.04 ± 26% sched_debug.cpu.nr_uninterruptible.stddev
Disclaimer:
Results have been estimated based on internal Intel analysis and are provided
for informational purposes only. Any difference in system hardware or software
design or configuration may affect actual performance.
--
0-DAY CI Kernel Test Service
https://github.com/intel/lkp-tests/wiki
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC PATCH v2] locking/rwsem: Block reader stealing during writer phase
2026-05-22 8:55 ` Peter Zijlstra
2026-05-22 9:14 ` Peter Zijlstra
2026-05-22 10:06 ` Peng Wang
@ 2026-06-04 14:14 ` Peng Wang
2 siblings, 0 replies; 8+ messages in thread
From: Peng Wang @ 2026-06-04 14:14 UTC (permalink / raw)
To: peterz; +Cc: Waiman Long, mingo, will, boqun, dbueso, linux-kernel
Reader optimistic lock stealing lets a slowpath reader acquire the
lock without queueing when WRITER_LOCKED and HANDOFF are clear. It
works well when writers are rare, but it can starve writers when
readers hold the rwsem across slow operations.
A common case is Direct I/O on inode->i_rwsem: both an O_DIRECT
appending writer and an O_DIRECT tailing reader hold inode->i_rwsem
across iomap_dio_rw(). After the last reader releases the lock the
queued writer is woken, but a newly arriving reader can win the
atomic fast path before the writer is scheduled. The stolen read
lock is then held for another DIO, so the writer repeatedly pays for
a reader I/O until the handoff timeout fires.
The handoff timeout already moves the lock briefly to a writer/
no-steal state, but only for one writer; as soon as it completes,
reader stealing is permitted again and the next queued writer is
starved by the same race. Make the writer phase explicit so that it
persists across consecutive queued writers and ends only when a
reader is granted. Use a new bit RWSEM_FLAG_WRITER_PHASE in the
existing reserved range of rw_semaphore->count, so struct
rw_semaphore is unchanged.
RWSEM_FLAG_WRITER_PHASE is set when a writer becomes the head of the
wait queue (in rwsem_mark_wake() when waking a queued writer head,
or in rwsem_down_write_slowpath() when enqueueing onto an empty
list), and cleared when the queue drains or transitions to a reader
phase (in rwsem_mark_wake() when waking a reader or emptying the
list; in rwsem_try_write_lock() when the acquiring writer is the
only remaining waiter; in rwsem_del_waiter() when removing the last
waiter).
While the bit is set, the reader fast path falls through via
RWSEM_READ_FAILED_MASK, slowpath stealing is disabled, and multiple
consecutive queued writers stay protected. Reader stealing is only
blocked in the precise condition that produces starvation, not when
the queue head is itself a reader.
The HANDOFF bit and its post-CAS invariant are not modified.
RWSEM_FLAG_WRITER_PHASE is only set or cleared under wait_lock,
never inside the rwsem_try_write_lock() compare-exchange that
handles HANDOFF.
Tested with one O_DIRECT appending writer and one O_DIRECT tailing
reader on ext4 over NVMe:
before patch after patch
writer-only baseline: ~390 MB/s, 20 us ~390 MB/s, 20 us
mixed write throughput: ~10 MB/s, ~970 us ~190 MB/s, ~50 us
Suggested-by: Peter Zijlstra <peterz@infradead.org>
Signed-off-by: Peng Wang <peng_wang@linux.alibaba.com>
---
kernel/locking/rwsem.c | 41 ++++++++++++++++++++++++++++-------------
1 file changed, 28 insertions(+), 13 deletions(-)
diff --git a/kernel/locking/rwsem.c b/kernel/locking/rwsem.c
index bf647097369c..f04112d2a336 100644
--- a/kernel/locking/rwsem.c
+++ b/kernel/locking/rwsem.c
@@ -85,7 +85,8 @@
* Bit 0 - writer locked bit
* Bit 1 - waiters present bit
* Bit 2 - lock handoff bit
- * Bits 3-7 - reserved
+ * Bit 3 - writer phase bit
+ * Bits 4-7 - reserved
* Bits 8-62 - 55-bit reader count
* Bit 63 - read fail bit
*
@@ -94,7 +95,8 @@
* Bit 0 - writer locked bit
* Bit 1 - waiters present bit
* Bit 2 - lock handoff bit
- * Bits 3-7 - reserved
+ * Bit 3 - writer phase bit
+ * Bits 4-7 - reserved
* Bits 8-30 - 23-bit reader count
* Bit 31 - read fail bit
*
@@ -106,10 +108,11 @@
* atomic_long_fetch_add() is used to obtain reader lock, whereas
* atomic_long_cmpxchg() will be used to obtain writer lock.
*
- * There are three places where the lock handoff bit may be set or cleared.
- * 1) rwsem_mark_wake() for readers -- set, clear
- * 2) rwsem_try_write_lock() for writers -- set, clear
- * 3) rwsem_del_waiter() -- clear
+ * There are three places where the lock handoff bit and writer phase bit
+ * may be set or cleared.
+ * 1) rwsem_mark_wake() -- set, clear
+ * 2) rwsem_try_write_lock() -- clear
+ * 3) rwsem_del_waiter() -- clear
*
* For all the above cases, wait_lock will be held. A writer must also
* be the first one in the wait_list to be eligible for setting the handoff
@@ -118,6 +121,7 @@
#define RWSEM_WRITER_LOCKED (1UL << 0)
#define RWSEM_FLAG_WAITERS (1UL << 1)
#define RWSEM_FLAG_HANDOFF (1UL << 2)
+#define RWSEM_FLAG_WRITER_PHASE (1UL << 3)
#define RWSEM_FLAG_READFAIL (1UL << (BITS_PER_LONG - 1))
#define RWSEM_READER_SHIFT 8
@@ -126,7 +130,9 @@
#define RWSEM_WRITER_MASK RWSEM_WRITER_LOCKED
#define RWSEM_LOCK_MASK (RWSEM_WRITER_MASK|RWSEM_READER_MASK)
#define RWSEM_READ_FAILED_MASK (RWSEM_WRITER_MASK|RWSEM_FLAG_WAITERS|\
- RWSEM_FLAG_HANDOFF|RWSEM_FLAG_READFAIL)
+ RWSEM_FLAG_HANDOFF |\
+ RWSEM_FLAG_WRITER_PHASE |\
+ RWSEM_FLAG_READFAIL)
/*
* All writes to owner are protected by WRITE_ONCE() to make sure that
@@ -396,7 +402,8 @@ rwsem_del_waiter(struct rw_semaphore *sem, struct rwsem_waiter *waiter)
lockdep_assert_held(&sem->wait_lock);
if (__rwsem_del_waiter(sem, waiter))
return true;
- atomic_long_andnot(RWSEM_FLAG_HANDOFF | RWSEM_FLAG_WAITERS, &sem->count);
+ atomic_long_andnot(RWSEM_FLAG_HANDOFF | RWSEM_FLAG_WAITERS |
+ RWSEM_FLAG_WRITER_PHASE, &sem->count);
return false;
}
@@ -444,12 +451,13 @@ static void rwsem_mark_wake(struct rw_semaphore *sem,
if (waiter->type == RWSEM_WAITING_FOR_WRITE) {
if (wake_type == RWSEM_WAKE_ANY) {
+ atomic_long_or(RWSEM_FLAG_WRITER_PHASE, &sem->count);
/*
* Mark writer at the front of the queue for wakeup.
* Until the task is actually later awoken later by
* the caller, other writers are able to steal it.
* Readers, on the other hand, will block as they
- * will notice the queued writer.
+ * will notice the writer phase.
*/
wake_q_add(wake_q, waiter->task);
lockevent_inc(rwsem_wake_writer);
@@ -554,13 +562,17 @@ static void rwsem_mark_wake(struct rw_semaphore *sem,
adjustment -= RWSEM_FLAG_WAITERS;
if (oldcount & RWSEM_FLAG_HANDOFF)
adjustment -= RWSEM_FLAG_HANDOFF;
+ if (oldcount & RWSEM_FLAG_WRITER_PHASE)
+ adjustment -= RWSEM_FLAG_WRITER_PHASE;
} else if (woken) {
/*
* When we've woken a reader, we no longer need to force
- * writers to give up the lock and we can clear HANDOFF.
+ * writers to give up the lock and we can clear writer phase.
*/
if (oldcount & RWSEM_FLAG_HANDOFF)
adjustment -= RWSEM_FLAG_HANDOFF;
+ if (oldcount & RWSEM_FLAG_WRITER_PHASE)
+ adjustment -= RWSEM_FLAG_WRITER_PHASE;
}
if (adjustment)
@@ -663,7 +675,8 @@ static inline bool rwsem_try_write_lock(struct rw_semaphore *sem,
new &= ~RWSEM_FLAG_HANDOFF;
if (list_empty(&first->list))
- new &= ~RWSEM_FLAG_WAITERS;
+ new &= ~(RWSEM_FLAG_WAITERS |
+ RWSEM_FLAG_WRITER_PHASE);
}
} while (!atomic_long_try_cmpxchg_acquire(&sem->count, &count, new));
@@ -1033,7 +1046,8 @@ rwsem_down_read_slowpath(struct rw_semaphore *sem, long count, unsigned int stat
/*
* Reader optimistic lock stealing.
*/
- if (!(count & (RWSEM_WRITER_LOCKED | RWSEM_FLAG_HANDOFF))) {
+ if (!(count & (RWSEM_WRITER_LOCKED | RWSEM_FLAG_HANDOFF |
+ RWSEM_FLAG_WRITER_PHASE))) {
rwsem_set_reader_owned(sem);
lockevent_inc(rwsem_rlock_steal);
@@ -1175,7 +1189,8 @@ rwsem_down_write_slowpath(struct rw_semaphore *sem, int state)
} else {
INIT_LIST_HEAD(&waiter.list);
sem->first_waiter = &waiter;
- atomic_long_or(RWSEM_FLAG_WAITERS, &sem->count);
+ atomic_long_or(RWSEM_FLAG_WAITERS | RWSEM_FLAG_WRITER_PHASE,
+ &sem->count);
}
/* wait until we successfully acquire the lock */
--
2.47.3
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-06-04 14:15 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-05-21 9:59 [PATCH] locking/rwsem: Remove reader optimistic lock stealing Peng Wang
2026-05-22 2:08 ` Waiman Long
2026-05-22 3:33 ` Peng Wang
2026-05-22 8:55 ` Peter Zijlstra
2026-05-22 9:14 ` Peter Zijlstra
2026-05-22 10:06 ` Peng Wang
2026-06-04 14:14 ` [RFC PATCH v2] locking/rwsem: Block reader stealing during writer phase Peng Wang
2026-05-31 7:27 ` [PATCH] locking/rwsem: Remove reader optimistic lock stealing kernel test robot
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome