mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Mete Durlu <meted@linux.ibm.com>
To: Heiko Carstens <hca@linux.ibm.com>,
	Alexander Gordeev <agordeev@linux.ibm.com>,
	Sven Schnelle <svens@linux.ibm.com>,
	Vasily Gorbik <gor@linux.ibm.com>,
	Christian Borntraeger <borntraeger@linux.ibm.com>,
	Peter Zijlstra <peterz@infradead.org>,
	Mark Rutland <mark.rutland@arm.com>,
	Juergen Christ <jchrist@linux.ibm.com>,
	Ilya Leoshkevich <iii@linux.ibm.com>
Cc: linux-kernel@vger.kernel.org, linux-s390@vger.kernel.org
Subject: Re: [PATCH v4 00/11] s390: More this_cpu_*() changes
Date: Tue, 22 Sep 2026 17:37:14 +0200	[thread overview]
Message-ID: <dc01ef2f-3ed9-4b25-ae5f-1bd317df27cb@linux.ibm.com> (raw)
In-Reply-To: <20260921155806.2447506-1-hca@linux.ibm.com>

On 21/09/2026 17:57, Heiko Carstens wrote:
[..snip..]
> v1:
> Most of this is only about cleaning up the percpu code after preemptible
> this_cpu_*() operations have been implemented. The first ten patches are
> all more or less trivial cleanup patches trying to make the code shorter
> and more readable.
> 
> The only non-trivial patch is the last one, which converts s390's
> this_cpu_*() operations to use a similar scheme like Mark Rutland
> provided it for arm64 [1]. This allows to simplify the irq entry and exit
> path, however at the cost of slightly worse code for this_cpu_*()
> operations.
> 
> The simplified irq entry and exit code seems to be worth it. Usable
> performance numbers are not available yet, however I don't expect big
> difference to before.
> 
> [1] https://lore.kernel.org/all/20260904161758.376504-1-mark.rutland@arm.com/
> 
> Thanks,
> Heiko
> 
> Heiko Carstens (11):
>    s390/percpu: Fix comment typo
>    s390/percpu: Add sanity check to GEN_MVIY macro
>    s390/lowcore: Remove _AC() from LOWCORE_ALT_ADDRESS
>    s390/percpu: Let MVIY_PERCPU() calculate alternative displacement
>    s390/percpu/lowcore: Add and use LC_PERCPU lowcore offset defines
>    s390/percpu: Rename inline assembly symbolic names
>    s390/percpu: Use __PCPU_BEGIN() and __PCPU_END() for inline assemblies
>    s390/percpu: Use percpu code section for this_cpu_cmpxchg128()
>    s390/percpu: Use percpu code section for this_cpu_xchg()
>    s390/percpu: Use percpu code section for this_cpu_cmpxchg()
>    s390/percpu: Rework to simplify percpu_entry() and percpu_exit()
> 
>   arch/s390/include/asm/entry-percpu.h |  71 ++----
>   arch/s390/include/asm/lowcore.h      |   5 +-
>   arch/s390/include/asm/percpu.h       | 354 +++++++++++++++++----------
>   arch/s390/kernel/irq.c               |  10 +-
>   arch/s390/kernel/nmi.c               |   4 +-
>   arch/s390/kernel/traps.c             |   4 +-
>   6 files changed, 248 insertions(+), 200 deletions(-)
> 

Hi all,

I did some benchmarking for this patch series.

Used a shared LPAR with 24 CPUs (12 cores w SMT2).

Base commit: f0100363d8c3 ("Merge tag 'xfs-fixes-7.3-rc5' of
gitolite.kernel.org:/pub/scm/fs/xfs/xfs-linux")

Some of the benchmark runs show +-5% standard deviation which can
be attributed to s390 being a virtualized system. Each run has been
repeated five times to circumvent the effects of standard deviation.
(In the odd case where baseline has -5% and patched run has +5%
deviation, the comparison can show ~+10% improvements which is highly
misleading).

Overall, the benchmarks do not show any real difference between baseline
and patched kernels;

1-) Hackbench
========================================================================
Repeated hackbench runs with different group and fd count combinations;
[1, 2, 4, 8]

$ hackbench -T -p -l $loops -g $g -f $f

no real difference.

2-) Stress-ng
========================================================================
Repeated stress-ng runs with range of stressors: [6, $(nproc)]

# 3d matrix operations
$ stress-ng --matrix-3d $cpu --matrix-3d-method mult --timeout 10

# repeated mmap() and munmap() operations and also writing to the
# allocated memory.
$ stress-ng --vm $cpu --vm-bytes 128M --timeout 10

# starts workers that each fork off 32 child processes. Each child
# tries to allocate some memory. Child processes use madvise and memset
# to produce VM activity.
$ stress-ng --mmapfork $cpu --mmapfork-bytes 128M --timeout 10

no real difference.

3-) Openblas Benchmark
========================================================================
Repeated matrix operations with different size of 3d matrices.

$ make -j$(nproc) \
         USE_OPENMP=1 \
         NUM_THREADS=$(nproc)\
         NOFORTRAN=1 > /dev/null 2>&1

$ make -C benchmark sgemm.goto \
         USE_OPENMP=1 \
         NUM_THREADS=$(nproc) \
         NOFORTRAN=1 > /dev/null 2>&1

$ OPENBLAS_LOOPS=50 sgemm.goto

no real difference between runs.

4-) Compiling kernel
========================================================================
Linux kernel compilation

$ make clean
$ make mrproper
$ make defconfig
$ time make -j$(nproc)

no real difference between runs.

  parent reply	other threads:[~2026-09-22 15:37 UTC|newest]

Thread overview: 14+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-21 15:57 Heiko Carstens
2026-09-21 15:57 ` [PATCH v4 01/11] s390/percpu: Fix comment typo Heiko Carstens
2026-09-21 15:57 ` [PATCH v4 02/11] s390/percpu: Add sanity check to GEN_MVIY macro Heiko Carstens
2026-09-21 15:57 ` [PATCH v4 03/11] s390/lowcore: Remove _AC() from LOWCORE_ALT_ADDRESS Heiko Carstens
2026-09-21 15:57 ` [PATCH v4 04/11] s390/percpu: Let MVIY_PERCPU() calculate alternative displacement Heiko Carstens
2026-09-21 15:57 ` [PATCH v4 05/11] s390/percpu/lowcore: Add and use LC_PERCPU lowcore offset defines Heiko Carstens
2026-09-21 15:58 ` [PATCH v4 06/11] s390/percpu: Rename inline assembly symbolic names Heiko Carstens
2026-09-21 15:58 ` [PATCH v4 07/11] s390/percpu: Use __PCPU_BEGIN() and __PCPU_END() for inline assemblies Heiko Carstens
2026-09-21 15:58 ` [PATCH v4 08/11] s390/percpu: Use percpu code section for this_cpu_cmpxchg128() Heiko Carstens
2026-09-21 15:58 ` [PATCH v4 09/11] s390/percpu: Use percpu code section for this_cpu_xchg() Heiko Carstens
2026-09-21 15:58 ` [PATCH v4 10/11] s390/percpu: Use percpu code section for this_cpu_cmpxchg() Heiko Carstens
2026-09-21 15:58 ` [PATCH v4 11/11] s390/percpu: Rework to simplify percpu_entry() and percpu_exit() Heiko Carstens
2026-09-22 15:37 ` Mete Durlu [this message]
2026-09-22 18:22   ` [PATCH v4 00/11] s390: More this_cpu_*() changes Heiko Carstens

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=dc01ef2f-3ed9-4b25-ae5f-1bd317df27cb@linux.ibm.com \
    --to=meted@linux.ibm.com \
    --cc=agordeev@linux.ibm.com \
    --cc=borntraeger@linux.ibm.com \
    --cc=gor@linux.ibm.com \
    --cc=hca@linux.ibm.com \
    --cc=iii@linux.ibm.com \
    --cc=jchrist@linux.ibm.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=mark.rutland@arm.com \
    --cc=peterz@infradead.org \
    --cc=svens@linux.ibm.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®