* [PATCH] patch in prfm for copy_template if requested
@ 2017-01-11 1:27 Andrew Pinski
2017-01-11 1:27 ` [PATCH] arm64: lib: " Andrew Pinski
0 siblings, 1 reply; 2+ messages in thread
From: Andrew Pinski @ 2017-01-11 1:27 UTC (permalink / raw)
To: linux-arm-kernel, linux-kernel; +Cc: Andrew Pinski
As mentioned in http://lists.infradead.org/pipermail/linux-arm-kernel/2016-February/404146.html
copy_template was left alone at the time which mentions:
"since the template really deals with 64 bytes per iteration,
which would need changing". The problem is that there is not enough
registers available to do 128 bytes at a time. There is only enough
registers to do 96 bytes at a time. If we did not have to save
dst or keep x5 free (that is used by the exception case) or keep
around the count; then we would have enough caller saved registers
free to copy 128 bytes at a time. For user space, we will be using
the SIMD registers which allows for not using any callee saved
registers and get better performance.
So basically this is my old patch which just patches in the prfm
to copy_template updated for the new name of the define and for
the nop not needed to be there any more.
Andrew Pinski (1):
arm64: lib: patch in prfm for copy_template if requested
arch/arm64/lib/copy_template.S | 9 ++++++++-
arch/arm64/lib/memcpy.S | 3 +++
2 files changed, 11 insertions(+), 1 deletion(-)
--
2.7.4
^ permalink raw reply [flat|nested] 2+ messages in thread
* [PATCH] arm64: lib: patch in prfm for copy_template if requested
2017-01-11 1:27 [PATCH] patch in prfm for copy_template if requested Andrew Pinski
@ 2017-01-11 1:27 ` Andrew Pinski
0 siblings, 0 replies; 2+ messages in thread
From: Andrew Pinski @ 2017-01-11 1:27 UTC (permalink / raw)
To: linux-arm-kernel, linux-kernel; +Cc: Andrew Pinski
On ThunderX T88 pass 1 and pass 2, there is no hardware prefetching so
we need to patch in explicit software prefetching instructions.
This speeds up copy_to_user and copy_from_user for large size.
The main use of large sizes is I/O read/writes.
Signed-off-by: Andrew Pinski <apinski@cavium.com>
---
arch/arm64/lib/copy_template.S | 9 ++++++++-
arch/arm64/lib/memcpy.S | 3 +++
2 files changed, 11 insertions(+), 1 deletion(-)
diff --git a/arch/arm64/lib/copy_template.S b/arch/arm64/lib/copy_template.S
index 410fbdb..ef99f686a 100644
--- a/arch/arm64/lib/copy_template.S
+++ b/arch/arm64/lib/copy_template.S
@@ -1,5 +1,5 @@
/*
- * Copyright (C) 2013 ARM Ltd.
+ * Copfrigt (C) 2013 ARM Ltd.
* Copyright (C) 2013 Linaro.
*
* This code is based on glibc cortex strings work originally authored by Linaro
@@ -163,12 +163,19 @@ D_h .req x14
*/
.p2align L1_CACHE_SHIFT
.Lcpy_body_large:
+alternative_if ARM64_HAS_NO_HW_PREFETCH
+ prfm pldl1strm, [src, #128]
+ prfm pldl1strm, [src, #256]
+alternative_else_nop_endif
/* pre-get 64 bytes data. */
ldp1 A_l, A_h, src, #16
ldp1 B_l, B_h, src, #16
ldp1 C_l, C_h, src, #16
ldp1 D_l, D_h, src, #16
1:
+alternative_if ARM64_HAS_NO_HW_PREFETCH
+ prfm pldl1strm, [src, #384]
+alternative_else_nop_endif
/*
* interlace the load of next 64 bytes data block with store of the last
* loaded 64 bytes data.
diff --git a/arch/arm64/lib/memcpy.S b/arch/arm64/lib/memcpy.S
index 6761393..ee30fd5 100644
--- a/arch/arm64/lib/memcpy.S
+++ b/arch/arm64/lib/memcpy.S
@@ -25,6 +25,9 @@
#include <linux/linkage.h>
#include <asm/assembler.h>
#include <asm/cache.h>
+#include <asm/alternative.h>
+#include <asm/cpufeature.h>
+
/*
* Copy a buffer from src to dest (alignment handled by the hardware)
--
2.7.4
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2017-01-11 2:30 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2017-01-11 1:27 [PATCH] patch in prfm for copy_template if requested Andrew Pinski
2017-01-11 1:27 ` [PATCH] arm64: lib: " Andrew Pinski
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®