mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v4 0/5] spi: dw: use threaded interrupt
@ 2026-09-29 12:18 Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 1/5] spi: dw: use DW_SPI_ISR directly Jisheng Zhang
                   ` (4 more replies)
  0 siblings, 5 replies; 7+ messages in thread
From: Jisheng Zhang @ 2026-09-29 12:18 UTC (permalink / raw)
  To: Mark Brown, Joseph Steel; +Cc: linux-spi, linux-kernel

To avoid blocking for an excessive amount of time, eventually impacting
on system responsiveness, hard interrupt handlers should finish
executing in as little time as possible.

Use threaded interrupt and move the SPI transfer handling to an
interrupt thread when non-native CS and host only.

After that, since the dw_reader() and dw_writer() are called in
threaded ISR now, so we can delay the unmasking interrupts until no
rx and tx action is taken, thus reduce the interrupt numbers further.

Tested with below two cmds
./spidev_test -D /dev/spidev1.3 -s 30000000 -S 327680 -I 1

./spidev_test -D /dev/spidev1.3 -s 30000000 -S 327680 -I 1000
./rtla timerlat top -q -k -P f:95

The first cmd is to check the interrupt numbers optmizaion result, the
2nd cmd group is to check the threaded interrupt improvement.

Before the patch:
each 320KB spi spidev_test transfer triggers 33118 interrupts

spidev_test reports ~22090 kbps
and rtla reports:
                                     Timer Latency
  0 00:00:37   |          IRQ Timer Latency (us)        |         Thread Timer Latency (us)
CPU COUNT      |      cur       min       avg       max |      cur       min       avg       max
  0 #9958      |        1         0        67    103394 |        6         4      2198    105031
  1 #36902     |        1         0         1        18 |        5         4         5        29

After the patch:
each 320KB spi spidev_test transfer only triggers 1 interrupts
spidev_test reports ~23520 kbps
and now rtla reports:
                                    Timer Latency
  0 00:00:58   |          IRQ Timer Latency (us)        |         Thread Timer Latency (us)
CPU COUNT      |      cur       min       avg       max |      cur       min       avg       max
  0 #58362     |        1         0         0        29 |        6         3         4        56
  1 #58363     |        1         0         1        23 |        6         4         5        68

In summary:
before the patch	after the patch
33118 interrutps	1 interrupts		reduced by 33117 times!
103394 us max latency	29 us max latency	reduced by 3564 times!
22090 kbps rate		23520 kbps rate		improved by 6.5%

Since v3:
  - If native cs, don't use threaded interrupt
  - do one round of FIFO in the hardirq
  - add cond_resched in dw_spi_irq_thread_fn

Since v2:
  - rebase against latest version
  - remove the "spi: dw: use DW_SPI_INT_MASK instead of hardcoded 0xff"
  - add three more patches to clean up irq code introduced by recent "enhanced
    spi" support
  - Don't use threaded interrupt for target mode and the enhanced spi

Since v1:
  - rebase against latest version
  - drop two patches which have been merged
  - correct some performance numbers
  - don't move request irq code block so no changes for err handling code path
  - move spi_finalize_current_transfer to the end of threaded irq fn
  - don't rely on irq status in threaded fn, but try rx and tx as much
    as possible in the loop


Jisheng Zhang (5):
  spi: dw: use DW_SPI_ISR directly
  spi: dw: remove useless dws->transfer_handler check
  spi: dw: remove duplicated "!rx_len && !tx_len" handling from
    dw_spi_irq
  spi: dw: restore previous irq handling behavior when !ctlr->cur_msg
  spi: dw: use threaded interrupt and optimize the threaded ISR

 drivers/spi/spi-dw-core.c | 78 ++++++++++++++++++++++++++++++++-------
 1 file changed, 64 insertions(+), 14 deletions(-)

-- 
2.53.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH v4 1/5] spi: dw: use DW_SPI_ISR directly
  2026-09-29 12:18 [PATCH v4 0/5] spi: dw: use threaded interrupt Jisheng Zhang
@ 2026-09-29 12:18 ` Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 2/5] spi: dw: remove useless dws->transfer_handler check Jisheng Zhang
                   ` (3 subsequent siblings)
  4 siblings, 0 replies; 7+ messages in thread
From: Jisheng Zhang @ 2026-09-29 12:18 UTC (permalink / raw)
  To: Mark Brown, Joseph Steel; +Cc: linux-spi, linux-kernel

The DW_SPI_ISR register reports the masked interrupts, no need to mask
again.

Signed-off-by: Jisheng Zhang <jszhang@kernel.org>
---
 drivers/spi/spi-dw-core.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/spi/spi-dw-core.c b/drivers/spi/spi-dw-core.c
index 206d3f9dd83d..1c060f0ec50c 100644
--- a/drivers/spi/spi-dw-core.c
+++ b/drivers/spi/spi-dw-core.c
@@ -275,7 +275,7 @@ static irqreturn_t dw_spi_irq(int irq, void *dev_id)
 {
 	struct spi_controller *ctlr = dev_id;
 	struct dw_spi *dws = spi_controller_get_devdata(ctlr);
-	u16 irq_status = dw_readl(dws, DW_SPI_ISR) & DW_SPI_INT_MASK;
+	u16 irq_status = dw_readl(dws, DW_SPI_ISR);
 
 	if (!irq_status)
 		return IRQ_NONE;
-- 
2.53.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH v4 2/5] spi: dw: remove useless dws->transfer_handler check
  2026-09-29 12:18 [PATCH v4 0/5] spi: dw: use threaded interrupt Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 1/5] spi: dw: use DW_SPI_ISR directly Jisheng Zhang
@ 2026-09-29 12:18 ` Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 3/5] spi: dw: remove duplicated "!rx_len && !tx_len" handling from dw_spi_irq Jisheng Zhang
                   ` (2 subsequent siblings)
  4 siblings, 0 replies; 7+ messages in thread
From: Jisheng Zhang @ 2026-09-29 12:18 UTC (permalink / raw)
  To: Mark Brown, Joseph Steel; +Cc: linux-spi, linux-kernel

transfer_handler is initialized before any DW SPI interrupt is
unmasked. While the IRQ is registered but the handler is NULL,
dw_spi_hw_init() has all DW SPI interrupts disabled. Therefore
dw_spi_irq() cannot observe a DW SPI interrupt with a NULL
transfer_handler.

Signed-off-by: Jisheng Zhang <jszhang@kernel.org>
---
 drivers/spi/spi-dw-core.c | 3 +--
 1 file changed, 1 insertion(+), 2 deletions(-)

diff --git a/drivers/spi/spi-dw-core.c b/drivers/spi/spi-dw-core.c
index 1c060f0ec50c..f4c4e9dae25e 100644
--- a/drivers/spi/spi-dw-core.c
+++ b/drivers/spi/spi-dw-core.c
@@ -280,8 +280,7 @@ static irqreturn_t dw_spi_irq(int irq, void *dev_id)
 	if (!irq_status)
 		return IRQ_NONE;
 
-	if (!dws->transfer_handler ||
-	    (!ctlr->cur_msg && dws->transfer_handler == dw_spi_transfer_handler)) {
+	if (!ctlr->cur_msg && dws->transfer_handler == dw_spi_transfer_handler) {
 		dw_spi_mask_intr(dws, 0xff);
 		return IRQ_HANDLED;
 	}
-- 
2.53.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH v4 3/5] spi: dw: remove duplicated "!rx_len && !tx_len" handling from dw_spi_irq
  2026-09-29 12:18 [PATCH v4 0/5] spi: dw: use threaded interrupt Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 1/5] spi: dw: use DW_SPI_ISR directly Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 2/5] spi: dw: remove useless dws->transfer_handler check Jisheng Zhang
@ 2026-09-29 12:18 ` Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 4/5] spi: dw: restore previous irq handling behavior when !ctlr->cur_msg Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 5/5] spi: dw: use threaded interrupt and optimize the threaded ISR Jisheng Zhang
  4 siblings, 0 replies; 7+ messages in thread
From: Jisheng Zhang @ 2026-09-29 12:18 UTC (permalink / raw)
  To: Mark Brown, Joseph Steel; +Cc: linux-spi, linux-kernel

dw_spi_irq() need not handle enhanced-transfer's "!rx_len && !tx_len"
case as that state is already handled by dw_spi_enh_handler()

Signed-off-by: Jisheng Zhang <jszhang@kernel.org>
---
 drivers/spi/spi-dw-core.c | 6 ------
 1 file changed, 6 deletions(-)

diff --git a/drivers/spi/spi-dw-core.c b/drivers/spi/spi-dw-core.c
index f4c4e9dae25e..91de357f97f3 100644
--- a/drivers/spi/spi-dw-core.c
+++ b/drivers/spi/spi-dw-core.c
@@ -284,12 +284,6 @@ static irqreturn_t dw_spi_irq(int irq, void *dev_id)
 		dw_spi_mask_intr(dws, 0xff);
 		return IRQ_HANDLED;
 	}
-	if (dws->transfer_handler == dw_spi_enh_handler &&
-	    !dws->rx_len && !dws->tx_len) {
-		dw_spi_mask_intr(dws, 0xff);
-		spi_finalize_current_transfer(ctlr);
-		return IRQ_HANDLED;
-	}
 
 	return dws->transfer_handler(dws);
 }
-- 
2.53.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH v4 4/5] spi: dw: restore previous irq handling behavior when !ctlr->cur_msg
  2026-09-29 12:18 [PATCH v4 0/5] spi: dw: use threaded interrupt Jisheng Zhang
                   ` (2 preceding siblings ...)
  2026-09-29 12:18 ` [PATCH v4 3/5] spi: dw: remove duplicated "!rx_len && !tx_len" handling from dw_spi_irq Jisheng Zhang
@ 2026-09-29 12:18 ` Jisheng Zhang
  2026-09-29 12:18 ` [PATCH v4 5/5] spi: dw: use threaded interrupt and optimize the threaded ISR Jisheng Zhang
  4 siblings, 0 replies; 7+ messages in thread
From: Jisheng Zhang @ 2026-09-29 12:18 UTC (permalink / raw)
  To: Mark Brown, Joseph Steel; +Cc: linux-spi, linux-kernel

Recent "enhanced" support changes the dw_spi_irq() behavior a bit
when !!ctlr->cur_msg:

    if (!ctlr->cur_msg && dws->transfer_handler == dw_spi_transfer_handler) {
            dw_spi_mask_intr(dws, 0xff);
            return IRQ_HANDLED;
    }

But it misses the dma case, where the transfer_handler ==
dw_spi_dma_transfer_handler, so this changes the previous long time
working behavior, let's restore the previous handling by only filtering
out the dw_spi_enh_handler.

Signed-off-by: Jisheng Zhang <jszhang@kernel.org>
---
 drivers/spi/spi-dw-core.c | 2 +-
 1 file changed, 1 insertion(+), 1 deletion(-)

diff --git a/drivers/spi/spi-dw-core.c b/drivers/spi/spi-dw-core.c
index 91de357f97f3..da7d4872e513 100644
--- a/drivers/spi/spi-dw-core.c
+++ b/drivers/spi/spi-dw-core.c
@@ -280,7 +280,7 @@ static irqreturn_t dw_spi_irq(int irq, void *dev_id)
 	if (!irq_status)
 		return IRQ_NONE;
 
-	if (!ctlr->cur_msg && dws->transfer_handler == dw_spi_transfer_handler) {
+	if (!ctlr->cur_msg && dws->transfer_handler != dw_spi_enh_handler) {
 		dw_spi_mask_intr(dws, 0xff);
 		return IRQ_HANDLED;
 	}
-- 
2.53.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* [PATCH v4 5/5] spi: dw: use threaded interrupt and optimize the threaded ISR
  2026-09-29 12:18 [PATCH v4 0/5] spi: dw: use threaded interrupt Jisheng Zhang
                   ` (3 preceding siblings ...)
  2026-09-29 12:18 ` [PATCH v4 4/5] spi: dw: restore previous irq handling behavior when !ctlr->cur_msg Jisheng Zhang
@ 2026-09-29 12:18 ` Jisheng Zhang
  2026-09-29 14:19   ` Mark Brown
  4 siblings, 1 reply; 7+ messages in thread
From: Jisheng Zhang @ 2026-09-29 12:18 UTC (permalink / raw)
  To: Mark Brown, Joseph Steel; +Cc: linux-spi, linux-kernel

To avoid blocking for an excessive amount of time, eventually impacting
on system responsiveness, hard interrupt handlers should finish
executing in as little time as possible.

Use threaded interrupt and move the SPI transfer handling to an
interrupt thread when non-native CS and host only.

After that, since the dw_reader() and dw_writer() are called in
threaded ISR now, so we can delay the unmasking interrupts until no
rx and tx action is taken, thus reduce the interrupt numbers further.

Tested with below two cmds
./spidev_test -D /dev/spidev1.3 -s 30000000 -S 327680 -I 1

./spidev_test -D /dev/spidev1.3 -s 30000000 -S 327680 -I 1000
./rtla timerlat top -q -k -P f:95

The first cmd is to check the interrupt numbers optmizaion result, the
2nd cmd group is to check the threaded interrupt improvement.

Before the patch:
each 320KB spi spidev_test transfer triggers 33118 interrupts

spidev_test reports ~22090 kbps
and rtla reports:
                                     Timer Latency
  0 00:00:37   |          IRQ Timer Latency (us)        |         Thread Timer Latency (us)
CPU COUNT      |      cur       min       avg       max |      cur       min       avg       max
  0 #9958      |        1         0        67    103394 |        6         4      2198    105031
  1 #36902     |        1         0         1        18 |        5         4         5        29

After the patch:
each 320KB spi spidev_test transfer only triggers 1 interrupts
spidev_test reports ~23520 kbps
and now rtla reports:
                                    Timer Latency
  0 00:00:58   |          IRQ Timer Latency (us)        |         Thread Timer Latency (us)
CPU COUNT      |      cur       min       avg       max |      cur       min       avg       max
  0 #58362     |        1         0         0        29 |        6         3         4        56
  1 #58363     |        1         0         1        23 |        6         4         5        68

In summary:
before the patch	after the patch
33118 interrutps	1 interrupts		reduced by 33117 times!
103394 us max latency	29 us max latency	reduced by 3564 times!
22090 kbps rate		23520 kbps rate		improved by 6.5%

Signed-off-by: Jisheng Zhang <jszhang@kernel.org>
---
 drivers/spi/spi-dw-core.c | 67 ++++++++++++++++++++++++++++++++++++---
 1 file changed, 62 insertions(+), 5 deletions(-)

diff --git a/drivers/spi/spi-dw-core.c b/drivers/spi/spi-dw-core.c
index da7d4872e513..bd86cbe8fcfe 100644
--- a/drivers/spi/spi-dw-core.c
+++ b/drivers/spi/spi-dw-core.c
@@ -132,10 +132,11 @@ static inline u32 dw_spi_rx_max(struct dw_spi *dws)
 	return min_t(u32, dws->rx_len, dw_readl(dws, DW_SPI_RXFLR));
 }
 
-static void dw_writer(struct dw_spi *dws)
+static u32 dw_writer(struct dw_spi *dws)
 {
 	u32 max = dw_spi_tx_max(dws);
 	u32 txw = 0;
+	u32 tx = 0;
 
 	while (max--) {
 		if (dws->tx) {
@@ -150,13 +151,16 @@ static void dw_writer(struct dw_spi *dws)
 		}
 		dw_write_io_reg(dws, DW_SPI_DR, txw);
 		--dws->tx_len;
+		++tx;
 	}
+	return tx;
 }
 
-static void dw_reader(struct dw_spi *dws)
+static u32 dw_reader(struct dw_spi *dws)
 {
 	u32 max = dw_spi_rx_max(dws);
 	u32 rxw;
+	u32 rx = 0;
 
 	while (max--) {
 		rxw = dw_read_io_reg(dws, DW_SPI_DR);
@@ -171,7 +175,9 @@ static void dw_reader(struct dw_spi *dws)
 			dws->rx += dws->n_bytes;
 		}
 		--dws->rx_len;
+		++rx;
 	}
+	return rx;
 }
 
 int dw_spi_check_status(struct dw_spi *dws, bool raw)
@@ -210,6 +216,51 @@ int dw_spi_check_status(struct dw_spi *dws, bool raw)
 }
 EXPORT_SYMBOL_NS_GPL(dw_spi_check_status, "SPI_DW_CORE");
 
+static irqreturn_t dw_spi_irq_thread_fn(int irq, void *dev_id)
+{
+	struct spi_controller *ctlr = dev_id;
+	struct dw_spi *dws = spi_controller_get_devdata(ctlr);
+	u32 rx, tx, imask, mask = 0;
+	bool finalize = false;
+
+	do {
+		/*
+		 * Read data from the Rx FIFO every time we've got a chance executing
+		 * this method. If there is nothing left to receive, terminate the
+		 * procedure. Otherwise adjust the Rx FIFO Threshold level if it's a
+		 * final stage of the transfer. By doing so we'll get the next IRQ
+		 * right when the leftover incoming data is received.
+		 */
+		rx = dw_reader(dws);
+		if (!dws->rx_len) {
+			mask |= 0xff;
+			finalize = true;
+		} else if (dws->rx_len <= dw_readl(dws, DW_SPI_RXFTLR)) {
+			dw_writel(dws, DW_SPI_RXFTLR, dws->rx_len - 1);
+		}
+
+		/*
+		 * Send data out as much as possible. The Tx FIFO Empty IRQ will be
+		 * disabled after the data transmission is finished so not to
+		 * have the TXE IRQ flood at the final stage of the transfer.
+		 */
+		tx = dw_writer(dws);
+		if (!dws->tx_len)
+			mask |= DW_SPI_INT_TXEI;
+		cond_resched();
+	} while (rx != 0 || tx != 0);
+
+	imask = DW_SPI_INT_TXEI | DW_SPI_INT_TXOI |
+		DW_SPI_INT_RXUI | DW_SPI_INT_RXOI | DW_SPI_INT_RXFI;
+	imask &= ~mask;
+	dw_spi_umask_intr(dws, imask);
+
+	if (finalize)
+		spi_finalize_current_transfer(dws->ctlr);
+
+	return IRQ_HANDLED;
+}
+
 static irqreturn_t dw_spi_transfer_handler(struct dw_spi *dws)
 {
 	u16 irq_status = dw_readl(dws, DW_SPI_ISR);
@@ -245,7 +296,13 @@ static irqreturn_t dw_spi_transfer_handler(struct dw_spi *dws)
 			dw_spi_mask_intr(dws, DW_SPI_INT_TXEI);
 	}
 
-	return IRQ_HANDLED;
+	if (spi_controller_is_target(dws->ctlr) || !dws->rx_len || !dws->tx_len ||
+		!spi_get_csgpiod(dws->ctlr->cur_msg->spi, 0)) {
+		return IRQ_HANDLED;
+	} else {
+		dw_spi_mask_intr(dws, 0xff);
+		return IRQ_WAKE_THREAD;
+	}
 }
 
 static irqreturn_t dw_spi_enh_handler(struct dw_spi *dws)
@@ -1305,8 +1362,8 @@ int dw_spi_add_controller(struct device *dev, struct dw_spi *dws)
 	/* Basic HW init */
 	dw_spi_hw_init(dev, dws);
 
-	ret = request_irq(dws->irq, dw_spi_irq, IRQF_SHARED, dev_name(dev),
-			  ctlr);
+	ret = request_threaded_irq(dws->irq, dw_spi_irq, dw_spi_irq_thread_fn,
+				   IRQF_SHARED, dev_name(dev), ctlr);
 	if (ret < 0 && ret != -ENOTCONN) {
 		dev_err(dev, "can not request IRQ\n");
 		goto err_free_ctlr;
-- 
2.53.0


^ permalink raw reply	[flat|nested] 7+ messages in thread

* Re: [PATCH v4 5/5] spi: dw: use threaded interrupt and optimize the threaded ISR
  2026-09-29 12:18 ` [PATCH v4 5/5] spi: dw: use threaded interrupt and optimize the threaded ISR Jisheng Zhang
@ 2026-09-29 14:19   ` Mark Brown
  0 siblings, 0 replies; 7+ messages in thread
From: Mark Brown @ 2026-09-29 14:19 UTC (permalink / raw)
  To: Jisheng Zhang; +Cc: Joseph Steel, linux-spi, linux-kernel

[-- Attachment #1: Type: text/plain, Size: 1241 bytes --]

On Tue, Sep 29, 2026 at 08:18:17PM +0800, Jisheng Zhang wrote:
> To avoid blocking for an excessive amount of time, eventually impacting
> on system responsiveness, hard interrupt handlers should finish
> executing in as little time as possible.

>  static irqreturn_t dw_spi_transfer_handler(struct dw_spi *dws)
>  {
>  	u16 irq_status = dw_readl(dws, DW_SPI_ISR);
> @@ -245,7 +296,13 @@ static irqreturn_t dw_spi_transfer_handler(struct dw_spi *dws)
>  			dw_spi_mask_intr(dws, DW_SPI_INT_TXEI);
>  	}
>  
> -	return IRQ_HANDLED;
> +	if (spi_controller_is_target(dws->ctlr) || !dws->rx_len || !dws->tx_len ||
> +		!spi_get_csgpiod(dws->ctlr->cur_msg->spi, 0)) {
> +		return IRQ_HANDLED;
> +	} else {
> +		dw_spi_mask_intr(dws, 0xff);
> +		return IRQ_WAKE_THREAD;
> +	}
>  }

If we're handing a short transfer in the first interrupt then we will
already have completed the transfer at the point we check have which
would allow another CPU to start another transfer first.  If things go
very badly this might mean that the length checks look at the values
from a new transfer and wake the thread.  We should probably set a local
flag before we complete the transfer and use that to decide if we wake
the thread.

[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 488 bytes --]

^ permalink raw reply	[flat|nested] 7+ messages in thread

end of thread, other threads:[~2026-09-29 14:19 UTC | newest]

Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 12:18 [PATCH v4 0/5] spi: dw: use threaded interrupt Jisheng Zhang
2026-09-29 12:18 ` [PATCH v4 1/5] spi: dw: use DW_SPI_ISR directly Jisheng Zhang
2026-09-29 12:18 ` [PATCH v4 2/5] spi: dw: remove useless dws->transfer_handler check Jisheng Zhang
2026-09-29 12:18 ` [PATCH v4 3/5] spi: dw: remove duplicated "!rx_len && !tx_len" handling from dw_spi_irq Jisheng Zhang
2026-09-29 12:18 ` [PATCH v4 4/5] spi: dw: restore previous irq handling behavior when !ctlr->cur_msg Jisheng Zhang
2026-09-29 12:18 ` [PATCH v4 5/5] spi: dw: use threaded interrupt and optimize the threaded ISR Jisheng Zhang
2026-09-29 14:19   ` Mark Brown

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®