mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH] mm: Make swapoff interruptible when unusing mms/shmem
@ 2026-09-30 23:17 Chris Down
  2026-09-30 23:46 ` Andrew Morton
  2026-10-01 12:50 ` Vineeth Remanan Pillai
  0 siblings, 2 replies; 3+ messages in thread
From: Chris Down @ 2026-09-30 23:17 UTC (permalink / raw)
  To: Andrew Morton
  Cc: Hugh Dickins, Baolin Wang, Chris Li, Kairui Song, Kemeng Shi,
	Nhat Pham, Baoquan He, Barry Song, Youngjun Park, Ying Huang,
	Kelley Nielsen, Vineeth Pillai, linux-mm, linux-kernel,
	kernel-team

try_to_unuse() only checks for a pending signal between mms, and
shmem_unuse() doesn't check at all. That means that once swapoff gets to
a process or a shmem file with a lot swapped out, nothing can interrupt
it until every last page of it has been read back in.

Just as one example of where this can concretely show up, freezing tasks
for suspend or hibernation has to wait for swapoff to notice the
freezer's fake signal, and gives up after freeze_timeout_msecs (20
seconds by default).

Here's a facetious example where one swaps out 2GiB of one process to a
swap file on ext4, starts swapoff, and half a second later tries to
freeze with pm_test=freezer. Writing to /sys/power/state then fails with
EBUSY and this in dmesg:

    Freezing user space processes failed after 20.003 seconds (1 tasks refusing to freeze, wq_busy=0):
    task:swapoff         state:D stack:0     pid:3175  tgid:3175  ppid:2955   task_flags:0x400100 flags:0x00000419
    Call trace:
     [...]
     io_schedule+0x44/0x70
     folio_wait_bit_common+0x1ec/0x3d0
     __folio_lock+0x24/0x40
     unuse_pte_range+0x2d0/0x348
     unuse_vma+0x158/0x248
     unuse_mm+0xfc/0x150
     try_to_unuse+0x104/0x3f8
     __do_sys_swapoff+0x220/0x5d8
     [...]

The same goes for anything else that wants swapoff to stop, like an
admin hitting ^C in a panic, of course.

Prior to commit b56a2d8af914 ("mm: rid swapoff of quadratic complexity")
try_to_unuse() was driven by find_next_to_unuse() which checks for a
signal before every entry, so let's restore that behaviour.

Just as an example of the improvements, here's how long freezing takes
in the same test while swapoff is happening on my computer:

                      before                  after
    400MiB anon       8.925s                  0.028s
    400MiB shmem      1.639s                  0.003s
    2GiB anon         failed after 20.003s    0.011s

Fixes: b56a2d8af914 ("mm: rid swapoff of quadratic complexity")
Signed-off-by: Chris Down <chris@chrisdown.name>
---
 mm/shmem.c    | 4 ++++
 mm/swapfile.c | 2 ++
 2 files changed, 6 insertions(+)

diff --git a/mm/shmem.c b/mm/shmem.c
index ae08cff4500c..72c8a61db76f 100644
--- a/mm/shmem.c
+++ b/mm/shmem.c
@@ -1742,6 +1742,10 @@ static int shmem_unuse_inode(struct inode *inode, unsigned int type)
 		if (ret < 0)
 			break;
 
+		if (signal_pending(current)) {
+			ret = -EINTR;
+			break;
+		}
 		start = indices[folio_batch_count(&fbatch) - 1];
 	} while (true);
 
diff --git a/mm/swapfile.c b/mm/swapfile.c
index 254ce86fa923..c3288910b3e3 100644
--- a/mm/swapfile.c
+++ b/mm/swapfile.c
@@ -2689,6 +2689,8 @@ static inline int unuse_pmd_range(struct vm_area_struct *vma, pud_t *pud,
 	pmd = pmd_offset(pud, addr);
 	do {
 		cond_resched();
+		if (signal_pending(current))
+			return -EINTR;
 		next = pmd_addr_end(addr, end);
 		ret = unuse_pte_range(vma, pmd, addr, next, type);
 		if (ret)

base-commit: 2ddb90ee544ae97215afc4698dc223293997cc43
-- 
2.50.1


^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] mm: Make swapoff interruptible when unusing mms/shmem
  2026-09-30 23:17 [PATCH] mm: Make swapoff interruptible when unusing mms/shmem Chris Down
@ 2026-09-30 23:46 ` Andrew Morton
  2026-10-01 12:50 ` Vineeth Remanan Pillai
  1 sibling, 0 replies; 3+ messages in thread
From: Andrew Morton @ 2026-09-30 23:46 UTC (permalink / raw)
  To: Chris Down
  Cc: Hugh Dickins, Baolin Wang, Chris Li, Kairui Song, Kemeng Shi,
	Nhat Pham, Baoquan He, Barry Song, Youngjun Park, Ying Huang,
	Kelley Nielsen, Vineeth Pillai, linux-mm, linux-kernel,
	kernel-team, Rafael J. Wysocki

On Thu, 1 Oct 2026 01:17:40 +0200 Chris Down <chris@chrisdown.name> wrote:

> try_to_unuse() only checks for a pending signal between mms, and
> shmem_unuse() doesn't check at all. That means that once swapoff gets to
> a process or a shmem file with a lot swapped out, nothing can interrupt
> it until every last page of it has been read back in.
> 
> Just as one example of where this can concretely show up, freezing tasks
> for suspend or hibernation has to wait for swapoff to notice the
> freezer's fake signal, and gives up after freeze_timeout_msecs (20
> seconds by default).

(cc Rafael)

> Here's a facetious example where one swaps out 2GiB of one process to a
> swap file on ext4, starts swapoff, and half a second later tries to
> freeze with pm_test=freezer. Writing to /sys/power/state then fails with
> EBUSY and this in dmesg:
> 
>     Freezing user space processes failed after 20.003 seconds (1 tasks refusing to freeze, wq_busy=0):
>     task:swapoff         state:D stack:0     pid:3175  tgid:3175  ppid:2955   task_flags:0x400100 flags:0x00000419
>     Call trace:
>      [...]
>      io_schedule+0x44/0x70
>      folio_wait_bit_common+0x1ec/0x3d0
>      __folio_lock+0x24/0x40
>      unuse_pte_range+0x2d0/0x348
>      unuse_vma+0x158/0x248
>      unuse_mm+0xfc/0x150
>      try_to_unuse+0x104/0x3f8
>      __do_sys_swapoff+0x220/0x5d8
>      [...]

ugh.

> The same goes for anything else that wants swapoff to stop, like an
> admin hitting ^C in a panic, of course.
> 
> Prior to commit b56a2d8af914 ("mm: rid swapoff of quadratic complexity")
> try_to_unuse() was driven by find_next_to_unuse() which checks for a
> signal before every entry, so let's restore that behaviour.

So things were all good before that change?

> Just as an example of the improvements, here's how long freezing takes
> in the same test while swapoff is happening on my computer:
> 
>                       before                  after
>     400MiB anon       8.925s                  0.028s
>     400MiB shmem      1.639s                  0.003s
>     2GiB anon         failed after 20.003s    0.011s
> 
> Fixes: b56a2d8af914 ("mm: rid swapoff of quadratic complexity")
> Signed-off-by: Chris Down <chris@chrisdown.name>

This sounds like a significant usability regression.  Should we
backport this?

otoh, it's been this way since 2019, so presumably nobody cares much?

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [PATCH] mm: Make swapoff interruptible when unusing mms/shmem
  2026-09-30 23:17 [PATCH] mm: Make swapoff interruptible when unusing mms/shmem Chris Down
  2026-09-30 23:46 ` Andrew Morton
@ 2026-10-01 12:50 ` Vineeth Remanan Pillai
  1 sibling, 0 replies; 3+ messages in thread
From: Vineeth Remanan Pillai @ 2026-10-01 12:50 UTC (permalink / raw)
  To: Chris Down
  Cc: Andrew Morton, Hugh Dickins, Baolin Wang, Chris Li, Kairui Song,
	Kemeng Shi, Nhat Pham, Baoquan He, Barry Song, Youngjun Park,
	Ying Huang, Kelley Nielsen, linux-mm, linux-kernel, kernel-team

On Wed, Sep 30, 2026 at 7:17 PM Chris Down <chris@chrisdown.name> wrote:
>
> try_to_unuse() only checks for a pending signal between mms, and
> shmem_unuse() doesn't check at all. That means that once swapoff gets to
> a process or a shmem file with a lot swapped out, nothing can interrupt
> it until every last page of it has been read back in.
>
> Just as one example of where this can concretely show up, freezing tasks
> for suspend or hibernation has to wait for swapoff to notice the
> freezer's fake signal, and gives up after freeze_timeout_msecs (20
> seconds by default).
>
> Here's a facetious example where one swaps out 2GiB of one process to a
> swap file on ext4, starts swapoff, and half a second later tries to
> freeze with pm_test=freezer. Writing to /sys/power/state then fails with
> EBUSY and this in dmesg:
>
>     Freezing user space processes failed after 20.003 seconds (1 tasks refusing to freeze, wq_busy=0):
>     task:swapoff         state:D stack:0     pid:3175  tgid:3175  ppid:2955   task_flags:0x400100 flags:0x00000419
>     Call trace:
>      [...]
>      io_schedule+0x44/0x70
>      folio_wait_bit_common+0x1ec/0x3d0
>      __folio_lock+0x24/0x40
>      unuse_pte_range+0x2d0/0x348
>      unuse_vma+0x158/0x248
>      unuse_mm+0xfc/0x150
>      try_to_unuse+0x104/0x3f8
>      __do_sys_swapoff+0x220/0x5d8
>      [...]
>
> The same goes for anything else that wants swapoff to stop, like an
> admin hitting ^C in a panic, of course.
>
> Prior to commit b56a2d8af914 ("mm: rid swapoff of quadratic complexity")
> try_to_unuse() was driven by find_next_to_unuse() which checks for a
> signal before every entry, so let's restore that behaviour.
>
Nice find, thanks for tracking this down and fixing it. The fix looks
good to me.

Acked-by: Vineeth Pillai (Google) <vineeth@bitbyteword.org>

Thanks,
Vineeth

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-10-01 12:50 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 23:17 [PATCH] mm: Make swapoff interruptible when unusing mms/shmem Chris Down
2026-09-30 23:46 ` Andrew Morton
2026-10-01 12:50 ` Vineeth Remanan Pillai

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®