mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [BUG/RFC] mm/madvise: MADV_WILLNEED skips swapped file-backed COW pages
@ 2026-08-31 18:24 Mike Kaplinskiy
  2026-09-03  7:05 ` zhaozhengzhuo
  2026-09-03 19:03 ` Lorenzo Stoakes (ARM)
  0 siblings, 2 replies; 3+ messages in thread
From: Mike Kaplinskiy @ 2026-08-31 18:24 UTC (permalink / raw)
  To: linux-mm; +Cc: linux-kernel, akpm, liam, ljs, david, vbabka, jannh

[-- Attachment #1: Type: text/plain, Size: 874 bytes --]

Hi,

I'm seeing a strange behavior when using MADV_WILLNEED to schedule
swap page-in. It seems MADV_WILLNEED does not schedule swap reads for
swapped-out COW pages in an ordinary file-backed MAP_PRIVATE mapping.

I reproduced this on Linux 7.0.14 on aarch64, but I think this code
hasn't changed in a while. mm/madvise.c:madvise_willneed walks swap
PTEs only when vma->vm_file is NULL, and sends other file-backed
mappings to vfs_fadvise(POSIX_FADV_WILLNEED). This skips the case of
MAP_PRIVATE mappings with changes, which frequently happens for
libraries/binaries with relocations and/or writable globals.

The code is at the top of
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/madvise.c#n282
.

Minimal reproducer attached. Would there be interest in changing this
path to prefetch private copies in addition to the readahead?

Thanks,
Mike

[-- Attachment #2: madvise-willneed-cow-repro.c --]
[-- Type: text/x-csrc, Size: 1766 bytes --]

#define _GNU_SOURCE
#include <assert.h>
#include <fcntl.h>
#include <stdio.h>
#include <stdlib.h>
#include <string.h>
#include <sys/mman.h>
#include <unistd.h>

#define SIZE (64UL << 20)

static size_t swap_kib(void *mapping)
{
	FILE *f = fopen("/proc/self/smaps", "r");
	char line[256], perms[5];
	unsigned long start, end, value = 0;
	int ours = 0;

	assert(f);
	while (fgets(line, sizeof(line), f)) {
		if (sscanf(line, "%lx-%lx %4s", &start, &end, perms) == 3)
			ours = start == (unsigned long)mapping;
		else if (ours && sscanf(line, "Swap: %lu kB", &value) == 1)
			break;
	}
	fclose(f);
	return value;
}

static unsigned long pswpin(void)
{
	FILE *f = fopen("/proc/vmstat", "r");
	char name[32];
	unsigned long value = 0;

	assert(f);
	while (fscanf(f, "%31s %lu", name, &value) == 2)
		if (!strcmp(name, "pswpin"))
			break;
	fclose(f);
	return value;
}

int main(void)
{
	char path[] = "/tmp/madvise-cow-XXXXXX";
	int fd = mkstemp(path);
	char *p;
	unsigned long before;
	volatile unsigned long sum = 0;

	assert(fd >= 0 && !ftruncate(fd, SIZE));
	p = mmap(NULL, SIZE, PROT_READ | PROT_WRITE, MAP_PRIVATE, fd, 0);
	assert(p != MAP_FAILED);
	unlink(path);

	for (size_t i = 0; i < SIZE; i += getpagesize())
		p[i] = i / getpagesize() + 1; /* create private COW pages */
	assert(!madvise(p, SIZE, MADV_PAGEOUT));
	usleep(500000);
	printf("after PAGEOUT:  Swap=%zu kB\n", swap_kib(p));

	before = pswpin();
	assert(!madvise(p, SIZE, MADV_WILLNEED));
	sleep(2); /* WILLNEED is asynchronous */
	printf("after WILLNEED: Swap=%zu kB, pswpin=+%lu\n",
	       swap_kib(p), pswpin() - before);

	before = pswpin();
	for (size_t i = 0; i < SIZE; i += getpagesize())
		sum += p[i];
	printf("after touch:    pswpin=+%lu (sum=%lu)\n",
	       pswpin() - before, sum);
}

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [BUG/RFC] mm/madvise: MADV_WILLNEED skips swapped file-backed COW pages
  2026-08-31 18:24 [BUG/RFC] mm/madvise: MADV_WILLNEED skips swapped file-backed COW pages Mike Kaplinskiy
@ 2026-09-03  7:05 ` zhaozhengzhuo
  2026-09-03 19:03 ` Lorenzo Stoakes (ARM)
  1 sibling, 0 replies; 3+ messages in thread
From: zhaozhengzhuo @ 2026-09-03  7:05 UTC (permalink / raw)
  To: Mike Kaplinskiy
  Cc: linux-mm, linux-kernel, Andrew Morton, Liam R . Howlett,
	Lorenzo Stoakes, David Hildenbrand, Vlastimil Babka, Jann Horn

Hi Mike,

Thanks for the report. I reproduced this on x86-64 with the current
mm-unstable tree at:

  e3b5239afe1b ("mm: gup: cleanup the gup_fast_*() call chain")

The test used a 64 MiB ordinary file-backed MAP_PRIVATE mapping and
disk-backed swap. Since the test machine has 128 GiB of RAM, the pages
initially remained in swap cache after MADV_PAGEOUT. I therefore added
controlled memory pressure between MADV_PAGEOUT and MADV_WILLNEED so
that the swap-cache folios were actually reclaimed.

Across three runs, the pswpin deltas (in 4 KiB pages) were:

                        run 1     run 2     run 3
  after MADV_WILLNEED       2         0         0
  during subsequent touch 16327     15811     15676

The mapping still reported about 64 MiB in Swap after MADV_WILLNEED,
and almost all swap reads occurred only when the pages were touched.

This appears to confirm that madvise_willneed() skips the private swap
entries because the VMA still has vm_file and therefore takes the file
readahead path. vfs_fadvise(POSIX_FADV_WILLNEED) can prefetch the
underlying file contents, but not the modified private contents stored
in swap.

Are you already working on a fix? If testing or other help would be
useful, I would be happy to help.

Thanks,
zhaozhengzhuo

^ permalink raw reply	[flat|nested] 3+ messages in thread

* Re: [BUG/RFC] mm/madvise: MADV_WILLNEED skips swapped file-backed COW pages
  2026-08-31 18:24 [BUG/RFC] mm/madvise: MADV_WILLNEED skips swapped file-backed COW pages Mike Kaplinskiy
  2026-09-03  7:05 ` zhaozhengzhuo
@ 2026-09-03 19:03 ` Lorenzo Stoakes (ARM)
  1 sibling, 0 replies; 3+ messages in thread
From: Lorenzo Stoakes (ARM) @ 2026-09-03 19:03 UTC (permalink / raw)
  To: Mike Kaplinskiy; +Cc: linux-mm, linux-kernel, akpm, liam, david, vbabka, jannh

On Mon, Aug 31, 2026 at 11:24:48AM -0700, Mike Kaplinskiy wrote:
> Hi,
>
> I'm seeing a strange behavior when using MADV_WILLNEED to schedule
> swap page-in. It seems MADV_WILLNEED does not schedule swap reads for
> swapped-out COW pages in an ordinary file-backed MAP_PRIVATE mapping.
>
> I reproduced this on Linux 7.0.14 on aarch64, but I think this code
> hasn't changed in a while. mm/madvise.c:madvise_willneed walks swap

It's more a known limitation of MADV_WILLNEED that has existed forever.

The issue is that every single MAP_PRIVATE-file backed mapping would then
need to be walked twice, once via page tables -> swap cache and once via
readahead.

But you could figure out if it was CoW'd...

Something like:

diff --git a/mm/madvise.c b/mm/madvise.c
index bc6a7dc73021..33ff3ba69c7f 100644
--- a/mm/madvise.c
+++ b/mm/madvise.c
@@ -297,10 +297,12 @@ static long madvise_willneed(struct madvise_behavior *madv_behavior)
        loff_t offset;

 #ifdef CONFIG_SWAP
-       if (!file) {
+       if (!file || (vma_is_cow_mapping(vma) && vma->anon_vma))
                walk_page_range_vma(vma, start, end, &swapin_walk_ops, vma);
                lru_add_drain(); /* Push any new pages onto the LRU now */
-               return 0;
+               if (!file)
+                       return 0;
        }

        if (shmem_mapping(file->f_mapping)) {

(This also happens to fix a bug there with MAP_PRIVATE-/dev/zero though
that'll get fixed with my upcoming series anyway :)

The vma->anon_vma check ensures CoW'd pages have actually been mapped in.

But then you'd have to do two walks for every single CoW'd MAP_PRIVATE-file
backed mapping.

The majority of the anon walk would be a no-op also.

> PTEs only when vma->vm_file is NULL, and sends other file-backed
> mappings to vfs_fadvise(POSIX_FADV_WILLNEED). This skips the case of
> MAP_PRIVATE mappings with changes, which frequently happens for
> libraries/binaries with relocations and/or writable globals.
>
> The code is at the top of
> https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/linux.git/tree/mm/madvise.c#n282

That code is skipping non-swap entries for a swapped out shmem folio? I
think that's irrelevant to this.

> .
>
> Minimal reproducer attached. Would there be interest in changing this
> path to prefetch private copies in addition to the readahead?

I mean I'd like to hear from others, if OK I can send the patch above.

Are we concerned about the inefficiency of this?

Or dropping the mmap lock right after and faulting in the file-backed bits?

I guess if you're doing an MADV_WILLNEED you are fine with it taking a bit
of extra time to swap stuff in.

>
> Thanks,
> Mike

--
Cheers, Lorenzo

^ permalink raw reply	[flat|nested] 3+ messages in thread

end of thread, other threads:[~2026-09-03 19:03 UTC | newest]

Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-08-31 18:24 [BUG/RFC] mm/madvise: MADV_WILLNEED skips swapped file-backed COW pages Mike Kaplinskiy
2026-09-03  7:05 ` zhaozhengzhuo
2026-09-03 19:03 ` Lorenzo Stoakes (ARM)

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®