* [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio
@ 2026-10-07 4:10 Kyle Zeng
2026-10-07 10:21 ` David Hildenbrand (Arm)
2026-10-07 11:14 ` Zi Yan
0 siblings, 2 replies; 3+ messages in thread
From: Kyle Zeng @ 2026-10-07 4:10 UTC (permalink / raw)
To: linux-mm
Cc: linux-kernel, Andrew Morton, David Hildenbrand, Zi Yan,
Baolin Wang, outbounddisclosures, Kyle Zeng, stable
collapse_file() can fail its reference-count or dirty-folio check after
unmapping with TTU_BATCH_FLUSH. Both paths put back the isolated folio,
then unlock it and drop the lookup reference before reaching the common
try_to_unmap_flush().
The page-cache reference does not keep the folio stable once the lock is
released. Another collapse can replace and free it. With its PTEs
already gone, that collapse cannot flush the first task's per-task TLB
batch, and retract_page_tables() skips short or unaligned VMAs. A CPU
can therefore retain a user translation to the freed folio. This has
been reproduced with unprivileged MADV_COLLAPSE on a memfd.
Flush at out_unlock while the lookup reference and folio lock are still
held. The common flush continues to cover the accumulated pagelist on
both success and rollback, preserving batching on successful collapses.
Fixes: 6d9df8a5889c ("mm/thp: collapse_file() do try_to_unmap(TTU_BATCH_FLUSH)")
Cc: stable@vger.kernel.org
Assisted-by: LLM
Signed-off-by: Kyle Zeng <kylebot@openai.com>
---
Changes in v2:
- Use Assisted-by: LLM.
- Explain that the common flush is a no-op after out_unlock flushes.
mm/khugepaged.c | 10 +++++++---
1 file changed, 7 insertions(+), 3 deletions(-)
diff --git a/mm/khugepaged.c b/mm/khugepaged.c
index 75639298efc2..e1a5890818ad 100644
--- a/mm/khugepaged.c
+++ b/mm/khugepaged.c
@@ -2478,6 +2478,11 @@ static enum scan_result collapse_file(struct mm_struct *mm, unsigned long addr,
index += folio_nr_pages(folio);
continue;
out_unlock:
+ /*
+ * The folio may have been unmapped with TTU_BATCH_FLUSH.
+ * Flush before releasing the lock and our last reference.
+ */
+ try_to_unmap_flush();
folio_unlock(folio);
folio_put(folio);
goto xa_unlocked;
@@ -2488,9 +2493,8 @@ static enum scan_result collapse_file(struct mm_struct *mm, unsigned long addr,
xa_unlocked:
/*
- * If collapse is successful, flush must be done now before copying.
- * If collapse is unsuccessful, does flush actually need to be done?
- * Do it anyway, to clear the state.
+ * Flush before copying the folios, or releasing them in rollback.
+ * This is a no-op if out_unlock already flushed the batch.
*/
try_to_unmap_flush();
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio
2026-10-07 4:10 [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio Kyle Zeng
@ 2026-10-07 10:21 ` David Hildenbrand (Arm)
2026-10-07 11:14 ` Zi Yan
1 sibling, 0 replies; 3+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-07 10:21 UTC (permalink / raw)
To: Kyle Zeng, linux-mm
Cc: linux-kernel, Andrew Morton, Zi Yan, Baolin Wang,
outbounddisclosures, stable
On 10/7/26 06:10, Kyle Zeng wrote:
> collapse_file() can fail its reference-count or dirty-folio check after
> unmapping with TTU_BATCH_FLUSH. Both paths put back the isolated folio,
> then unlock it and drop the lookup reference before reaching the common
> try_to_unmap_flush().
>
> The page-cache reference does not keep the folio stable once the lock is
> released. Another collapse can replace and free it. With its PTEs
> already gone, that collapse cannot flush the first task's per-task TLB
> batch, and retract_page_tables() skips short or unaligned VMAs. A CPU
> can therefore retain a user translation to the freed folio. This has
> been reproduced with unprivileged MADV_COLLAPSE on a memfd.
>
> Flush at out_unlock while the lookup reference and folio lock are still
> held. The common flush continues to cover the accumulated pagelist on
> both success and rollback, preserving batching on successful collapses.
>
> Fixes: 6d9df8a5889c ("mm/thp: collapse_file() do try_to_unmap(TTU_BATCH_FLUSH)")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Kyle Zeng <kylebot@openai.com>
> ---
LGTM
Acked-by: David Hildenbrand (Arm) <david@kernel.org>
--
Cheers,
David
^ permalink raw reply [flat|nested] 3+ messages in thread
* Re: [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio
2026-10-07 4:10 [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio Kyle Zeng
2026-10-07 10:21 ` David Hildenbrand (Arm)
@ 2026-10-07 11:14 ` Zi Yan
1 sibling, 0 replies; 3+ messages in thread
From: Zi Yan @ 2026-10-07 11:14 UTC (permalink / raw)
To: Kyle Zeng, linux-mm
Cc: linux-kernel, Andrew Morton, David Hildenbrand, Baolin Wang,
outbounddisclosures, stable
On Wed Oct 7, 2026 at 12:10 AM EDT, Kyle Zeng wrote:
> collapse_file() can fail its reference-count or dirty-folio check after
> unmapping with TTU_BATCH_FLUSH. Both paths put back the isolated folio,
> then unlock it and drop the lookup reference before reaching the common
> try_to_unmap_flush().
>
> The page-cache reference does not keep the folio stable once the lock is
> released. Another collapse can replace and free it. With its PTEs
> already gone, that collapse cannot flush the first task's per-task TLB
> batch, and retract_page_tables() skips short or unaligned VMAs. A CPU
> can therefore retain a user translation to the freed folio. This has
> been reproduced with unprivileged MADV_COLLAPSE on a memfd.
>
> Flush at out_unlock while the lookup reference and folio lock are still
> held. The common flush continues to cover the accumulated pagelist on
> both success and rollback, preserving batching on successful collapses.
>
> Fixes: 6d9df8a5889c ("mm/thp: collapse_file() do try_to_unmap(TTU_BATCH_FLUSH)")
> Cc: stable@vger.kernel.org
> Assisted-by: LLM
> Signed-off-by: Kyle Zeng <kylebot@openai.com>
> ---
> Changes in v2:
> - Use Assisted-by: LLM.
> - Explain that the common flush is a no-op after out_unlock flushes.
>
> mm/khugepaged.c | 10 +++++++---
> 1 file changed, 7 insertions(+), 3 deletions(-)
>
LGTM. Thanks.
Reviewed-by: Zi Yan <ziy@nvidia.com>
--
Best Regards,
Yan, Zi
^ permalink raw reply [flat|nested] 3+ messages in thread
end of thread, other threads:[~2026-10-07 11:15 UTC | newest]
Thread overview: 3+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-07 4:10 [PATCH v2] mm/khugepaged: flush deferred unmaps before dropping a failed folio Kyle Zeng
2026-10-07 10:21 ` David Hildenbrand (Arm)
2026-10-07 11:14 ` Zi Yan
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®