* [PATCH v3] mm/memfd: fix hugetlb reservation accounting in error paths
@ 2026-09-27 9:47 Hongfu Li
2026-10-07 22:35 ` Andrew Morton
0 siblings, 1 reply; 2+ messages in thread
From: Hongfu Li @ 2026-09-27 9:47 UTC (permalink / raw)
To: hughd, baolin.wang, akpm, vivek.kasireddy
Cc: muchun.song, osalvador, david, linux-mm, linux-kernel, Hongfu Li, stable
From: Hongfu Li <lihongfu@kylinos.cn>
If hugetlb_add_to_page_cache() in memfd_alloc_folio() fails with -EEXIST,
a concurrent fault has already instantiated the folio in the page cache,
and the reservation now belongs to that folio. Calling
hugetlb_unreserve_pages() in that case incorrectly removes the region
backing the cached folio and releases one reservation more than it
should, leaving that folio in the page cache with no region recording it.
That leaves resv_huge_pages one page short until that folio is removed
from the page cache, so available_huge_pages() reports a page that is
not actually free and the pool can grant one reservation more than it
can back. Applications using hugetlb memfds can then fail to allocate
or fault in a page they already reserved. They see -ENOMEM or -ENOSPC
from the allocation or fault path even though the reservation and
HugePages_Free still look healthy.
So hold the hugetlb fault mutex from hugetlb_reserve_pages() until the
error-path unreserve completes to make the reserve, allocate and
instantiate steps atomic against concurrent faults. With the mutex held
from the start, a concurrent fault can no longer consume the reservation
between reserve and allocate/instantiate. If a fault completed before the
mutex was taken, it has already added the region for that index, so
hugetlb_reserve_pages() returns 0 and the error path leaves the region in
place.
Fixes: 717cf9357325 ("mm/memfd: reserve hugetlb folios before allocation")
Cc: stable@vger.kernel.org
Signed-off-by: Hongfu Li <lihongfu@kylinos.cn>
---
v3:
- Update commit message to describe the accounting damage and its
user-visible effect; no code changes.
- Add Cc: stable@vger.kernel.org.
v2:
- Take the hugetlb fault mutex before hugetlb_reserve_pages() and hold
it until the error-path unreserve completes.
- Update commit message
- Link to v1: https://lore.kernel.org/all/20260831090631.29227-1-hongfu.li@linux.dev/
---
mm/memfd.c | 31 ++++++++++++++++---------------
1 file changed, 16 insertions(+), 15 deletions(-)
diff --git a/mm/memfd.c b/mm/memfd.c
index c708d92533f4..0f6fff004f5e 100644
--- a/mm/memfd.c
+++ b/mm/memfd.c
@@ -82,22 +82,31 @@ struct folio *memfd_alloc_folio(struct file *memfd, pgoff_t idx)
struct hstate *h = hstate_file(memfd);
int err = -ENOMEM;
long nr_resv;
+ u32 hash;
gfp_mask = htlb_alloc_mask(h);
gfp_mask &= ~(__GFP_HIGHMEM | __GFP_MOVABLE);
idx >>= huge_page_order(h);
+ /*
+ * Serialize hugepage allocation and instantiation to prevent
+ * races with concurrent allocations, as required by all other
+ * callers of hugetlb_add_to_page_cache().
+ */
+ hash = hugetlb_fault_mutex_hash(memfd->f_mapping, idx);
+ mutex_lock(&hugetlb_fault_mutex_table[hash]);
+
nr_resv = hugetlb_reserve_pages(inode, idx, idx + 1, NULL, EMPTY_VMA_FLAGS);
- if (nr_resv < 0)
- return ERR_PTR(nr_resv);
+ if (nr_resv < 0) {
+ err = nr_resv;
+ goto out_unlock;
+ }
folio = alloc_hugetlb_folio_reserve(h,
numa_node_id(),
NULL,
gfp_mask);
if (folio) {
- u32 hash;
-
/*
* Zero the folio to prevent information leaks to userspace.
* Use folio_zero_user() which is optimized for huge/gigantic
@@ -112,20 +121,9 @@ struct folio *memfd_alloc_folio(struct file *memfd, pgoff_t idx)
*/
__folio_mark_uptodate(folio);
- /*
- * Serialize hugepage allocation and instantiation to prevent
- * races with concurrent allocations, as required by all other
- * callers of hugetlb_add_to_page_cache().
- */
- hash = hugetlb_fault_mutex_hash(memfd->f_mapping, idx);
- mutex_lock(&hugetlb_fault_mutex_table[hash]);
-
err = hugetlb_add_to_page_cache(folio,
memfd->f_mapping,
idx);
-
- mutex_unlock(&hugetlb_fault_mutex_table[hash]);
-
if (err) {
folio_put(folio);
goto err_unresv;
@@ -133,11 +131,14 @@ struct folio *memfd_alloc_folio(struct file *memfd, pgoff_t idx)
hugetlb_set_folio_subpool(folio, subpool_inode(inode));
folio_unlock(folio);
+ mutex_unlock(&hugetlb_fault_mutex_table[hash]);
return folio;
}
err_unresv:
if (nr_resv > 0)
hugetlb_unreserve_pages(inode, idx, idx + 1, 0);
+out_unlock:
+ mutex_unlock(&hugetlb_fault_mutex_table[hash]);
return ERR_PTR(err);
}
#endif
--
2.54.0
^ permalink raw reply [flat|nested] 2+ messages in thread
* Re: [PATCH v3] mm/memfd: fix hugetlb reservation accounting in error paths
2026-09-27 9:47 [PATCH v3] mm/memfd: fix hugetlb reservation accounting in error paths Hongfu Li
@ 2026-10-07 22:35 ` Andrew Morton
0 siblings, 0 replies; 2+ messages in thread
From: Andrew Morton @ 2026-10-07 22:35 UTC (permalink / raw)
To: Hongfu Li
Cc: hughd, baolin.wang, vivek.kasireddy, muchun.song, osalvador,
david, linux-mm, linux-kernel, Hongfu Li, stable
On Sun, 27 Sep 2026 17:47:57 +0800 Hongfu Li <hongfu.li@linux.dev> wrote:
> From: Hongfu Li <lihongfu@kylinos.cn>
>
> If hugetlb_add_to_page_cache() in memfd_alloc_folio() fails with -EEXIST,
> a concurrent fault has already instantiated the folio in the page cache,
> and the reservation now belongs to that folio. Calling
> hugetlb_unreserve_pages() in that case incorrectly removes the region
> backing the cached folio and releases one reservation more than it
> should, leaving that folio in the page cache with no region recording it.
>
> That leaves resv_huge_pages one page short until that folio is removed
> from the page cache, so available_huge_pages() reports a page that is
> not actually free and the pool can grant one reservation more than it
> can back. Applications using hugetlb memfds can then fail to allocate
> or fault in a page they already reserved. They see -ENOMEM or -ENOSPC
> from the allocation or fault path even though the reservation and
> HugePages_Free still look healthy.
>
> So hold the hugetlb fault mutex from hugetlb_reserve_pages() until the
> error-path unreserve completes to make the reserve, allocate and
> instantiate steps atomic against concurrent faults. With the mutex held
> from the start, a concurrent fault can no longer consume the reservation
> between reserve and allocate/instantiate. If a fault completed before the
> mutex was taken, it has already added the region for that index, so
> hugetlb_reserve_pages() returns 0 and the error path leaves the region in
> place.
>
> Fixes: 717cf9357325 ("mm/memfd: reserve hugetlb folios before allocation")
> Cc: stable@vger.kernel.org
> Signed-off-by: Hongfu Li <lihongfu@kylinos.cn>
Can we please have review of this cc:stable regression fix?
> v3:
> - Update commit message to describe the accounting damage and its
> user-visible effect; no code changes.
> - Add Cc: stable@vger.kernel.org.
> v2:
> - Take the hugetlb fault mutex before hugetlb_reserve_pages() and hold
> it until the error-path unreserve completes.
> - Update commit message
> - Link to v1: https://lore.kernel.org/all/20260831090631.29227-1-hongfu.li@linux.dev/
> ---
> mm/memfd.c | 31 ++++++++++++++++---------------
> 1 file changed, 16 insertions(+), 15 deletions(-)
>
> diff --git a/mm/memfd.c b/mm/memfd.c
> index c708d92533f4..0f6fff004f5e 100644
> --- a/mm/memfd.c
> +++ b/mm/memfd.c
> @@ -82,22 +82,31 @@ struct folio *memfd_alloc_folio(struct file *memfd, pgoff_t idx)
> struct hstate *h = hstate_file(memfd);
> int err = -ENOMEM;
> long nr_resv;
> + u32 hash;
>
> gfp_mask = htlb_alloc_mask(h);
> gfp_mask &= ~(__GFP_HIGHMEM | __GFP_MOVABLE);
> idx >>= huge_page_order(h);
>
> + /*
> + * Serialize hugepage allocation and instantiation to prevent
> + * races with concurrent allocations, as required by all other
> + * callers of hugetlb_add_to_page_cache().
> + */
> + hash = hugetlb_fault_mutex_hash(memfd->f_mapping, idx);
> + mutex_lock(&hugetlb_fault_mutex_table[hash]);
> +
> nr_resv = hugetlb_reserve_pages(inode, idx, idx + 1, NULL, EMPTY_VMA_FLAGS);
> - if (nr_resv < 0)
> - return ERR_PTR(nr_resv);
> + if (nr_resv < 0) {
> + err = nr_resv;
> + goto out_unlock;
> + }
>
> folio = alloc_hugetlb_folio_reserve(h,
> numa_node_id(),
> NULL,
> gfp_mask);
> if (folio) {
> - u32 hash;
> -
> /*
> * Zero the folio to prevent information leaks to userspace.
> * Use folio_zero_user() which is optimized for huge/gigantic
> @@ -112,20 +121,9 @@ struct folio *memfd_alloc_folio(struct file *memfd, pgoff_t idx)
> */
> __folio_mark_uptodate(folio);
>
> - /*
> - * Serialize hugepage allocation and instantiation to prevent
> - * races with concurrent allocations, as required by all other
> - * callers of hugetlb_add_to_page_cache().
> - */
> - hash = hugetlb_fault_mutex_hash(memfd->f_mapping, idx);
> - mutex_lock(&hugetlb_fault_mutex_table[hash]);
> -
> err = hugetlb_add_to_page_cache(folio,
> memfd->f_mapping,
> idx);
> -
> - mutex_unlock(&hugetlb_fault_mutex_table[hash]);
> -
> if (err) {
> folio_put(folio);
> goto err_unresv;
> @@ -133,11 +131,14 @@ struct folio *memfd_alloc_folio(struct file *memfd, pgoff_t idx)
>
> hugetlb_set_folio_subpool(folio, subpool_inode(inode));
> folio_unlock(folio);
> + mutex_unlock(&hugetlb_fault_mutex_table[hash]);
> return folio;
> }
> err_unresv:
> if (nr_resv > 0)
> hugetlb_unreserve_pages(inode, idx, idx + 1, 0);
> +out_unlock:
> + mutex_unlock(&hugetlb_fault_mutex_table[hash]);
> return ERR_PTR(err);
> }
> #endif
> --
> 2.54.0
>
^ permalink raw reply [flat|nested] 2+ messages in thread
end of thread, other threads:[~2026-10-07 22:35 UTC | newest]
Thread overview: 2+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-27 9:47 [PATCH v3] mm/memfd: fix hugetlb reservation accounting in error paths Hongfu Li
2026-10-07 22:35 ` Andrew Morton
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®