From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from out-170.mta0.migadu.com (out-170.mta0.migadu.com [91.218.175.170]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 00CD43ED5A8 for ; Fri, 6 Mar 2026 16:15:13 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=91.218.175.170 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772813716; cv=none; b=jpeCviOg2+DrdAvyueM3X6MtjqW+mvkIA581qdknXbOxv/INPmDxb4PIZnIAxT1GDoiDlJm/Avv4MWcna5HG9JGJEOUel5Rc58TmVybnI63sIhtStpp2oqRS8L9kD8y/6oE9f8klDH0E3Wad/eHW6Kuvohz8ua59y5isj9bhb+w= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772813716; c=relaxed/simple; bh=tPjB3ZpJKCyjG+Mgm0U/0oxejuTnlqvTG/wOcUGER2g=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=D/P1ApvHqBfTUq9NU04BxXpccGZ9n9cWIsp8h/Ww3+F47Li0tMWQmZVn+PDlTMq319JYPZCmG9FBRC3VZLdiNzFuhMtkUSDeU4EmS4zg4iTrc+19JSRtrhEDoTChsmUhU/wmXA7VE4f42Bcy6bukk46vGSowM4FuX8Sl2a8LQYM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev; spf=pass smtp.mailfrom=linux.dev; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b=ilyx6wsK; arc=none smtp.client-ip=91.218.175.170 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=linux.dev Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=linux.dev Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux.dev header.i=@linux.dev header.b="ilyx6wsK" Message-ID: <28e48b47-f215-4e4a-b55a-01dbf293ff35@linux.dev> DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux.dev; s=key1; t=1772813711; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=F46BSWUVzTSNMVtPbFqNmwug0wRx85MgEhXeZ2WefSg=; b=ilyx6wsK0ezghYuXNj+LGWBzh5gjqHMVDVuo4AxTgmeFbqG88SfGgWNXWl8GSj02nrFkHD /KanIQP0HSHOvrQZ6kfuzHwAXhRtPy7RUg/q+Z2U5jYPfaW2W0c1aCMX+3OU1I2k6eGseM TPmwPA6wGXmBBvt+RIi1t9yozapqjyU= Date: Fri, 6 Mar 2026 19:15:07 +0300 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Subject: Re: [PATCH] mm: migrate: requeue destination folio on deferred split queue Content-Language: en-GB To: Zi Yan Cc: "David Hildenbrand (Arm)" , Andrew Morton , npache@redhat.com, linux-mm@kvack.org, matthew.brost@intel.com, joshua.hahnjy@gmail.com, hannes@cmpxchg.org, rakie.kim@sk.com, byungchul@sk.com, gourry@gourry.net, ying.huang@linux.alibaba.com, apopple@nvidia.com, linux-kernel@vger.kernel.org, kernel-team@meta.com References: <20260306133556.2051251-1-usama.arif@linux.dev> <64051a59-680f-40ae-b291-b884aeb7c77b@linux.dev> <993B37FF-29E1-41F5-A1E8-F38B9CD24478@nvidia.com> X-Report-Abuse: Please report any abuse attempt to abuse@migadu.com and include these headers. From: Usama Arif In-Reply-To: <993B37FF-29E1-41F5-A1E8-F38B9CD24478@nvidia.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit X-Migadu-Flow: FLOW_OUT On 06/03/2026 14:46, Zi Yan wrote: > On 6 Mar 2026, at 9:12, Usama Arif wrote: > >> On 06/03/2026 13:49, David Hildenbrand (Arm) wrote: >>> On 3/6/26 14:35, Usama Arif wrote: >>>> During folio migration, __folio_migrate_mapping() removes the source >>>> folio from the deferred split queue, but the destination folio is never >>>> re-queued. This causes underutilized THPs to escape the shrinker after >>>> NUMA migration, since they silently drop off the deferred split list. >>>> >>>> Fix this by calling deferred_split_folio() on the destination folio >>>> after a successful migration, for large rmappable folios. >>>> >>>> Reported-by: Johannes Weiner >>>> Fixes: dafff3f4c850 ("mm: split underused THPs") >>>> Signed-off-by: Usama Arif >>>> --- >>>> mm/migrate.c | 11 +++++++++++ >>>> 1 file changed, 11 insertions(+) >>>> >>>> diff --git a/mm/migrate.c b/mm/migrate.c >>>> index ece77ccb2ec0..98d0a594f7b7 100644 >>>> --- a/mm/migrate.c >>>> +++ b/mm/migrate.c >>>> @@ -1393,6 +1393,17 @@ static int migrate_folio_move(free_folio_t put_new_folio, unsigned long private, >>>> if (old_page_state & PAGE_WAS_MAPPED) >>>> remove_migration_ptes(src, dst, 0); >>>> >>>> + /* >>>> + * Requeue the destination folio on the deferred split queue if >>>> + * the source was a large folio that was on the queue. Without >>>> + * this, NUMA migration causes underutilized THPs to escape >>>> + * the shrinker since the source is unqueued in >>>> + * __folio_migrate_mapping() and the destination is never >>>> + * re-queued. >>>> + */ >>>> + if (folio_test_large(dst) && folio_test_large_rmappable(dst)) >>>> + deferred_split_folio(dst, false); >>> >>> Doesn't that mean that you will readd any large folios, even if already >>> previously taken off the list after scanning? >>> >>> So I am not sure if your "if the source was a large folio that was on >>> the queue." comment is accurate? >>> >> >> Yes you are right. How about something like below? We also won't need to check >> for anon and non-device folios with this as we only set the the flag if it was >> already on deferred_split list. > > BTW, migrate_pages() tries to split partially mapped folios before migration[1], > so what remains in the deferred_list would be: > > 1. partially mapped but with a pin, > 2. fully mapped but potentially underused. > Yes, thats right. > I wonder if you want to do an underused scan before migration and try to split > underused THPs. hmm, I think we should keep THPs as is if there is no memory pressure (proactive or otherwise). Scanning THPs for zeros has a cost and we would also lose the benefit of THPs when we dont need memory. > Or to avoid this additional scan, find a way of detecting > zero pages at page copy time and split it after migration. > Yeah but I think we lose the benefits of THPs after migration when we dont need additional memory? > Anyway, it seems that all large folios are in this deferred_list. Maybe, like > David suggested in his LSFMM proposal, we should scan large folios on LRU lists > at reclaim time instead, since there is not much difference between deferred_list > and LRU lists right now. > Yeah the THP shrinker is a very basic implementation and there are a lot of > > [1] https://elixir.bootlin.com/linux/v6.19.3/source/mm/migrate.c#L1840 > Also Johannes pointed out its not great storing this information in page flags, we can just keep it as local variable. This is what the patch would look like: diff --git a/mm/migrate.c b/mm/migrate.c index ece77ccb2ec0..48a972f158ab 100644 --- a/mm/migrate.c +++ b/mm/migrate.c @@ -1360,6 +1360,7 @@ static int migrate_folio_move(free_folio_t put_new_folio, unsigned long private, int rc; int old_page_state = 0; struct anon_vma *anon_vma = NULL; + bool src_deferred_split = false; struct list_head *prev; __migrate_folio_extract(dst, &old_page_state, &anon_vma); @@ -1373,6 +1374,10 @@ static int migrate_folio_move(free_folio_t put_new_folio, unsigned long private, goto out_unlock_both; } + if (folio_test_large(src) && folio_test_large_rmappable(src) && + !data_race(list_empty(&src->_deferred_list))) + src_deferred_split = true; + rc = move_to_new_folio(dst, src, mode); if (rc) goto out; @@ -1393,6 +1398,15 @@ static int migrate_folio_move(free_folio_t put_new_folio, unsigned long private, if (old_page_state & PAGE_WAS_MAPPED) remove_migration_ptes(src, dst, 0); + /* + * Requeue the destination folio on the deferred split queue if + * the source was on the queue. The source is unqueued in + * __folio_migrate_mapping(), so we recorded the state from + * before move_to_new_folio(). + */ + if (src_deferred_split) + deferred_split_folio(dst, false); + out_unlock_both: folio_unlock(dst); folio_set_owner_migrate_reason(dst, reason);