From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from m16.mail.126.com (m16.mail.126.com [117.135.210.8]) by smtp.subspace.kernel.org (Postfix) with ESMTP id 5557A179A7 for ; Tue, 14 Jan 2025 02:52:19 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=117.135.210.8 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736823143; cv=none; b=Zm5hff+bA1GFH9SeN6SS/XD9bzSLZIzWvcj6vdIe2lS9P0ea+CBESh1quDEEz8TGLwUO0sLPwA1KzBEcxo6w0qqlvE0Zcdr8VLddbU4Ge4mGtFhE4zyUfuLCBaPjWEktYfO+6467KvrtQJsvfdQKXDC8ya5t+F5/bFuO1skv/zE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736823143; c=relaxed/simple; bh=ScyXe22Qp4FkQgDDS1XvJqsulhP/Zhi0MuM97wh+Jhk=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=fUx6EwzAAEN+4jliXyP9IrN+EuelkBLiYNtbt+3ja0vzodR4X4c1F7aUvu+RiyuhmLqv7ijuTBdU++m36UKOHSheiRvMcxDx+Wdrzi+xhA2S3uqYq8aXDL381ivTRdx+BXaDqyx9jSFDMCBxPBbhkuHI2x8YuqDKYjoMtPOiDck= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=126.com; spf=pass smtp.mailfrom=126.com; dkim=pass (1024-bit key) header.d=126.com header.i=@126.com header.b=Btco0AOp; arc=none smtp.client-ip=117.135.210.8 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=126.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=126.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=126.com header.i=@126.com header.b="Btco0AOp" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=126.com; s=s110527; h=Message-ID:Date:MIME-Version:Subject:From: Content-Type; bh=C/cPjmlA+t6jlHSj6/8PTaHoULJ6MSmV+P7CNNzIowc=; b=Btco0AOpUSq6HvzSeoup2jqATNDsc8A3440ohN3BjTbro6FSaU8if3E6sdVaKR cD6qBc0TBdmin/ITDkuYhvi0N8c1wYqWQseyhuJGFvUqGFC7WT4vE6tnU5B7RbK9 aHcFS7rcokuDgEavx4jvFg/q0LulG6ulC4SVBYx8gbiKA= Received: from [172.19.20.199] (unknown []) by gzga-smtp-mtada-g0-3 (Coremail) with SMTP id _____wDnf3dI0YVnlvOLBA--.522S2; Tue, 14 Jan 2025 10:51:53 +0800 (CST) Message-ID: Date: Tue, 14 Jan 2025 10:51:52 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH V3] mm: compaction: skip memory compaction when there are not enough migratable pages To: Johannes Weiner Cc: akpm@linux-foundation.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org, 21cnbao@gmail.com, david@redhat.com, baolin.wang@linux.alibaba.com, liuzixing@hygon.cn, Vlastimil Babka References: <1736335854-548-1-git-send-email-yangge1116@126.com> <20250113154657.GA829144@cmpxchg.org> From: Ge Yang In-Reply-To: <20250113154657.GA829144@cmpxchg.org> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit X-CM-TRANSID:_____wDnf3dI0YVnlvOLBA--.522S2 X-Coremail-Antispam: 1Uf129KBjvJXoWxKF1ktFWkGFy7GF1xtrWDJwb_yoW7Ar43pF yxGFn8Kr4DXFZFkr1Iq3Wv9F92y3yftFW8JFySyryfu3ZI9FySya1DtFyDCF1DZr1jqw4Y qFWq9wnruws8Za7anT9S1TB71UUUUU7qnTZGkaVYY2UrUUUUjbIjqfuFe4nvWSU5nxnvy2 9KBjDUYxBIdaVFxhVjvjDU0xZFpf9x07jbHUDUUUUU= X-CM-SenderInfo: 51dqwwjhrrila6rslhhfrp/1tbiOhLUG2eFxZQF1QABsF 在 2025/1/13 23:46, Johannes Weiner 写道: > CC Vlastimil > > On Wed, Jan 08, 2025 at 07:30:54PM +0800, yangge1116@126.com wrote: >> From: yangge >> >> There are 4 NUMA nodes on my machine, and each NUMA node has 32GB >> of memory. I have configured 16GB of CMA memory on each NUMA node, >> and starting a 32GB virtual machine with device passthrough is >> extremely slow, taking almost an hour. >> >> During the start-up of the virtual machine, it will call >> pin_user_pages_remote(..., FOLL_LONGTERM, ...) to allocate memory. >> Long term GUP cannot allocate memory from CMA area, so a maximum of >> 16 GB of no-CMA memory on a NUMA node can be used as virtual machine >> memory. There is 16GB of free CMA memory on a NUMA node, which is >> sufficient to pass the order-0 watermark check, causing the >> __compaction_suitable() function to consistently return true. >> However, if there aren't enough migratable pages available, performing >> memory compaction is also meaningless. Besides checking whether >> the order-0 watermark is met, __compaction_suitable() also needs >> to determine whether there are sufficient migratable pages available >> for memory compaction. >> >> For costly allocations, because __compaction_suitable() always >> returns true, __alloc_pages_slowpath() can't exit at the appropriate >> place, resulting in excessively long virtual machine startup times. >> Call trace: >> __alloc_pages_slowpath >> if (compact_result == COMPACT_SKIPPED || >> compact_result == COMPACT_DEFERRED) >> goto nopage; // should exit __alloc_pages_slowpath() from here >> >> When the 16G of non-CMA memory on a single node is exhausted, we will >> fallback to allocating memory on other nodes. In order to quickly >> fallback to remote nodes, we should skip memory compaction when >> migratable pages are insufficient. After this fix, it only takes a >> few tens of seconds to start a 32GB virtual machine with device >> passthrough functionality. >> >> Signed-off-by: yangge >> --- >> >> V3: >> - fix build error >> >> V2: >> - consider unevictable folios >> >> mm/compaction.c | 20 ++++++++++++++++++++ >> 1 file changed, 20 insertions(+) >> >> diff --git a/mm/compaction.c b/mm/compaction.c >> index 07bd227..a9f1261 100644 >> --- a/mm/compaction.c >> +++ b/mm/compaction.c >> @@ -2383,7 +2383,27 @@ static bool __compaction_suitable(struct zone *zone, int order, >> int highest_zoneidx, >> unsigned long wmark_target) >> { >> + pg_data_t __maybe_unused *pgdat = zone->zone_pgdat; >> + unsigned long sum, nr_pinned; >> unsigned long watermark; >> + >> + sum = node_page_state(pgdat, NR_INACTIVE_FILE) + >> + node_page_state(pgdat, NR_INACTIVE_ANON) + >> + node_page_state(pgdat, NR_ACTIVE_FILE) + >> + node_page_state(pgdat, NR_ACTIVE_ANON) + >> + node_page_state(pgdat, NR_UNEVICTABLE); > > What about PAGE_MAPPING_MOVABLE pages that aren't on this list? For > example, zsmalloc backend pages can be a large share of allocated > memory, and they are compactable. You would give up on compaction > prematurely and cause unnecessary allocation failures. > Yes, indeed, there are pages that are not in the LRU list but support migration. Currently, technologies such as balloon, z3fold, and zsmalloc are utilizing such pages. I feel that we could add an item to node_stat_item to keep statistics on these pages. > That scenario is way more common than the one you're trying to fix. > > I think trying to make this list complete, and maintaining it, is > painstaking and error prone. And errors are hard to detect: they will > just manifest as spurious failures in higher order requests that you'd > need to catch with tracing enabled in the right moments. > > So I'm not a fan of this approach. > > Compaction is already skipped when previous runs were not successful. > See defer_compaction() and compaction_deferred(). Why is this not > helping here? if (prio != COMPACT_PRIO_ASYNC && (status == COMPACT_COMPLETE || status == COMPACT_PARTIAL_SKIPPED)) defer_compaction(zone, order); When prio != COMPACT_PRIO_ASYNC, defer_compaction(zone, order) will be executed. In the __alloc_page_slowpath() function, during the first execution of __alloc_pages_direct_compact(), prio is equal to COMPACT_PRIO_ASYNC, and therefore defer_compaction(zone, order) will not be executed. Instead, it will eventually proceed to the time-consuming __alloc_pages_direct_reclaim(). This can be avoided in scenarios where memory compaction is not suitable. > >> + nr_pinned = node_page_state(pgdat, NR_FOLL_PIN_ACQUIRED) - >> + node_page_state(pgdat, NR_FOLL_PIN_RELEASED); > > Likewise, as Barry notes, not all pinned pages are necessarily LRU > pages. remap_vmalloc_range() pages come to mind. You can't do subset > math on potentially disjunct sets. Indeed, some problem scenarios are unsolvable currently, but there are some scenarios that can be resolved through this approach. Currently, we haven't come up with a better solution yet. > >> + /* >> + * Gup-pinned pages are non-migratable. After subtracting these pages, >> + * we need to check if the remaining pages are sufficient for memory >> + * compaction. >> + */ >> + if ((sum - nr_pinned) < (1 << order)) >> + return false; >> +