From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-dl2-f41.google.com (mail-dl2-f41.google.com [74.125.229.169]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C3D7952BE47 for ; Thu, 1 Oct 2026 15:34:20 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.229.169 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790868862; cv=none; b=Y6peJYklT50HdCnSteInqjit/EmoukNJzEMHwJhsQOQMSXC/6WE6FgQw+4vrXJAykLDaFge/UQOx4kXJcj0OJeogPyfBj0Ekkwl9s87tM5A0cMu/lIkH5TFe23Jo7nTEz1W027YL4lx30b/mB4AU0D1/13bUGGEfMHVXgg45+ms= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790868862; c=relaxed/simple; bh=zfoGhZ9tkatQAPTvMhsHChhPWCw6Z8eEJLu/axoyvVE=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=GBVgSCS8FHQH6TZd2cz2pT87a2gNIBJJ4JO1JXl4IMTpJDbdYrMnjfhk4bzhphuGrQMX6/jbfwkts0y61jiILxjd1paUxKoxihckAtdK2+dho5sET1JyEyLqRiT8SBddR/d6uyEjor+wS0bLn1m5uyNf98gaFjkJbgwIOeROl1w= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=K4HyrvoH; arc=none smtp.client-ip=74.125.229.169 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="K4HyrvoH" Received: by mail-dl2-f41.google.com with SMTP id a92af1059eb24-1474c6b7742so5025121c88.1 for ; Thu, 01 Oct 2026 08:34:20 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1790868859; x=1791473659; darn=vger.kernel.org; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:from:to:cc:subject:date :message-id:reply-to:content-type; bh=UYLCjPpLfBewvYdjBDkNxTVbyP/bcB0jfIhL2wBACGM=; b=K4HyrvoHiwX8tLPo+4vior3S0EMYyxp/gpIhJSbLb1l3Lf6kYSSmvLmignKZdHTDDo 1+ac/DDls+UQO0aG4SHQ0VwLizJMVnGWgiCVJwAS1hd1MYRPwLaSx3oTUJtCHiKugQuc 1j1cnEGOS+Tq/mPmzAiPdaRw7AdX96UlrdUo6Q0GmtP5FdCgHYbhOgozOktoyY9D9/80 wdt8IjxsO9qtTzrFYRM//HDZnLWl6ozOad3mPps3lAOrP99OseKFDGRdxE+cpUDkMVTO zgdz/Ma3IDq1UE5mocAi/ODqNcQgP7oV7v8kbozPWLs5ofsq0TQ9by06qbysfG5kf4dw tj1g== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790868859; x=1791473659; h=cc:to:in-reply-to:references:message-id:content-transfer-encoding :content-type:mime-version:subject:date:from:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=UYLCjPpLfBewvYdjBDkNxTVbyP/bcB0jfIhL2wBACGM=; b=zzJT45Bly0GorvkFmO56TrSGb8KDnAP+wGtEC0d8mp5Y32B4CAiQoOm+IYTCVVMrZd Z81wQRy8IrqGj/Oi9Y5+m6HtNj8OHXfKT/U2pufZmKhdheE4OBFuyOl3eytY8EArNHcL obQs4gnhi3IrqFRqP2vgbDfpEvto8BsnAnu6exTxWvmd0t1+6bXBzwMTe/NkJ78nP5LH 5PDdlkhKh7NhkBcYKRr8+eKWacE+l7MMX0l2Br8fTrx7tJiGjr4Mm0XSYaOOVCEzcHrR 3tZ/OhMncd7umzOvmjGG5zQBDxh3s9A9YK72JbkqIdch7rn0lSZY7hB1Zj8F1zTHsXVJ QP1Q== X-Forwarded-Encrypted: i=1; AKwUvByubiTaKH21IwEXtgqk9rISL8DtA16Ugg7uvSjx1Gif2pvSGaQTKNZWA/i//mdxqB7/gTgqcFkYRfxXVhg=@vger.kernel.org X-Gm-Message-State: AFuF++nLUnp29zQ3AoDvx92FmWnHET82uy3x1N9gbIfbKrnelRfElC2y JZDlwitUGQ7K3dpVM4uH3YXf0R+dkK5Q9lV4t6a4Ea0jrJXyDs5u8HbeTY8nQK2M X-Gm-Gg: AYBFou1sBg+XJ+RPY8/NIF1uJqFqOsGEHQohpaW1FBii64cdE8kktwTqYGaXZIG+gm6 tmjcmkAIB+b2ro/qlWyi/AkWtKlplCTCDlQJhg0NkpDR/VkLANH7qGdOg0VqtEqGKQ9DWSvmvIE 7KENvhpY8QLJyJRRIc20CXJn50t7bSwJhEiZZxZvOZ/f7rv1/O/C8DGfWJVcYlOS8yidHLucu0O O7CWGl7pOZcgX5wT9rWQGYSIqNNum0btU34SRRSlSsjHrVR/dMkjDN0Gfdz6dgqCZQegcjE9mkK ec0ZbX1r+Y2nYgky3mGhPYiEEpG65acAYTBD+lMaBojO7s0ZMDN6/3F7NMc+qpd/bUnY1A/dt5Z RqX3fheY0xypSbzvXZWuSZe+iNOq4QnYsMqMiWZEKREtYeARu0/Nm1oyS3bmqVXrR2ozDjXzYW6 UhzKhnpAFC+rlvf5Pst3OxtJHGxFrisCFDtsjntddVG74S3sJsW/6pventyLOYj3t2Nw== X-Received: by 2002:a05:7022:e1b:b0:14e:906f:9c8e with SMTP id a92af1059eb24-14e906f9dc4mr2454634c88.20.1790868859216; Thu, 01 Oct 2026 08:34:19 -0700 (PDT) Received: from [127.0.1.1] ([23.254.208.9]) by smtp.gmail.com with ESMTPSA id a92af1059eb24-14f266d1156sm146683c88.6.2026.10.01.08.34.09 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 01 Oct 2026 08:34:18 -0700 (PDT) From: Qiliang Yuan Date: Thu, 01 Oct 2026 23:33:49 +0800 Subject: [PATCH 2/2] mm/compaction: defer failed async direct compaction Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20261001-bug-mm-thp-async-compact-defer-v1-2-0174c7923430@gmail.com> References: <20261001-bug-mm-thp-async-compact-defer-v1-0-0174c7923430@gmail.com> In-Reply-To: <20261001-bug-mm-thp-async-compact-defer-v1-0-0174c7923430@gmail.com> To: Andrew Morton , Kairui Song , Qi Zheng , Shakeel Butt , Barry Song , Axel Rasmussen , Yuanchu Xie , Wei Xu , Baoquan He , Baolin Wang , David Hildenbrand , Lorenzo Stoakes , "Liam R. Howlett" , Vlastimil Babka , Mike Rapoport , Suren Baghdasaryan , Michal Hocko , Steven Rostedt , Masami Hiramatsu , Mathieu Desnoyers , Brendan Jackman , Johannes Weiner , Zi Yan Cc: linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-trace-kernel@vger.kernel.org, Qiliang Yuan X-Mailer: b4 0.13.0 A THP fault on a MADV_HUGEPAGE VMA first tries the local node only, with __GFP_THISNODE | __GFP_NORETRY, and runs a single round of async direct compaction there before falling back to other nodes. A failed async run is never deferred. When the local node has plenty of free memory but no free pageblock, and what sits between the free pages can't be migrated, such as memory long-term pinned for RDMA, each of these faults scans the zone again and fails. Populating a large buffer pays for one failed compaction per 2M fault. A KV-cache store that registers hundreds of GiB for RDMA reports registration growing from tens of seconds to tens of minutes, and works around it with MPOL_INTERLEAVE. Defer the async state when an async run fails, and check it before further async runs. Sync compaction keeps its own state and behaves as before, and a successful compaction or allocation resets both. Keep resetting the pageblock skip hints only when sync compaction restarts, as async relies on them. Populating a 4 GiB MADV_HUGEPAGE buffer from node 0 of a two-node VM, with node 0 (24 GiB) fragmented by long-term pinning every other page, median of 3 runs: before after direct compactions 735 15 time in compaction 122.7 ms 3.6 ms population time 380 ms 273 ms THPs on node 0 1313 1126 The THPs that no longer land on node 0 come from node 1. With the same fragmentation but movable memory, compaction still succeeds and the median run still gets all 2048 THPs on node 0, as before. Signed-off-by: Qiliang Yuan --- mm/compaction.c | 21 ++++++++++++++------- 1 file changed, 14 insertions(+), 7 deletions(-) diff --git a/mm/compaction.c b/mm/compaction.c index 7f8845d1990aa..25f758bc54f06 100644 --- a/mm/compaction.c +++ b/mm/compaction.c @@ -2600,7 +2600,9 @@ compact_zone(struct compact_control *cc, struct capture_control *capc) /* * Clear pageblock skip if there were failures recently and compaction - * is about to be retried after being deferred. + * is about to be retried after being deferred. Only do it when sync + * compaction restarts: async compaction relies on the skip hints, and + * clearing them on every async retry would rescan the whole zone. */ if (compaction_restarting(cc->zone, cc->order)) __reset_isolation_suitable(cc->zone); @@ -2858,8 +2860,10 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order, !__cpuset_zone_allowed(zone, gfp_mask)) continue; - if (prio > MIN_COMPACT_PRIORITY - && compaction_deferred(zone, order, true)) { + if (prio > MIN_COMPACT_PRIORITY && + (compaction_deferred(zone, order, true) || + (prio == COMPACT_PRIO_ASYNC && + compaction_deferred(zone, order, false)))) { rc = max_t(enum compact_result, COMPACT_DEFERRED, rc); continue; } @@ -2890,14 +2894,17 @@ enum compact_result try_to_compact_pages(gfp_t gfp_mask, unsigned int order, break; } - if (prio != COMPACT_PRIO_ASYNC && (status == COMPACT_COMPLETE || - status == COMPACT_PARTIAL_SKIPPED)) + if (status == COMPACT_COMPLETE || + status == COMPACT_PARTIAL_SKIPPED) /* * We think that allocation won't succeed in this zone * so we defer compaction there. If it ends up - * succeeding after all, it will be reset. + * succeeding after all, it will be reset. A failed + * async run only defers further async runs, as sync + * compaction may succeed on pageblocks it skipped. */ - defer_compaction(zone, order, true); + defer_compaction(zone, order, + prio != COMPACT_PRIO_ASYNC); /* * We might have stopped compacting due to need_resched() in -- 2.43.0