From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f73.google.com (mail-wm1-f73.google.com [209.85.128.73]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 84E933859D4 for ; Fri, 12 Jun 2026 14:15:52 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.73 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781273757; cv=none; b=s5XLCa1wwWYtzDFXE6qG8HxxF9gF3wCqe4AvnVjgEppjfD+ynRFIMDdc+Qbz0fAodOUQFtGltP2TNwMH4JP6+1tCq045shOrDH36xCkfDBKGJVOOYXWsLhcHrRKMaOW51HXh8Q8pFfHq9Fhmt/9B6ogWlOTgFmvF76tT9bZgDFE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1781273757; c=relaxed/simple; bh=LxOR7X7j8q0QanoB2nAHMNRngg+bbK59z6izKhQjVyY=; h=Date:Mime-Version:Message-ID:Subject:From:To:Cc:Content-Type; b=rAMPAYScorO47D9pmats/ITw4XkrNukAKYFCp9c41uI/NslMrzbiM0CBrnSLvCWDAgnd1cCzEaPKTdXjaaP8gMiXnT5L6w0nWDruQZCuwrXAjlNJ3AVKjRNe4ki2fcaPrmkcukAFo8CfLAFeTqG96PjsbCKh/W0yZXThb8ldd70= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jackmanb.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=RQzICvBW; arc=none smtp.client-ip=209.85.128.73 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jackmanb.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="RQzICvBW" Received: by mail-wm1-f73.google.com with SMTP id 5b1f17b1804b1-490b37e1f48so7496755e9.0 for ; Fri, 12 Jun 2026 07:15:52 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1781273751; x=1781878551; darn=vger.kernel.org; h=cc:to:from:subject:message-id:mime-version:date:from:to:cc:subject :date:message-id:reply-to; bh=3aqjAT0fRHYdS66UFRNsM8EkSHwMJb62tgtmp8RP6ZI=; b=RQzICvBW/3Mb6oC13iSEb2oh403OtoXisQdrrc1z9fPSDVDLlIfOn1gjt5A38RCtPU tvui9J0TbGorszrTI7yezhTNfAOM0OsPDxhiu9ZUqgfEyEFJTtHtnDsFDmorNHuLr0YS qYUlE+4i1iWas5KoO5186xdjJK9iCVQ6UR4yKrA8qWoNgHPDv9DY51LVyARPD7mHwP+w Ds9unSgTqQ99/9qlBo9APtxAif/rywOAdX9ET5K0YNhL4Xgy5Gz8dug9uHAvJSw3hqRy LTF+CeFwaZdRdS36G9ezrlwan9pEeDYpIq4JYIXI3woeQYn0BcvQbtddOCjo5pl1Dsy+ JXHA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1781273751; x=1781878551; h=cc:to:from:subject:message-id:mime-version:date:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=3aqjAT0fRHYdS66UFRNsM8EkSHwMJb62tgtmp8RP6ZI=; b=Cm0O5NhDq/DtTquMnMJo6z7qUzbgldJItXBm/l0vj58Lg6x9E33W1jBZ+TQ3PxMJQ+ Lc9mJlaSfLAAzBEb1U86/18m/zvSBRj1yPBbeyikudBEZ+5MggB0GfTqv5Q50V7bxDpj sG11+bsNqqKxvwoLoAgvVwmFnzHDkopo7bTe1RpHkClwgC8xlDK7QcP0lsZQH3H1MRop vNOd0i9FUwQ1hloGq3YWa+Rco/uQoNdzQ6Jio+aj/RZxROsK47gDyeUwHS1Pk9k6gCI0 D1jHV8767n97rFr9JBjFL5Pz+eFWCc3j1sPvFHvsIrhl1EUfXQEo5KFpTtHyPE/2QBWN jGFA== X-Forwarded-Encrypted: i=1; AFNElJ9y5bBMwDLURwEajF4i2GJzboC70rAqP1ZRVJF4HSy4afschFfeJ9aglS8xwO1eWSqjvKPsVqCE7mLVL4Q=@vger.kernel.org X-Gm-Message-State: AOJu0YzFCig6vXnsXLtnpgsGIENJ/P76J4/GWsS4ZAtJLQf4YqUey45L RwZ6MfeV/ofXnIqwCmZGEjqA4imXl90W+hVDnIb0D8YRFW1dUOpKaZSkq8BN5ceMkWe1fZGesps x89XvW7rNqy9kmw== X-Received: from wmpe36.prod.google.com ([2002:a05:600c:4ba4:b0:490:1b6b:9818]) (user=jackmanb job=prod-delivery.src-stubby-dispatcher) by 2002:a05:600c:3586:b0:490:c024:2ec8 with SMTP id 5b1f17b1804b1-490ec339ec0mr48700225e9.0.1781273749232; Fri, 12 Jun 2026 07:15:49 -0700 (PDT) Date: Fri, 12 Jun 2026 14:15:44 +0000 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 X-B4-Tracking: v=1; b=H4sIAJAULGoC/x2MQQqAMAzAviI9O3BFRfyKeFi1ag/qWEUE2d8tH gNJXlBOwgp98ULiW1TOw8CXBUxbOFZ2MhsDVthWrUe3LtFFVpVdNFxmO8KmCzVRwzSDdTHxIs/ /HMacP0vTMpljAAAA X-Change-Id: 20260612-gfp-pessimisation-b258a4bb5ebd X-Mailer: b4 0.14.3 Message-ID: <20260612-gfp-pessimisation-v1-1-936eb04202e7@google.com> Subject: [PATCH] mm/page_alloc: drop flag-conversion "optimisation" From: Brendan Jackman To: Andrew Morton , Vlastimil Babka , Suren Baghdasaryan , Michal Hocko , Johannes Weiner , Zi Yan Cc: "Harry Yoo (Oracle)" , Gregory Price , linux-mm@kvack.org, linux-kernel@vger.kernel.org, Brendan Jackman Content-Type: text/plain; charset="utf-8" This code uses flag equivalences to try to optimise conversion from GFP_ to ALLOC_ but there's no clear reason to believe it makes things faster. Even if it gets rid of conditional branches, it just trades them for a data dependency. CPUs are pretty good at conditional branches. But, in my GCC x86 build it doesn't look like there are any branches anyway, the compiler found some conditional instruction tricks. (Caveat: This was extracted & annotated by Gemini AI, I did not actually read the disasm myself) Old code: ae50: 8b 04 24 mov (%rsp),%eax # Load gfp_mask ... ae5d: 41 89 c4 mov %eax,%r12d ae64: 41 81 e4 20 08 00 00 and $0x820,%r12d # Mask both flags at once ... ae6f: 44 89 e1 mov %r12d,%ecx ae77: 83 c9 40 or $0x40,%ecx # OR with ALLOC_CPUSET (0x40) ae7a: 89 4c 24 60 mov %ecx,0x60(%rsp) # Store to alloc_flags New code: For __GFP_HIGH ( 0x20 ): It uses the Carry Flag (via sbb ) to conditionally add 0x20 to the base 0x40 ( ALLOC_CPUSET ) flag: ae63: 83 e0 20 and $0x20,%eax # Test __GFP_HIGH ... ae6a: 83 f8 01 cmp $0x1,%eax # Set carry flag if 0 ae6f: 45 19 e4 sbb %r12d,%r12d # %r12d = (gfp & 0x20) ? 0 : -1 ae80: 41 83 e4 e0 and $0xffffffe0,%r12d # %r12d = (gfp & 0x20) ? 0 : -32 ae87: 41 83 c4 60 add $0x60,%r12d # %r12d = (gfp & 0x20) ? 0x60 : 0x40 For __GFP_KSWAPD_RECLAIM ( 0x800 ): It uses a conditional move ( cmov ) later in the function to set the ALLOC_KSWAPD ( 0x800 ) bit: ae72: 25 00 08 00 00 and $0x800,%eax # Test __GFP_KSWAPD_RECLAIM ae77: 89 44 24 30 mov %eax,0x30(%rsp) # Store result ... af2c: 80 cf 08 or $0x8,%bh # Set ALLOC_KSWAPD (0x800) in temp reg af2f: 45 85 c9 test %r9d,%r9d # Check if __GFP_KSWAPD_RECLAIM was set af32: 0f 44 d8 cmove %eax,%ebx # If not, revert to flags without it Testing with a modified version[0] of lib/free_pages_test.c (adding printks with timing)... [0] https://github.com/bjackman/aethelred/blob/2ccdc84ef087c2a631914f58e106e99e19bd3b98/page-alloc-test/page-alloc-test.c Old results from a Sapphire Rapids consumer CPU: [ 67.157118] page_alloc_test: Testing with GFP_KERNEL [ 67.157122] page_alloc_test: Starting 1,000,000 allocations... [ 70.704446] page_alloc_test: Completed. Time: 3543002 us (Avg: 3543.00 ns per alloc+free loop) [ 70.704456] page_alloc_test: Testing with GFP_KERNEL | __GFP_COMP [ 70.704460] page_alloc_test: Starting 1,000,000 allocations... [ 70.944672] page_alloc_test: Completed. Time: 239980 us (Avg: 239.98 ns per alloc+free loop) [ 70.944675] page_alloc_test: Test completed New results: [ 70.079015] page_alloc_test: Testing with GFP_KERNEL [ 70.079020] page_alloc_test: Starting 1,000,000 allocations... [ 73.669396] page_alloc_test: Completed. Time: 3586954 us (Avg: 3586.95 ns per alloc+free loop) [ 73.669402] page_alloc_test: Testing with GFP_KERNEL | __GFP_COMP [ 73.669405] page_alloc_test: Starting 1,000,000 allocations... [ 73.905084] page_alloc_test: Completed. Time: 235496 us (Avg: 235.49 ns per alloc+free loop) [ 73.905086] page_alloc_test: Test completed Seems like a wash. So, drop the flag value coupling here and let the compiler and CPU do their job. Superscalar CPUs are pretty neat after all. (Used AI for the disasm but the rest is all manual). Signed-off-by: Brendan Jackman --- mm/page_alloc.c | 14 ++++---------- 1 file changed, 4 insertions(+), 10 deletions(-) diff --git a/mm/page_alloc.c b/mm/page_alloc.c index ee902a468c2f5..9e1949ea13a6d 100644 --- a/mm/page_alloc.c +++ b/mm/page_alloc.c @@ -4478,22 +4478,16 @@ gfp_to_alloc_flags(gfp_t gfp_mask, unsigned int order) { unsigned int alloc_flags = ALLOC_WMARK_MIN | ALLOC_CPUSET; - /* - * __GFP_HIGH is assumed to be the same as ALLOC_MIN_RESERVE - * and __GFP_KSWAPD_RECLAIM is assumed to be the same as ALLOC_KSWAPD - * to save two branches. - */ - BUILD_BUG_ON(__GFP_HIGH != (__force gfp_t) ALLOC_MIN_RESERVE); - BUILD_BUG_ON(__GFP_KSWAPD_RECLAIM != (__force gfp_t) ALLOC_KSWAPD); - /* * The caller may dip into page reserves a bit more if the caller * cannot run direct reclaim, or if the caller has realtime scheduling * policy or is asking for __GFP_HIGH memory. GFP_ATOMIC requests will * set both ALLOC_NON_BLOCK and ALLOC_MIN_RESERVE(__GFP_HIGH). */ - alloc_flags |= (__force int) - (gfp_mask & (__GFP_HIGH | __GFP_KSWAPD_RECLAIM)); + if (gfp_mask & __GFP_HIGH) + alloc_flags |= ALLOC_MIN_RESERVE; + if (gfp_mask & __GFP_KSWAPD_RECLAIM) + alloc_flags |= ALLOC_KSWAPD; if (!(gfp_mask & __GFP_DIRECT_RECLAIM)) { /* --- base-commit: ca2351ac6da277a470d4fcf122b53267e02b2716 change-id: 20260612-gfp-pessimisation-b258a4bb5ebd Best regards, -- Brendan Jackman