From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pj1-f46.google.com (mail-pj1-f46.google.com [209.85.216.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 8AC4C18E75A for ; Wed, 29 Jan 2025 17:53:18 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.216.46 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1738173200; cv=none; b=SrjwlqCFKd5mClOriez9GaPG+up5mHTdK0gXDnhb9nVFQmcTMBdspepwb/5BF2jpFAUEVK0YPGPk1tgR3EPpqf3WhLGfbbJvKvefnJeQaE3Y2XesdgBrR7aCOuDiblscgb7iAHbZ43fOeEDR7HtXXnfAPTezN2ndtpiABbTmC3g= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1738173200; c=relaxed/simple; bh=PoTwTghGg0+4nalsAX2QGDktONbaK1ZHMqdVpgG2K6M=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=D2Gj01XWIfos8vC+MRyQ1TZKNzjn7nroeMxssKyy7UUFrJzTmVmzkegaZIX9vjgKwb1yqvD3MJ2/D95olvWvFp0MNzaREFU0BmlwggCpsFlGUvxfw3zxRiVu60w9ak7wbScY17U+Npy4oex6i92CVsuSWAU1q+rIRrD4fim3jpk= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com; spf=pass smtp.mailfrom=bytedance.com; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b=UnxD5EET; arc=none smtp.client-ip=209.85.216.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=bytedance.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=bytedance.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=bytedance.com header.i=@bytedance.com header.b="UnxD5EET" Received: by mail-pj1-f46.google.com with SMTP id 98e67ed59e1d1-2ef748105deso9631846a91.1 for ; Wed, 29 Jan 2025 09:53:18 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=bytedance.com; s=google; t=1738173198; x=1738777998; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:from:references:cc:to :content-language:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=KdtDciE7wia/Lnf64s9wqrIFVKAtIQ+OPsXndNgHzhU=; b=UnxD5EETstNZP6k1/L7xMZcub4zqxL0JaI0KHD9qBEOklEVnXUjgjt2lkL2payBr93 D0n28nyKnYz6viLl7kLm9qWa2gmAGaBQB4edNyH0Oe54Kz14VWcqZMKWMVfnBOOre8Ha OOFq+vDAXH2h9B+S+qpxoKrkWZdA8YgRZdczDDq9/oj63IgkpKCSmU2WKR5VKlWXb3HO jQk/Tj7LVp6iYJUBXTJxceIaWaB0r8VuH14xEZ2R3k+jRZk1Xr4MPTw/KNtDIvCd67id tigtdveLsrL6kkys3Kik2XrQBL2z2Pij88yhTbfEjzboDzbHXiC1DfLUCJfhJvq0NOe1 fY8Q== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1738173198; x=1738777998; h=content-transfer-encoding:in-reply-to:from:references:cc:to :content-language:subject:user-agent:mime-version:date:message-id :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=KdtDciE7wia/Lnf64s9wqrIFVKAtIQ+OPsXndNgHzhU=; b=A6qovqVkzkTXZnS6NI0asHYqZJaSO1/aY4WXR9k4unB8Apo42JqQaIkdYEShR/Bf4S fk/x678mLsiTpFe77CU5I7G3+uoTHMm1dIouXI0KuwqW8xA9zwth148SWOKO700fqIlh WeGwRvYVxGnncga4SbVQX3OoZ19AA1c/wsEqkM+3hRuNpIj1a9Lbp/Likv9SsPN4meIv Vew4CgMn8woziW7h/h9FELHKbQVHZeB4fJcoW+URc5pD4ZnfoQI/Rled9bZ4a7Kg+RV5 QKeR+bzEuuJEL415yiHfxCQkwXloq2sUfb3oQXggZRjbPWUgjJLy+gRk7Fkyopxo33hE OuqA== X-Forwarded-Encrypted: i=1; AJvYcCVjtGvqaxMrTzEmbyP78Z9lLTMZqGZs1ZV0QJwY/YKrTNwjMkQACfeX0q63fj28cgy6SEBDNqivhTvoD/I=@vger.kernel.org X-Gm-Message-State: AOJu0Yw1umfRwgvCKpGdzybJ2Yatr/nA9nHIXP2rswRoEDTsJCw0BnlC xEyzPzedHnF99nYxO8Lm8eynkvQFMhdjgMk5HD5rvgf1Kqke0hkO3XQNhJdyluo= X-Gm-Gg: ASbGncvIo9TKX+tXVkD5O6usZX3y7a4+VDxFcSk3Vc6Bll8LbIUaQjOb3QEtIM+z/zG cKZE6/sAtpH0Rypg8apHnB0hmEgNTRAOosnuFKObrn+LyGdDKphIrd5L3SFgzfFuMebVykpAJh/ 6x+6eDtrQ7qY56uDNsQSukz4Am0PNGlO8tywQ+9qVxT4LEh0aLYb6Cs2P9tM264a912eaQdIQcy TLOHdU8WigDoV+KCMPT57Sks/vPFKnWURgMxe9wCHSch0WihbOFMyU0JUi7hS5F0fxUpA84/7wO Bria2fw5achbQ+FtGfGCyPheNqOVgR1vJM7O7naBsP2xLF1k3uumNc39ZbyGyKthJcqEfVJCjxE 8YAet4bW+ZNcBt9A= X-Google-Smtp-Source: AGHT+IGoJvasbXxXpHQCiqhGeXCLcysAaQlK/QCdxvK3cdUTVFD8UHXCU5mbBYZY56hWe0Z2Pnb2og== X-Received: by 2002:a17:90b:2e8b:b0:2ee:9d57:243 with SMTP id 98e67ed59e1d1-2f83abbd864mr5497816a91.1.1738173197782; Wed, 29 Jan 2025 09:53:17 -0800 (PST) Received: from ?IPV6:240e:370:8b10:5140:d8c5:7248:cf0a:d072? ([240e:370:8b10:5140:d8c5:7248:cf0a:d072]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-2f83bc9823dsm2058859a91.5.2025.01.29.09.53.10 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Wed, 29 Jan 2025 09:53:17 -0800 (PST) Message-ID: <92fe6c2b-ef8d-4ed4-bf96-27f0760329b7@bytedance.com> Date: Thu, 30 Jan 2025 01:53:06 +0800 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [linus:master] [x86] 4817f70c25: stress-ng.mmapaddr.ops_per_sec 63.0% regression Content-Language: en-US To: Rik van Riel Cc: "Paul E. McKenney" , Matthew Wilcox , Peter Zijlstra , David Hildenbrand , kernel test robot , oe-lkp@lists.linux.dev, lkp@intel.com, linux-kernel@vger.kernel.org, Andrew Morton , Dave Hansen , Andy Lutomirski , Catalin Marinas , David Rientjes , Hugh Dickins , Jann Horn , Lorenzo Stoakes , Mel Gorman , Muchun Song , Peter Xu , Will Deacon , Zach O'Keefe , Dan Carpenter , Frederic Weisbecker , Neeraj Upadhyay References: <20250128113100.GB7145@noisy.programming.kicks-ass.net> <2c599f1f-f63b-4504-a026-46bf73954124@redhat.com> <20250128132847.GB505@noisy.programming.kicks-ass.net> <18172465-37e3-48a9-9d67-0d44fdcb93bd@redhat.com> <8b661116-0a85-4928-91ed-3c01ebbf8d39@bytedance.com> <2212111cad3180948cf388a7e5c8689df0fdda08.camel@surriel.com> <14159fb4-0c59-4653-9265-73f415e70063@bytedance.com> <20250129105920.7a4bffa1@fangorn> <488401565cac6b5f9e3232d0ca481055876c919b.camel@surriel.com> <05da0ae9-073e-4578-b65c-f837c28eead8@paulmck-laptop> <20250129115320.1334ad5f@fangorn> <96f55f36-5cb7-4226-a8ea-374a7891ecff@bytedance.com> From: Qi Zheng In-Reply-To: <96f55f36-5cb7-4226-a8ea-374a7891ecff@bytedance.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 8bit On 2025/1/30 01:33, Qi Zheng wrote: > > > On 2025/1/30 00:53, Rik van Riel wrote: >> On Wed, 29 Jan 2025 08:36:12 -0800 >> "Paul E. McKenney" wrote: >>> On Wed, Jan 29, 2025 at 11:14:29AM -0500, Rik van Riel wrote: >> >>>> Paul, does this look like it could do the trick, >>>> or do we need something else to make RCU freeing >>>> happy again? >>> >>> I don't claim to fully understand the issue, but this would prevent >>> any RCU grace periods starting subsequently from completing.  It would >>> not prevent RCU callbacks from being invoked for RCU grace periods that >>> started earlier. >>> >>> So it won't prevent RCU callbacks from being invoked. >> >> That makes things clear! I guess we need a different approach. >> >> Qi, does the patch below resolve the regression for you? >> >> ---8<--- >> >>  From 5de4fa686fca15678a7e0a186852f921166854a3 Mon Sep 17 00:00:00 2001 >> From: Rik van Riel >> Date: Wed, 29 Jan 2025 10:51:51 -0500 >> Subject: [PATCH 2/2] mm,rcu: prevent RCU callbacks from running with >> pcp lock >>   held >> >> Enabling MMU_GATHER_RCU_TABLE_FREE can create contention on the >> zone->lock.  This turns out to be because in some configurations >> RCU callbacks are called when IRQs are re-enabled inside >> rmqueue_bulk, while the CPU is still holding the per-cpu pages lock. >> >> That results in the RCU callbacks being unable to grab the >> PCP lock, and taking the slow path with the zone->lock for >> each item freed. >> >> Speed things up by blocking RCU callbacks while holding the >> PCP lock. >> >> Signed-off-by: Rik van Riel >> Suggested-by: Paul McKenney >> Reported-by: Qi Zheng >> --- >>   mm/page_alloc.c | 10 +++++++--- >>   1 file changed, 7 insertions(+), 3 deletions(-) >> >> diff --git a/mm/page_alloc.c b/mm/page_alloc.c >> index 6e469c7ef9a4..73e334f403fd 100644 >> --- a/mm/page_alloc.c >> +++ b/mm/page_alloc.c >> @@ -94,11 +94,15 @@ static DEFINE_MUTEX(pcp_batch_high_lock); >>   #if defined(CONFIG_SMP) || defined(CONFIG_PREEMPT_RT) >>   /* >> - * On SMP, spin_trylock is sufficient protection. >> + * On SMP, spin_trylock is sufficient protection against recursion. >>    * On PREEMPT_RT, spin_trylock is equivalent on both SMP and UP. >> + * >> + * Block softirq execution to prevent RCU frees from running in softirq >> + * context while this CPU holds the PCP lock, which could result in a >> whole >> + * bunch of frees contending on the zone->lock. >>    */ >> -#define pcp_trylock_prepare(flags)    do { } while (0) >> -#define pcp_trylock_finish(flag)    do { } while (0) >> +#define pcp_trylock_prepare(flags)    local_bh_disable() >> +#define pcp_trylock_finish(flag)    local_bh_enable() > > I just tested this, and it doesn't seem to improve much: > > root@debian:~# stress-ng --timeout 60 --times --verify --metrics > --no-rand-seed --mmapaddr 64 > stress-ng: info:  [671] dispatching hogs: 64 mmapaddr > stress-ng: info:  [671] successful run completed in 60.07s (1 min, 0.07 > secs) > stress-ng: info:  [671] stressor       bogo ops real time  usr time  sys > time   bogo ops/s   bogo ops/s > stress-ng: info:  [671]                           (secs)    (secs) > (secs)   (real time) (usr+sys time) > stress-ng: info:  [671] mmapaddr       19803127     60.01    235.20 > 1146.76    330007.29     14329.74 > stress-ng: info:  [671] for a 60.07s run time: > stress-ng: info:  [671]    1441.59s available CPU time > stress-ng: info:  [671]     235.57s user time   ( 16.34%) > stress-ng: info:  [671]    1147.20s system time ( 79.58%) > stress-ng: info:  [671]    1382.77s total time  ( 95.92%) > stress-ng: info:  [671] load average: 41.42 11.91 4.10 > > The _raw_spin_unlock_irqrestore hotspot still exists: > >   15.87%  [kernel]  [k] _raw_spin_unlock_irqrestore >    9.18%  [kernel]  [k] clear_page_rep >    7.03%  [kernel]  [k] do_syscall_64 >    3.67%  [kernel]  [k] _raw_spin_lock >    3.28%  [kernel]  [k] __slab_free >    2.03%  [kernel]  [k] rcu_cblist_dequeue >    1.98%  [kernel]  [k] flush_tlb_mm_range >    1.88%  [kernel]  [k] lruvec_stat_mod_folio.part.131 >    1.85%  [kernel]  [k] get_page_from_freelist >    1.64%  [kernel]  [k] kmem_cache_alloc_noprof >    1.61%  [kernel]  [k] tlb_remove_table_rcu >    1.39%  [kernel]  [k] mtree_range_walk >    1.36%  [kernel]  [k] __alloc_frozen_pages_noprof >    1.27%  [kernel]  [k] pmd_install >    1.24%  [kernel]  [k] memcpy_orig >    1.23%  [kernel]  [k] __call_rcu_common.constprop.77 >    1.17%  [kernel]  [k] free_pgd_range >    1.15%  [kernel]  [k] pte_alloc_one > > The call stack is as follows: > > bpftrace -e 'k:_raw_spin_unlock_irqrestore {@[kstack,comm]=count();} > interval:s:1 {exit();}' > > @[ >     _raw_spin_unlock_irqrestore+5 >     hrtimer_interrupt+289 >     __sysvec_apic_timer_interrupt+85 >     sysvec_apic_timer_interrupt+108 >     asm_sysvec_apic_timer_interrupt+26 >     tlb_remove_table_rcu+48 >     rcu_do_batch+424 >     rcu_core+401 >     handle_softirqs+204 >     irq_exit_rcu+208 >     sysvec_apic_timer_interrupt+61 >     asm_sysvec_apic_timer_interrupt+26 > , stress-ng-mmapa]: 8 > > The tlb_remove_table_rcu() is called very rarely, so I guess the > PCP cache is basically empty at this time, resulting in the following > call stack: > But I think this may be just an extreme test scenario, because my test machine has no other load at this time. Under normal workload, page table pages should only occupy a small part of the PCP cache, and delayed freeing should not have much impact on the PCP cache. Thanks!