From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ej2-f11.google.com (mail-ej2-f11.google.com [74.125.228.139]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 3BC472D12EC for ; Tue, 18 Aug 2026 14:05:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.228.139 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787061924; cv=none; b=dSKwkDUs7zdSLxyf+jG8ruAZ/dfDsNu8f+glf0+E3qrmqqavcmxj9UEof+7JMx0VXfFT3PSJUrhctCdr55hGDeM5oKGYN6K3nw5JPoGScDJOAhoxRCv4ugC01nMqk389b42a9q7KHapDfq9NVYcwlDmS2kIflCHa/wIUvBNCJ5k= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787061924; c=relaxed/simple; bh=6lN7JRMhV++fUWBRY/mOyEZ6HBCxjfX02lNtSVCMpUg=; h=Mime-Version:Content-Type:Date:Message-Id:Cc:Subject:From:To: References:In-Reply-To; b=CO/97Zy+sJuGdB91JcS4sM76sFppk1MelZZmiVeY/W8PyeF1+qEsofj4Zbt4quqXQZx6JVgDGKqX1CJvV0rehkiwevKETfQ/EytV1AjylOYIGdM6OhWTjE93OK25vhNMdNqyu+gCEvZPQJLTDVVxJbSiza5DROWpiqDpp34kFHE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=hWnAAPQv; arc=none smtp.client-ip=74.125.228.139 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="hWnAAPQv" Received: by mail-ej2-f11.google.com with SMTP id a640c23a62f3a-c202c6c0475so108690566b.1 for ; Tue, 18 Aug 2026 07:05:23 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1787061921; x=1787666721; darn=vger.kernel.org; h=in-reply-to:references:to:from:subject:cc:message-id:date :content-type:content-transfer-encoding:mime-version:from:to:cc :subject:date:message-id:reply-to:content-type; bh=8J8uVRc1aTCENx69srR/YmAYw9UpjNQQND8eHx6fKtQ=; b=hWnAAPQvsn8q4ieYhxBLS0l2dlDwcusBC6NyPuLYLeULqSOZEP7S052MJiyCjTBpuG 7kYCgpwVN1exOwGGYy5ym0Dd3Wd8BPEbAVXvyLyIsSaMNAo2VG8Ndghf9Q1XmD/B2bxA /2lLyfglwM3x91nRW+zJnkcJSjNtGSa9/HYtWqBdpJXcaynD8ln7HHPCHtAb8uZ8msIS H8VXxSvbMBgJTbi6jGQg9scKcFkfuslB42UxkNPq2Gk1mPu4/PBCUAMwMxQNgzug/ZUA wAhTVwDCwYSV3b21LNYCafNtq2LnAgDu9DQm7lRWpWwpjp26wxC5wMm2jKmfBW5NzV9F cSOg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787061921; x=1787666721; h=in-reply-to:references:to:from:subject:cc:message-id:date :content-type:content-transfer-encoding:mime-version:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=8J8uVRc1aTCENx69srR/YmAYw9UpjNQQND8eHx6fKtQ=; b=VCyg2bl701bDMfq+uHCgJ8OhYPkzuUeDQXqM2ecAV6QKWkSLgma+lZjg4xbp+Vrkmf 6Ub/Ac0tPGfNBz9IOfaEXnZQKbPDteXr05jf3UR6/MWrLMhFYlEle/P44i3nr+TEk91+ Kkuz1dxjnvH+uqF4JI9QjpQ05ZWSUHifS1LnIYGAxfDen1IfMCHVnfe6bY56T2WSu/kQ Wc+mmuPg1oCFN6gPtdIB9eC1y6CY0buTcn4K1V/Wu7YDVA9CSYT+GbepWCPXkJYRiQZr UD6aYSmwqbcXeB0XJJ3uX5OMnSFx1JbhaIkG8Y8RftEEyFjOKoIaMsbbYvpzShjkwJkZ iGfA== X-Forwarded-Encrypted: i=1; AHgh+RoSFIR4uf6R37CzUXFN3LdWNMNDfHFV5S0Mdbdqu7SZc60k5jiZswRAzDFb33cdiZGUdr8PkJ7FTTDgt+c=@vger.kernel.org X-Gm-Message-State: AOJu0YztBRR0tywELZgfnneNbGysCu6MMzhl8g7zuFLWsWXc+EEfAzfk lCfRDzIHHZW7ZHTr8hme3SleYvu1OpAAWBd56ass1+Aj8e+xvmXLo248 X-Gm-Gg: AR+sD10qo6sST1q8fOMVmKczQnOIBXDPOJ7ke4wvJN8cl+3harb9wv38STDcQAiPY62 yjcpdIJSsmJaiQ6QLXTCJ0dlee7JpJlfH79CXiMP6VvV3bFWGWUmgzOXxEyiUiBSBUqZjGAj7hJ yWfRy7xHGeN9cS3OpaIxd+gUPeA6xe8dVeMQE22YqFkc85BXGcojSv9rS0Mx96jy2i2yiQCauqH 4Q830TX+nGeSYYXBxiFR2v7IpCh9z6N5Ltv6uMbr35hYyUf2F/t5ev24PXKDCQBG4faXxjpk1pl htrdh2q53t6nAOYW9iGO72jDRnLjP1fzbj540fXm9JqTaEcadzWLuAaSugro6YvmSzC7eVEgn1j gZaen2EC5SkzPfCovPxOWJly2s8vReGyjwJDLNOCBzre+uNjhulBUDcNG2aq5hugHwDd3nroQcF SHTjmXRNw6osttUVIEAWT2XOx5ll7qoXRhbu9EQNa7Tqc2hXlTPaRlUnnzh53QRa0RYdxm1/iWZ 1B9Uf9s6P5VCjra3LloN3uiRw7N3gf/w8TrMisTgFSu6hrdaoHbZ9jW9JKLIr15q1fImqoG6hUU owEvi4T8zbxuTQ814sDmrjo6vbA= X-Received: by 2002:a17:906:c147:b0:c16:785:cfb6 with SMTP id a640c23a62f3a-c218d9ba609mr329965666b.3.1787061921093; Tue, 18 Aug 2026 07:05:21 -0700 (PDT) Received: from localhost (nat-icclus-192-26-29-3.epfl.ch. [192.26.29.3]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-c2187d269dbsm144722766b.7.2026.08.18.07.05.20 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 18 Aug 2026 07:05:20 -0700 (PDT) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Transfer-Encoding: quoted-printable Content-Type: text/plain; charset=UTF-8 Date: Tue, 18 Aug 2026 16:05:20 +0200 Message-Id: Cc: , , "Vincent Li" , Subject: Re: [PATCH v1] LoongArch: Fix __smp_mb__{before,after}_atomic() From: "Kumar Kartikeya Dwivedi" To: "Tiezhu Yang" , "Huacai Chen" X-Mailer: aerc 0.21.0 References: <20260817035908.460-1-yangtiezhu@loongson.cn> <3103b3b4-5c0d-ed87-2e15-31a4d2289ec5@loongson.cn> <48ade10d-dd1f-ac6f-e827-d76a81183486@loongson.cn> In-Reply-To: <48ade10d-dd1f-ac6f-e827-d76a81183486@loongson.cn> On Tue Aug 18, 2026 at 3:50 PM CEST, Tiezhu Yang wrote: > On 2026/8/18 =E4=B8=8B=E5=8D=889:35, Kumar Kartikeya Dwivedi wrote: >> On Tue Aug 18, 2026 at 3:28 PM CEST, Tiezhu Yang wrote: >>> On 2026/8/17 =E4=B8=8A=E5=8D=8811:59, Tiezhu Yang wrote: >>>> When testing the BPF selftest "sudo ./test_progs -t timer_lockup", the= re >>>> is a kernel lockup and panic: >>> >>> ... >>> >>>> With this patch, the lockless "store-before-load" ordering is enforced= by >>>> the DBAR instruction. The BPF timer_lockup selftest was stressed for 5= 000 >>>> consecutive loops on a physical LoongArch machine without encountering= any >>>> further lockups or warnings: >>>> >>>> for i in {1..5000}; do sudo ./test_progs -t timer_lockup; done >>> >>> ... >>> >>>> diff --git a/arch/loongarch/include/asm/barrier.h b/arch/loongarch/inc= lude/asm/barrier.h >>>> index 4b663f197706..adfe343dfa65 100644 >>>> --- a/arch/loongarch/include/asm/barrier.h >>>> +++ b/arch/loongarch/include/asm/barrier.h >>>> @@ -57,8 +57,8 @@ >>>> #define __WEAK_LLSC_MB " \n" >>>> #endif >>>> >>>> -#define __smp_mb__before_atomic() barrier() >>>> -#define __smp_mb__after_atomic() barrier() >>>> +#define __smp_mb__before_atomic() __smp_mb() >>>> +#define __smp_mb__after_atomic() __smp_mb() >>>> >>>> /** >>>> * array_index_mask_nospec() - generate a ~0 mask when index < size= , 0 otherwise >>> >>> Hi all, >>> >>> To address the soft lockup while avoiding the performance overhead of >>> executing two separate instructions, I think there is a more elegant >>> way to fix this directly in the BPF core helper: >>> >>> We can replace atomic_inc() and smp_mb__after_atomic() with a single >>> atomic_fetch_add() in bpf_timer_cancel(). This consolidates the logic >>> into a single native atomic operation with full ordering, which not >>> only guarantees the store-load order to eliminate the deadlock but >>> also improves the performance for weak memory model architectures. >>> >>> Any thoughts on this approach? >>> >>> ``` >>> diff --git a/kernel/bpf/helpers.c b/kernel/bpf/helpers.c >>> index c18f1e16edee..ee2b3a4dcc05 100644 >>> --- a/kernel/bpf/helpers.c >>> +++ b/kernel/bpf/helpers.c >>> @@ -1591,9 +1591,7 @@ BPF_CALL_1(bpf_timer_cancel, struct bpf_async_ker= n >>> *, async) >>> */ >>> if (!cur_t) >>> goto drop; >>> - atomic_inc(&t->cancelling); >>> - /* Need full barrier after relaxed atomic_inc */ >>> - smp_mb__after_atomic(); >>> + atomic_fetch_add(1, &t->cancelling); >>> inc =3D true; >>> if (atomic_read(&cur_t->cancelling)) { >>> /* We're cancelling timer t, while some other timer >>> callback is >>> ``` >>> >>> I tested the above diff, the BPF timer_lockup selftest was stressed >>> for 5000 consecutive loops on a physical LoongArch machine without >>> encountering any further lockups or warnings. >> >> There is a full barrier in both cases. I don't know what performance imp= rovement >> will be achieved by changing this. > > We performed benchmarking using UnixBench on a physical LoongArch > machine, the UnixBench score is different (amadd.w + dbar < amadd_db.w). > I mean, this function is already quite heavy. We have multiple fully ordere= d atomics (refcount bumps/drops, xchg(), etc.) spread across the operation. D= o you observe any measurable speedup in the throughput of this function? The sec= ond question is whether timer cancellation is really frequent. In practice, I d= on't think that is the case. The actual hrtimer_cancel() in itself is heavy and waits synchronously for the callback to finish. I'm not opposed to it or anything, I just don't think it's worth it in this case. If the latency of cancel is a problem there are other bigger opportun= ities to pursue than this. >> Don't you need to fix the lowering for >> smp_mb__after_atomic() anyway? It's used in several other places in the = kernel. > > The LoongArch architecture-level fix for smp_mb__after_atomic() is > a fundamental bug fix that will be pushed separately to the LoongArch > tree. This BPF-layer change is intended as a generic, cross-architecture > performance optimization. > > Thanks, > Tiezhu