From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id EE775259490 for ; Wed, 8 Jan 2025 03:38:11 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736307493; cv=none; b=mWbrv4wG/ECvLP9wDIlapZX1cntvlXsThaaU8AMVLlOxD/2f5l0N8Rk7HKslW9D8hVp+JNYldRKfkzamvFm2z+i++kElgeKbBLkGeKQ+t27uCCSFaCI8BKgzgd3DgWVCiGJB1SI1pyzcUoAVOBBzoxETsy7sZ0oF5ZemSBxIUDk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736307493; c=relaxed/simple; bh=HuBkXU4WGAIKikDiA4v9faPr62D42N2iT3vf6U8PU7g=; h=From:Message-ID:Date:MIME-Version:Subject:To:Cc:References: In-Reply-To:Content-Type; b=JzKP9MNqlplPaQlaedPjWLyDnbkl3qVZIVxXLROHYFlF0Db0mXmKmyvBu2z+UksUPxOvsfeYHa9b2x+UPO/MlCWkCC2UjX+2LoXU0sA/1bMb6is5kG4RYYbZclR4KXPSrsgiHt7NloXdyWvf7P3v7lpeJfF+ft5+AElMcBS8BbI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=fX8k3UMq; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="fX8k3UMq" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1736307490; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=q1T2+dFg8czdV4/r0tUFcPs7edWiEyCb3Pvfi0CFBS4=; b=fX8k3UMqQJv8/NSEtVtuWyiB4EbUtT7YXP/aOtKnFBojP316humbaIGSwzuJiCukkLxpfe sfv7g/SpTbjdoKQ7UGPLI9hMReMvXEJl74aHaRiz76MbDwjGm8JjRg4wH/OB+YgQZ8LGRf pScMg41URIVkJH4e9q9uRjL9IB/YK/o= Received: from mail-qk1-f200.google.com (mail-qk1-f200.google.com [209.85.222.200]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-570-G8T52gMNMhGMjD8jaWyfxg-1; Tue, 07 Jan 2025 22:38:09 -0500 X-MC-Unique: G8T52gMNMhGMjD8jaWyfxg-1 X-Mimecast-MFC-AGG-ID: G8T52gMNMhGMjD8jaWyfxg Received: by mail-qk1-f200.google.com with SMTP id af79cd13be357-7b864496708so4711912585a.2 for ; Tue, 07 Jan 2025 19:38:09 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1736307489; x=1736912289; h=content-transfer-encoding:in-reply-to:content-language:references :cc:to:subject:user-agent:mime-version:date:message-id:from :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=q1T2+dFg8czdV4/r0tUFcPs7edWiEyCb3Pvfi0CFBS4=; b=D+ZdmV00LZM69zinfCCLqJms3S7ih//aLQWrbq0Ec37B8GC/h/CWw0vUqlwTpQ3NZ2 fxqEviqBAWHpcegQ/Xrl4b7aEYjQrTon14h2Z3hsFP07oWz7xUuB18M+YGAfSfehj8IS AAGaOG1EXcxmVgv8LSxaTNezaDyJIX3S+D3lOTBPyIEWZyoXwh551XVmahYs2+TJ6xcV o9nq8fP42hpqw8cop6dMrPEBFLAfauJwPjFEGDIZZatDIGM9yuRnGwxqgT9Q4MJ5V94K DZp5VmFOXcTW+cAyBi+szbUI65dPI3dY49Xrq/rNYM3mwM+ZWyOeCUUcHPV2bmGhf/M2 G/RQ== X-Forwarded-Encrypted: i=1; AJvYcCWBFItsWCKLri3yifatRH2MSqbuKGbC3xNRMxVlQLCjg08fEDaiGdKGOrcNbCgwITiFqu9T6Y+bcy4I99o=@vger.kernel.org X-Gm-Message-State: AOJu0YzKOUImX4JnNIkf8Q9q/HZ0e+xuXr6f3zrswytjssILI9qvJFBR qi8Xvp+bFM7ryJZv/UMEa3MCb68dhpZ19ESF0/sbWyleedUzoAHBknmMMtWbjzEN1ViEKsn/Nrf H/x1ach2sa55DDJ69JvvyRAbvtgiEtHQHD3g9Q8SHXRlmm1b9RKUZh5CyeRhLDw== X-Gm-Gg: ASbGnctcvWuyY6cVbIoOWgBqM27RzuhtXyQFA8R2bfp8ibW25oDvRht/PK7gUp2nrHv N/ZZSUQKT7zRFOmMBJUkWO3U45NNK03TUXbHsFYBGMiZ50kspQTxt/yI/W/ZNWS8YfndyW5UBYs GGucD7dAXJ0jN6NUHCvVx0DXZKH5xYggOWTY9uU/ZDzR7fNopa44unJkEOBdwa0NbI+gRCHljKd iXqrhBqcU8RpEvUUTNtu7QFCewVg7wrBYz4eN09NYrdxdD/Brx5KLsHtXCUz+7n49m3lO5Bsi7P B/0o5nNnk1u0X67bIEV8B8x+ X-Received: by 2002:a05:620a:1aa7:b0:7b6:e20d:2b55 with SMTP id af79cd13be357-7bcd97b1aa0mr201432985a.41.1736307489221; Tue, 07 Jan 2025 19:38:09 -0800 (PST) X-Google-Smtp-Source: AGHT+IF2jp1gzMsOIs/upwLQ9lfk8KRV5eRMFJzkXxfha13aTa4I5WcjU0JLvHMEgpdV7UefW2UHow== X-Received: by 2002:a05:620a:1aa7:b0:7b6:e20d:2b55 with SMTP id af79cd13be357-7bcd97b1aa0mr201430385a.41.1736307488888; Tue, 07 Jan 2025 19:38:08 -0800 (PST) Received: from ?IPV6:2601:188:ca00:a00:f844:fad5:7984:7bd7? ([2601:188:ca00:a00:f844:fad5:7984:7bd7]) by smtp.gmail.com with ESMTPSA id af79cd13be357-7b9ac478e59sm1648658485a.78.2025.01.07.19.38.06 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Tue, 07 Jan 2025 19:38:08 -0800 (PST) From: Waiman Long X-Google-Original-From: Waiman Long Message-ID: Date: Tue, 7 Jan 2025 22:38:06 -0500 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH bpf-next v1 09/22] rqspinlock: Protect waiters in queue from stalls To: Kumar Kartikeya Dwivedi , bpf@vger.kernel.org, linux-kernel@vger.kernel.org Cc: Barret Rhoden , Linus Torvalds , Peter Zijlstra , Waiman Long , Alexei Starovoitov , Andrii Nakryiko , Daniel Borkmann , Martin KaFai Lau , Eduard Zingerman , "Paul E. McKenney" , Tejun Heo , Josh Don , Dohyun Kim , kernel-team@meta.com References: <20250107140004.2732830-1-memxor@gmail.com> <20250107140004.2732830-10-memxor@gmail.com> Content-Language: en-US In-Reply-To: <20250107140004.2732830-10-memxor@gmail.com> Content-Type: text/plain; charset=UTF-8; format=flowed Content-Transfer-Encoding: 7bit On 1/7/25 8:59 AM, Kumar Kartikeya Dwivedi wrote: > Implement the wait queue cleanup algorithm for rqspinlock. There are > three forms of waiters in the original queued spin lock algorithm. The > first is the waiter which acquires the pending bit and spins on the lock > word without forming a wait queue. The second is the head waiter that is > the first waiter heading the wait queue. The third form is of all the > non-head waiters queued behind the head, waiting to be signalled through > their MCS node to overtake the responsibility of the head. > > In this commit, we are concerned with the second and third kind. First, > we augment the waiting loop of the head of the wait queue with a > timeout. When this timeout happens, all waiters part of the wait queue > will abort their lock acquisition attempts. This happens in three steps. > First, the head breaks out of its loop waiting for pending and locked > bits to turn to 0, and non-head waiters break out of their MCS node spin > (more on that later). Next, every waiter (head or non-head) attempts to > check whether they are also the tail waiter, in such a case they attempt > to zero out the tail word and allow a new queue to be built up for this > lock. If they succeed, they have no one to signal next in the queue to > stop spinning. Otherwise, they signal the MCS node of the next waiter to > break out of its spin and try resetting the tail word back to 0. This > goes on until the tail waiter is found. In case of races, the new tail > will be responsible for performing the same task, as the old tail will > then fail to reset the tail word and wait for its next pointer to be > updated before it signals the new tail to do the same. > > Lastly, all of these waiters release the rqnode and return to the > caller. This patch underscores the point that rqspinlock's timeout does > not apply to each waiter individually, and cannot be relied upon as an > upper bound. It is possible for the rqspinlock waiters to return early > from a failed lock acquisition attempt as soon as stalls are detected. > > The head waiter cannot directly WRITE_ONCE the tail to zero, as it may > race with a concurrent xchg and a non-head waiter linking its MCS node > to the head's MCS node through 'prev->next' assignment. > > Reviewed-by: Barret Rhoden > Signed-off-by: Kumar Kartikeya Dwivedi > --- > kernel/locking/rqspinlock.c | 42 +++++++++++++++++++++++++++++--- > kernel/locking/rqspinlock.h | 48 +++++++++++++++++++++++++++++++++++++ > 2 files changed, 87 insertions(+), 3 deletions(-) > create mode 100644 kernel/locking/rqspinlock.h > > diff --git a/kernel/locking/rqspinlock.c b/kernel/locking/rqspinlock.c > index dd305573db13..f712fe4b1f38 100644 > --- a/kernel/locking/rqspinlock.c > +++ b/kernel/locking/rqspinlock.c > @@ -77,6 +77,8 @@ struct rqspinlock_timeout { > u16 spin; > }; > > +#define RES_TIMEOUT_VAL 2 > + > static noinline int check_timeout(struct rqspinlock_timeout *ts) > { > u64 time = ktime_get_mono_fast_ns(); > @@ -305,12 +307,18 @@ int __lockfunc resilient_queued_spin_lock_slowpath(struct qspinlock *lock, u32 v > * head of the waitqueue. > */ > if (old & _Q_TAIL_MASK) { > + int val; > + > prev = decode_tail(old, qnodes); > > /* Link @node into the waitqueue. */ > WRITE_ONCE(prev->next, node); > > - arch_mcs_spin_lock_contended(&node->locked); > + val = arch_mcs_spin_lock_contended(&node->locked); > + if (val == RES_TIMEOUT_VAL) { > + ret = -EDEADLK; > + goto waitq_timeout; > + } > > /* > * While waiting for the MCS lock, the next pointer may have > @@ -334,7 +342,35 @@ int __lockfunc resilient_queued_spin_lock_slowpath(struct qspinlock *lock, u32 v > * sequentiality; this is because the set_locked() function below > * does not imply a full barrier. > */ > - val = atomic_cond_read_acquire(&lock->val, !(VAL & _Q_LOCKED_PENDING_MASK)); > + RES_RESET_TIMEOUT(ts); > + val = atomic_cond_read_acquire(&lock->val, !(VAL & _Q_LOCKED_PENDING_MASK) || > + RES_CHECK_TIMEOUT(ts, ret)); This has the same wfe problem for arm64. Cheers, Longman