From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0a-00069f02.pphosted.com (mx0a-00069f02.pphosted.com [205.220.165.32]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id B478E3BCD33 for ; Tue, 11 Aug 2026 15:53:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=205.220.165.32 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786463641; cv=none; b=HwKerj6QYjZRQNfwwqexjOmpAK2ibbRVHmx1qb8iU03KjnzApQ5AVxZm/y7G6qjtVGRyUCoBvlVNfVnaY9AneeYnJZLE6BDlLq2Ebon0p+SLvqSSqgEZXoLH82SnLN0YgkWqLLJs0KXYNtv4eOeY86omsccayS6Diqo1JqMFMCQ= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786463641; c=relaxed/simple; bh=Z0rP/ONyO2C17ijv2MYIZSfIyfkcRAQXcfsauGiv/M8=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=M9oRuAsbAjKgA+hzkEBzQzEHE9sMi5+TcVPDPGh+0JcDmfcWCI+MSmis7pWP7HL8XwiWzPqPPKLPiEnIWQp/74vJlTIB0XNtijIHvMaod1QucmuE+tINhdj/rfqGZkN+fRoPexyvtyJY87hQ2mB8/VeHW5rgWRaHkW+I0AmIQJU= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oracle.com; spf=pass smtp.mailfrom=oracle.com; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b=GY1LoD1A; arc=none smtp.client-ip=205.220.165.32 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oracle.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=oracle.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="GY1LoD1A" Received: from pps.filterd (m0246617.ppops.net [127.0.0.1]) by mx0b-00069f02.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 67BFMaCG2599665; Tue, 11 Aug 2026 15:53:48 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s= corp-2025-04-25; bh=11tT7Yz82H+hTijMLxUSJNYTFrPYfi0zfItrXM8qiAg=; b= GY1LoD1ApI4DXDnrErmcdwvjnIYe+OWIUJQg5flCy9v9J0lj/k+VJcMa8Elxq0MR GbnrZADUhPsxCktUpWhmbtUxpUNwI7bw5Sh0DS1QkycLYmba0kRUpItjNJ2ufJS2 88TuMi5rPP0JOpEJCgq21xQWJ1eOrXMsfLJ3wI/1Rw8PuUwLYkkgrKSThcO0xiot wpJS2Gz30c1DXyK0FJFW0t2EzynHD0pMVML1vQLGo3U529k65nX/ZsAKjgPPkhEa 8coWqg8V2+zUFp3h0wzE7pnmI3bN/HnvR5riDqizBe/jHiFOvKm4UXyJj2TxAN/T AkzTASenx4O7kf0jZ6LWJA== Received: from phxpaimrmta01.imrmtpd1.prodappphxaev1.oraclevcn.com (phxpaimrmta01.appoci.oracle.com [138.1.114.2]) by mx0b-00069f02.pphosted.com (PPS) with ESMTPS id 4fww0sw69x-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Tue, 11 Aug 2026 15:53:48 +0000 (GMT) Received: from pps.filterd (phxpaimrmta01.imrmtpd1.prodappphxaev1.oraclevcn.com [127.0.0.1]) by phxpaimrmta01.imrmtpd1.prodappphxaev1.oraclevcn.com (8.18.1.7/8.18.1.7) with ESMTP id 67BFjC9G033068; Tue, 11 Aug 2026 15:53:47 GMT Received: from pps.reinject (localhost [127.0.0.1]) by phxpaimrmta01.imrmtpd1.prodappphxaev1.oraclevcn.com (PPS) with ESMTPS id 4fwtwdqw0y-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Tue, 11 Aug 2026 15:53:47 +0000 (GMT) Received: from phxpaimrmta01.imrmtpd1.prodappphxaev1.oraclevcn.com (phxpaimrmta01.imrmtpd1.prodappphxaev1.oraclevcn.com [127.0.0.1]) by pps.reinject (8.18.1.12/8.18.1.12) with ESMTP id 67BFrhMH037679; Tue, 11 Aug 2026 15:53:46 GMT Received: from lab61.no.oracle.com (lab61.no.oracle.com [10.172.144.82]) by phxpaimrmta01.imrmtpd1.prodappphxaev1.oraclevcn.com (PPS) with ESMTP id 4fwtwdqvwn-2; Tue, 11 Aug 2026 15:53:46 +0000 (GMT) From: =?UTF-8?q?H=C3=A5kon=20Bugge?= To: linux-kernel@vger.kernel.org, Peter Zijlstra , Ingo Molnar , Will Deacon , Boqun Feng , Waiman Long , Chris Wilson Cc: John Stultz , Bradley Morgan , Tejun Heo , =?UTF-8?q?H=C3=A5kon=20Bugge?= , Ingo Molnar Subject: [PATCH v3 2/2] test-ww_mutex: Fix deadlock in test_cycle_work Date: Tue, 11 Aug 2026 17:53:38 +0200 Message-ID: <20260811155340.3867610-2-haakon.bugge@oracle.com> X-Mailer: git-send-email 2.43.5 In-Reply-To: <20260811155340.3867610-1-haakon.bugge@oracle.com> References: <20260811155340.3867610-1-haakon.bugge@oracle.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-11_03,2026-08-10_03,2025-10-01_01 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 mlxscore=0 adultscore=0 mlxlogscore=999 lowpriorityscore=0 suspectscore=0 malwarescore=0 phishscore=0 bulkscore=0 spamscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.19.0-2606160000 definitions=main-2608110132 X-Proofpoint-Spam-Info: AW1haW4tMjYwODExMDEzMyBTYWx0ZWRfX3cHOUE6VnYz9 Hp4Eut6zidASyMF8SvtXPxOgRMP72SkLDz5CRSXUEkIRRwiZD/BTYjFmFBb9JOZOSnkS72zD2mb ZjVmbLnt/7J13ipxx8QrvZieRSIc2ihmbw4VxxMvZBjx2EOwuMkE X-Proofpoint-ORIG-GUID: PF9x_Ui7aUhP6XQ-EEKEwGcB7YeGYMwC X-Proofpoint-GUID: PF9x_Ui7aUhP6XQ-EEKEwGcB7YeGYMwC X-Authority-Analysis: v=2.4 cv=FcAHAp+6 c=1 sm=1 tr=0 ts=6a7b458c cx=c_pps a=XiAAW1AwiKB2Y8Wsi+sD2Q==:117 a=XiAAW1AwiKB2Y8Wsi+sD2Q==:17 a=IkcTkHD0fZMA:10 a=Sv0fKeRqtYgA:10 a=M51BFTxLslgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=jiCTI4zE5U7BLdzWsZGv:22 a=7Gl3-_t3PgB9XO-mQDs3:22 a=yPCof4ZbAAAA:8 a=jcFIXsoQAAAA:8 a=WAdc_XZaImYrliwInUAA:9 a=3ZKOabzyN94A:10 a=QEXdDO2ut3YA:10 a=M1-Q_cRM6PFdmZSamPpU:22 X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODExMDEzMyBTYWx0ZWRfX3rCzOsZRvSrc vRWsDRFHXp7PN3pxg8wF0qLMDX0lR6DtTG415OKBaTiBKY5PLMUiz7ed29K0MNLOISvMKh/tWGE Oldxbn2ZLCBUyZfnn0i64ccXdfB3g//GFIqUtQUST+AsugCB9cfgt8hhKchGJFudHSa5c1IAqVQ PhPMlDT2CXM23YD2a6fcnUd1FRR3XdE1ocWmYa3FYhcqFGK1n0GL5vda5YHXaLNSEHVFiKBjy2M kDaUVPLhZwUMpUJk7H6S9CSs/yQTDXbpnruZpzmJHoNNiGIJzttbHagOr/qc5w1D5Q4rg/2DPTo /jT1d95b3K7szZ+hWtdhJXQ/VEJFZyM7ghRbGKu3PnxSTPUpAx8ksz4OSiXnpSVumU5To3N5M60 EgyMFPuG6YdEgMHjHEjU5IvjgYzTjllsqxbC1tVLgStQSta+o90MObK88i/Ps9tIXCeTNGAr8eE ZdrBypZrg6WIEiwyR0w== When running with N online CPUs, where N is fairly large, let's say 512, a deadlock may happen in test_cycle_work() when running with N + 1 kernel threads. The reason is that the min_active is too small in order to let N + 1 worker threads run concurrently. The following is my analyzes. The test sets up a circular dependency with ww_mutexes and completions: work[1] completes signal[0] work[2] completes signal[1] ... work[511] completes signal[510] work[512] would complete signal[511] work[0] completes signal[512] Therefore: work[0..510] -> blocked in ww_mutex_lock(b_mutex) work[511] -> blocked in wait_for_completion(signal[511]) work[512] -> inactive; test_cycle_work() has not started Because the last worker thread has not started, worker 511 hangs forever in wait_for_completion(). This bug produces the following splats (slightly edited for better brevity). This for the worker threads hung in ww_mutex_lock(): INFO: task kworker/u2066:1:13658 blocked for more than 123 seconds. Workqueue: test-ww_mutex test_cycle_work [test_ww_mutex] Call Trace: __schedule+0x28e/0x670 schedule+0x27/0xa0 schedule_preempt_disabled+0x15/0x30 __ww_mutex_lock.constprop.0+0x841/0xe00 test_cycle_work+0x9d/0x150 [test_ww_mutex] process_one_work+0x196/0x370 worker_thread+0x1af/0x320 kthread+0xe3/0x120 ret_from_fork+0x19e/0x260 ret_from_fork_asm+0x1a/0x30 And this for the single worker thread 511, hung in wait_for_completion(): Workqueue: test-ww_mutex test_cycle_work [test_ww_mutex] Call Trace: __schedule+0x28e/0x670 schedule+0x27/0xa0 schedule_timeout+0xac/0xf0 __wait_for_common+0x97/0x1b0 test_cycle_work+0x80/0x160 [test_ww_mutex] process_one_work+0x196/0x370 worker_thread+0x1af/0x320 kthread+0xe3/0x120 ret_from_fork+0x19e/0x260 ret_from_fork_asm+0x1a/0x30 We fix this by adjusting the wq's min_active parameter. When num_online_cpus() has been sampled in run_tests(), we adjust the {min,max}_active values of the wq, to make sure both are set to ncpus + 1 in order to avoid the above deadlock. Note that this fix is invariant to num_online_cpus() changing after it has been sampled, because the RC here is the number of runnable worker threads vs. threads created, not per se the number of online CPUs. Also, in order to avoid exceeding WQ_MAX_ACTIVE, we create cycle_ncpus and clamp it. We do not want to change ncpus for the other tests, not affected by this bug. Fixes: d1b42b800e5d ("locking/ww_mutex: Add kselftests for resolving ww_mutex cyclic deadlocks") Signed-off-by: HÃ¥kon Bugge Reviewed-by: Bradley Morgan --- v2 -> v3: * No changes v1 -> v2: * Added Bradley's r-b * Reworded comment in run_tests() --- kernel/locking/test-ww_mutex.c | 13 ++++++++++++- 1 file changed, 12 insertions(+), 1 deletion(-) diff --git a/kernel/locking/test-ww_mutex.c b/kernel/locking/test-ww_mutex.c index 47e016a4f4fea..e11f6b869b4d4 100644 --- a/kernel/locking/test-ww_mutex.c +++ b/kernel/locking/test-ww_mutex.c @@ -676,6 +676,7 @@ static int stress(struct ww_class *class, int nlocks, int nthreads, unsigned int static int run_tests(struct ww_class *class) { int ncpus = num_online_cpus(); + int cycle_ncpus = min_t(int, ncpus, WQ_MAX_ACTIVE - 1); int ret, i; ret = test_mutex(class); @@ -696,7 +697,17 @@ static int run_tests(struct ww_class *class) return ret; } - ret = test_cycle(class, ncpus); + /* + * test_cycle_work() has a linear dependency which requires + * all kernel threads to be run-able at once. With N CPUs and + * N + 1 worker threads, deadlock may happen. Hence, adjust + * min_active. Raise max first, min_active is clamped to it. + * Cap N so that N + 1 doesn't exceed WQ_MAX_ACTIVE. + */ + workqueue_set_max_active(wq, cycle_ncpus + 1); + workqueue_set_min_active(wq, cycle_ncpus + 1); + + ret = test_cycle(class, cycle_ncpus); if (ret) return ret; -- 2.43.5