From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mx0a-00069f02.pphosted.com (mx0a-00069f02.pphosted.com [205.220.165.32]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CC0D142AF9F for ; Mon, 10 Aug 2026 16:35:00 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=205.220.165.32 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786379702; cv=none; b=VztuumQZf7+vsnfbyoOYT/NVlVjlDUVDDTEDYpXLvlVVVmHyjK+JGTkjlck6Aozb+awAWuBGZcbXxK3DRBhJlgyZYxuZ62UojTatHUMz8UVXQGFLJtHp3Hv+9+2TcPbPi0vRFUBpCDmaxctiRgJE2zfEvNaac1yk2Tc1H+qgR80= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786379702; c=relaxed/simple; bh=d847JEF6UzlMJD89+z/kGUy1I8HaZYHi2dL8tUOSYXQ=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=dqtSuLVqwBTpL+mRntMF2Xxm70B+6gLMzxBlYYN0niYQz48SdURTyS9i7d0CpG/QPbPXRltxfnAhGMZhUHEgQN6IbxPBR36HOAzBFBnsk9IBTTXo0bbDLfbsCx4DCUWAmxMPmtLUCJ0Clj7hj14PhW01qtj1sN5kOGHatBAluj0= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oracle.com; spf=pass smtp.mailfrom=oracle.com; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b=nybBNX12; arc=none smtp.client-ip=205.220.165.32 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=oracle.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=oracle.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=oracle.com header.i=@oracle.com header.b="nybBNX12" Received: from pps.filterd (m0333521.ppops.net [127.0.0.1]) by mx0b-00069f02.pphosted.com (8.18.1.11/8.18.1.11) with ESMTP id 67AFjDjQ559926; Mon, 10 Aug 2026 16:34:41 GMT DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=oracle.com; h=cc :content-transfer-encoding:content-type:date:from:in-reply-to :message-id:mime-version:references:subject:to; s= corp-2025-04-25; bh=9qgx/Gk6tspchyfv0/DujKaVHJv0n7mHq0aRMVuLfUo=; b= nybBNX12Gdh2ve9waqn2c3JxVMFL7ppxTgJZVFRksunCoU8HLTivDTNJ6mHjcFpC WfY7f7NW3mRNQlR5ut+ADmWk8OX1SP+D5vkbRqbUh0rsghmBQfmT8eX0IzTEcghQ GXoymOmRGRQrF99YLbytB8//Das8fLBVFrwfy87/blw1QWdEYdqpYGWkDi6uln/T TruGFWCSNL/wbInS3QU9dIeu61HuIEdZmlL6fryJy01tiJrsL/UhtFj4+zIgWUa7 n1WyTgl0zyfpnkoiSI0HwNuMRoPN4OV/a4KD2nyAWfJNaDvEHWIDbTby/Zs1H59a Z+yljlXEYDvTveLBm/zL9A== Received: from iadpaimrmta01.imrmtpd1.prodappiadaev1.oraclevcn.com (iadpaimrmta01.appoci.oracle.com [130.35.100.223]) by mx0b-00069f02.pphosted.com (PPS) with ESMTPS id 4fwvestwss-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Mon, 10 Aug 2026 16:34:40 +0000 (GMT) Received: from pps.filterd (iadpaimrmta01.imrmtpd1.prodappiadaev1.oraclevcn.com [127.0.0.1]) by iadpaimrmta01.imrmtpd1.prodappiadaev1.oraclevcn.com (8.18.1.7/8.18.1.7) with ESMTP id 67AGUbnS020310; Mon, 10 Aug 2026 16:34:39 GMT Received: from pps.reinject (localhost [127.0.0.1]) by iadpaimrmta01.imrmtpd1.prodappiadaev1.oraclevcn.com (PPS) with ESMTPS id 4fwtwpf788-1 (version=TLSv1.2 cipher=ECDHE-RSA-AES256-GCM-SHA384 bits=256 verify=OK); Mon, 10 Aug 2026 16:34:39 +0000 (GMT) Received: from iadpaimrmta01.imrmtpd1.prodappiadaev1.oraclevcn.com (iadpaimrmta01.imrmtpd1.prodappiadaev1.oraclevcn.com [127.0.0.1]) by pps.reinject (8.18.1.12/8.18.1.12) with ESMTP id 67AGYaaU038843; Mon, 10 Aug 2026 16:34:38 GMT Received: from lab61.no.oracle.com (lab61.no.oracle.com [10.172.144.82]) by iadpaimrmta01.imrmtpd1.prodappiadaev1.oraclevcn.com (PPS) with ESMTP id 4fwtwpf761-2; Mon, 10 Aug 2026 16:34:38 +0000 (GMT) From: =?UTF-8?q?H=C3=A5kon=20Bugge?= To: linux-kernel@vger.kernel.org, Peter Zijlstra , Ingo Molnar , Will Deacon , Boqun Feng , Waiman Long , Chris Wilson Cc: John Stultz , Bradley Morgan , =?UTF-8?q?H=C3=A5kon=20Bugge?= , Ingo Molnar Subject: [PATCH v2 2/2] test-ww_mutex: Fix deadlock in test_cycle_work Date: Mon, 10 Aug 2026 18:34:31 +0200 Message-ID: <20260810163433.3765919-2-haakon.bugge@oracle.com> X-Mailer: git-send-email 2.43.5 In-Reply-To: <20260810163433.3765919-1-haakon.bugge@oracle.com> References: <20260810163433.3765919-1-haakon.bugge@oracle.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit X-Proofpoint-Virus-Version: vendor=baseguard engine=ICAP:2.0.293,Aquarius:18.0.1176,Hydra:6.1.134,FMLib:17.12.100.49 definitions=2026-08-10_04,2026-08-10_02,2025-10-01_01 X-Proofpoint-Spam-Details: rule=notspam policy=default score=0 lowpriorityscore=0 mlxscore=0 suspectscore=0 mlxlogscore=999 spamscore=0 phishscore=0 malwarescore=0 adultscore=0 bulkscore=0 classifier=spam adjust=0 reason=mlx scancount=1 engine=8.19.0-2606160000 definitions=main-2608100142 X-Authority-Analysis: v=2.4 cv=R8wz39RX c=1 sm=1 tr=0 ts=6a79fda0 b=1 cx=c_pps a=zPCbziy225d3KhSqZt3L1A==:117 a=zPCbziy225d3KhSqZt3L1A==:17 a=IkcTkHD0fZMA:10 a=Sv0fKeRqtYgA:10 a=M51BFTxLslgA:10 a=VkNPw1HP01LnGYTKEx00:22 a=jiCTI4zE5U7BLdzWsZGv:22 a=x0eKOSpe3m1H3M0S9YoZ:22 a=yPCof4ZbAAAA:8 a=jcFIXsoQAAAA:8 a=WAdc_XZaImYrliwInUAA:9 a=3ZKOabzyN94A:10 a=QEXdDO2ut3YA:10 a=M1-Q_cRM6PFdmZSamPpU:22 a=5yU3S35YU4bGjq-dph-N:22 a=Bho9c0fBagfJEIQBS7DQ:22 cc=ntf awl=host:12100 X-Proofpoint-GUID: c-qpJTq9g6KPxa8PHpL4kYKqK--b_9vu X-Proofpoint-Spam-Details-Enc: AW1haW4tMjYwODEwMDE0MiBTYWx0ZWRfXydQjuXXXZbfn 6xNpjQoa+9aj2RCsOPPg+Gg7jIkZBAadKRiV37GmAV89Wn6FYVYoNYA/bZ0ZCSMgJUx5FI+m6tB 6oRycIsqjq3Dock7qOnhS2LVAF4L+WmQOhk0u0L8zxPAL3Jp6FRaluTlg23W5VcY3dbj4XMgIQX YD9zLzlzCtiZgTiZ8s9MDs8dVbNxuWULukNuDayLb9pfECaV/HkdqGZs+hqRCMuwRnZXKXtpR7R ELpe22CUacjI+g/B+hxN+eSHX3IHgNzujAESHrHVz2pCFMlQhXq8JHQ6jI1BWtUXIIOraHYi+2A B60vg8hn9G570ZXcNdDPYyE59N6U2CQNf+cRSsbRyY5zxX3MHTaYt7BtCuR8x2Nuz+luoVxNv4c q35e/AoxvDshos0Uuc/ZVbazvDd+PAcMAJnvQWKp5VHxDjMvquaNIm6fenfrk3AGqPynNWy09Zc juStZJnYwc3llSU25OjKcYOvzFh4ERZaFycLaqyc= X-Proofpoint-Spam-Info: AW1haW4tMjYwODEwMDE0MiBTYWx0ZWRfX6fS6/oasPCDk TrZoEPxk6y6YOZk7LysL8ySF9vZQH4KHOe2MD43l10WQIT+SiyffZ7BPn/Y8zHuy3I7MrexmTBD cN2qIPYurS+t/sjuD+e0UiwiP+2WUbI+koAkbFsaE5+OHHAghFRN X-Proofpoint-ORIG-GUID: c-qpJTq9g6KPxa8PHpL4kYKqK--b_9vu When running with N online CPUs, where N is fairly large, let's say 512, a deadlock may happen in test_cycle_work() when running with N + 1 kernel threads. The reason is that the min_active is too small in order to let N + 1 worker threads run concurrently. The following is my analyzes. The test sets up a circular dependency with ww_mutexes and completions: work[1] completes signal[0] work[2] completes signal[1] ... work[511] completes signal[510] work[512] would complete signal[511] work[0] completes signal[512] Therefore: work[0..510] -> blocked in ww_mutex_lock(b_mutex) work[511] -> blocked in wait_for_completion(signal[511]) work[512] -> inactive; test_cycle_work() has not started Because the last worker thread has not started, worker 511 hangs forever in wait_for_completion(). This bug produces the following splats (slightly edited for better brevity). This for the worker threads hung in ww_mutex_lock(): INFO: task kworker/u2066:1:13658 blocked for more than 123 seconds. Workqueue: test-ww_mutex test_cycle_work [test_ww_mutex] Call Trace: __schedule+0x28e/0x670 schedule+0x27/0xa0 schedule_preempt_disabled+0x15/0x30 __ww_mutex_lock.constprop.0+0x841/0xe00 test_cycle_work+0x9d/0x150 [test_ww_mutex] process_one_work+0x196/0x370 worker_thread+0x1af/0x320 kthread+0xe3/0x120 ret_from_fork+0x19e/0x260 ret_from_fork_asm+0x1a/0x30 And this for the single worker thread 511, hung in wait_for_completion(): Workqueue: test-ww_mutex test_cycle_work [test_ww_mutex] Call Trace: __schedule+0x28e/0x670 schedule+0x27/0xa0 schedule_timeout+0xac/0xf0 __wait_for_common+0x97/0x1b0 test_cycle_work+0x80/0x160 [test_ww_mutex] process_one_work+0x196/0x370 worker_thread+0x1af/0x320 kthread+0xe3/0x120 ret_from_fork+0x19e/0x260 ret_from_fork_asm+0x1a/0x30 We fix this by adjusting the wq's min_active parameter. When num_online_cpus() has been sampled in run_tests(), we adjust the {min,max}_active values of the wq, to make sure both are set to ncpus + 1 in order to avoid the above deadlock. Note that this fix is invariant to num_online_cpus() changing after it has been sampled, because the RC here is the number of runnable worker threads vs. threads created, not per se the number of online CPUs. Also, in order to avoid exceeding WQ_MAX_ACTIVE, we create cycle_ncpus and clamp it. We do not want to change ncpus for the other tests, not affected by this bug. Fixes: d1b42b800e5d ("locking/ww_mutex: Add kselftests for resolving ww_mutex cyclic deadlocks") Signed-off-by: HÃ¥kon Bugge Reviewed-by: Bradley Morgan --- v1 -> v2: * Added Bradley's r-b * Reworded comment in run_tests() --- kernel/locking/test-ww_mutex.c | 13 ++++++++++++- 1 file changed, 12 insertions(+), 1 deletion(-) diff --git a/kernel/locking/test-ww_mutex.c b/kernel/locking/test-ww_mutex.c index 47e016a4f4fea..e11f6b869b4d4 100644 --- a/kernel/locking/test-ww_mutex.c +++ b/kernel/locking/test-ww_mutex.c @@ -676,6 +676,7 @@ static int stress(struct ww_class *class, int nlocks, int nthreads, unsigned int static int run_tests(struct ww_class *class) { int ncpus = num_online_cpus(); + int cycle_ncpus = min_t(int, ncpus, WQ_MAX_ACTIVE - 1); int ret, i; ret = test_mutex(class); @@ -696,7 +697,17 @@ static int run_tests(struct ww_class *class) return ret; } - ret = test_cycle(class, ncpus); + /* + * test_cycle_work() has a linear dependency which requires + * all kernel threads to be run-able at once. With N CPUs and + * N + 1 worker threads, deadlock may happen. Hence, adjust + * min_active. Raise max first, min_active is clamped to it. + * Cap N so that N + 1 doesn't exceed WQ_MAX_ACTIVE. + */ + workqueue_set_max_active(wq, cycle_ncpus + 1); + workqueue_set_min_active(wq, cycle_ncpus + 1); + + ret = test_cycle(class, cycle_ncpus); if (ret) return ret; -- 2.43.5