From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from stravinsky.debian.org (stravinsky.debian.org [82.195.75.108]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 46FCA421A19 for ; Tue, 26 May 2026 18:08:16 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=82.195.75.108 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779818898; cv=none; b=B4eYD8wkCDiWdZBVsCjQfZwKRR3CH+VNoJHa+1ljTbxlv/xYT+PSz+9jw1K7GTowQ9NAtpm/vmccaSio99KL6UkxIvdkuQjJWI8PPF18hWMlaqRN3aR5W1hhRWNIzKgXESolWseQSbDvMNZNB7/h4heAVNx7Qe3L4tA7NymZcJI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1779818898; c=relaxed/simple; bh=OHFpmGg6eK2vffd/eDnZJyfgvVD70ayi3QDPz+fZYSo=; h=From:Date:Subject:MIME-Version:Content-Type:Message-Id:References: In-Reply-To:To:Cc; b=hnkQ1GLCCiB72ceCV4Py02RiaiBhLPAHSmhugRma7pLH04UayZLRYbHQv15Zk6vhbjX6fBi+56KGDmSQU0uYJXaOBQhVdoFgDqOfubE6BNNuZkRTCaRxrw7vzbFF5sP3gZpEwo7lyqX3VBmsx5fzpRJyQ7Lw/Nd2d9Jptcf6m28= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org; spf=pass smtp.mailfrom=debian.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b=QMrHnuzL; arc=none smtp.client-ip=82.195.75.108 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=debian.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b="QMrHnuzL" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=debian.org; s=smtpauto.stravinsky; h=X-Debian-User:Cc:To:In-Reply-To:References: Message-Id:Content-Transfer-Encoding:Content-Type:MIME-Version:Subject:Date: From:Reply-To:Content-ID:Content-Description; bh=p6PRWoqmsxSGKSyXhbzegHEoF9nWLbHQq8uk2FbvmCI=; b=QMrHnuzLbun0FFk0ZfJQGHpjwZ DXqHOdugwf2NLrifwrULOFwwIeoiaOAMg6BKohEbZXsjdQ31MC3u/L5boCuWJbWlMVa42kHGwanvv o2HFdOGuRda9asN/0QtKIOrqFuCj1NlZFi1a7cvN6+TTvSydF0KlYmjsGM8jDzml5UgMHKKi0Jxkh /VXPgvDGCr/7Ct834uNFFnIwM9mO/Y3pPKmoqBNwgOaUDNv0qu8Rw5rvR7UXyo2t6AIWnZyjOkwrd rGIMxfn864pY8WfYWXVrLEIf+gqdKBgcK4te6fJI6RQb7lYWTbZ3GA6XC4OxzO7SY03SWIhK92l9f 6RVHcHAQ==; Received: from authenticated-user by stravinsky.debian.org with esmtpsa (TLS1.3:ECDHE_SECP256R1__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim 4.96) (envelope-from ) id 1wRwCJ-002Wty-1I; Tue, 26 May 2026 18:08:11 +0000 From: Breno Leitao Date: Tue, 26 May 2026 14:08:06 -0400 Subject: [PATCH 2/2] workqueue: defer wake_up_process() outside pool->lock on hot paths Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit Message-Id: <20260526-fastwake-v1-2-e69ad86923e6@debian.org> References: <20260526-fastwake-v1-0-e69ad86923e6@debian.org> In-Reply-To: <20260526-fastwake-v1-0-e69ad86923e6@debian.org> To: Tejun Heo , Lai Jiangshan Cc: linux-kernel@vger.kernel.org, marco.crivellari@suse.com, frederic@kernel.org, bigeasy@linutronix.de, Breno Leitao , kernel-team@meta.com X-Mailer: b4 0.16-dev-d5d98 X-Developer-Signature: v=1; a=openpgp-sha256; l=4350; i=leitao@debian.org; h=from:subject:message-id; bh=OHFpmGg6eK2vffd/eDnZJyfgvVD70ayi3QDPz+fZYSo=; b=owEBbQKS/ZANAwAKATWjk5/8eHdtAcsmYgBqFeGHp2tCYDsdnUfbqKDq18COGd6ap/Qs5n5v3 Xk0EvJHODmJAjMEAAEKAB0WIQSshTmm6PRnAspKQ5s1o5Of/Hh3bQUCahXhhwAKCRA1o5Of/Hh3 bbVKD/97brpwlAm+HuM8f7tITpHvHr75pPebuGC8YQLZeHXAkda77TMMTyQsr6lQIOrOgWTDhMs r42ZkZCba7AchaBwDzoL0G4bpgP315eN4VrhZXGUzYx+1372lmaUIg4Zj6uT7XxUloz65nXpjBD uJI3vnaHBmT7zhwqq8LRBqlMpU9qAmT3sqwyVjMaP9ALgnCVflYVZoSd07oriZ43qocBl2Bf1m5 dL6rphcOGBHVFvEn91ACU5GvIAtQ/GfRnrW/4sTEypmrA2LWCLsgzurvYWGROekeow5yOWVEsvk hqUY8FqciOIIYBe3ZhG0vFejA54gTDAB0uUtkQvqF6lTyUK+CcIHho6LpdmoXRDiVH/S1ICwFrB /eQhNAr0Lfk6tD2LI/teOZ5hmFnyTbytXQhOtAIfIbNFMbK5mBPoMqKloaYZuJsyZa6AaE9aJ9/ bXiLGwyg6xGFM0ckfQ85BsVmAYorNb8elOOYotUDhDIG1bIBJlok3r29G+e2/rGa87pgiAvQZI7 FGkY3p9p6bwS+NKri85QygWx+v130wdfTHZmpxtFEziL/bBb0kYzZyN8H2Ow4j4pIeGNYbjYhvU E1ytnLi563kBEwR73omi9PQRIursSe6eazOMNZE8yHjKECye80QKY3J8FocEOGj8FZ71gCJhkI2 mv6jRtmpRJCfUNQ== X-Developer-Key: i=leitao@debian.org; a=openpgp; fpr=AC8539A6E8F46702CA4A439B35A3939FFC78776D X-Debian-User: leitao Both __queue_work() (enqueue) and process_one_work() (per-work chain kick on unbound/CPU_INTENSIVE pools) call kick_pool() while holding pool->lock. kick_pool() ends in wake_up_process(), which takes the target task's rq->lock. Holding pool->lock across that runqueue lock acquisition lengthens the locked region on the two hottest paths of a contended unbound workqueue. Use the new kick_pool_pick() helper to select the worker to wake while holding pool->lock, then call wake_up_process() after pool->lock is released. All state that requires pool->lock (worker selection, wake_cpu adjustment, BH-pool fast path) is still done under the lock; only the unrelated rq->lock acquisition is moved out. Measured on a CONFIG_SMP arm64 VM (8 vCPUs) with the test_workqueue benchmark (lib/test_workqueue.c) using a batched-submit mode (8 producer kthreads, 200000 work items each, WQ_UNBOUND). Averages of five runs per scope: affinity_scope baseline (items/s) patched (items/s) gain -------------- ------------------ ----------------- ---- cpu 1,419,973 1,413,896 -0.4 % (no contention) smt 1,442,921 1,437,164 -0.4 % (no contention) cache_shard 1,184,058 1,279,184 +8.0 % cache 1,167,603 1,271,341 +8.9 % numa 1,163,617 1,285,427 +10.5 % system 1,175,933 1,255,227 +6.7 % Enqueue latency on the contended scopes also drops (p50 ~2875 -> ~2625 ns, p99 ~5000 -> ~4200 ns). The cpu/smt scopes use per-CPU pools with no producer/consumer contention, so as expected they are unchanged. Signed-off-by: Breno Leitao --- kernel/workqueue.c | 21 +++++++++++++++++++-- 1 file changed, 19 insertions(+), 2 deletions(-) diff --git a/kernel/workqueue.c b/kernel/workqueue.c index b788d7c44ac0..1403a4b195a3 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -2301,6 +2301,7 @@ static void __queue_work(int cpu, struct workqueue_struct *wq, { struct pool_workqueue *pwq; struct worker_pool *last_pool, *pool; + struct task_struct *wake_p = NULL; unsigned int work_flags; unsigned int req_cpu = cpu; @@ -2415,7 +2416,7 @@ static void __queue_work(int cpu, struct workqueue_struct *wq, trace_workqueue_activate_work(work); insert_work(pwq, work, &pool->worklist, work_flags); - kick_pool(pool); + wake_p = kick_pool_pick(pool); } else { work_flags |= WORK_STRUCT_INACTIVE; insert_work(pwq, work, &pwq->inactive_works, work_flags); @@ -2423,6 +2424,15 @@ static void __queue_work(int cpu, struct workqueue_struct *wq, out: raw_spin_unlock(&pool->lock); + /* + * Issue the wakeup after dropping pool->lock to shorten the + * locked region on this hot enqueue path. kick_pool_pick() did all + * of the work that required the lock (worker selection and + * wake_cpu setup); the wake_up_process() itself only needs to + * take the target rq->lock. + */ + if (wake_p) + wake_up_process(wake_p); rcu_read_unlock(); } @@ -3243,6 +3253,7 @@ __acquires(&pool->lock) { struct pool_workqueue *pwq = get_work_pwq(work); struct worker_pool *pool = worker->pool; + struct task_struct *wake_p; unsigned long work_data; int lockdep_start_depth, rcu_start_depth; bool bh_draining = pool->flags & POOL_BH_DRAINING; @@ -3296,8 +3307,11 @@ __acquires(&pool->lock) * since nr_running would always be >= 1 at this point. This is used to * chain execution of the pending work items for WORKER_NOT_RUNNING * workers such as the UNBOUND and CPU_INTENSIVE ones. + * + * Select the worker to wake while holding pool->lock, but defer the + * actual wake_up_process() until after the lock is dropped below. */ - kick_pool(pool); + wake_p = kick_pool_pick(pool); /* * Record the last pool and clear PENDING which should be the last @@ -3310,6 +3324,9 @@ __acquires(&pool->lock) pwq->stats[PWQ_STAT_STARTED]++; raw_spin_unlock_irq(&pool->lock); + if (wake_p) + wake_up_process(wake_p); + rcu_start_depth = rcu_preempt_depth(); lockdep_start_depth = lockdep_depth(current); /* see drain_dead_softirq_workfn() */ -- 2.51.0