From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from stravinsky.debian.org (stravinsky.debian.org [82.195.75.108]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 22FB22D249B for ; Wed, 24 Jun 2026 11:47:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=82.195.75.108 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782301682; cv=none; b=BRebOiH9ifTImoq9KYNeXGSMnWQCa+zodSelfek55DKdc/zpzmEd+PYrL6hndhJ7JpaHkx/pwwp1NcFL4J4xRxIFcIY2eY0e31ghYJetbwslT+V3Z/N6ZPsKndwE9NXQPl78QcscDXt1+KUALRrckvznJ08BacEL8OIEz5UrPIU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1782301682; c=relaxed/simple; bh=ZB+/SGtaP4oIekVdW4IuE97mmm4U2LgnSOtFc58j5Po=; h=From:Subject:Date:Message-Id:MIME-Version:Content-Type:To:Cc; b=C0fnE1zUdgIFB6f1IJC31TLjn7V5z/CvwxVn7wb6IuXBto/bDmAKe4Wz53MZ4H7oOD02Y7meHrjHLkmeGTu53aeaD4uYvjEix16OdPtvWPc8hNQtTAFGpybr6o97ld4plyRyo257JsVw+yVJGZf86VfPQRcrIcpEWZtq5EDVask= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org; spf=pass smtp.mailfrom=debian.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b=Z8N07PRG; arc=none smtp.client-ip=82.195.75.108 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=debian.org Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=debian.org Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=debian.org header.i=@debian.org header.b="Z8N07PRG" DKIM-Signature: v=1; a=rsa-sha256; q=dns/txt; c=relaxed/relaxed; d=debian.org; s=smtpauto.stravinsky; h=X-Debian-User:Cc:To:Content-Transfer-Encoding: Content-Type:MIME-Version:Message-Id:Date:Subject:From:Reply-To:Content-ID: Content-Description:In-Reply-To:References; bh=2VOCY2PN3cFckRg+bwmXwQVLu0bQk8GN1FwzZZx56c0=; b=Z8N07PRGFfXx7VPO6C4Ybg20o+ zaMC4AD5wClacF7ypTj4quIMm8MaZD16NO5ys6cV+F0R7/qRWG4KCncKjVuY6s3ZaNL9WY9mt3qOt 7jcnDjVwB0fEifv69jKaoigkjhsAIqw3UwXKvr1TejuVz57NCdA7aDdS3rxbgzwV4iV7N8HDxJENL YHJEGwAjxGJ4M09uDUc1loBeQnMAMzP0tYIqVDHiGCR2XYRs9GyUJU8OztY2h88IXl+PkSoY8CNF9 mcJhaFO6P+GXnmiLBFqxl3IOsKVGlSPjsDRRxTa9Ruz2/pTr3dP/KUHKkFhK2LpyEPVRdf+plW2Zd h32OK1dw==; Received: from authenticated-user by stravinsky.debian.org with esmtpsa (TLS1.3:ECDHE_X25519__RSA_PSS_RSAE_SHA256__AES_256_GCM:256) (Exim 4.96) (envelope-from ) id 1wcM5A-002Oxd-2C; Wed, 24 Jun 2026 11:47:53 +0000 From: Breno Leitao Subject: [PATCH v4 0/3] workqueue: Shrink the lock time Date: Wed, 24 Jun 2026 04:47:38 -0700 Message-Id: <20260624-fastwake-v4-0-7b6d7b494a44@debian.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="utf-8" Content-Transfer-Encoding: 7bit X-B4-Tracking: v=1; b=H4sIANrDO2oC/1WOQQ6CMBREr9L8tTXlIy1l5T2Mi0J/oZqAabFqC Hc3wEJcTjLvzUwQKXiKULEJAiUf/dBDxU4HBk1n+pa4t1AxQIFSFCi5M3F8mTtxgbpEZ6XMM4Q Dg0cg59+r6nLdcnzWN2rGhV8anY/jED7rVsKlt2mlyH/ahFxw1EoVGTpSzpwt1d70xyG0sHhTv iOz3aGUc8GVtibTrrGidH/kPM9fVoD9l+0AAAA= X-Change-ID: 20260526-fastwake-02982fd66312 To: Tejun Heo , Lai Jiangshan , bigeasy@linutronix.de Cc: linux-kernel@vger.kernel.org, marco.crivellari@suse.com, frederic@kernel.org, bigeasy@linutronix.de, Hillf Danton , Breno Leitao , kernel-team@meta.com, kmagar@redhat.com, psuriset@redhat.com, david.dai@linux.dev X-Mailer: b4 0.14.3 X-Developer-Signature: v=1; a=openpgp-sha256; l=3804; i=leitao@debian.org; h=from:subject:message-id; bh=ZB+/SGtaP4oIekVdW4IuE97mmm4U2LgnSOtFc58j5Po=; b=owEBbQKS/ZANAwAIATWjk5/8eHdtAcsmYgBqO8Pko9EHuVW1pTOJU0FgWQagyRcvn2hOTQqlu UeYbSYAkzKJAjMEAAEIAB0WIQSshTmm6PRnAspKQ5s1o5Of/Hh3bQUCajvD5AAKCRA1o5Of/Hh3 bd3LD/sFcxROk8MDmOzMJ7Hwe6Fyvoglo0aJ1TuWTNcLkUrpUBszzLnfCfXrzL3/I7MyVBdHn+k exPfZrSfegB0HXYY+MLPFLOoBshmO8RhGJTILF8T059BWczGFPo6KiH5uPjXO28KCmRkKl4s1l2 KUKBqA2slzHj2xDjIazTBWTjqeR6L1mi6J7VZHdXTxUu4JyLVvc0NP6Vah3/kWNrcQqSZeitSxB jVr7mcSg+CjPdzqrEQy5sjwSP71pFgEaq0m6f6eYUh8X43eMcLMROFsr7EF+zH0fJP9FQZ4IzNy 8YEZ3uUY1LKpo8d2Mx/gMCvP50seUXM8YWlSRdpVwI4U2BkUbtoWIMjpWog0dx4dS1FQpbGyEbA ZutdsjRHJH6NjqdYFvngTYVmhsiEt/cZvYVhLvNTh7UW5OS8jjkbxU8lq7pidGCRj+twFpQAV20 LfaJCNoxB5HeCGTuQ5+ZcdAsdLLyoxbEeKFHPmbdTv9X5DutZBVzXuvMa2gac9A59fvHGLuc3AD py2J7s59ZfyLYmgPDEdsSnhwsTTDDwMcRhMUkgLvb8wSxWawZMy/ktIBkOncRai/Hi6QhS5EHX7 e5KeZFpntKm83olxtu+19hHAi+6dSmegsbJjNJzdwuF30WKPyWdRKeKbp9IHI5FSTny6oEnI8uv 7j0fV5X9IL09M6g== X-Developer-Key: i=leitao@debian.org; a=openpgp; fpr=AC8539A6E8F46702CA4A439B35A3939FFC78776D X-Debian-User: leitao The goal of this patchset is to decrease the time spent under the workqueue pool->lock. Currently the worker process is woken up inside pool->lock. The wakeup ends in wake_up_process(), which takes the target task's rq->lock, so rq->lock nests under pool->lock on the two hottest paths of a contended unbound workqueue (__queue_work() enqueue and process_one_work() chain kick). On some architectures the wakeup is even more expensive: on arm64 waking a CPU that is idle (in wfi) issues an IPI. Doing all of that while holding pool->lock lengthens the locked region and hurts throughput on contended unbound pools. This series shortens the locked region by selecting and claiming the worker to wake under pool->lock, but issuing the actual wakeup after the lock is dropped, using the wake_q machinery (wake_q_add() under the lock, wake_up_q() after). Because the win is a shorter pool->lock hold time, it shows up most clearly as lower enqueue latency under contention. Performance numbers (based on in-kernel workqueue microbenchmark) VMs and arm64 (Grace) is where this series is meant to pay off -- waking an idle CPU sitting in wfi costs an IPI (on arm; similar type of operation on VMs), so doing it under pool->lock lengthens the critical section. The arm64 bare-metal numbers match what the x86-or-arm64 VM showed: affinity_scope baseline patched tput p95 (items/s) (items/s) gain drop -------------- --------- --------- ------ ------ cpu 2,569,880 3,029,740 +17.9% -13.6% smt 2,586,485 3,044,788 +17.7% -14.0% cache_shard 572,055 797,621 +39.4% -37.1% cache 538,132 724,997 +34.7% -30.1% numa 528,673 658,215 +24.5% -20.5% system 524,287 614,486 +17.2% -21.1% (p95 drop = change in p95 enqueue latency; negative is better.) (tput gain = number of requests enqueued per sec; bigger is better.) Patch 1 is a pure refactor introducing kick_pool_pick(). Patch 2 defers the wakeup on the enqueue path (__queue_work()). Patch 3 defers the wakeup on the per-work chain-kick path (process_one_work()). Signed-off-by: Breno Leitao Changes in v4: - replace raw_spin_unlock_wake() with a standard raw_spin_unlock() + wake_up_q() (Sebastian Andrzej Siewior) - Link to v3: https://lore.kernel.org/r/20260616-fastwake-v3-0-79da19fcd08f@debian.org Changes in v3: - Drop the "park kicked worker on pool->kicked_list" patch (v2 1/4). * That is a fix that is independent of this patch, in case we want to revamp it, it can be sent separately. - Link to v2: https://lore.kernel.org/r/20260603-fastwake-v2-0-2977512fe7fa@debian.org Changes in v2: - Close the idle_cull_fn() vs kicked-worker race by parking the kicked worker on a new pool->kicked_list under pool->lock (new patch 1). Reported by Hillf Danton. - Use the wake_q machinery (wake_q_add() / wake_up_q() via raw_spin_unlock_wake()) instead of plumbing a task_struct out of the helper by hand. Suggested by Sebastian Andrzej Siewior. - Link to v1: https://lore.kernel.org/r/20260526-fastwake-v1-0-e69ad86923e6@debian.org --- Breno Leitao (3): workqueue: split kick_pool() into kick_pool_pick() + wake_up_q() workqueue: defer the worker wakeup outside pool->lock in __queue_work() workqueue: defer the worker wakeup outside pool->lock in process_one_work() kernel/workqueue.c | 40 +++++++++++++++++++++++++++++++++------- 1 file changed, 33 insertions(+), 7 deletions(-) --- base-commit: 8d6dbbbe3ba62de0a63e962ee004afb848c8e3ac change-id: 20260526-fastwake-02982fd66312 Best regards, -- Breno Leitao