mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH wq/for-7.4 v3 0/3] workqueue: Keep max_active and percpu_max_active in sync
@ 2026-10-09 12:18 Breno Leitao
  2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 1/3] workqueue: Move the initial max_active setup into a helper Breno Leitao
                   ` (2 more replies)
  0 siblings, 3 replies; 4+ messages in thread
From: Breno Leitao @ 2026-10-09 12:18 UTC (permalink / raw)
  To: Tejun Heo, Lai Jiangshan
  Cc: linux-kernel, marco.crivellari, Breno Leitao, kernel-team

Commit 27db9dd7f84f3a ("workqueue: Give percpu workqueues their own
max_active") gave the percpu backend a limit of its own, so a workqueue
now carries both percpu_max_active and max_active.

__alloc_workqueue() only ever sets the pair matching the backend the
workqueue is created on, so the other stays at zero for its lifetime.

Nothing reads "the other scope" max_active today, so this is groundwork
rather than a fix.

Three patches. The first moves the initial limit setup out of
__alloc_workqueue() into wq_init_max_active().

The second sets both domains rather than one, according to tejun's
recommendation:

  U = unbound max_active, the system-wide target
  L = unbound min_active, the floor on each node
  P = percpu_max_active, the limit on each CPU
  N = online CPUs in the effective unbound mask, at least 1

  Setting P:
    U = min(P * N, WQ_MAX_ACTIVE)
    L = P

  Setting U:
    L = min(L, U)
    P = max(DIV_ROUND_UP(U, N), L)

  Setting L:
    L = clamp(L, 0, U)
    P = max(DIV_ROUND_UP(U, N), L)

The third rederives both domains once CPU topology is known. Workqueues
created before smp_init() brings up the secondary CPUs compute N as 1,
and nothing revisits that afterwards.

This shouldn't change the behaviour of workqueue for now, but keeps the
groundwork correct for the PERCPU affinity scope work it's meant to
enable.

Signed-off-by: Breno Leitao <leitao@debian.org>
---
Changes in v3:
- Dropped the silly percpu_concurrency_managed attribute patch.
- Fixed N to come from the workqueue's own unbound_effective_cpumask()
- Added a patch rederiving both domains once workqueue_init_topology()
  knows the real CPU count.
- Handled BH in one place: wq_init_max_active() sets the limits of a BH
  workqueue through wq_init_bh_max_active() and returns, instead of
  giving BH its own branch further down (Tejun)
- Set min_active to percpu_max_active when deriving from a percpu limit
  instead of min(percpu_max_active, max_active), as max_active can never
  be lower (Tejun)
- Changelog of patch 2 now refers to the new PERCPU affinity scope
  rather than WQ_AFFN_CPU (Tejun)
- Link to v2: https://patch.msgid.link/20260925-wq_final-v2-0-860ee052169e@debian.org

Changes in v2:
- Cut the series down to the three patches that change nothing
  observable. The runtime scaling, the affinity scopes, the backend
  rework, the concurrency management attribute and the sysfs files are
  all held back.
- Set both max_active domains instead of holding them equal, deriving
  one from the other rather than duplicating the value (Tejun)
- Link to v1:
  https://patch.msgid.link/20260918-wq_final-v1-0-5c43c08a26bc@debian.org

---
Breno Leitao (3):
      workqueue: Move the initial max_active setup into a helper
      workqueue: Keep the limits of both max_active domains current
      workqueue: Rescale max_active domains once CPU topology is known

 kernel/workqueue.c | 168 ++++++++++++++++++++++++++++++++++++++++-------------
 1 file changed, 129 insertions(+), 39 deletions(-)
---
base-commit: 9d4815c14f7faf789aeaa63515024168daf0390c
change-id: 20260914-wq_final-bcfbcd3b0f79

Best regards,
--  
Breno Leitao <leitao@debian.org>


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH wq/for-7.4 v3 1/3] workqueue: Move the initial max_active setup into a helper
  2026-10-09 12:18 [PATCH wq/for-7.4 v3 0/3] workqueue: Keep max_active and percpu_max_active in sync Breno Leitao
@ 2026-10-09 12:18 ` Breno Leitao
  2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 2/3] workqueue: Keep the limits of both max_active domains current Breno Leitao
  2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 3/3] workqueue: Rescale max_active domains once CPU topology is known Breno Leitao
  2 siblings, 0 replies; 4+ messages in thread
From: Breno Leitao @ 2026-10-09 12:18 UTC (permalink / raw)
  To: Tejun Heo, Lai Jiangshan
  Cc: linux-kernel, marco.crivellari, Breno Leitao, kernel-team

Extract the max/min active assignment into a helper, simplifying
__alloc_workqueue() and keeping the next patch easier to review and reason
about.

PS: Eventually, wq_init_max_active() and wq_adjust_max_active() can share
some code, but, keep this as a simple code movement for now, and then we
can come back to it.

Signed-off-by: Breno Leitao <leitao@debian.org>
---
 kernel/workqueue.c | 46 +++++++++++++++++++++++++++-------------------
 1 file changed, 27 insertions(+), 19 deletions(-)

diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index e618108c6127da..97d2c9c0a7d6ab 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -6031,6 +6031,32 @@ static int init_rescuer(struct workqueue_struct *wq)
 	return 0;
 }
 
+static void wq_init_max_active(struct workqueue_struct *wq, int max_active)
+{
+	int effective_max_active;
+
+	if (wq->flags & WQ_BH) {
+		/*
+		 * BH workqueues always share a single execution context per CPU
+		 * and don't impose any max_active limit.
+		 */
+		effective_max_active = INT_MAX;
+	} else {
+		effective_max_active = max_active ?: WQ_DFL_ACTIVE;
+		effective_max_active = wq_clamp_max_active(effective_max_active,
+							   wq->flags, wq->name);
+	}
+
+	if (wq->flags & WQ_UNBOUND) {
+		wq->max_active = effective_max_active;
+		wq->min_active = min(effective_max_active, WQ_DFL_MIN_ACTIVE);
+		wq->saved_min_active = wq->min_active;
+	} else {
+		wq->percpu_max_active = effective_max_active;
+	}
+	wq->saved_max_active = effective_max_active;
+}
+
 /**
  * wq_adjust_max_active - update a wq's max_active to the current setting
  * @wq: target workqueue
@@ -6156,27 +6182,9 @@ static struct workqueue_struct *__alloc_workqueue(const char *fmt,
 		flags &= ~WQ_PERCPU;
 	}
 
-	if (flags & WQ_BH) {
-		/*
-		 * BH workqueues always share a single execution context per CPU
-		 * and don't impose any max_active limit.
-		 */
-		max_active = INT_MAX;
-	} else {
-		max_active = max_active ?: WQ_DFL_ACTIVE;
-		max_active = wq_clamp_max_active(max_active, flags, wq->name);
-	}
-
 	/* init wq */
 	wq->flags = flags;
-	if (flags & WQ_UNBOUND) {
-		wq->max_active = max_active;
-		wq->min_active = min(max_active, WQ_DFL_MIN_ACTIVE);
-		wq->saved_min_active = wq->min_active;
-	} else {
-		wq->percpu_max_active = max_active;
-	}
-	wq->saved_max_active = max_active;
+	wq_init_max_active(wq, max_active);
 	mutex_init(&wq->mutex);
 	atomic_set(&wq->nr_pwqs_to_flush, 0);
 	INIT_LIST_HEAD(&wq->pwqs);

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH wq/for-7.4 v3 2/3] workqueue: Keep the limits of both max_active domains current
  2026-10-09 12:18 [PATCH wq/for-7.4 v3 0/3] workqueue: Keep max_active and percpu_max_active in sync Breno Leitao
  2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 1/3] workqueue: Move the initial max_active setup into a helper Breno Leitao
@ 2026-10-09 12:18 ` Breno Leitao
  2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 3/3] workqueue: Rescale max_active domains once CPU topology is known Breno Leitao
  2 siblings, 0 replies; 4+ messages in thread
From: Breno Leitao @ 2026-10-09 12:18 UTC (permalink / raw)
  To: Tejun Heo, Lai Jiangshan
  Cc: linux-kernel, marco.crivellari, Breno Leitao, kernel-team

Commit 27db9dd7f84f3a ("workqueue: Give percpu workqueues their own
max_active") gave the percpu backend a limit of its own, so a workqueue
now carries limits for both scopes, but just one is used today.

Given that we are introducing a new PERCPU affinity scope which will be
able to do CM, we want to keep both values set, so, the transition from
scopes would be able to succeed. This cannot be done under the flip, given
this is RCU protected. See discussion at [1]

The two numbers cannot simply be kept equal. percpu_max_active caps every
CPU on its own, while max_active is a workqueue wide budget distributed
among the nodes with min_active as the per-node floor. A per-CPU limit of
16 on a 64 CPU machine is not a workqueue wide limit of 16.

Use the scale function suggested by Tejun in [1].

N is the online CPUs the unbound backend of the workqueue may actually
use. For a workqueue whose own cpumask has been narrowed, that is
unbound_effective_cpumask(), not the global wq_unbound_cpumask -- the
latter overcounts and makes percpu_max_active come out too low.
wq_update_node_max_active() already distributes max_active the same way;
match it.

Suggested-by: Tejun Heo <tj@kernel.org>
Link: https://lore.kernel.org/all/ec6c22eddc8ce0b1ef2434ed62acd560@kernel.org/ [1]
Signed-off-by: Breno Leitao <leitao@debian.org>
---
 kernel/workqueue.c | 126 +++++++++++++++++++++++++++++++++++++++++------------
 1 file changed, 99 insertions(+), 27 deletions(-)

diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index 97d2c9c0a7d6ab..244a199f2dd61a 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -393,6 +393,7 @@ struct workqueue_struct {
 	int			min_active;	/* WO: min active works */
 	int			saved_max_active; /* WQ: saved max_active */
 	int			saved_min_active; /* WQ: saved min_active */
+	int			saved_percpu_max_active; /* WQ: percpu limit */
 
 	struct workqueue_attrs	*attrs;	/* PW: workqueue attributes */
 	struct pool_workqueue __rcu *dfl_pwq;   /* PW: only for unbound wqs */
@@ -421,6 +422,14 @@ struct workqueue_struct {
 	struct wq_node_nr_active *node_nr_active[]; /* I: per-node nr_active */
 };
 
+static int wq_user_max_active(struct workqueue_struct *wq)
+{
+	if (wq->flags & WQ_UNBOUND)
+		return READ_ONCE(wq->saved_max_active);
+
+	return READ_ONCE(wq->saved_percpu_max_active);
+}
+
 /*
  * Each pod type describes how CPUs should be grouped for unbound workqueues.
  * See the comment above workqueue_attrs->affn_scope.
@@ -4549,7 +4558,7 @@ static bool start_flush_work(struct work_struct *work, struct wq_barrier *barr,
 	 * workqueues the deadlock happens when the rescuer stalls, blocking
 	 * forward progress.
 	 */
-	if (!from_cancel && (wq->saved_max_active == 1 || wq->rescuer))
+	if (!from_cancel && (wq_user_max_active(wq) == 1 || wq->rescuer))
 		touch_wq_lockdep_map(wq);
 
 	rcu_read_unlock();
@@ -6031,74 +6040,121 @@ static int init_rescuer(struct workqueue_struct *wq)
 	return 0;
 }
 
+static int wq_nr_online_cpus(const struct cpumask *mask)
+{
+	return max(cpumask_weight_and(mask, cpu_online_mask), 1);
+}
+
+static int wq_nr_effective_cpus(struct workqueue_struct *wq)
+{
+	lockdep_assert_held(&wq->mutex);
+
+	if (!installed_pwq(wq, -1))
+		return wq_nr_online_cpus(wq_unbound_cpumask);
+
+	return wq_nr_online_cpus(unbound_effective_cpumask(wq));
+}
+
+static int wq_scale_to_unbound(int percpu_max, int nr_cpus)
+{
+	return min_t(s64, (s64)percpu_max * nr_cpus, WQ_MAX_ACTIVE);
+}
+
+static int wq_scale_to_percpu(int max_active, int min_active, int nr_cpus)
+{
+	return max(DIV_ROUND_UP(max_active, nr_cpus), min_active);
+}
+
+static void wq_init_bh_max_active(struct workqueue_struct *wq)
+{
+	wq->max_active = INT_MAX;
+	wq->min_active = INT_MAX;
+	wq->percpu_max_active = INT_MAX;
+	wq->saved_max_active = INT_MAX;
+	wq->saved_min_active = INT_MAX;
+	wq->saved_percpu_max_active = INT_MAX;
+}
+
 static void wq_init_max_active(struct workqueue_struct *wq, int max_active)
 {
 	int effective_max_active;
+	int nr_cpus = wq_nr_online_cpus(wq_unbound_cpumask);
 
 	if (wq->flags & WQ_BH) {
 		/*
 		 * BH workqueues always share a single execution context per CPU
 		 * and don't impose any max_active limit.
 		 */
-		effective_max_active = INT_MAX;
-	} else {
-		effective_max_active = max_active ?: WQ_DFL_ACTIVE;
-		effective_max_active = wq_clamp_max_active(effective_max_active,
-							   wq->flags, wq->name);
+		wq_init_bh_max_active(wq);
+		return;
 	}
 
+	effective_max_active = max_active ?: WQ_DFL_ACTIVE;
+	effective_max_active = wq_clamp_max_active(effective_max_active,
+						   wq->flags, wq->name);
+
 	if (wq->flags & WQ_UNBOUND) {
 		wq->max_active = effective_max_active;
 		wq->min_active = min(effective_max_active, WQ_DFL_MIN_ACTIVE);
-		wq->saved_min_active = wq->min_active;
+		wq->percpu_max_active =
+			wq_scale_to_percpu(effective_max_active, wq->min_active,
+					   nr_cpus);
 	} else {
 		wq->percpu_max_active = effective_max_active;
+		wq->max_active =
+			wq_scale_to_unbound(effective_max_active, nr_cpus);
+		wq->min_active = effective_max_active;
 	}
-	wq->saved_max_active = effective_max_active;
+
+	wq->saved_max_active = wq->max_active;
+	wq->saved_min_active = wq->min_active;
+	wq->saved_percpu_max_active = wq->percpu_max_active;
 }
 
 /**
  * wq_adjust_max_active - update a wq's max_active to the current setting
  * @wq: target workqueue
  *
- * If @wq isn't freezing, set the limit that applies to @wq's backing to the
- * saved_max_active and activate inactive work items accordingly. If @wq is
- * freezing, clear it to zero.
+ * If @wq isn't freezing, publish the saved limits of both domains and activate
+ * inactive work items accordingly. If @wq is freezing, clear them to zero.
  */
 static void wq_adjust_max_active(struct workqueue_struct *wq)
 {
+	int new_max, new_min, new_percpu_max;
 	bool activated;
-	int new_max, new_min;
 
 	lockdep_assert_held(&wq->mutex);
 
 	if ((wq->flags & WQ_FREEZABLE) && workqueue_freezing) {
 		new_max = 0;
 		new_min = 0;
+		new_percpu_max = 0;
 	} else {
 		new_max = wq->saved_max_active;
 		new_min = wq->saved_min_active;
+		new_percpu_max = wq->saved_percpu_max_active;
 	}
 
 	/*
-	 * Update the limit and then kick inactive work items if more active
+	 * Update both domains and then kick inactive work items if more active
 	 * work items are allowed. This doesn't break work item ordering
 	 * because new work items are always queued behind existing inactive
 	 * work items if there are any.
+	 *
+	 * Which domain a pwq honours follows its pool, see
+	 * pwq_tryinc_nr_active(). Keeping both current means a pwq is never
+	 * metered against a limit nobody set.
 	 */
-	if (wq->flags & WQ_UNBOUND) {
-		if (wq->max_active == new_max && wq->min_active == new_min)
-			return;
+	if (wq->max_active == new_max && wq->min_active == new_min &&
+	    wq->percpu_max_active == new_percpu_max)
+		return;
 
-		WRITE_ONCE(wq->max_active, new_max);
-		WRITE_ONCE(wq->min_active, new_min);
-		wq_update_node_max_active(wq, -1);
-	} else {
-		if (wq->percpu_max_active == new_max)
-			return;
+	WRITE_ONCE(wq->max_active, new_max);
+	WRITE_ONCE(wq->min_active, new_min);
+	WRITE_ONCE(wq->percpu_max_active, new_percpu_max);
 
-		WRITE_ONCE(wq->percpu_max_active, new_max);
-	}
+	if (wq->flags & WQ_UNBOUND)
+		wq_update_node_max_active(wq, -1);
 
 	if (new_max == 0)
 		return;
@@ -6448,6 +6504,8 @@ EXPORT_SYMBOL_GPL(destroy_workqueue);
  */
 void workqueue_set_max_active(struct workqueue_struct *wq, int max_active)
 {
+	int nr_cpus;
+
 	/* max_active doesn't mean anything for BH workqueues */
 	if (WARN_ON(wq->flags & WQ_BH))
 		return;
@@ -6459,9 +6517,20 @@ void workqueue_set_max_active(struct workqueue_struct *wq, int max_active)
 
 	mutex_lock(&wq->mutex);
 
-	wq->saved_max_active = max_active;
-	if (wq->flags & WQ_UNBOUND)
+	nr_cpus = wq_nr_effective_cpus(wq);
+
+	if (wq->flags & WQ_UNBOUND) {
+		wq->saved_max_active = max_active;
 		wq->saved_min_active = min(wq->saved_min_active, max_active);
+		wq->saved_percpu_max_active =
+			wq_scale_to_percpu(max_active, wq->saved_min_active,
+					   nr_cpus);
+	} else {
+		wq->saved_percpu_max_active = max_active;
+		wq->saved_max_active = wq_scale_to_unbound(max_active,
+							   nr_cpus);
+		wq->saved_min_active = max_active;
+	}
 
 	wq_adjust_max_active(wq);
 
@@ -6492,6 +6561,9 @@ void workqueue_set_min_active(struct workqueue_struct *wq, int min_active)
 
 	mutex_lock(&wq->mutex);
 	wq->saved_min_active = clamp(min_active, 0, wq->saved_max_active);
+	wq->saved_percpu_max_active =
+		wq_scale_to_percpu(wq->saved_max_active, wq->saved_min_active,
+				   wq_nr_effective_cpus(wq));
 	wq_adjust_max_active(wq);
 	mutex_unlock(&wq->mutex);
 }
@@ -7563,7 +7635,7 @@ static ssize_t max_active_show(struct device *dev,
 {
 	struct workqueue_struct *wq = dev_to_wq(dev);
 
-	return scnprintf(buf, PAGE_SIZE, "%d\n", wq->saved_max_active);
+	return scnprintf(buf, PAGE_SIZE, "%d\n", wq_user_max_active(wq));
 }
 
 static ssize_t max_active_store(struct device *dev,

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 4+ messages in thread

* [PATCH wq/for-7.4 v3 3/3] workqueue: Rescale max_active domains once CPU topology is known
  2026-10-09 12:18 [PATCH wq/for-7.4 v3 0/3] workqueue: Keep max_active and percpu_max_active in sync Breno Leitao
  2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 1/3] workqueue: Move the initial max_active setup into a helper Breno Leitao
  2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 2/3] workqueue: Keep the limits of both max_active domains current Breno Leitao
@ 2026-10-09 12:18 ` Breno Leitao
  2 siblings, 0 replies; 4+ messages in thread
From: Breno Leitao @ 2026-10-09 12:18 UTC (permalink / raw)
  To: Tejun Heo, Lai Jiangshan
  Cc: linux-kernel, marco.crivellari, Breno Leitao, kernel-team

wq_init_max_active() derives the domain a workqueue's backend doesn't use
(percpu_max_active for an unbound workqueue, max_active for a percpu one)
from the number of online CPUs at creation time. Workqueues created by
workqueue_init_early() or an early initcall run before smp_init() brings
up the secondary CPUs, so that count is 1 for them, and nothing
recomputes it afterwards.

Redo wq_init_max_active() once workqueue_init_topology() has the real
CPU count, in the same pass that already walks every workqueue to update
per-node max_active. With one CPU online both domains were equal, so
saved_max_active is still the max_active the workqueue was created with
and can be passed back as is.

Signed-off-by: Breno Leitao <leitao@debian.org>
---
 kernel/workqueue.c | 10 ++++++++++
 1 file changed, 10 insertions(+)

diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index 244a199f2dd61a..dd52a7d0999d17 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -8870,6 +8870,16 @@ void __init workqueue_init_topology(void)
 	list_for_each_entry(wq, &workqueues, list) {
 		for_each_online_cpu(cpu)
 			unbound_wq_update_pwq(wq, cpu);
+
+		/*
+		 * wq_init_max_active() ran before smp_init() for these, so
+		 * saved_max_active is still their max_active argument. Redo it
+		 * now that all CPUs are online.
+		 */
+		mutex_lock(&wq->mutex);
+		wq_init_max_active(wq, wq->saved_max_active);
+		mutex_unlock(&wq->mutex);
+
 		if (wq->flags & WQ_UNBOUND) {
 			mutex_lock(&wq->mutex);
 			wq_update_node_max_active(wq, -1);

-- 
2.53.0-Meta


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-09 12:22 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-09 12:18 [PATCH wq/for-7.4 v3 0/3] workqueue: Keep max_active and percpu_max_active in sync Breno Leitao
2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 1/3] workqueue: Move the initial max_active setup into a helper Breno Leitao
2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 2/3] workqueue: Keep the limits of both max_active domains current Breno Leitao
2026-10-09 12:18 ` [PATCH wq/for-7.4 v3 3/3] workqueue: Rescale max_active domains once CPU topology is known Breno Leitao

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®