* [PATCH wq/for-7.4 v2 0/3] workqueue: Keep max_active and percpu_max_active in sync
@ 2026-09-25 13:09 Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 1/3] workqueue: Move the initial max_active setup into a helper Breno Leitao
` (2 more replies)
0 siblings, 3 replies; 8+ messages in thread
From: Breno Leitao @ 2026-09-25 13:09 UTC (permalink / raw)
To: Tejun Heo, Lai Jiangshan
Cc: linux-kernel, marco.crivellari, Breno Leitao, kernel-team
Commit 27db9dd7f84f3a ("workqueue: Give percpu workqueues their own
max_active") gave the percpu backend a limit of its own, so a workqueue
now carries both percpu_max_active and max cpu.
__alloc_workqueue() only ever sets the pair matching the backend the
workqueue is created on, so the other stays at zero for its lifetime.
Nothing reads "the other scope" max_active today, so this is groundwork
rather than a fix
Two patches. The first moves the initial limit setup out of
__alloc_workqueue() into wq_init_max_active()
The second sets both domains rather than one, according to tejun's
recommendation:
U = unbound max_active, the system-wide target
L = unbound min_active, the floor on each node
P = percpu_max_active, the limit on each CPU
N = online CPUs in the effective unbound mask, at least 1
Setting P:
U = min(P * N, WQ_MAX_ACTIVE)
L = P
Setting U:
L = min(L, U)
P = max(DIV_ROUND_UP(U, N), L)
Setting L:
L = clamp(L, 0, U)
P = max(DIV_ROUND_UP(U, N), L)
Also, add the flag for CM in the workqueue attributes, as discussed on
v1. This will help will enable the introduction of PERCPU affinity scope
later.
This shouldn't change the behaviour of workqueue for now, but, enable
the new PERCPU introduction.
Signed-off-by: Breno Leitao <leitao@debian.org>
---
Changes in v2:
- Cut the series down to the three patches that change nothing
observable. The runtime scaling, the affinity scopes, the backend
rework, the concurrency management attribute and the sysfs files are
all held back.
- Set both max_active domains instead of holding them equal, deriving
one from the other rather than duplicating the value (Tejun)
- Link to v1:
https://patch.msgid.link/20260918-wq_final-v1-0-5c43c08a26bc@debian.org
---
Breno Leitao (3):
workqueue: Move the initial max_active setup into a helper
workqueue: Keep the limits of both max_active domains current
workqueue: add a percpu_concurrency_managed workqueue attribute
include/linux/workqueue.h | 9 +++
kernel/workqueue.c | 173 ++++++++++++++++++++++++++++++++--------------
2 files changed, 132 insertions(+), 50 deletions(-)
---
base-commit: 9d4815c14f7faf789aeaa63515024168daf0390c
change-id: 20260914-wq_final-bcfbcd3b0f79
Best regards,
--
Breno Leitao <leitao@debian.org>
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH wq/for-7.4 v2 1/3] workqueue: Move the initial max_active setup into a helper
2026-09-25 13:09 [PATCH wq/for-7.4 v2 0/3] workqueue: Keep max_active and percpu_max_active in sync Breno Leitao
@ 2026-09-25 13:09 ` Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 2/3] workqueue: Keep the limits of both max_active domains current Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 3/3] workqueue: add a percpu_concurrency_managed workqueue attribute Breno Leitao
2 siblings, 0 replies; 8+ messages in thread
From: Breno Leitao @ 2026-09-25 13:09 UTC (permalink / raw)
To: Tejun Heo, Lai Jiangshan
Cc: linux-kernel, marco.crivellari, Breno Leitao, kernel-team
Extract the max/min active assignment into a helper, simplifying
__alloc_workqueue() and keeping the next patch easier to review and reason
about.
Signed-off-by: Breno Leitao <leitao@debian.org>
---
kernel/workqueue.c | 46 +++++++++++++++++++++++++++-------------------
1 file changed, 27 insertions(+), 19 deletions(-)
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index e618108c6127da..97d2c9c0a7d6ab 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -6031,6 +6031,32 @@ static int init_rescuer(struct workqueue_struct *wq)
return 0;
}
+static void wq_init_max_active(struct workqueue_struct *wq, int max_active)
+{
+ int effective_max_active;
+
+ if (wq->flags & WQ_BH) {
+ /*
+ * BH workqueues always share a single execution context per CPU
+ * and don't impose any max_active limit.
+ */
+ effective_max_active = INT_MAX;
+ } else {
+ effective_max_active = max_active ?: WQ_DFL_ACTIVE;
+ effective_max_active = wq_clamp_max_active(effective_max_active,
+ wq->flags, wq->name);
+ }
+
+ if (wq->flags & WQ_UNBOUND) {
+ wq->max_active = effective_max_active;
+ wq->min_active = min(effective_max_active, WQ_DFL_MIN_ACTIVE);
+ wq->saved_min_active = wq->min_active;
+ } else {
+ wq->percpu_max_active = effective_max_active;
+ }
+ wq->saved_max_active = effective_max_active;
+}
+
/**
* wq_adjust_max_active - update a wq's max_active to the current setting
* @wq: target workqueue
@@ -6156,27 +6182,9 @@ static struct workqueue_struct *__alloc_workqueue(const char *fmt,
flags &= ~WQ_PERCPU;
}
- if (flags & WQ_BH) {
- /*
- * BH workqueues always share a single execution context per CPU
- * and don't impose any max_active limit.
- */
- max_active = INT_MAX;
- } else {
- max_active = max_active ?: WQ_DFL_ACTIVE;
- max_active = wq_clamp_max_active(max_active, flags, wq->name);
- }
-
/* init wq */
wq->flags = flags;
- if (flags & WQ_UNBOUND) {
- wq->max_active = max_active;
- wq->min_active = min(max_active, WQ_DFL_MIN_ACTIVE);
- wq->saved_min_active = wq->min_active;
- } else {
- wq->percpu_max_active = max_active;
- }
- wq->saved_max_active = max_active;
+ wq_init_max_active(wq, max_active);
mutex_init(&wq->mutex);
atomic_set(&wq->nr_pwqs_to_flush, 0);
INIT_LIST_HEAD(&wq->pwqs);
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH wq/for-7.4 v2 2/3] workqueue: Keep the limits of both max_active domains current
2026-09-25 13:09 [PATCH wq/for-7.4 v2 0/3] workqueue: Keep max_active and percpu_max_active in sync Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 1/3] workqueue: Move the initial max_active setup into a helper Breno Leitao
@ 2026-09-25 13:09 ` Breno Leitao
2026-09-28 20:15 ` Tejun Heo
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 3/3] workqueue: add a percpu_concurrency_managed workqueue attribute Breno Leitao
2 siblings, 1 reply; 8+ messages in thread
From: Breno Leitao @ 2026-09-25 13:09 UTC (permalink / raw)
To: Tejun Heo, Lai Jiangshan
Cc: linux-kernel, marco.crivellari, Breno Leitao, kernel-team
Commit 27db9dd7f84f3a ("workqueue: Give percpu workqueues their own
max_active") gave the percpu backend a limit of its own, so a workqueue
now carries limits for two both scopes, but just one is used today.
Given that we are introducing a WQ_AFFN_CPU which will be able to do CM,
we want to keep both values set, so, the transition from scopes would be
able to succeed. This cannot be done under the flip, given this is RCU
protected. See discussion at [1]
The two numbers cannot simply be kept equal. percpu_max_active caps every
CPU on its own, while max_active is a workqueue wide budget distributed
among the nodes with min_active as the per-node floor. A per-CPU limit of
16 on a 64 CPU machine is not a workqueue wide limit of 16.
Use the scale function suggested by Tejun in [1].
Suggested-by: Tejun Heo <tj@kernel.org>
Link: https://lore.kernel.org/all/ec6c22eddc8ce0b1ef2434ed62acd560@kernel.org/ [1]
Signed-off-by: Breno Leitao <leitao@debian.org>
---
kernel/workqueue.c | 97 +++++++++++++++++++++++++++++++++++++++++-------------
1 file changed, 75 insertions(+), 22 deletions(-)
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index 97d2c9c0a7d6ab..545332d6118159 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -393,6 +393,7 @@ struct workqueue_struct {
int min_active; /* WO: min active works */
int saved_max_active; /* WQ: saved max_active */
int saved_min_active; /* WQ: saved min_active */
+ int saved_percpu_max_active; /* WQ: percpu limit */
struct workqueue_attrs *attrs; /* PW: workqueue attributes */
struct pool_workqueue __rcu *dfl_pwq; /* PW: only for unbound wqs */
@@ -421,6 +422,13 @@ struct workqueue_struct {
struct wq_node_nr_active *node_nr_active[]; /* I: per-node nr_active */
};
+static int wq_user_max_active(struct workqueue_struct *wq)
+{
+ if (wq->flags & WQ_UNBOUND)
+ return READ_ONCE(wq->saved_max_active);
+ return READ_ONCE(wq->saved_percpu_max_active);
+}
+
/*
* Each pod type describes how CPUs should be grouped for unbound workqueues.
* See the comment above workqueue_attrs->affn_scope.
@@ -4549,7 +4557,7 @@ static bool start_flush_work(struct work_struct *work, struct wq_barrier *barr,
* workqueues the deadlock happens when the rescuer stalls, blocking
* forward progress.
*/
- if (!from_cancel && (wq->saved_max_active == 1 || wq->rescuer))
+ if (!from_cancel && (wq_user_max_active(wq) == 1 || wq->rescuer))
touch_wq_lockdep_map(wq);
rcu_read_unlock();
@@ -6031,9 +6039,25 @@ static int init_rescuer(struct workqueue_struct *wq)
return 0;
}
+static int wq_nr_online_cpus(const struct cpumask *mask)
+{
+ return max(cpumask_weight_and(mask, cpu_online_mask), 1);
+}
+
+static int wq_scale_to_unbound(int percpu_max, int nr_cpus)
+{
+ return min_t(s64, (s64)percpu_max * nr_cpus, WQ_MAX_ACTIVE);
+}
+
+static int wq_scale_to_percpu(int max_active, int min_active, int nr_cpus)
+{
+ return max(DIV_ROUND_UP(max_active, nr_cpus), min_active);
+}
+
static void wq_init_max_active(struct workqueue_struct *wq, int max_active)
{
int effective_max_active;
+ int nr_cpus = wq_nr_online_cpus(wq_unbound_cpumask);
if (wq->flags & WQ_BH) {
/*
@@ -6050,55 +6074,69 @@ static void wq_init_max_active(struct workqueue_struct *wq, int max_active)
if (wq->flags & WQ_UNBOUND) {
wq->max_active = effective_max_active;
wq->min_active = min(effective_max_active, WQ_DFL_MIN_ACTIVE);
- wq->saved_min_active = wq->min_active;
+ wq->percpu_max_active = wq_scale_to_percpu(effective_max_active,
+ wq->min_active,
+ nr_cpus);
+ } else if (wq->flags & WQ_BH) { /* BH implies !WQ_UNBOUND */
+ wq->max_active = effective_max_active;
+ wq->min_active = effective_max_active;
+ wq->percpu_max_active = effective_max_active;
} else {
wq->percpu_max_active = effective_max_active;
+ wq->max_active = wq_scale_to_unbound(effective_max_active,
+ nr_cpus);
+ wq->min_active = min(effective_max_active, wq->max_active);
}
- wq->saved_max_active = effective_max_active;
+
+ wq->saved_max_active = wq->max_active;
+ wq->saved_min_active = wq->min_active;
+ wq->saved_percpu_max_active = wq->percpu_max_active;
}
/**
* wq_adjust_max_active - update a wq's max_active to the current setting
* @wq: target workqueue
*
- * If @wq isn't freezing, set the limit that applies to @wq's backing to the
- * saved_max_active and activate inactive work items accordingly. If @wq is
- * freezing, clear it to zero.
+ * If @wq isn't freezing, publish the saved limits of both domains and activate
+ * inactive work items accordingly. If @wq is freezing, clear them to zero.
*/
static void wq_adjust_max_active(struct workqueue_struct *wq)
{
+ int new_max, new_min, new_percpu_max;
bool activated;
- int new_max, new_min;
lockdep_assert_held(&wq->mutex);
if ((wq->flags & WQ_FREEZABLE) && workqueue_freezing) {
new_max = 0;
new_min = 0;
+ new_percpu_max = 0;
} else {
new_max = wq->saved_max_active;
new_min = wq->saved_min_active;
+ new_percpu_max = wq->saved_percpu_max_active;
}
/*
- * Update the limit and then kick inactive work items if more active
+ * Update both domains and then kick inactive work items if more active
* work items are allowed. This doesn't break work item ordering
* because new work items are always queued behind existing inactive
* work items if there are any.
+ *
+ * Which domain a pwq honours follows its pool, see
+ * pwq_tryinc_nr_active(). Keeping both current means a pwq is never
+ * metered against a limit nobody set.
*/
- if (wq->flags & WQ_UNBOUND) {
- if (wq->max_active == new_max && wq->min_active == new_min)
- return;
+ if (wq->max_active == new_max && wq->min_active == new_min &&
+ wq->percpu_max_active == new_percpu_max)
+ return;
- WRITE_ONCE(wq->max_active, new_max);
- WRITE_ONCE(wq->min_active, new_min);
- wq_update_node_max_active(wq, -1);
- } else {
- if (wq->percpu_max_active == new_max)
- return;
+ WRITE_ONCE(wq->max_active, new_max);
+ WRITE_ONCE(wq->min_active, new_min);
+ WRITE_ONCE(wq->percpu_max_active, new_percpu_max);
- WRITE_ONCE(wq->percpu_max_active, new_max);
- }
+ if (wq->flags & WQ_UNBOUND)
+ wq_update_node_max_active(wq, -1);
if (new_max == 0)
return;
@@ -6448,6 +6486,8 @@ EXPORT_SYMBOL_GPL(destroy_workqueue);
*/
void workqueue_set_max_active(struct workqueue_struct *wq, int max_active)
{
+ int nr_cpus;
+
/* max_active doesn't mean anything for BH workqueues */
if (WARN_ON(wq->flags & WQ_BH))
return;
@@ -6456,12 +6496,22 @@ void workqueue_set_max_active(struct workqueue_struct *wq, int max_active)
return;
max_active = wq_clamp_max_active(max_active, wq->flags, wq->name);
+ nr_cpus = wq_nr_online_cpus(wq_unbound_cpumask);
mutex_lock(&wq->mutex);
- wq->saved_max_active = max_active;
- if (wq->flags & WQ_UNBOUND)
+ if (wq->flags & WQ_UNBOUND) {
+ wq->saved_max_active = max_active;
wq->saved_min_active = min(wq->saved_min_active, max_active);
+ wq->saved_percpu_max_active =
+ wq_scale_to_percpu(max_active, wq->saved_min_active,
+ nr_cpus);
+ } else {
+ wq->saved_percpu_max_active = max_active;
+ wq->saved_max_active = wq_scale_to_unbound(max_active,
+ nr_cpus);
+ wq->saved_min_active = min(max_active, wq->saved_max_active);
+ }
wq_adjust_max_active(wq);
@@ -6492,6 +6542,9 @@ void workqueue_set_min_active(struct workqueue_struct *wq, int min_active)
mutex_lock(&wq->mutex);
wq->saved_min_active = clamp(min_active, 0, wq->saved_max_active);
+ wq->saved_percpu_max_active =
+ wq_scale_to_percpu(wq->saved_max_active, wq->saved_min_active,
+ wq_nr_online_cpus(wq_unbound_cpumask));
wq_adjust_max_active(wq);
mutex_unlock(&wq->mutex);
}
@@ -7563,7 +7616,7 @@ static ssize_t max_active_show(struct device *dev,
{
struct workqueue_struct *wq = dev_to_wq(dev);
- return scnprintf(buf, PAGE_SIZE, "%d\n", wq->saved_max_active);
+ return scnprintf(buf, PAGE_SIZE, "%d\n", wq_user_max_active(wq));
}
static ssize_t max_active_store(struct device *dev,
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH wq/for-7.4 v2 3/3] workqueue: add a percpu_concurrency_managed workqueue attribute
2026-09-25 13:09 [PATCH wq/for-7.4 v2 0/3] workqueue: Keep max_active and percpu_max_active in sync Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 1/3] workqueue: Move the initial max_active setup into a helper Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 2/3] workqueue: Keep the limits of both max_active domains current Breno Leitao
@ 2026-09-25 13:09 ` Breno Leitao
2026-09-28 21:31 ` Tejun Heo
2 siblings, 1 reply; 8+ messages in thread
From: Breno Leitao @ 2026-09-25 13:09 UTC (permalink / raw)
To: Tejun Heo, Lai Jiangshan
Cc: linux-kernel, marco.crivellari, Breno Leitao, kernel-team
Introduce percpu_concurrency_managed, similar to the discussion in [1],
and use it when deciding to do concurrency management, instead of
relying on WQ_PERCPU (the only WQ type that does CM as of now).
Do not set it for WQ_BH. BH uses static per-cpu pools too, but it forces
max_active to INT_MAX and does not impose the normal max-active
throttle.
The attr is workqueue-only. wqattrs_clear_for_pool() clears it before
attrs are stored in a worker_pool. Pool selection still uses WQ_PERCPU;
attrs do not switch a workqueue onto static per-cpu pools yet.
Link: https://lore.kernel.org/all/8a1437751a6da2c166284b1ca562fb0a@kernel.org/ [1]
Signed-off-by: Breno Leitao <leitao@debian.org>
---
include/linux/workqueue.h | 9 +++++++++
kernel/workqueue.c | 40 ++++++++++++++++++++++++++--------------
2 files changed, 35 insertions(+), 14 deletions(-)
diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h
index a283766a192aaf..32747f797c4449 100644
--- a/include/linux/workqueue.h
+++ b/include/linux/workqueue.h
@@ -204,6 +204,15 @@ struct workqueue_attrs {
*/
enum wq_affn_scope affn_scope;
+ /**
+ * @percpu_concurrency_managed: use per-cpu concurrency accounting
+ *
+ * Workqueues with this set meter active work per CPU against
+ * percpu_max_active. Workqueues with this clear either use unbound
+ * max_active accounting or do not impose a max_active limit.
+ */
+ bool percpu_concurrency_managed;
+
/**
* @ordered: work items must be executed one by one in queueing order
*/
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index 545332d6118159..7059fc1bd9b090 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -424,9 +424,9 @@ struct workqueue_struct {
static int wq_user_max_active(struct workqueue_struct *wq)
{
- if (wq->flags & WQ_UNBOUND)
- return READ_ONCE(wq->saved_max_active);
- return READ_ONCE(wq->saved_percpu_max_active);
+ if (READ_ONCE(wq->attrs->percpu_concurrency_managed))
+ return READ_ONCE(wq->saved_percpu_max_active);
+ return READ_ONCE(wq->saved_max_active);
}
/*
@@ -1861,11 +1861,13 @@ static bool pwq_tryinc_nr_active(struct pool_workqueue *pwq, bool fill)
lockdep_assert_held(&pool->lock);
- /*
- * A concurrency-managed per-cpu pool accounts nr_active per pwq, so
- * pwq->nr_active against wq->percpu_max_active is sufficient.
- */
- if (is_percpu_pool(pool)) {
+ /* BH workqueues use per-cpu pools but do not throttle max_active. */
+ if (pool->flags & POOL_BH) {
+ obtained = true;
+ goto out;
+ }
+
+ if (READ_ONCE(wq->attrs->percpu_concurrency_managed)) {
obtained = pwq->nr_active < READ_ONCE(wq->percpu_max_active);
goto out;
}
@@ -2091,6 +2093,7 @@ static void node_activate_pending_pwq(struct wq_node_nr_active *nna,
*/
static void pwq_dec_nr_active(struct pool_workqueue *pwq)
{
+ struct workqueue_struct *wq = pwq->wq;
struct worker_pool *pool = pwq->pool;
struct wq_node_nr_active *nna;
@@ -2102,16 +2105,15 @@ static void pwq_dec_nr_active(struct pool_workqueue *pwq)
*/
pwq->nr_active--;
- /*
- * A concurrency-managed per-cpu pool only needs to kick the first
- * inactive work item on @pwq itself.
- */
- if (is_percpu_pool(pool)) {
+ if (pool->flags & POOL_BH)
+ return;
+
+ if (READ_ONCE(wq->attrs->percpu_concurrency_managed)) {
pwq_activate_first_inactive(pwq, false);
return;
}
- nna = wq_node_nr_active(pwq->wq, pool->node);
+ nna = wq_node_nr_active(wq, pool->node);
/*
* If @pwq is for an unbound workqueue, it's more complicated because
@@ -5034,6 +5036,7 @@ static void copy_workqueue_attrs(struct workqueue_attrs *to,
* get_unbound_pool() explicitly clears the fields.
*/
to->affn_scope = from->affn_scope;
+ to->percpu_concurrency_managed = from->percpu_concurrency_managed;
to->ordered = from->ordered;
}
@@ -5044,6 +5047,7 @@ static void copy_workqueue_attrs(struct workqueue_attrs *to,
static void wqattrs_clear_for_pool(struct workqueue_attrs *attrs)
{
attrs->affn_scope = WQ_AFFN_NR_TYPES;
+ attrs->percpu_concurrency_managed = false;
attrs->ordered = false;
if (attrs->affn_strict)
cpumask_copy(attrs->cpumask, cpu_possible_mask);
@@ -5851,6 +5855,10 @@ int apply_workqueue_attrs(struct workqueue_struct *wq,
if (WARN_ON(!(wq->flags & WQ_UNBOUND)))
return -EINVAL;
+ /* Attrs cannot switch a workqueue to per-cpu pools yet. */
+ if (WARN_ON(attrs->percpu_concurrency_managed))
+ return -EINVAL;
+
mutex_lock(&wq_pool_mutex);
ret = apply_workqueue_attrs_locked(wq, attrs);
mutex_unlock(&wq_pool_mutex);
@@ -5942,6 +5950,10 @@ static struct workqueue_attrs *alloc_wq_std_attrs(struct workqueue_struct *wq)
if (wq->flags & __WQ_ORDERED)
attrs->ordered = true;
+ /* BH uses per-cpu pools but not max_active concurrency management. */
+ if ((wq->flags & WQ_PERCPU) && !(wq->flags & WQ_BH))
+ attrs->percpu_concurrency_managed = true;
+
return attrs;
}
--
2.53.0-Meta
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH wq/for-7.4 v2 2/3] workqueue: Keep the limits of both max_active domains current
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 2/3] workqueue: Keep the limits of both max_active domains current Breno Leitao
@ 2026-09-28 20:15 ` Tejun Heo
2026-09-29 14:40 ` Breno Leitao
0 siblings, 1 reply; 8+ messages in thread
From: Tejun Heo @ 2026-09-28 20:15 UTC (permalink / raw)
To: Breno Leitao; +Cc: Lai Jiangshan, Marco Crivellari, linux-kernel, kernel-team
Hello, Breno.
On Fri, Sep 25, 2026 at 06:09:51AM -0700, Breno Leitao wrote:
> Given that we are introducing a WQ_AFFN_CPU which will be able to do CM,
> we want to keep both values set, so, the transition from scopes would be
> able to succeed.
This would be the new PERCPU scope rather than WQ_AFFN_CPU.
> + int nr_cpus = wq_nr_online_cpus(wq_unbound_cpumask);
Workqueues from workqueue_init_early() and early initcalls are created
before smp_init(), when only the boot CPU is online, so nr_cpus is 1 for
them. For example, system_dfl_wq ends up with percpu_max_active 2048 where
32 would be right on a 64 CPU machine, and nothing rederives it once all
CPUs are up. Maybe rederive the other domain in workqueue_init_topology(),
which already walks all workqueues to update node max_active?
> + } else if (wq->flags & WQ_BH) { /* BH implies !WQ_UNBOUND */
> + wq->max_active = effective_max_active;
> + wq->min_active = effective_max_active;
> + wq->percpu_max_active = effective_max_active;
Maybe set these in the WQ_BH branch at the top so that BH is handled in one
place?
> } else {
> wq->percpu_max_active = effective_max_active;
> + wq->max_active = wq_scale_to_unbound(effective_max_active,
> + nr_cpus);
> + wq->min_active = min(effective_max_active, wq->max_active);
max_active can't be lower than percpu_max_active here, so min_active can
just be percpu_max_active. Same in workqueue_set_max_active().
> max_active = wq_clamp_max_active(max_active, wq->flags, wq->name);
> + nr_cpus = wq_nr_online_cpus(wq_unbound_cpumask);
For unbound workqueues, nr_cpus should match what
wq_update_node_max_active() distributes max_active over, the online CPUs in
unbound_effective_cpumask(). With wq_unbound_cpumask, percpu_max_active
comes out too low for a workqueue with a narrowed cpumask, e.g. 8 instead of
64 for max_active 256 on 4 of 64 CPUs. Same in workqueue_set_min_active().
wq_unbound_cpumask is right for percpu workqueues.
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH wq/for-7.4 v2 3/3] workqueue: add a percpu_concurrency_managed workqueue attribute
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 3/3] workqueue: add a percpu_concurrency_managed workqueue attribute Breno Leitao
@ 2026-09-28 21:31 ` Tejun Heo
2026-09-29 13:51 ` Breno Leitao
0 siblings, 1 reply; 8+ messages in thread
From: Tejun Heo @ 2026-09-28 21:31 UTC (permalink / raw)
To: Breno Leitao; +Cc: Lai Jiangshan, Marco Crivellari, linux-kernel, kernel-team
Hello, Breno.
On Fri, Sep 25, 2026 at 06:09:52AM -0700, Breno Leitao wrote:
> Introduce percpu_concurrency_managed, similar to the discussion in [1],
> and use it when deciding to do concurrency management, instead of
> relying on WQ_PERCPU (the only WQ type that does CM as of now).
Once PERCPU is its own scope, the scope picks the backend, and concurrency
management and the per-cpu max_active follow from that. Opting out of
concurrency management on the per-cpu pools is what WQ_CPU_INTENSIVE already
does, and I don't see a need to change that at runtime. How about dropping
this patch for now?
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH wq/for-7.4 v2 3/3] workqueue: add a percpu_concurrency_managed workqueue attribute
2026-09-28 21:31 ` Tejun Heo
@ 2026-09-29 13:51 ` Breno Leitao
0 siblings, 0 replies; 8+ messages in thread
From: Breno Leitao @ 2026-09-29 13:51 UTC (permalink / raw)
To: Tejun Heo; +Cc: Lai Jiangshan, Marco Crivellari, linux-kernel, kernel-team
Hello Tejun,
On Mon, Sep 28, 2026 at 11:31:11AM -1000, Tejun Heo wrote:
> On Fri, Sep 25, 2026 at 06:09:52AM -0700, Breno Leitao wrote:
> > Introduce percpu_concurrency_managed, similar to the discussion in [1],
> > and use it when deciding to do concurrency management, instead of
> > relying on WQ_PERCPU (the only WQ type that does CM as of now).
>
> Once PERCPU is its own scope, the scope picks the backend, and concurrency
> management and the per-cpu max_active follow from that. Opting out of
> concurrency management on the per-cpu pools is what WQ_CPU_INTENSIVE already
> does, and I don't see a need to change that at runtime. How about dropping
> this patch for now?
Ack, I will continue to decide about concurrency management by looking
at the affinity scope other than a new knob.
Thanks for the review,
--breno
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH wq/for-7.4 v2 2/3] workqueue: Keep the limits of both max_active domains current
2026-09-28 20:15 ` Tejun Heo
@ 2026-09-29 14:40 ` Breno Leitao
0 siblings, 0 replies; 8+ messages in thread
From: Breno Leitao @ 2026-09-29 14:40 UTC (permalink / raw)
To: Tejun Heo; +Cc: Lai Jiangshan, Marco Crivellari, linux-kernel, kernel-team
Hello Tejun,
On Mon, Sep 28, 2026 at 10:15:31AM -1000, Tejun Heo wrote:
> On Fri, Sep 25, 2026 at 06:09:51AM -0700, Breno Leitao wrote:
> > Given that we are introducing a WQ_AFFN_CPU which will be able to do CM,
> > we want to keep both values set, so, the transition from scopes would be
> > able to succeed.
>
> This would be the new PERCPU scope rather than WQ_AFFN_CPU.
yes, sorry, that is waht I meant.
> > + int nr_cpus = wq_nr_online_cpus(wq_unbound_cpumask);
>
> Workqueues from workqueue_init_early() and early initcalls are created
> before smp_init(), when only the boot CPU is online, so nr_cpus is 1 for
> them. For example, system_dfl_wq ends up with percpu_max_active 2048 where
> 32 would be right on a 64 CPU machine, and nothing rederives it once all
> CPUs are up. Maybe rederive the other domain in workqueue_init_topology(),
> which already walks all workqueues to update node max_active?
Very good point, and I've just tested it and reproduce the behaviour you
raised.
>
> > + } else if (wq->flags & WQ_BH) { /* BH implies !WQ_UNBOUND */
> > + wq->max_active = effective_max_active;
> > + wq->min_active = effective_max_active;
> > + wq->percpu_max_active = effective_max_active;
>
> Maybe set these in the WQ_BH branch at the top so that BH is handled in one
> place?
Ack!
> > } else {
> > wq->percpu_max_active = effective_max_active;
> > + wq->max_active = wq_scale_to_unbound(effective_max_active,
> > + nr_cpus);
> > + wq->min_active = min(effective_max_active, wq->max_active);
>
> max_active can't be lower than percpu_max_active here, so min_active can
> just be percpu_max_active. Same in workqueue_set_max_active().
Ack!
> > max_active = wq_clamp_max_active(max_active, wq->flags, wq->name);
> > + nr_cpus = wq_nr_online_cpus(wq_unbound_cpumask);
>
> For unbound workqueues, nr_cpus should match what
> wq_update_node_max_active() distributes max_active over, the online CPUs in
> unbound_effective_cpumask(). With wq_unbound_cpumask, percpu_max_active
> comes out too low for a workqueue with a narrowed cpumask, e.g. 8 instead of
> 64 for max_active 256 on 4 of 64 CPUs. Same in workqueue_set_min_active().
> wq_unbound_cpumask is right for percpu workqueues.
Right, thanks for catching that. I'll change both setters to use
unbound_effective_cpumask() once the workqueue has a pwq installed, and
only fall back to wq_unbound_cpumask before that (percpu workqueues, or
unbound ones still being set up), so it lines up with what
wq_update_node_max_active() already does.
Thanks for the review,
--breno
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-09-29 14:41 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-25 13:09 [PATCH wq/for-7.4 v2 0/3] workqueue: Keep max_active and percpu_max_active in sync Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 1/3] workqueue: Move the initial max_active setup into a helper Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 2/3] workqueue: Keep the limits of both max_active domains current Breno Leitao
2026-09-28 20:15 ` Tejun Heo
2026-09-29 14:40 ` Breno Leitao
2026-09-25 13:09 ` [PATCH wq/for-7.4 v2 3/3] workqueue: add a percpu_concurrency_managed workqueue attribute Breno Leitao
2026-09-28 21:31 ` Tejun Heo
2026-09-29 13:51 ` Breno Leitao
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®