* [RFC v6 0/3] Realtime workqueues and panthor realtime submission
@ 2026-10-01 16:07 Tvrtko Ursulin
2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
` (2 more replies)
0 siblings, 3 replies; 8+ messages in thread
From: Tvrtko Ursulin @ 2026-10-01 16:07 UTC (permalink / raw)
To: linux-kernel
Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon,
Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo
This is a continuation of the previous discussion which was here:
https://lore.kernel.org/dri-devel/20260702143745.79293-1-tvrtko.ursulin@igalia.com/
Work is now converted to a much simpler approach by adding real-time scheduling
workqueues based on Tejun's feedback
To re-cap, when an userspace thread submits GPU work, due how the DRM scheduler
uses workqueues to feed the GPU, and regardless of the GPU rendering context
priority, or the CPU scheduling priority of the userspace thread itself, the
use of workqueues can add significant latency to the submit path.
When CPU is busy with enough backround load this translates to severe latency
spikes measured as time between userspace submitting work and GPU actually being
given that work to execute.
With the panthor workqueue upgraded to use WQ_HIGHPRI and varying the CPU
priority of the submit thread, the test program from
https://gitlab.freedesktop.org/panfrost/linux/-/work_items/49 reproduces these
kind of latencies:
. N RT
M 27 28 32
95% 163 246 809
98% 924 991 1882
Legend:
M = Median submit latency in us
95% = Percentile latency in us
. = Userspace submit thread SCHED_OTHER
N = -||= nice -1
RT = -||- FIFO 1
Upgrading the panthor workqueues so that the realtime GPU priority queue uses
WQ_RT, submit latency becomes completely controlled with the median of 14us and
95 and 98-th percentiles at 23us and 25us respectively.
Important to note is that VK_QUEUE_GLOBAL_PRIORITY_REALTIME already required
the userspace to have CAP_SYS_NICE, meaning access to real-time workqueues is
effectively also guared behind this capability.
v2:
* See patch 1 changelog.
v3:
* See patch 1 changelog + apologies for v3 following so quickly after v2.
I have found a race condition as I expanded the testing to a second platform.
v4:
* See patch 1 changelog.
v5:
* Rebased on top of tj/for-7.4 branch.
* New patch in series simplifies unbound sysfs attribute handling.
* WQ_RTPRI reworked to fold RT marker into attr->nice.
v6:
* Fixed accidental sysfs attribute rename (patch 1).
* Various review feedback as per patch 2 changelog.
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
Tvrtko Ursulin (3):
workqueue: Simplify unbound sysfs attribute registration
workqueue: Add support for real-time workers
drm/panthor: Create per queue priority workqueues
Documentation/core-api/workqueue.rst | 8 +
drivers/gpu/drm/panthor/panthor_sched.c | 37 ++++-
include/linux/workqueue.h | 5 +-
kernel/workqueue.c | 199 +++++++++++++++---------
tools/workqueue/wq_dump.py | 9 +-
5 files changed, 175 insertions(+), 83 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 8+ messages in thread* [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration 2026-10-01 16:07 [RFC v6 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin @ 2026-10-01 16:07 ` Tvrtko Ursulin 2026-10-02 19:30 ` Tejun Heo 2026-10-01 16:07 ` [RFC v6 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin 2026-10-01 16:07 ` [RFC v6 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin 2 siblings, 1 reply; 8+ messages in thread From: Tvrtko Ursulin @ 2026-10-01 16:07 UTC (permalink / raw) To: linux-kernel Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon, Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo Instead of manually registering each attribute we can put them in an attribute group with a visibility check and device core will handle the rest, which simplifies the registration and error unwind. Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com> Cc: Boris Brezillon <boris.brezillon@collabora.com> Cc: Bradley Morgan <include@grrlz.net> Cc: Chia-I Wu <olv@google.com> Cc: Liviu Dudau <liviu.dudau@arm.com> Cc: Matthew Brost <matthew.brost@intel.com> Cc: Steven Price <steven.price@arm.com> Cc: Tejun Heo <tj@kernel.org> --- v2: * Fix accidental rename of affinity_scope. --- kernel/workqueue.c | 98 ++++++++++++++++++++++++---------------------- 1 file changed, 52 insertions(+), 46 deletions(-) diff --git a/kernel/workqueue.c b/kernel/workqueue.c index e618108c6127..c83d68d7d0ee 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -7599,8 +7599,8 @@ static const struct attribute_group wq_sysfs_group = { }; __ATTRIBUTE_GROUPS(wq_sysfs); -static ssize_t wq_nice_show(struct device *dev, struct device_attribute *attr, - char *buf) +static ssize_t nice_show(struct device *dev, struct device_attribute *attr, + char *buf) { struct workqueue_struct *wq = dev_to_wq(dev); int written; @@ -7627,8 +7627,8 @@ static struct workqueue_attrs *wq_sysfs_prep_attrs(struct workqueue_struct *wq) return attrs; } -static ssize_t wq_nice_store(struct device *dev, struct device_attribute *attr, - const char *buf, size_t count) +static ssize_t nice_store(struct device *dev, struct device_attribute *attr, + const char *buf, size_t count) { struct workqueue_struct *wq = dev_to_wq(dev); struct workqueue_attrs *attrs; @@ -7652,8 +7652,8 @@ static ssize_t wq_nice_store(struct device *dev, struct device_attribute *attr, return ret ?: count; } -static ssize_t wq_cpumask_show(struct device *dev, - struct device_attribute *attr, char *buf) +static ssize_t unbound_cpumask_show(struct device *dev, + struct device_attribute *attr, char *buf) { struct workqueue_struct *wq = dev_to_wq(dev); int written; @@ -7665,9 +7665,9 @@ static ssize_t wq_cpumask_show(struct device *dev, return written; } -static ssize_t wq_cpumask_store(struct device *dev, - struct device_attribute *attr, - const char *buf, size_t count) +static ssize_t unbound_cpumask_store(struct device *dev, + struct device_attribute *attr, + const char *buf, size_t count) { struct workqueue_struct *wq = dev_to_wq(dev); struct workqueue_attrs *attrs; @@ -7689,8 +7689,8 @@ static ssize_t wq_cpumask_store(struct device *dev, return ret ?: count; } -static ssize_t wq_affn_scope_show(struct device *dev, - struct device_attribute *attr, char *buf) +static ssize_t affinity_scope_show(struct device *dev, + struct device_attribute *attr, char *buf) { struct workqueue_struct *wq = dev_to_wq(dev); int written; @@ -7708,9 +7708,9 @@ static ssize_t wq_affn_scope_show(struct device *dev, return written; } -static ssize_t wq_affn_scope_store(struct device *dev, - struct device_attribute *attr, - const char *buf, size_t count) +static ssize_t affinity_scope_store(struct device *dev, + struct device_attribute *attr, + const char *buf, size_t count) { struct workqueue_struct *wq = dev_to_wq(dev); struct workqueue_attrs *attrs; @@ -7731,8 +7731,8 @@ static ssize_t wq_affn_scope_store(struct device *dev, return ret ?: count; } -static ssize_t wq_affinity_strict_show(struct device *dev, - struct device_attribute *attr, char *buf) +static ssize_t affinity_strict_show(struct device *dev, + struct device_attribute *attr, char *buf) { struct workqueue_struct *wq = dev_to_wq(dev); @@ -7740,9 +7740,9 @@ static ssize_t wq_affinity_strict_show(struct device *dev, wq->attrs->affn_strict); } -static ssize_t wq_affinity_strict_store(struct device *dev, - struct device_attribute *attr, - const char *buf, size_t count) +static ssize_t affinity_strict_store(struct device *dev, + struct device_attribute *attr, + const char *buf, size_t count) { struct workqueue_struct *wq = dev_to_wq(dev); struct workqueue_attrs *attrs; @@ -7762,14 +7762,40 @@ static ssize_t wq_affinity_strict_store(struct device *dev, return ret ?: count; } -static struct device_attribute wq_sysfs_unbound_attrs[] = { - __ATTR(nice, 0644, wq_nice_show, wq_nice_store), - __ATTR(cpumask, 0644, wq_cpumask_show, wq_cpumask_store), - __ATTR(affinity_scope, 0644, wq_affn_scope_show, wq_affn_scope_store), - __ATTR(affinity_strict, 0644, wq_affinity_strict_show, wq_affinity_strict_store), - __ATTR_NULL, +static DEVICE_ATTR_RW(nice); +static DEVICE_ATTR_RW(affinity_scope); +static DEVICE_ATTR_RW(affinity_strict); +/* Avoid naming clash with the other cpumask */ +static struct device_attribute dev_attr_unbound_cpumask = + __ATTR(cpumask, 0644, unbound_cpumask_show, unbound_cpumask_store); + +static struct attribute *wq_sysfs_unbound_attrs[] = { + &dev_attr_nice.attr, + &dev_attr_unbound_cpumask.attr, + &dev_attr_affinity_scope.attr, + &dev_attr_affinity_strict.attr, + NULL, }; +static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj, + struct attribute *attr, int n) +{ + struct device *dev = kobj_to_dev(kobj); + struct workqueue_struct *wq = dev_to_wq(dev); + + if (!(wq->flags & WQ_UNBOUND)) + return SYSFS_GROUP_INVISIBLE; + + return attr->mode; +} + +static const struct attribute_group wq_sysfs_unbound_group = { + .is_visible = wq_sysfs_unbound_group_visible, + .attrs = wq_sysfs_unbound_attrs, +}; + +__ATTRIBUTE_GROUPS(wq_sysfs_unbound); + static const struct bus_type wq_subsys = { .name = "workqueue", .dev_groups = wq_sysfs_groups, @@ -7907,14 +7933,9 @@ int workqueue_sysfs_register(struct workqueue_struct *wq) wq_dev->wq = wq; wq_dev->dev.bus = &wq_subsys; wq_dev->dev.release = wq_device_release; + wq_dev->dev.groups = wq_sysfs_unbound_groups; dev_set_name(&wq_dev->dev, "%s", wq->name); - /* - * attrs are created separately. Suppress uevent until - * everything is ready. - */ - dev_set_uevent_suppress(&wq_dev->dev, true); - ret = device_register(&wq_dev->dev); if (ret) { put_device(&wq_dev->dev); @@ -7922,21 +7943,6 @@ int workqueue_sysfs_register(struct workqueue_struct *wq) return ret; } - if (wq->flags & WQ_UNBOUND) { - struct device_attribute *attr; - - for (attr = wq_sysfs_unbound_attrs; attr->attr.name; attr++) { - ret = device_create_file(&wq_dev->dev, attr); - if (ret) { - device_unregister(&wq_dev->dev); - wq->wq_dev = NULL; - return ret; - } - } - } - - dev_set_uevent_suppress(&wq_dev->dev, false); - kobject_uevent(&wq_dev->dev.kobj, KOBJ_ADD); return 0; } -- 2.55.0 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration 2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin @ 2026-10-02 19:30 ` Tejun Heo 0 siblings, 0 replies; 8+ messages in thread From: Tejun Heo @ 2026-10-02 19:30 UTC (permalink / raw) To: Tvrtko Ursulin Cc: linux-kernel, dri-devel, kernel-dev, Boris Brezillon, Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Lai Jiangshan, Breno Leitao Hello, The following is a Claude-generated review. On Thu, Oct 01, 2026 at 05:07:09PM +0100, Tvrtko Ursulin wrote: > +static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj, > + struct attribute *attr, int n) > +{ > + struct device *dev = kobj_to_dev(kobj); > + struct workqueue_struct *wq = dev_to_wq(dev); > + > + if (!(wq->flags & WQ_UNBOUND)) > + return SYSFS_GROUP_INVISIBLE; SYSFS_GROUP_INVISIBLE is documented for named groups only. For an unnamed group, fs/sysfs/group.c masks the bit off and the leftover 0 is what hides the attribute. Can you return 0 here like wq_sysfs_is_visible() does? Thanks. -- tejun ^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC v6 2/3] workqueue: Add support for real-time workers 2026-10-01 16:07 [RFC v6 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin 2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin @ 2026-10-01 16:07 ` Tvrtko Ursulin 2026-10-01 18:48 ` [RFC v6.1 " Tvrtko Ursulin 2026-10-01 16:07 ` [RFC v6 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin 2 siblings, 1 reply; 8+ messages in thread From: Tvrtko Ursulin @ 2026-10-01 16:07 UTC (permalink / raw) To: linux-kernel Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon, Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo For use cases such as the DRM scheduler submitting work to the GPU on behalf of low latency userspace applications, where latter have sufficient privileges to have had successfully obtained realtime Vulkan global priority, competing with random background CPU load can create large latency spikes which gets in the way of a smooth user experience. For these situations the existing WQ_HIGHPRI does not bring a noticeable improvement and a stronger hint is needed. Lets add WQ_RT which creates workers with a SCHED_FIFO scheduling class to improve this. We use a minimum priority level since we only care about winning the contest against normal background CPU load. Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com> Cc: Boris Brezillon <boris.brezillon@collabora.com> Cc: Bradley Morgan <include@grrlz.net> Cc: Chia-I Wu <olv@google.com> Cc: Liviu Dudau <liviu.dudau@arm.com> Cc: Matthew Brost <matthew.brost@intel.com> Cc: Steven Price <steven.price@arm.com> Cc: Tejun Heo <tj@kernel.org> --- v2: * Limit WQ_RTPRI to unbound workqueues and make it have strict CPU affinitity. (Tejun) * Fixed commit message typos. (AI) * Fixed sysfs handling, max_active setting and user modified nice application. (AI) v3: * Fix worker->pool null pointer dereference race by moving the global decrement to detach_dying_workers(). * Rebase for upstream changes. v4: * Fixed onion unwind. * Moved affinity setting to default attributes. v5: * Dropped global and local limits. * Documented in workqueue.rst. * Added NR_WQ_ATTRIBUTES. * Reverted BH handling changes. v6: * Dropped separate attr->prio in favour of RTPRI_NICE_LEVEL checks. (Tejun) * Reworked on top of tj/for-7.4. v7: * Convert to attrs->prio encoded analoguous to task_struct->prio. * Rename flag to WQ_PRIO and do not re-order enums. * Forbid WQ_RT affinity modifications via sysfs. * Added wq_dump.py support. --- Documentation/core-api/workqueue.rst | 8 +++ include/linux/workqueue.h | 5 +- kernel/workqueue.c | 101 +++++++++++++++++++-------- tools/workqueue/wq_dump.py | 9 ++- 4 files changed, 90 insertions(+), 33 deletions(-) diff --git a/Documentation/core-api/workqueue.rst b/Documentation/core-api/workqueue.rst index bb770f556568..d699c3832b19 100644 --- a/Documentation/core-api/workqueue.rst +++ b/Documentation/core-api/workqueue.rst @@ -225,6 +225,14 @@ resources, scheduled and executed. each other. Each maintains its separate pool of workers and implements concurrency management among its workers. +``WQ_RT`` + Real-time priority workqueues must be created as unbound and will be + configured with the strict CPU affinity set. Their worker threads use the FIFO + scheduling policy with the lowest applicable priority. + + To be used sparingly for use cases such as the real-time GPU rendering + contexts accessible to privileged clients. + ``WQ_CPU_INTENSIVE`` Work items of a CPU intensive wq do not contribute to the concurrency level. In other words, runnable CPU intensive diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h index a283766a192a..ebec9dcc9e5f 100644 --- a/include/linux/workqueue.h +++ b/include/linux/workqueue.h @@ -147,9 +147,9 @@ enum wq_affn_scope { */ struct workqueue_attrs { /** - * @nice: nice level + * @prio: priority encoded analoguous to task_struct->prio. */ - int nice; + int prio; /** * @cpumask: allowed CPUs @@ -404,6 +404,7 @@ enum wq_flags { */ WQ_POWER_EFFICIENT = 1 << 7, WQ_PERCPU = 1 << 8, /* bound to a specific cpu */ + WQ_RT = 1 << 9, /* real-time priority, valid only with WQ_UNBOUND */ __WQ_DESTROYING = 1 << 15, /* internal: workqueue is destroying */ __WQ_DRAINING = 1 << 16, /* internal: workqueue is draining */ diff --git a/kernel/workqueue.c b/kernel/workqueue.c index c83d68d7d0ee..afe39a18ad9e 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -47,6 +47,7 @@ #include <linux/jhash.h> #include <linux/hashtable.h> #include <linux/rculist.h> +#include <linux/sched/rt.h> #include <linux/nodemask.h> #include <linux/moduleparam.h> #include <linux/uaccess.h> @@ -126,7 +127,8 @@ enum wq_internal_consts { * all cpus. Give MIN_NICE. */ RESCUER_NICE_LEVEL = MIN_NICE, - HIGHPRI_NICE_LEVEL = MIN_NICE, + HIGHPRI_PRIORITY = NICE_TO_PRIO(MIN_NICE), + RT_PRIORITY = MAX_PRIO, WQ_NAME_LEN = 32, WORKER_ID_LEN = 10 + WQ_NAME_LEN, /* "kworker/R-" + WQ_NAME_LEN */ @@ -1275,7 +1277,7 @@ static bool assign_work(struct work_struct *work, struct worker *worker, static struct irq_work *bh_pool_irq_work(struct worker_pool *pool) { - int high = pool->attrs->nice == HIGHPRI_NICE_LEVEL ? 1 : 0; + int high = pool->attrs->prio == HIGHPRI_PRIORITY ? 1 : 0; return &per_cpu(bh_pool_irq_works, pool->cpu)[high]; } @@ -1290,7 +1292,7 @@ static void kick_bh_pool(struct worker_pool *pool) return; } #endif - if (pool->attrs->nice == HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio == HIGHPRI_PRIORITY) raise_softirq_irqoff(HI_SOFTIRQ); else raise_softirq_irqoff(TASKLET_SOFTIRQ); @@ -2959,7 +2961,8 @@ static int format_worker_id(char *buf, size_t size, struct worker *worker, if (pool->cpu >= 0) return scnprintf(buf, size, "kworker/%d:%d%s", pool->cpu, worker->id, - pool->attrs->nice < 0 ? "H" : ""); + pool->attrs->prio < NICE_TO_PRIO(0) ? + "H" : ""); else return scnprintf(buf, size, "kworker/u%d:%d", pool->id, worker->id); @@ -3018,7 +3021,12 @@ static struct worker *create_worker(struct worker_pool *pool) goto fail; } - set_user_nice(worker->task, pool->attrs->nice); + if (rt_prio(pool->attrs->prio)) + sched_set_fifo_low(worker->task); + else + set_user_nice(worker->task, + PRIO_TO_NICE(pool->attrs->prio)); + kthread_bind_mask(worker->task, pool_allowed_cpus(pool)); } @@ -3910,7 +3918,7 @@ static void bh_worker(struct worker *worker) if (budget_exhausted) trace_workqueue_bh_budget_yield(pool, restarts, timeout, - pool->attrs->nice == HIGHPRI_NICE_LEVEL); + pool->attrs->prio == HIGHPRI_PRIORITY); } /* @@ -3969,7 +3977,7 @@ static void drain_dead_softirq_workfn(struct work_struct *work) * don't hog this CPU's BH. */ if (repeat) { - if (pool->attrs->nice == HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio == HIGHPRI_PRIORITY) queue_work(system_bh_highpri_wq, work); else queue_work(system_bh_wq, work); @@ -4001,7 +4009,7 @@ void workqueue_softirq_dead(unsigned int cpu) dead_work.pool = pool; init_completion(&dead_work.done); - if (pool->attrs->nice == HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio == HIGHPRI_PRIORITY) queue_work(system_bh_highpri_wq, &dead_work.work); else queue_work(system_bh_wq, &dead_work.work); @@ -5015,7 +5023,7 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(void) static void copy_workqueue_attrs(struct workqueue_attrs *to, const struct workqueue_attrs *from) { - to->nice = from->nice; + to->prio = from->prio; cpumask_copy(to->cpumask, from->cpumask); cpumask_copy(to->__pod_cpumask, from->__pod_cpumask); to->affn_strict = from->affn_strict; @@ -5046,7 +5054,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs *attrs) { u32 hash = 0; - hash = jhash_1word(attrs->nice, hash); + hash = jhash_1word(attrs->prio, hash); hash = jhash_1word(attrs->affn_strict, hash); hash = jhash(cpumask_bits(attrs->__pod_cpumask), BITS_TO_LONGS(nr_cpumask_bits) * sizeof(long), hash); @@ -5060,7 +5068,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs *attrs) static bool wqattrs_equal(const struct workqueue_attrs *a, const struct workqueue_attrs *b) { - if (a->nice != b->nice) + if (a->prio != b->prio) return false; if (a->affn_strict != b->affn_strict) return false; @@ -5928,8 +5936,19 @@ static struct workqueue_attrs *alloc_wq_std_attrs(struct workqueue_struct *wq) if (!attrs) return NULL; - if (wq->flags & WQ_HIGHPRI) - attrs->nice = HIGHPRI_NICE_LEVEL; + if (wq->flags & WQ_RT) { + attrs->prio = RT_PRIORITY; + /* + * RT workqueues have strict CPU affinity for low + * latency execution. + */ + attrs->affn_scope = WQ_AFFN_CPU; + attrs->affn_strict = true; + } else if (wq->flags & WQ_HIGHPRI) { + attrs->prio = HIGHPRI_PRIORITY; + } else { + attrs->prio = DEFAULT_PRIO; + } if (wq->flags & __WQ_ORDERED) attrs->ordered = true; @@ -6115,6 +6134,12 @@ static struct workqueue_struct *__alloc_workqueue(const char *fmt, return NULL; } + if (flags & WQ_RT) { + if (WARN_ON_ONCE((flags & (WQ_HIGHPRI | WQ_UNBOUND)) != + WQ_UNBOUND)) + return NULL; + } + /* see the comment above the definition of WQ_POWER_EFFICIENT */ if ((flags & WQ_POWER_EFFICIENT) && wq_power_efficient) flags = (flags & ~WQ_PERCPU) | WQ_UNBOUND; @@ -6671,9 +6696,9 @@ static void pr_cont_pool_info(struct worker_pool *pool) pr_cont(" flags=0x%x", pool->flags); if (pool->flags & POOL_BH) pr_cont(" bh%s", - pool->attrs->nice == HIGHPRI_NICE_LEVEL ? "-hi" : ""); + pool->attrs->prio == HIGHPRI_PRIORITY ? "-hi" : ""); else - pr_cont(" nice=%d", pool->attrs->nice); + pr_cont(" nice=%d", PRIO_TO_NICE(pool->attrs->prio)); } static void pr_cont_worker_id(struct worker *worker) @@ -6682,7 +6707,7 @@ static void pr_cont_worker_id(struct worker *worker) if (pool->flags & POOL_BH) pr_cont("bh%s", - pool->attrs->nice == HIGHPRI_NICE_LEVEL ? "-hi" : ""); + pool->attrs->prio == HIGHPRI_PRIORITY ? "-hi" : ""); else pr_cont("%d%s", task_pid_nr(worker->task), worker->rescue_wq ? "(RESCUER)" : ""); @@ -7606,7 +7631,11 @@ static ssize_t nice_show(struct device *dev, struct device_attribute *attr, int written; mutex_lock(&wq->mutex); - written = scnprintf(buf, PAGE_SIZE, "%d\n", wq->attrs->nice); + if (wq->attrs->prio == RT_PRIORITY) + written = scnprintf(buf, PAGE_SIZE, "rt\n"); + else + written = scnprintf(buf, PAGE_SIZE, "%d\n", + PRIO_TO_NICE(wq->attrs->prio)); mutex_unlock(&wq->mutex); return written; @@ -7632,19 +7661,21 @@ static ssize_t nice_store(struct device *dev, struct device_attribute *attr, { struct workqueue_struct *wq = dev_to_wq(dev); struct workqueue_attrs *attrs; - int ret = -ENOMEM; + int ret, nice = 0; + + if (sscanf(buf, "%d", &nice) != 1 || nice < MIN_NICE || nice > MAX_NICE) + return -EINVAL; mutex_lock(&wq_pool_mutex); attrs = wq_sysfs_prep_attrs(wq); - if (!attrs) + if (!attrs) { + ret = -ENOMEM; goto out_unlock; + } - if (sscanf(buf, "%d", &attrs->nice) == 1 && - attrs->nice >= MIN_NICE && attrs->nice <= MAX_NICE) - ret = apply_workqueue_attrs_locked(wq, attrs); - else - ret = -EINVAL; + attrs->prio = NICE_TO_PRIO(nice); + ret = apply_workqueue_attrs_locked(wq, attrs); out_unlock: mutex_unlock(&wq_pool_mutex); @@ -7716,6 +7747,10 @@ static ssize_t affinity_scope_store(struct device *dev, struct workqueue_attrs *attrs; int affn, ret = -ENOMEM; + /* Do not allow affinity changes for RT workers. */ + if (wq->flags & WQ_RT) + return -EINVAL; + affn = parse_affn_scope(buf); if (affn < 0) return affn; @@ -7748,6 +7783,10 @@ static ssize_t affinity_strict_store(struct device *dev, struct workqueue_attrs *attrs; int v, ret = -ENOMEM; + /* Do not allow affinity changes for RT workers. */ + if (wq->flags & WQ_RT) + return -EINVAL; + if (sscanf(buf, "%d", &v) != 1) return -EINVAL; @@ -7786,6 +7825,10 @@ static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj, if (!(wq->flags & WQ_UNBOUND)) return SYSFS_GROUP_INVISIBLE; + /* Do not allow priority changes for RT workers. */ + if ((wq->flags & WQ_RT) && !strcmp(attr->name, "nice")) + return 0444; + return attr->mode; } @@ -8310,13 +8353,13 @@ static void __init restrict_unbound_cpumask(const char *name, const struct cpuma cpumask_and(wq_unbound_cpumask, wq_unbound_cpumask, mask); } -static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int nice) +static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int prio) { BUG_ON(init_worker_pool(pool)); pool->cpu = cpu; cpumask_copy(pool->attrs->cpumask, cpumask_of(cpu)); cpumask_copy(pool->attrs->__pod_cpumask, cpumask_of(cpu)); - pool->attrs->nice = nice; + pool->attrs->prio = prio; pool->attrs->affn_strict = true; pool->node = cpu_to_node(cpu); @@ -8339,7 +8382,7 @@ static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int n void __init workqueue_init_early(void) { struct wq_pod_type *pt = &wq_pod_types[WQ_AFFN_SYSTEM]; - int std_nice[NR_STD_WORKER_POOLS] = { 0, HIGHPRI_NICE_LEVEL }; + int std_prio[NR_STD_WORKER_POOLS] = { DEFAULT_PRIO, HIGHPRI_PRIORITY }; void (*irq_work_fns[NR_STD_WORKER_POOLS])(struct irq_work *) = { bh_pool_kick_normal, bh_pool_kick_highpri }; int i, cpu; @@ -8391,7 +8434,7 @@ void __init workqueue_init_early(void) i = 0; for_each_bh_worker_pool(pool, cpu) { - init_cpu_worker_pool(pool, cpu, std_nice[i]); + init_cpu_worker_pool(pool, cpu, std_prio[i]); pool->flags |= POOL_BH; init_irq_work(bh_pool_irq_work(pool), irq_work_fns[i]); i++; @@ -8399,7 +8442,7 @@ void __init workqueue_init_early(void) i = 0; for_each_cpu_worker_pool(pool, cpu) - init_cpu_worker_pool(pool, cpu, std_nice[i++]); + init_cpu_worker_pool(pool, cpu, std_prio[i++]); } system_wq = alloc_workqueue("events", WQ_PERCPU | __WQ_DEPRECATED, 0); diff --git a/tools/workqueue/wq_dump.py b/tools/workqueue/wq_dump.py index 9313ebe0c525..371601b086ca 100644 --- a/tools/workqueue/wq_dump.py +++ b/tools/workqueue/wq_dump.py @@ -24,7 +24,7 @@ Worker Pools Lists all worker pools indexed by their ID. For each pool: ref number of pool_workqueue's associated with this pool - nice nice value of the worker threads in the pool + prio priority of the worker threads in the pool idle number of idle workers workers number of all workers cpu CPU the pool is associated with (per-cpu pool) @@ -122,6 +122,8 @@ POOL_BH = prog['POOL_BH'] WQ_NAME_LEN = prog['WQ_NAME_LEN'].value_() cpumask_str_len = len(cpumask_str(wq_unbound_cpumask)) +rt_prio = prog.constant('RT_PRIORITY', filename='kernel/workqueue.c') + print('Affinity Scopes') print('===============') @@ -163,7 +165,10 @@ for pi, pool in idr_for_each(worker_pool_idr): for pi, pool in idr_for_each(worker_pool_idr): pool = drgn.Object(prog, 'struct worker_pool', address=pool) - print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} nice={pool.attrs.nice.value_():3} ', end='') + prio = pool.attrs.prio.value_() + if prio == rt_prio: + prio = 'rt' + print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} prio={prio:3} ', end='') print(f'idle/workers={pool.nr_idle.value_():3}/{pool.nr_workers.value_():3} ', end='') if pool.cpu >= 0: print(f'cpu={pool.cpu.value_():3}', end='') -- 2.55.0 ^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC v6.1 2/3] workqueue: Add support for real-time workers 2026-10-01 16:07 ` [RFC v6 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin @ 2026-10-01 18:48 ` Tvrtko Ursulin 2026-10-02 19:30 ` Tejun Heo 0 siblings, 1 reply; 8+ messages in thread From: Tvrtko Ursulin @ 2026-10-01 18:48 UTC (permalink / raw) To: linux-kernel Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon, Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo For use cases such as the DRM scheduler submitting work to the GPU on behalf of low latency userspace applications, where latter have sufficient privileges to have had successfully obtained realtime Vulkan global priority, competing with random background CPU load can create large latency spikes which gets in the way of a smooth user experience. For these situations the existing WQ_HIGHPRI does not bring a noticeable improvement and a stronger hint is needed. Lets add WQ_RT which creates workers with a SCHED_FIFO scheduling class to improve this. We use a minimum priority level since we only care about winning the contest against normal background CPU load. Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com> Cc: Boris Brezillon <boris.brezillon@collabora.com> Cc: Bradley Morgan <include@grrlz.net> Cc: Chia-I Wu <olv@google.com> Cc: Liviu Dudau <liviu.dudau@arm.com> Cc: Matthew Brost <matthew.brost@intel.com> Cc: Steven Price <steven.price@arm.com> Cc: Tejun Heo <tj@kernel.org> --- v2: * Limit WQ_RTPRI to unbound workqueues and make it have strict CPU affinitity. (Tejun) * Fixed commit message typos. (AI) * Fixed sysfs handling, max_active setting and user modified nice application. (AI) v3: * Fix worker->pool null pointer dereference race by moving the global decrement to detach_dying_workers(). * Rebase for upstream changes. v4: * Fixed onion unwind. * Moved affinity setting to default attributes. v5: * Dropped global and local limits. * Documented in workqueue.rst. * Added NR_WQ_ATTRIBUTES. * Reverted BH handling changes. v6: * Dropped separate attr->prio in favour of RTPRI_NICE_LEVEL checks. (Tejun) * Reworked on top of tj/for-7.4. v7: * Convert to attrs->prio encoded analoguous to task_struct->prio. * Rename flag to WQ_PRIO and do not re-order enums. * Forbid WQ_RT affinity modifications via sysfs. * Added wq_dump.py support. v8: * Fix max rt prio confusion. * Initialize priority also in alloc_workqueue_attrs_noprof. * Restrict rt cpumask modifications. --- Documentation/core-api/workqueue.rst | 8 ++ include/linux/workqueue.h | 5 +- kernel/workqueue.c | 106 +++++++++++++++++++-------- tools/workqueue/wq_dump.py | 9 ++- 4 files changed, 95 insertions(+), 33 deletions(-) diff --git a/Documentation/core-api/workqueue.rst b/Documentation/core-api/workqueue.rst index bb770f556568..d699c3832b19 100644 --- a/Documentation/core-api/workqueue.rst +++ b/Documentation/core-api/workqueue.rst @@ -225,6 +225,14 @@ resources, scheduled and executed. each other. Each maintains its separate pool of workers and implements concurrency management among its workers. +``WQ_RT`` + Real-time priority workqueues must be created as unbound and will be + configured with the strict CPU affinity set. Their worker threads use the FIFO + scheduling policy with the lowest applicable priority. + + To be used sparingly for use cases such as the real-time GPU rendering + contexts accessible to privileged clients. + ``WQ_CPU_INTENSIVE`` Work items of a CPU intensive wq do not contribute to the concurrency level. In other words, runnable CPU intensive diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h index a283766a192a..ebec9dcc9e5f 100644 --- a/include/linux/workqueue.h +++ b/include/linux/workqueue.h @@ -147,9 +147,9 @@ enum wq_affn_scope { */ struct workqueue_attrs { /** - * @nice: nice level + * @prio: priority encoded analoguous to task_struct->prio. */ - int nice; + int prio; /** * @cpumask: allowed CPUs @@ -404,6 +404,7 @@ enum wq_flags { */ WQ_POWER_EFFICIENT = 1 << 7, WQ_PERCPU = 1 << 8, /* bound to a specific cpu */ + WQ_RT = 1 << 9, /* real-time priority, valid only with WQ_UNBOUND */ __WQ_DESTROYING = 1 << 15, /* internal: workqueue is destroying */ __WQ_DRAINING = 1 << 16, /* internal: workqueue is draining */ diff --git a/kernel/workqueue.c b/kernel/workqueue.c index c83d68d7d0ee..31f3ea0dd62e 100644 --- a/kernel/workqueue.c +++ b/kernel/workqueue.c @@ -47,6 +47,7 @@ #include <linux/jhash.h> #include <linux/hashtable.h> #include <linux/rculist.h> +#include <linux/sched/rt.h> #include <linux/nodemask.h> #include <linux/moduleparam.h> #include <linux/uaccess.h> @@ -126,7 +127,8 @@ enum wq_internal_consts { * all cpus. Give MIN_NICE. */ RESCUER_NICE_LEVEL = MIN_NICE, - HIGHPRI_NICE_LEVEL = MIN_NICE, + HIGHPRI_PRIORITY = NICE_TO_PRIO(MIN_NICE), + RT_PRIORITY = MAX_RT_PRIO - 1, WQ_NAME_LEN = 32, WORKER_ID_LEN = 10 + WQ_NAME_LEN, /* "kworker/R-" + WQ_NAME_LEN */ @@ -1275,7 +1277,7 @@ static bool assign_work(struct work_struct *work, struct worker *worker, static struct irq_work *bh_pool_irq_work(struct worker_pool *pool) { - int high = pool->attrs->nice == HIGHPRI_NICE_LEVEL ? 1 : 0; + int high = pool->attrs->prio == HIGHPRI_PRIORITY ? 1 : 0; return &per_cpu(bh_pool_irq_works, pool->cpu)[high]; } @@ -1290,7 +1292,7 @@ static void kick_bh_pool(struct worker_pool *pool) return; } #endif - if (pool->attrs->nice == HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio == HIGHPRI_PRIORITY) raise_softirq_irqoff(HI_SOFTIRQ); else raise_softirq_irqoff(TASKLET_SOFTIRQ); @@ -2959,7 +2961,8 @@ static int format_worker_id(char *buf, size_t size, struct worker *worker, if (pool->cpu >= 0) return scnprintf(buf, size, "kworker/%d:%d%s", pool->cpu, worker->id, - pool->attrs->nice < 0 ? "H" : ""); + pool->attrs->prio < NICE_TO_PRIO(0) ? + "H" : ""); else return scnprintf(buf, size, "kworker/u%d:%d", pool->id, worker->id); @@ -3018,7 +3021,12 @@ static struct worker *create_worker(struct worker_pool *pool) goto fail; } - set_user_nice(worker->task, pool->attrs->nice); + if (rt_prio(pool->attrs->prio)) + sched_set_fifo_low(worker->task); + else + set_user_nice(worker->task, + PRIO_TO_NICE(pool->attrs->prio)); + kthread_bind_mask(worker->task, pool_allowed_cpus(pool)); } @@ -3910,7 +3918,7 @@ static void bh_worker(struct worker *worker) if (budget_exhausted) trace_workqueue_bh_budget_yield(pool, restarts, timeout, - pool->attrs->nice == HIGHPRI_NICE_LEVEL); + pool->attrs->prio == HIGHPRI_PRIORITY); } /* @@ -3969,7 +3977,7 @@ static void drain_dead_softirq_workfn(struct work_struct *work) * don't hog this CPU's BH. */ if (repeat) { - if (pool->attrs->nice == HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio == HIGHPRI_PRIORITY) queue_work(system_bh_highpri_wq, work); else queue_work(system_bh_wq, work); @@ -4001,7 +4009,7 @@ void workqueue_softirq_dead(unsigned int cpu) dead_work.pool = pool; init_completion(&dead_work.done); - if (pool->attrs->nice == HIGHPRI_NICE_LEVEL) + if (pool->attrs->prio == HIGHPRI_PRIORITY) queue_work(system_bh_highpri_wq, &dead_work.work); else queue_work(system_bh_wq, &dead_work.work); @@ -5005,6 +5013,7 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(void) goto fail; cpumask_copy(attrs->cpumask, cpu_possible_mask); + attrs->prio = DEFAULT_PRIO; attrs->affn_scope = WQ_AFFN_DFL; return attrs; fail: @@ -5015,7 +5024,7 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(void) static void copy_workqueue_attrs(struct workqueue_attrs *to, const struct workqueue_attrs *from) { - to->nice = from->nice; + to->prio = from->prio; cpumask_copy(to->cpumask, from->cpumask); cpumask_copy(to->__pod_cpumask, from->__pod_cpumask); to->affn_strict = from->affn_strict; @@ -5046,7 +5055,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs *attrs) { u32 hash = 0; - hash = jhash_1word(attrs->nice, hash); + hash = jhash_1word(attrs->prio, hash); hash = jhash_1word(attrs->affn_strict, hash); hash = jhash(cpumask_bits(attrs->__pod_cpumask), BITS_TO_LONGS(nr_cpumask_bits) * sizeof(long), hash); @@ -5060,7 +5069,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs *attrs) static bool wqattrs_equal(const struct workqueue_attrs *a, const struct workqueue_attrs *b) { - if (a->nice != b->nice) + if (a->prio != b->prio) return false; if (a->affn_strict != b->affn_strict) return false; @@ -5928,8 +5937,19 @@ static struct workqueue_attrs *alloc_wq_std_attrs(struct workqueue_struct *wq) if (!attrs) return NULL; - if (wq->flags & WQ_HIGHPRI) - attrs->nice = HIGHPRI_NICE_LEVEL; + if (wq->flags & WQ_RT) { + attrs->prio = RT_PRIORITY; + /* + * RT workqueues have strict CPU affinity for low + * latency execution. + */ + attrs->affn_scope = WQ_AFFN_CPU; + attrs->affn_strict = true; + } else if (wq->flags & WQ_HIGHPRI) { + attrs->prio = HIGHPRI_PRIORITY; + } else { + attrs->prio = DEFAULT_PRIO; + } if (wq->flags & __WQ_ORDERED) attrs->ordered = true; @@ -6115,6 +6135,12 @@ static struct workqueue_struct *__alloc_workqueue(const char *fmt, return NULL; } + if (flags & WQ_RT) { + if (WARN_ON_ONCE((flags & (WQ_HIGHPRI | WQ_UNBOUND)) != + WQ_UNBOUND)) + return NULL; + } + /* see the comment above the definition of WQ_POWER_EFFICIENT */ if ((flags & WQ_POWER_EFFICIENT) && wq_power_efficient) flags = (flags & ~WQ_PERCPU) | WQ_UNBOUND; @@ -6671,9 +6697,9 @@ static void pr_cont_pool_info(struct worker_pool *pool) pr_cont(" flags=0x%x", pool->flags); if (pool->flags & POOL_BH) pr_cont(" bh%s", - pool->attrs->nice == HIGHPRI_NICE_LEVEL ? "-hi" : ""); + pool->attrs->prio == HIGHPRI_PRIORITY ? "-hi" : ""); else - pr_cont(" nice=%d", pool->attrs->nice); + pr_cont(" nice=%d", PRIO_TO_NICE(pool->attrs->prio)); } static void pr_cont_worker_id(struct worker *worker) @@ -6682,7 +6708,7 @@ static void pr_cont_worker_id(struct worker *worker) if (pool->flags & POOL_BH) pr_cont("bh%s", - pool->attrs->nice == HIGHPRI_NICE_LEVEL ? "-hi" : ""); + pool->attrs->prio == HIGHPRI_PRIORITY ? "-hi" : ""); else pr_cont("%d%s", task_pid_nr(worker->task), worker->rescue_wq ? "(RESCUER)" : ""); @@ -7606,7 +7632,11 @@ static ssize_t nice_show(struct device *dev, struct device_attribute *attr, int written; mutex_lock(&wq->mutex); - written = scnprintf(buf, PAGE_SIZE, "%d\n", wq->attrs->nice); + if (wq->attrs->prio == RT_PRIORITY) + written = scnprintf(buf, PAGE_SIZE, "rt\n"); + else + written = scnprintf(buf, PAGE_SIZE, "%d\n", + PRIO_TO_NICE(wq->attrs->prio)); mutex_unlock(&wq->mutex); return written; @@ -7632,19 +7662,21 @@ static ssize_t nice_store(struct device *dev, struct device_attribute *attr, { struct workqueue_struct *wq = dev_to_wq(dev); struct workqueue_attrs *attrs; - int ret = -ENOMEM; + int ret, nice = 0; + + if (sscanf(buf, "%d", &nice) != 1 || nice < MIN_NICE || nice > MAX_NICE) + return -EINVAL; mutex_lock(&wq_pool_mutex); attrs = wq_sysfs_prep_attrs(wq); - if (!attrs) + if (!attrs) { + ret = -ENOMEM; goto out_unlock; + } - if (sscanf(buf, "%d", &attrs->nice) == 1 && - attrs->nice >= MIN_NICE && attrs->nice <= MAX_NICE) - ret = apply_workqueue_attrs_locked(wq, attrs); - else - ret = -EINVAL; + attrs->prio = NICE_TO_PRIO(nice); + ret = apply_workqueue_attrs_locked(wq, attrs); out_unlock: mutex_unlock(&wq_pool_mutex); @@ -7673,6 +7705,10 @@ static ssize_t unbound_cpumask_store(struct device *dev, struct workqueue_attrs *attrs; int ret = -ENOMEM; + /* Do not allow cpumask changes for RT workers. */ + if (wq->flags & WQ_RT) + return -EINVAL; + mutex_lock(&wq_pool_mutex); attrs = wq_sysfs_prep_attrs(wq); @@ -7716,6 +7752,10 @@ static ssize_t affinity_scope_store(struct device *dev, struct workqueue_attrs *attrs; int affn, ret = -ENOMEM; + /* Do not allow affinity changes for RT workers. */ + if (wq->flags & WQ_RT) + return -EINVAL; + affn = parse_affn_scope(buf); if (affn < 0) return affn; @@ -7748,6 +7788,10 @@ static ssize_t affinity_strict_store(struct device *dev, struct workqueue_attrs *attrs; int v, ret = -ENOMEM; + /* Do not allow affinity changes for RT workers. */ + if (wq->flags & WQ_RT) + return -EINVAL; + if (sscanf(buf, "%d", &v) != 1) return -EINVAL; @@ -7786,6 +7830,10 @@ static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj, if (!(wq->flags & WQ_UNBOUND)) return SYSFS_GROUP_INVISIBLE; + /* Do not allow priority changes for RT workers. */ + if ((wq->flags & WQ_RT) && !strcmp(attr->name, "nice")) + return 0444; + return attr->mode; } @@ -8310,13 +8358,13 @@ static void __init restrict_unbound_cpumask(const char *name, const struct cpuma cpumask_and(wq_unbound_cpumask, wq_unbound_cpumask, mask); } -static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int nice) +static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int prio) { BUG_ON(init_worker_pool(pool)); pool->cpu = cpu; cpumask_copy(pool->attrs->cpumask, cpumask_of(cpu)); cpumask_copy(pool->attrs->__pod_cpumask, cpumask_of(cpu)); - pool->attrs->nice = nice; + pool->attrs->prio = prio; pool->attrs->affn_strict = true; pool->node = cpu_to_node(cpu); @@ -8339,7 +8387,7 @@ static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int n void __init workqueue_init_early(void) { struct wq_pod_type *pt = &wq_pod_types[WQ_AFFN_SYSTEM]; - int std_nice[NR_STD_WORKER_POOLS] = { 0, HIGHPRI_NICE_LEVEL }; + int std_prio[NR_STD_WORKER_POOLS] = { DEFAULT_PRIO, HIGHPRI_PRIORITY }; void (*irq_work_fns[NR_STD_WORKER_POOLS])(struct irq_work *) = { bh_pool_kick_normal, bh_pool_kick_highpri }; int i, cpu; @@ -8391,7 +8439,7 @@ void __init workqueue_init_early(void) i = 0; for_each_bh_worker_pool(pool, cpu) { - init_cpu_worker_pool(pool, cpu, std_nice[i]); + init_cpu_worker_pool(pool, cpu, std_prio[i]); pool->flags |= POOL_BH; init_irq_work(bh_pool_irq_work(pool), irq_work_fns[i]); i++; @@ -8399,7 +8447,7 @@ void __init workqueue_init_early(void) i = 0; for_each_cpu_worker_pool(pool, cpu) - init_cpu_worker_pool(pool, cpu, std_nice[i++]); + init_cpu_worker_pool(pool, cpu, std_prio[i++]); } system_wq = alloc_workqueue("events", WQ_PERCPU | __WQ_DEPRECATED, 0); diff --git a/tools/workqueue/wq_dump.py b/tools/workqueue/wq_dump.py index 9313ebe0c525..371601b086ca 100644 --- a/tools/workqueue/wq_dump.py +++ b/tools/workqueue/wq_dump.py @@ -24,7 +24,7 @@ Worker Pools Lists all worker pools indexed by their ID. For each pool: ref number of pool_workqueue's associated with this pool - nice nice value of the worker threads in the pool + prio priority of the worker threads in the pool idle number of idle workers workers number of all workers cpu CPU the pool is associated with (per-cpu pool) @@ -122,6 +122,8 @@ POOL_BH = prog['POOL_BH'] WQ_NAME_LEN = prog['WQ_NAME_LEN'].value_() cpumask_str_len = len(cpumask_str(wq_unbound_cpumask)) +rt_prio = prog.constant('RT_PRIORITY', filename='kernel/workqueue.c') + print('Affinity Scopes') print('===============') @@ -163,7 +165,10 @@ for pi, pool in idr_for_each(worker_pool_idr): for pi, pool in idr_for_each(worker_pool_idr): pool = drgn.Object(prog, 'struct worker_pool', address=pool) - print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} nice={pool.attrs.nice.value_():3} ', end='') + prio = pool.attrs.prio.value_() + if prio == rt_prio: + prio = 'rt' + print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} prio={prio:3} ', end='') print(f'idle/workers={pool.nr_idle.value_():3}/{pool.nr_workers.value_():3} ', end='') if pool.cpu >= 0: print(f'cpu={pool.cpu.value_():3}', end='') -- 2.55.0 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [RFC v6.1 2/3] workqueue: Add support for real-time workers 2026-10-01 18:48 ` [RFC v6.1 " Tvrtko Ursulin @ 2026-10-02 19:30 ` Tejun Heo 0 siblings, 0 replies; 8+ messages in thread From: Tejun Heo @ 2026-10-02 19:30 UTC (permalink / raw) To: Tvrtko Ursulin Cc: linux-kernel, dri-devel, kernel-dev, Boris Brezillon, Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Lai Jiangshan, Breno Leitao Hello, The following is a Claude-generated review. On Thu, Oct 01, 2026 at 07:48:48PM +0100, Tvrtko Ursulin wrote: > For use cases such as the DRM scheduler submitting work to the GPU on > behalf of low latency userspace applications, where latter have sufficient > privileges to have had successfully obtained realtime Vulkan global > priority, competing with random background CPU load can create large > latency spikes which gets in the way of a smooth user experience. Can you add the compositor / DRM master usage model from the v5 discussion here? The cover letter also still says CAP_SYS_NICE is required while group_priority_permit() accepts DRM master too. > + HIGHPRI_PRIORITY = NICE_TO_PRIO(MIN_NICE), > + RT_PRIORITY = MAX_RT_PRIO - 1, MAX_RT_PRIO - 1 is what rt_priority 0 would map to, while sched_set_fifo_low() gives the workers prio 98. Nothing depends on it today, but maybe MAX_RT_PRIO - 2 with a comment tying it to sched_set_fifo_low(), or soften the kerneldoc which says the encoding matches task_struct->prio? Also, s/analoguous/analogous/ there. > + if (wq->flags & WQ_RT) { > + attrs->prio = RT_PRIORITY; > + /* > + * RT workqueues have strict CPU affinity for low > + * latency execution. > + */ > + attrs->affn_scope = WQ_AFFN_CPU; > + attrs->affn_strict = true; apply_workqueue_attrs() is unrestricted, so a caller can hand a WQ_RT wq the attrs from alloc_workqueue_attrs() and turn it into a normal non-strict wq while the flag stays set. Maybe apply_wqattrs_prepare() should force prio and affinity for WQ_RT instead of restricting only the sysfs side? > + if (flags & WQ_RT) { > + if (WARN_ON_ONCE((flags & (WQ_HIGHPRI | WQ_UNBOUND)) != > + WQ_UNBOUND)) > + return NULL; > + } alloc_ordered_workqueue() with WQ_RT passes this and ends up with a single pool spanning all CPUs because ordered wqs use dfl_pwq everywhere, while the doc and sysfs say strict per-CPU. Should __WQ_ORDERED be rejected too? The nested ifs can also be a single condition. > - pr_cont(" nice=%d", pool->attrs->nice); > + pr_cont(" nice=%d", PRIO_TO_NICE(pool->attrs->prio)); This prints nice=-21 for RT pools. Can you show "rt" here like nice_show()? > + /* Do not allow cpumask changes for RT workers. */ > + if (wq->flags & WQ_RT) > + return -EINVAL; The three stores check WQ_RT and return -EINVAL while nice goes read-only through is_visible below. Can all four go through wq_sysfs_unbound_group_visible() returning 0444 for WQ_RT? That drops the three store checks. Note that the global cpumask still applies to WQ_RT wqs through workqueue_apply_unbound_cpumask(), so the comment overstates a bit. The interface comment at the top of the sysfs section also still says nice is RW int. > + /* Do not allow priority changes for RT workers. */ > + if ((wq->flags & WQ_RT) && !strcmp(attr->name, "nice")) > + return 0444; attr == &dev_attr_nice.attr would avoid the strcmp. > + prio = pool.attrs.prio.value_() > + if prio == rt_prio: > + prio = 'rt' > + print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} prio={prio:3} ', end='') Can we keep printing nice for non-RT pools? prio=120 is the internal encoding, sysfs and the pool dumps print nice, and the example output in workqueue.rst would go stale. prog['RT_PRIORITY'] would match the rest of the file too. One more thing which isn't in the diff. The rescuer of a WQ_RT | WQ_MEM_RECLAIM wq, which is what panthor-drm-rt is, still runs at nice -20, so under memory pressure the RT wq's work items run as CFS. Should rescuer_thread() use sched_set_fifo_low() for WQ_RT? Thanks. -- tejun ^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC v6 3/3] drm/panthor: Create per queue priority workqueues 2026-10-01 16:07 [RFC v6 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin 2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin 2026-10-01 16:07 ` [RFC v6 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin @ 2026-10-01 16:07 ` Tvrtko Ursulin 2026-10-02 19:30 ` Tejun Heo 2 siblings, 1 reply; 8+ messages in thread From: Tvrtko Ursulin @ 2026-10-01 16:07 UTC (permalink / raw) To: linux-kernel Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon, Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo Split the single workqueue shared between the driver internal logic and DRM scheduler use into separate ones, where the DRM scheduler one is created per GPU priority level using the appropriate mapping to workqueue priorities. Low and medium GPU priority are served by a normal workqueue, high is server by a WQ_HIGHPRI instance, while realtime GPU priority is using the newly added WQ_RTPRI flag for lowest possible latency. These workqueues are device global and for all three we set the maximum concurrency to two in order to keep the GPU optimally fed with work. Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com> Cc: Boris Brezillon <boris.brezillon@collabora.com> Cc: Chia-I Wu <olv@google.com> Cc: Liviu Dudau <liviu.dudau@arm.com> Cc: Matthew Brost <matthew.brost@intel.com> Cc: Steven Price <steven.price@arm.com> Cc: Tejun Heo <tj@kernel.org> --- drivers/gpu/drm/panthor/panthor_sched.c | 37 ++++++++++++++++++++++--- 1 file changed, 33 insertions(+), 4 deletions(-) diff --git a/drivers/gpu/drm/panthor/panthor_sched.c b/drivers/gpu/drm/panthor/panthor_sched.c index 5b34032deff8..257895e5e48b 100644 --- a/drivers/gpu/drm/panthor/panthor_sched.c +++ b/drivers/gpu/drm/panthor/panthor_sched.c @@ -152,11 +152,18 @@ struct panthor_scheduler { * * Used for the scheduler tick, group update or other kind of FW * event processing that can't be handled in the threaded interrupt - * path. Also passed to the drm_gpu_scheduler instances embedded - * in panthor_queue. + * path. */ struct workqueue_struct *wq; + /** + * @submit_wq: Per priority workqueues for the DRM scheduler + * + * Passed to the drm_gpu_scheduler instances embedded + * in panthor_queue based on the queue priority. + */ + struct workqueue_struct *submit_wq[PANTHOR_CSG_PRIORITY_COUNT]; + /** * @heap_alloc_wq: Workqueue used to schedule tiler_oom works. * @@ -3582,8 +3589,14 @@ group_create_queue(struct panthor_group *group, goto err_free_queue; } + if (group->priority >= ARRAY_SIZE(group->ptdev->scheduler->submit_wq) || + !group->ptdev->scheduler->submit_wq[group->priority]) { + ret = -EINVAL; + goto err_free_queue; + } + sched_args.name = queue->name; - + sched_args.submit_wq = group->ptdev->scheduler->submit_wq[group->priority]; ret = drm_sched_init(&queue->scheduler, &sched_args); if (ret) goto err_free_queue; @@ -4073,6 +4086,15 @@ static void panthor_sched_fini(struct drm_device *ddev, void *res) if (!sched || !sched->csg_slot_count) return; + if (sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM]) + destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM]); + + if (sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH]) + destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH]); + + if (sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]) + destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]); + if (sched->wq) destroy_workqueue(sched->wq); @@ -4174,7 +4196,14 @@ int panthor_sched_init(struct panthor_device *ptdev) */ sched->heap_alloc_wq = alloc_workqueue("panthor-heap-alloc", WQ_UNBOUND, 0); sched->wq = alloc_workqueue("panthor-csf-sched", WQ_MEM_RECLAIM | WQ_UNBOUND, 0); - if (!sched->wq || !sched->heap_alloc_wq) { + sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] = alloc_workqueue("panthor-drm", WQ_MEM_RECLAIM | WQ_UNBOUND, 2); + sched->submit_wq[PANTHOR_CSG_PRIORITY_LOW] = sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM]; + sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] = alloc_workqueue("panthor-drm-high", WQ_HIGHPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2); + sched->submit_wq[PANTHOR_CSG_PRIORITY_RT] = alloc_workqueue("panthor-drm-rt", WQ_RT | WQ_MEM_RECLAIM | WQ_UNBOUND, 2); + if (!sched->wq || !sched->heap_alloc_wq || + !sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] || + !sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] || + !sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]) { panthor_sched_fini(&ptdev->base, sched); drm_err(&ptdev->base, "Failed to allocate the workqueues"); return -ENOMEM; -- 2.55.0 ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [RFC v6 3/3] drm/panthor: Create per queue priority workqueues 2026-10-01 16:07 ` [RFC v6 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin @ 2026-10-02 19:30 ` Tejun Heo 0 siblings, 0 replies; 8+ messages in thread From: Tejun Heo @ 2026-10-02 19:30 UTC (permalink / raw) To: Tvrtko Ursulin Cc: linux-kernel, dri-devel, kernel-dev, Boris Brezillon, Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price Hello, The following is a Claude-generated review. On Thu, Oct 01, 2026 at 05:07:11PM +0100, Tvrtko Ursulin wrote: > Low and medium GPU priority are served by a normal workqueue, > high is server by a WQ_HIGHPRI instance, while realtime GPU priority is > using the newly added WQ_RTPRI flag for lowest possible latency. s/server/served/ and the flag is WQ_RT now. > + if (group->priority >= ARRAY_SIZE(group->ptdev->scheduler->submit_wq) || > + !group->ptdev->scheduler->submit_wq[group->priority]) { > + ret = -EINVAL; > + goto err_free_queue; > + } panthor_group_create() already rejects priorities >= PANTHOR_CSG_PRIORITY_COUNT and all slots are populated once panthor_sched_init() succeeded, so this can't fire. > + sched_args.submit_wq = group->ptdev->scheduler->submit_wq[group->priority]; The .submit_wq = sched->wq initializer at the top of the function is still there and gets overwritten here. > + if (sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]) > + destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]); > + > if (sched->wq) > destroy_workqueue(sched->wq); Work on sched->wq signals job fences whose callbacks queue onto the submit wqs, so destroying sched->wq first would be the safer order. > + sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] = alloc_workqueue("panthor-drm", WQ_MEM_RECLAIM | WQ_UNBOUND, 2); > + sched->submit_wq[PANTHOR_CSG_PRIORITY_LOW] = sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM]; > + sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] = alloc_workqueue("panthor-drm-high", WQ_HIGHPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2); > + sched->submit_wq[PANTHOR_CSG_PRIORITY_RT] = alloc_workqueue("panthor-drm-rt", WQ_RT | WQ_MEM_RECLAIM | WQ_UNBOUND, 2); For an unbound wq max_active applies to the whole wq, so this is two in-flight items per priority level across all the queues on the device, where the shared wq before had no effective limit. queue_run_job() blocks on sched->lock which tick_work() holds across FW round trips, so two blocked run_jobs would stall every other queue's run and free work at that level. What's the reason for 2? Also, these lines are over 100 columns. Thanks. -- tejun ^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-10-02 19:30 UTC | newest] Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2026-10-01 16:07 [RFC v6 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin 2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin 2026-10-02 19:30 ` Tejun Heo 2026-10-01 16:07 ` [RFC v6 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin 2026-10-01 18:48 ` [RFC v6.1 " Tvrtko Ursulin 2026-10-02 19:30 ` Tejun Heo 2026-10-01 16:07 ` [RFC v6 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin 2026-10-02 19:30 ` Tejun Heo
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®