* [RFC v6 0/3] Realtime workqueues and panthor realtime submission
@ 2026-10-01 16:07 Tvrtko Ursulin
2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
` (2 more replies)
0 siblings, 3 replies; 8+ messages in thread
From: Tvrtko Ursulin @ 2026-10-01 16:07 UTC (permalink / raw)
To: linux-kernel
Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon,
Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo
This is a continuation of the previous discussion which was here:
https://lore.kernel.org/dri-devel/20260702143745.79293-1-tvrtko.ursulin@igalia.com/
Work is now converted to a much simpler approach by adding real-time scheduling
workqueues based on Tejun's feedback
To re-cap, when an userspace thread submits GPU work, due how the DRM scheduler
uses workqueues to feed the GPU, and regardless of the GPU rendering context
priority, or the CPU scheduling priority of the userspace thread itself, the
use of workqueues can add significant latency to the submit path.
When CPU is busy with enough backround load this translates to severe latency
spikes measured as time between userspace submitting work and GPU actually being
given that work to execute.
With the panthor workqueue upgraded to use WQ_HIGHPRI and varying the CPU
priority of the submit thread, the test program from
https://gitlab.freedesktop.org/panfrost/linux/-/work_items/49 reproduces these
kind of latencies:
. N RT
M 27 28 32
95% 163 246 809
98% 924 991 1882
Legend:
M = Median submit latency in us
95% = Percentile latency in us
. = Userspace submit thread SCHED_OTHER
N = -||= nice -1
RT = -||- FIFO 1
Upgrading the panthor workqueues so that the realtime GPU priority queue uses
WQ_RT, submit latency becomes completely controlled with the median of 14us and
95 and 98-th percentiles at 23us and 25us respectively.
Important to note is that VK_QUEUE_GLOBAL_PRIORITY_REALTIME already required
the userspace to have CAP_SYS_NICE, meaning access to real-time workqueues is
effectively also guared behind this capability.
v2:
* See patch 1 changelog.
v3:
* See patch 1 changelog + apologies for v3 following so quickly after v2.
I have found a race condition as I expanded the testing to a second platform.
v4:
* See patch 1 changelog.
v5:
* Rebased on top of tj/for-7.4 branch.
* New patch in series simplifies unbound sysfs attribute handling.
* WQ_RTPRI reworked to fold RT marker into attr->nice.
v6:
* Fixed accidental sysfs attribute rename (patch 1).
* Various review feedback as per patch 2 changelog.
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
Tvrtko Ursulin (3):
workqueue: Simplify unbound sysfs attribute registration
workqueue: Add support for real-time workers
drm/panthor: Create per queue priority workqueues
Documentation/core-api/workqueue.rst | 8 +
drivers/gpu/drm/panthor/panthor_sched.c | 37 ++++-
include/linux/workqueue.h | 5 +-
kernel/workqueue.c | 199 +++++++++++++++---------
tools/workqueue/wq_dump.py | 9 +-
5 files changed, 175 insertions(+), 83 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration
2026-10-01 16:07 [RFC v6 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin
@ 2026-10-01 16:07 ` Tvrtko Ursulin
2026-10-02 19:30 ` Tejun Heo
2026-10-01 16:07 ` [RFC v6 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin
2026-10-01 16:07 ` [RFC v6 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin
2 siblings, 1 reply; 8+ messages in thread
From: Tvrtko Ursulin @ 2026-10-01 16:07 UTC (permalink / raw)
To: linux-kernel
Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon,
Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost,
Steven Price, Tejun Heo
Instead of manually registering each attribute we can put them in an
attribute group with a visibility check and device core will handle the
rest, which simplifies the registration and error unwind.
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Bradley Morgan <include@grrlz.net>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
---
v2:
* Fix accidental rename of affinity_scope.
---
kernel/workqueue.c | 98 ++++++++++++++++++++++++----------------------
1 file changed, 52 insertions(+), 46 deletions(-)
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index e618108c6127..c83d68d7d0ee 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -7599,8 +7599,8 @@ static const struct attribute_group wq_sysfs_group = {
};
__ATTRIBUTE_GROUPS(wq_sysfs);
-static ssize_t wq_nice_show(struct device *dev, struct device_attribute *attr,
- char *buf)
+static ssize_t nice_show(struct device *dev, struct device_attribute *attr,
+ char *buf)
{
struct workqueue_struct *wq = dev_to_wq(dev);
int written;
@@ -7627,8 +7627,8 @@ static struct workqueue_attrs *wq_sysfs_prep_attrs(struct workqueue_struct *wq)
return attrs;
}
-static ssize_t wq_nice_store(struct device *dev, struct device_attribute *attr,
- const char *buf, size_t count)
+static ssize_t nice_store(struct device *dev, struct device_attribute *attr,
+ const char *buf, size_t count)
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
@@ -7652,8 +7652,8 @@ static ssize_t wq_nice_store(struct device *dev, struct device_attribute *attr,
return ret ?: count;
}
-static ssize_t wq_cpumask_show(struct device *dev,
- struct device_attribute *attr, char *buf)
+static ssize_t unbound_cpumask_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
{
struct workqueue_struct *wq = dev_to_wq(dev);
int written;
@@ -7665,9 +7665,9 @@ static ssize_t wq_cpumask_show(struct device *dev,
return written;
}
-static ssize_t wq_cpumask_store(struct device *dev,
- struct device_attribute *attr,
- const char *buf, size_t count)
+static ssize_t unbound_cpumask_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t count)
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
@@ -7689,8 +7689,8 @@ static ssize_t wq_cpumask_store(struct device *dev,
return ret ?: count;
}
-static ssize_t wq_affn_scope_show(struct device *dev,
- struct device_attribute *attr, char *buf)
+static ssize_t affinity_scope_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
{
struct workqueue_struct *wq = dev_to_wq(dev);
int written;
@@ -7708,9 +7708,9 @@ static ssize_t wq_affn_scope_show(struct device *dev,
return written;
}
-static ssize_t wq_affn_scope_store(struct device *dev,
- struct device_attribute *attr,
- const char *buf, size_t count)
+static ssize_t affinity_scope_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t count)
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
@@ -7731,8 +7731,8 @@ static ssize_t wq_affn_scope_store(struct device *dev,
return ret ?: count;
}
-static ssize_t wq_affinity_strict_show(struct device *dev,
- struct device_attribute *attr, char *buf)
+static ssize_t affinity_strict_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
{
struct workqueue_struct *wq = dev_to_wq(dev);
@@ -7740,9 +7740,9 @@ static ssize_t wq_affinity_strict_show(struct device *dev,
wq->attrs->affn_strict);
}
-static ssize_t wq_affinity_strict_store(struct device *dev,
- struct device_attribute *attr,
- const char *buf, size_t count)
+static ssize_t affinity_strict_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t count)
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
@@ -7762,14 +7762,40 @@ static ssize_t wq_affinity_strict_store(struct device *dev,
return ret ?: count;
}
-static struct device_attribute wq_sysfs_unbound_attrs[] = {
- __ATTR(nice, 0644, wq_nice_show, wq_nice_store),
- __ATTR(cpumask, 0644, wq_cpumask_show, wq_cpumask_store),
- __ATTR(affinity_scope, 0644, wq_affn_scope_show, wq_affn_scope_store),
- __ATTR(affinity_strict, 0644, wq_affinity_strict_show, wq_affinity_strict_store),
- __ATTR_NULL,
+static DEVICE_ATTR_RW(nice);
+static DEVICE_ATTR_RW(affinity_scope);
+static DEVICE_ATTR_RW(affinity_strict);
+/* Avoid naming clash with the other cpumask */
+static struct device_attribute dev_attr_unbound_cpumask =
+ __ATTR(cpumask, 0644, unbound_cpumask_show, unbound_cpumask_store);
+
+static struct attribute *wq_sysfs_unbound_attrs[] = {
+ &dev_attr_nice.attr,
+ &dev_attr_unbound_cpumask.attr,
+ &dev_attr_affinity_scope.attr,
+ &dev_attr_affinity_strict.attr,
+ NULL,
};
+static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj,
+ struct attribute *attr, int n)
+{
+ struct device *dev = kobj_to_dev(kobj);
+ struct workqueue_struct *wq = dev_to_wq(dev);
+
+ if (!(wq->flags & WQ_UNBOUND))
+ return SYSFS_GROUP_INVISIBLE;
+
+ return attr->mode;
+}
+
+static const struct attribute_group wq_sysfs_unbound_group = {
+ .is_visible = wq_sysfs_unbound_group_visible,
+ .attrs = wq_sysfs_unbound_attrs,
+};
+
+__ATTRIBUTE_GROUPS(wq_sysfs_unbound);
+
static const struct bus_type wq_subsys = {
.name = "workqueue",
.dev_groups = wq_sysfs_groups,
@@ -7907,14 +7933,9 @@ int workqueue_sysfs_register(struct workqueue_struct *wq)
wq_dev->wq = wq;
wq_dev->dev.bus = &wq_subsys;
wq_dev->dev.release = wq_device_release;
+ wq_dev->dev.groups = wq_sysfs_unbound_groups;
dev_set_name(&wq_dev->dev, "%s", wq->name);
- /*
- * attrs are created separately. Suppress uevent until
- * everything is ready.
- */
- dev_set_uevent_suppress(&wq_dev->dev, true);
-
ret = device_register(&wq_dev->dev);
if (ret) {
put_device(&wq_dev->dev);
@@ -7922,21 +7943,6 @@ int workqueue_sysfs_register(struct workqueue_struct *wq)
return ret;
}
- if (wq->flags & WQ_UNBOUND) {
- struct device_attribute *attr;
-
- for (attr = wq_sysfs_unbound_attrs; attr->attr.name; attr++) {
- ret = device_create_file(&wq_dev->dev, attr);
- if (ret) {
- device_unregister(&wq_dev->dev);
- wq->wq_dev = NULL;
- return ret;
- }
- }
- }
-
- dev_set_uevent_suppress(&wq_dev->dev, false);
- kobject_uevent(&wq_dev->dev.kobj, KOBJ_ADD);
return 0;
}
--
2.55.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC v6 2/3] workqueue: Add support for real-time workers
2026-10-01 16:07 [RFC v6 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin
2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
@ 2026-10-01 16:07 ` Tvrtko Ursulin
2026-10-01 18:48 ` [RFC v6.1 " Tvrtko Ursulin
2026-10-01 16:07 ` [RFC v6 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin
2 siblings, 1 reply; 8+ messages in thread
From: Tvrtko Ursulin @ 2026-10-01 16:07 UTC (permalink / raw)
To: linux-kernel
Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon,
Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost,
Steven Price, Tejun Heo
For use cases such as the DRM scheduler submitting work to the GPU on
behalf of low latency userspace applications, where latter have sufficient
privileges to have had successfully obtained realtime Vulkan global
priority, competing with random background CPU load can create large
latency spikes which gets in the way of a smooth user experience.
For these situations the existing WQ_HIGHPRI does not bring a noticeable
improvement and a stronger hint is needed.
Lets add WQ_RT which creates workers with a SCHED_FIFO scheduling class to
improve this.
We use a minimum priority level since we only care about winning the
contest against normal background CPU load.
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Bradley Morgan <include@grrlz.net>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
---
v2:
* Limit WQ_RTPRI to unbound workqueues and make it have strict CPU
affinitity. (Tejun)
* Fixed commit message typos. (AI)
* Fixed sysfs handling, max_active setting and user modified nice
application. (AI)
v3:
* Fix worker->pool null pointer dereference race by moving the
global decrement to detach_dying_workers().
* Rebase for upstream changes.
v4:
* Fixed onion unwind.
* Moved affinity setting to default attributes.
v5:
* Dropped global and local limits.
* Documented in workqueue.rst.
* Added NR_WQ_ATTRIBUTES.
* Reverted BH handling changes.
v6:
* Dropped separate attr->prio in favour of RTPRI_NICE_LEVEL checks. (Tejun)
* Reworked on top of tj/for-7.4.
v7:
* Convert to attrs->prio encoded analoguous to task_struct->prio.
* Rename flag to WQ_PRIO and do not re-order enums.
* Forbid WQ_RT affinity modifications via sysfs.
* Added wq_dump.py support.
---
Documentation/core-api/workqueue.rst | 8 +++
include/linux/workqueue.h | 5 +-
kernel/workqueue.c | 101 +++++++++++++++++++--------
tools/workqueue/wq_dump.py | 9 ++-
4 files changed, 90 insertions(+), 33 deletions(-)
diff --git a/Documentation/core-api/workqueue.rst b/Documentation/core-api/workqueue.rst
index bb770f556568..d699c3832b19 100644
--- a/Documentation/core-api/workqueue.rst
+++ b/Documentation/core-api/workqueue.rst
@@ -225,6 +225,14 @@ resources, scheduled and executed.
each other. Each maintains its separate pool of workers and
implements concurrency management among its workers.
+``WQ_RT``
+ Real-time priority workqueues must be created as unbound and will be
+ configured with the strict CPU affinity set. Their worker threads use the FIFO
+ scheduling policy with the lowest applicable priority.
+
+ To be used sparingly for use cases such as the real-time GPU rendering
+ contexts accessible to privileged clients.
+
``WQ_CPU_INTENSIVE``
Work items of a CPU intensive wq do not contribute to the
concurrency level. In other words, runnable CPU intensive
diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h
index a283766a192a..ebec9dcc9e5f 100644
--- a/include/linux/workqueue.h
+++ b/include/linux/workqueue.h
@@ -147,9 +147,9 @@ enum wq_affn_scope {
*/
struct workqueue_attrs {
/**
- * @nice: nice level
+ * @prio: priority encoded analoguous to task_struct->prio.
*/
- int nice;
+ int prio;
/**
* @cpumask: allowed CPUs
@@ -404,6 +404,7 @@ enum wq_flags {
*/
WQ_POWER_EFFICIENT = 1 << 7,
WQ_PERCPU = 1 << 8, /* bound to a specific cpu */
+ WQ_RT = 1 << 9, /* real-time priority, valid only with WQ_UNBOUND */
__WQ_DESTROYING = 1 << 15, /* internal: workqueue is destroying */
__WQ_DRAINING = 1 << 16, /* internal: workqueue is draining */
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index c83d68d7d0ee..afe39a18ad9e 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -47,6 +47,7 @@
#include <linux/jhash.h>
#include <linux/hashtable.h>
#include <linux/rculist.h>
+#include <linux/sched/rt.h>
#include <linux/nodemask.h>
#include <linux/moduleparam.h>
#include <linux/uaccess.h>
@@ -126,7 +127,8 @@ enum wq_internal_consts {
* all cpus. Give MIN_NICE.
*/
RESCUER_NICE_LEVEL = MIN_NICE,
- HIGHPRI_NICE_LEVEL = MIN_NICE,
+ HIGHPRI_PRIORITY = NICE_TO_PRIO(MIN_NICE),
+ RT_PRIORITY = MAX_PRIO,
WQ_NAME_LEN = 32,
WORKER_ID_LEN = 10 + WQ_NAME_LEN, /* "kworker/R-" + WQ_NAME_LEN */
@@ -1275,7 +1277,7 @@ static bool assign_work(struct work_struct *work, struct worker *worker,
static struct irq_work *bh_pool_irq_work(struct worker_pool *pool)
{
- int high = pool->attrs->nice == HIGHPRI_NICE_LEVEL ? 1 : 0;
+ int high = pool->attrs->prio == HIGHPRI_PRIORITY ? 1 : 0;
return &per_cpu(bh_pool_irq_works, pool->cpu)[high];
}
@@ -1290,7 +1292,7 @@ static void kick_bh_pool(struct worker_pool *pool)
return;
}
#endif
- if (pool->attrs->nice == HIGHPRI_NICE_LEVEL)
+ if (pool->attrs->prio == HIGHPRI_PRIORITY)
raise_softirq_irqoff(HI_SOFTIRQ);
else
raise_softirq_irqoff(TASKLET_SOFTIRQ);
@@ -2959,7 +2961,8 @@ static int format_worker_id(char *buf, size_t size, struct worker *worker,
if (pool->cpu >= 0)
return scnprintf(buf, size, "kworker/%d:%d%s",
pool->cpu, worker->id,
- pool->attrs->nice < 0 ? "H" : "");
+ pool->attrs->prio < NICE_TO_PRIO(0) ?
+ "H" : "");
else
return scnprintf(buf, size, "kworker/u%d:%d",
pool->id, worker->id);
@@ -3018,7 +3021,12 @@ static struct worker *create_worker(struct worker_pool *pool)
goto fail;
}
- set_user_nice(worker->task, pool->attrs->nice);
+ if (rt_prio(pool->attrs->prio))
+ sched_set_fifo_low(worker->task);
+ else
+ set_user_nice(worker->task,
+ PRIO_TO_NICE(pool->attrs->prio));
+
kthread_bind_mask(worker->task, pool_allowed_cpus(pool));
}
@@ -3910,7 +3918,7 @@ static void bh_worker(struct worker *worker)
if (budget_exhausted)
trace_workqueue_bh_budget_yield(pool, restarts, timeout,
- pool->attrs->nice == HIGHPRI_NICE_LEVEL);
+ pool->attrs->prio == HIGHPRI_PRIORITY);
}
/*
@@ -3969,7 +3977,7 @@ static void drain_dead_softirq_workfn(struct work_struct *work)
* don't hog this CPU's BH.
*/
if (repeat) {
- if (pool->attrs->nice == HIGHPRI_NICE_LEVEL)
+ if (pool->attrs->prio == HIGHPRI_PRIORITY)
queue_work(system_bh_highpri_wq, work);
else
queue_work(system_bh_wq, work);
@@ -4001,7 +4009,7 @@ void workqueue_softirq_dead(unsigned int cpu)
dead_work.pool = pool;
init_completion(&dead_work.done);
- if (pool->attrs->nice == HIGHPRI_NICE_LEVEL)
+ if (pool->attrs->prio == HIGHPRI_PRIORITY)
queue_work(system_bh_highpri_wq, &dead_work.work);
else
queue_work(system_bh_wq, &dead_work.work);
@@ -5015,7 +5023,7 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(void)
static void copy_workqueue_attrs(struct workqueue_attrs *to,
const struct workqueue_attrs *from)
{
- to->nice = from->nice;
+ to->prio = from->prio;
cpumask_copy(to->cpumask, from->cpumask);
cpumask_copy(to->__pod_cpumask, from->__pod_cpumask);
to->affn_strict = from->affn_strict;
@@ -5046,7 +5054,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs *attrs)
{
u32 hash = 0;
- hash = jhash_1word(attrs->nice, hash);
+ hash = jhash_1word(attrs->prio, hash);
hash = jhash_1word(attrs->affn_strict, hash);
hash = jhash(cpumask_bits(attrs->__pod_cpumask),
BITS_TO_LONGS(nr_cpumask_bits) * sizeof(long), hash);
@@ -5060,7 +5068,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs *attrs)
static bool wqattrs_equal(const struct workqueue_attrs *a,
const struct workqueue_attrs *b)
{
- if (a->nice != b->nice)
+ if (a->prio != b->prio)
return false;
if (a->affn_strict != b->affn_strict)
return false;
@@ -5928,8 +5936,19 @@ static struct workqueue_attrs *alloc_wq_std_attrs(struct workqueue_struct *wq)
if (!attrs)
return NULL;
- if (wq->flags & WQ_HIGHPRI)
- attrs->nice = HIGHPRI_NICE_LEVEL;
+ if (wq->flags & WQ_RT) {
+ attrs->prio = RT_PRIORITY;
+ /*
+ * RT workqueues have strict CPU affinity for low
+ * latency execution.
+ */
+ attrs->affn_scope = WQ_AFFN_CPU;
+ attrs->affn_strict = true;
+ } else if (wq->flags & WQ_HIGHPRI) {
+ attrs->prio = HIGHPRI_PRIORITY;
+ } else {
+ attrs->prio = DEFAULT_PRIO;
+ }
if (wq->flags & __WQ_ORDERED)
attrs->ordered = true;
@@ -6115,6 +6134,12 @@ static struct workqueue_struct *__alloc_workqueue(const char *fmt,
return NULL;
}
+ if (flags & WQ_RT) {
+ if (WARN_ON_ONCE((flags & (WQ_HIGHPRI | WQ_UNBOUND)) !=
+ WQ_UNBOUND))
+ return NULL;
+ }
+
/* see the comment above the definition of WQ_POWER_EFFICIENT */
if ((flags & WQ_POWER_EFFICIENT) && wq_power_efficient)
flags = (flags & ~WQ_PERCPU) | WQ_UNBOUND;
@@ -6671,9 +6696,9 @@ static void pr_cont_pool_info(struct worker_pool *pool)
pr_cont(" flags=0x%x", pool->flags);
if (pool->flags & POOL_BH)
pr_cont(" bh%s",
- pool->attrs->nice == HIGHPRI_NICE_LEVEL ? "-hi" : "");
+ pool->attrs->prio == HIGHPRI_PRIORITY ? "-hi" : "");
else
- pr_cont(" nice=%d", pool->attrs->nice);
+ pr_cont(" nice=%d", PRIO_TO_NICE(pool->attrs->prio));
}
static void pr_cont_worker_id(struct worker *worker)
@@ -6682,7 +6707,7 @@ static void pr_cont_worker_id(struct worker *worker)
if (pool->flags & POOL_BH)
pr_cont("bh%s",
- pool->attrs->nice == HIGHPRI_NICE_LEVEL ? "-hi" : "");
+ pool->attrs->prio == HIGHPRI_PRIORITY ? "-hi" : "");
else
pr_cont("%d%s", task_pid_nr(worker->task),
worker->rescue_wq ? "(RESCUER)" : "");
@@ -7606,7 +7631,11 @@ static ssize_t nice_show(struct device *dev, struct device_attribute *attr,
int written;
mutex_lock(&wq->mutex);
- written = scnprintf(buf, PAGE_SIZE, "%d\n", wq->attrs->nice);
+ if (wq->attrs->prio == RT_PRIORITY)
+ written = scnprintf(buf, PAGE_SIZE, "rt\n");
+ else
+ written = scnprintf(buf, PAGE_SIZE, "%d\n",
+ PRIO_TO_NICE(wq->attrs->prio));
mutex_unlock(&wq->mutex);
return written;
@@ -7632,19 +7661,21 @@ static ssize_t nice_store(struct device *dev, struct device_attribute *attr,
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
- int ret = -ENOMEM;
+ int ret, nice = 0;
+
+ if (sscanf(buf, "%d", &nice) != 1 || nice < MIN_NICE || nice > MAX_NICE)
+ return -EINVAL;
mutex_lock(&wq_pool_mutex);
attrs = wq_sysfs_prep_attrs(wq);
- if (!attrs)
+ if (!attrs) {
+ ret = -ENOMEM;
goto out_unlock;
+ }
- if (sscanf(buf, "%d", &attrs->nice) == 1 &&
- attrs->nice >= MIN_NICE && attrs->nice <= MAX_NICE)
- ret = apply_workqueue_attrs_locked(wq, attrs);
- else
- ret = -EINVAL;
+ attrs->prio = NICE_TO_PRIO(nice);
+ ret = apply_workqueue_attrs_locked(wq, attrs);
out_unlock:
mutex_unlock(&wq_pool_mutex);
@@ -7716,6 +7747,10 @@ static ssize_t affinity_scope_store(struct device *dev,
struct workqueue_attrs *attrs;
int affn, ret = -ENOMEM;
+ /* Do not allow affinity changes for RT workers. */
+ if (wq->flags & WQ_RT)
+ return -EINVAL;
+
affn = parse_affn_scope(buf);
if (affn < 0)
return affn;
@@ -7748,6 +7783,10 @@ static ssize_t affinity_strict_store(struct device *dev,
struct workqueue_attrs *attrs;
int v, ret = -ENOMEM;
+ /* Do not allow affinity changes for RT workers. */
+ if (wq->flags & WQ_RT)
+ return -EINVAL;
+
if (sscanf(buf, "%d", &v) != 1)
return -EINVAL;
@@ -7786,6 +7825,10 @@ static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj,
if (!(wq->flags & WQ_UNBOUND))
return SYSFS_GROUP_INVISIBLE;
+ /* Do not allow priority changes for RT workers. */
+ if ((wq->flags & WQ_RT) && !strcmp(attr->name, "nice"))
+ return 0444;
+
return attr->mode;
}
@@ -8310,13 +8353,13 @@ static void __init restrict_unbound_cpumask(const char *name, const struct cpuma
cpumask_and(wq_unbound_cpumask, wq_unbound_cpumask, mask);
}
-static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int nice)
+static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int prio)
{
BUG_ON(init_worker_pool(pool));
pool->cpu = cpu;
cpumask_copy(pool->attrs->cpumask, cpumask_of(cpu));
cpumask_copy(pool->attrs->__pod_cpumask, cpumask_of(cpu));
- pool->attrs->nice = nice;
+ pool->attrs->prio = prio;
pool->attrs->affn_strict = true;
pool->node = cpu_to_node(cpu);
@@ -8339,7 +8382,7 @@ static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int n
void __init workqueue_init_early(void)
{
struct wq_pod_type *pt = &wq_pod_types[WQ_AFFN_SYSTEM];
- int std_nice[NR_STD_WORKER_POOLS] = { 0, HIGHPRI_NICE_LEVEL };
+ int std_prio[NR_STD_WORKER_POOLS] = { DEFAULT_PRIO, HIGHPRI_PRIORITY };
void (*irq_work_fns[NR_STD_WORKER_POOLS])(struct irq_work *) =
{ bh_pool_kick_normal, bh_pool_kick_highpri };
int i, cpu;
@@ -8391,7 +8434,7 @@ void __init workqueue_init_early(void)
i = 0;
for_each_bh_worker_pool(pool, cpu) {
- init_cpu_worker_pool(pool, cpu, std_nice[i]);
+ init_cpu_worker_pool(pool, cpu, std_prio[i]);
pool->flags |= POOL_BH;
init_irq_work(bh_pool_irq_work(pool), irq_work_fns[i]);
i++;
@@ -8399,7 +8442,7 @@ void __init workqueue_init_early(void)
i = 0;
for_each_cpu_worker_pool(pool, cpu)
- init_cpu_worker_pool(pool, cpu, std_nice[i++]);
+ init_cpu_worker_pool(pool, cpu, std_prio[i++]);
}
system_wq = alloc_workqueue("events", WQ_PERCPU | __WQ_DEPRECATED, 0);
diff --git a/tools/workqueue/wq_dump.py b/tools/workqueue/wq_dump.py
index 9313ebe0c525..371601b086ca 100644
--- a/tools/workqueue/wq_dump.py
+++ b/tools/workqueue/wq_dump.py
@@ -24,7 +24,7 @@ Worker Pools
Lists all worker pools indexed by their ID. For each pool:
ref number of pool_workqueue's associated with this pool
- nice nice value of the worker threads in the pool
+ prio priority of the worker threads in the pool
idle number of idle workers
workers number of all workers
cpu CPU the pool is associated with (per-cpu pool)
@@ -122,6 +122,8 @@ POOL_BH = prog['POOL_BH']
WQ_NAME_LEN = prog['WQ_NAME_LEN'].value_()
cpumask_str_len = len(cpumask_str(wq_unbound_cpumask))
+rt_prio = prog.constant('RT_PRIORITY', filename='kernel/workqueue.c')
+
print('Affinity Scopes')
print('===============')
@@ -163,7 +165,10 @@ for pi, pool in idr_for_each(worker_pool_idr):
for pi, pool in idr_for_each(worker_pool_idr):
pool = drgn.Object(prog, 'struct worker_pool', address=pool)
- print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} nice={pool.attrs.nice.value_():3} ', end='')
+ prio = pool.attrs.prio.value_()
+ if prio == rt_prio:
+ prio = 'rt'
+ print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} prio={prio:3} ', end='')
print(f'idle/workers={pool.nr_idle.value_():3}/{pool.nr_workers.value_():3} ', end='')
if pool.cpu >= 0:
print(f'cpu={pool.cpu.value_():3}', end='')
--
2.55.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC v6 3/3] drm/panthor: Create per queue priority workqueues
2026-10-01 16:07 [RFC v6 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin
2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
2026-10-01 16:07 ` [RFC v6 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin
@ 2026-10-01 16:07 ` Tvrtko Ursulin
2026-10-02 19:30 ` Tejun Heo
2 siblings, 1 reply; 8+ messages in thread
From: Tvrtko Ursulin @ 2026-10-01 16:07 UTC (permalink / raw)
To: linux-kernel
Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon,
Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo
Split the single workqueue shared between the driver internal logic and
DRM scheduler use into separate ones, where the DRM scheduler one is
created per GPU priority level using the appropriate mapping to
workqueue priorities.
Low and medium GPU priority are served by a normal workqueue,
high is server by a WQ_HIGHPRI instance, while realtime GPU priority is
using the newly added WQ_RTPRI flag for lowest possible latency.
These workqueues are device global and for all three we set the maximum
concurrency to two in order to keep the GPU optimally fed with work.
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
---
drivers/gpu/drm/panthor/panthor_sched.c | 37 ++++++++++++++++++++++---
1 file changed, 33 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/panthor/panthor_sched.c b/drivers/gpu/drm/panthor/panthor_sched.c
index 5b34032deff8..257895e5e48b 100644
--- a/drivers/gpu/drm/panthor/panthor_sched.c
+++ b/drivers/gpu/drm/panthor/panthor_sched.c
@@ -152,11 +152,18 @@ struct panthor_scheduler {
*
* Used for the scheduler tick, group update or other kind of FW
* event processing that can't be handled in the threaded interrupt
- * path. Also passed to the drm_gpu_scheduler instances embedded
- * in panthor_queue.
+ * path.
*/
struct workqueue_struct *wq;
+ /**
+ * @submit_wq: Per priority workqueues for the DRM scheduler
+ *
+ * Passed to the drm_gpu_scheduler instances embedded
+ * in panthor_queue based on the queue priority.
+ */
+ struct workqueue_struct *submit_wq[PANTHOR_CSG_PRIORITY_COUNT];
+
/**
* @heap_alloc_wq: Workqueue used to schedule tiler_oom works.
*
@@ -3582,8 +3589,14 @@ group_create_queue(struct panthor_group *group,
goto err_free_queue;
}
+ if (group->priority >= ARRAY_SIZE(group->ptdev->scheduler->submit_wq) ||
+ !group->ptdev->scheduler->submit_wq[group->priority]) {
+ ret = -EINVAL;
+ goto err_free_queue;
+ }
+
sched_args.name = queue->name;
-
+ sched_args.submit_wq = group->ptdev->scheduler->submit_wq[group->priority];
ret = drm_sched_init(&queue->scheduler, &sched_args);
if (ret)
goto err_free_queue;
@@ -4073,6 +4086,15 @@ static void panthor_sched_fini(struct drm_device *ddev, void *res)
if (!sched || !sched->csg_slot_count)
return;
+ if (sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM])
+ destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM]);
+
+ if (sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH])
+ destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH]);
+
+ if (sched->submit_wq[PANTHOR_CSG_PRIORITY_RT])
+ destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]);
+
if (sched->wq)
destroy_workqueue(sched->wq);
@@ -4174,7 +4196,14 @@ int panthor_sched_init(struct panthor_device *ptdev)
*/
sched->heap_alloc_wq = alloc_workqueue("panthor-heap-alloc", WQ_UNBOUND, 0);
sched->wq = alloc_workqueue("panthor-csf-sched", WQ_MEM_RECLAIM | WQ_UNBOUND, 0);
- if (!sched->wq || !sched->heap_alloc_wq) {
+ sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] = alloc_workqueue("panthor-drm", WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
+ sched->submit_wq[PANTHOR_CSG_PRIORITY_LOW] = sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM];
+ sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] = alloc_workqueue("panthor-drm-high", WQ_HIGHPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
+ sched->submit_wq[PANTHOR_CSG_PRIORITY_RT] = alloc_workqueue("panthor-drm-rt", WQ_RT | WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
+ if (!sched->wq || !sched->heap_alloc_wq ||
+ !sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] ||
+ !sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] ||
+ !sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]) {
panthor_sched_fini(&ptdev->base, sched);
drm_err(&ptdev->base, "Failed to allocate the workqueues");
return -ENOMEM;
--
2.55.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* [RFC v6.1 2/3] workqueue: Add support for real-time workers
2026-10-01 16:07 ` [RFC v6 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin
@ 2026-10-01 18:48 ` Tvrtko Ursulin
2026-10-02 19:30 ` Tejun Heo
0 siblings, 1 reply; 8+ messages in thread
From: Tvrtko Ursulin @ 2026-10-01 18:48 UTC (permalink / raw)
To: linux-kernel
Cc: dri-devel, kernel-dev, Tvrtko Ursulin, Boris Brezillon,
Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost,
Steven Price, Tejun Heo
For use cases such as the DRM scheduler submitting work to the GPU on
behalf of low latency userspace applications, where latter have sufficient
privileges to have had successfully obtained realtime Vulkan global
priority, competing with random background CPU load can create large
latency spikes which gets in the way of a smooth user experience.
For these situations the existing WQ_HIGHPRI does not bring a noticeable
improvement and a stronger hint is needed.
Lets add WQ_RT which creates workers with a SCHED_FIFO scheduling class to
improve this.
We use a minimum priority level since we only care about winning the
contest against normal background CPU load.
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Bradley Morgan <include@grrlz.net>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
---
v2:
* Limit WQ_RTPRI to unbound workqueues and make it have strict CPU
affinitity. (Tejun)
* Fixed commit message typos. (AI)
* Fixed sysfs handling, max_active setting and user modified nice
application. (AI)
v3:
* Fix worker->pool null pointer dereference race by moving the
global decrement to detach_dying_workers().
* Rebase for upstream changes.
v4:
* Fixed onion unwind.
* Moved affinity setting to default attributes.
v5:
* Dropped global and local limits.
* Documented in workqueue.rst.
* Added NR_WQ_ATTRIBUTES.
* Reverted BH handling changes.
v6:
* Dropped separate attr->prio in favour of RTPRI_NICE_LEVEL checks. (Tejun)
* Reworked on top of tj/for-7.4.
v7:
* Convert to attrs->prio encoded analoguous to task_struct->prio.
* Rename flag to WQ_PRIO and do not re-order enums.
* Forbid WQ_RT affinity modifications via sysfs.
* Added wq_dump.py support.
v8:
* Fix max rt prio confusion.
* Initialize priority also in alloc_workqueue_attrs_noprof.
* Restrict rt cpumask modifications.
---
Documentation/core-api/workqueue.rst | 8 ++
include/linux/workqueue.h | 5 +-
kernel/workqueue.c | 106 +++++++++++++++++++--------
tools/workqueue/wq_dump.py | 9 ++-
4 files changed, 95 insertions(+), 33 deletions(-)
diff --git a/Documentation/core-api/workqueue.rst b/Documentation/core-api/workqueue.rst
index bb770f556568..d699c3832b19 100644
--- a/Documentation/core-api/workqueue.rst
+++ b/Documentation/core-api/workqueue.rst
@@ -225,6 +225,14 @@ resources, scheduled and executed.
each other. Each maintains its separate pool of workers and
implements concurrency management among its workers.
+``WQ_RT``
+ Real-time priority workqueues must be created as unbound and will be
+ configured with the strict CPU affinity set. Their worker threads use the FIFO
+ scheduling policy with the lowest applicable priority.
+
+ To be used sparingly for use cases such as the real-time GPU rendering
+ contexts accessible to privileged clients.
+
``WQ_CPU_INTENSIVE``
Work items of a CPU intensive wq do not contribute to the
concurrency level. In other words, runnable CPU intensive
diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h
index a283766a192a..ebec9dcc9e5f 100644
--- a/include/linux/workqueue.h
+++ b/include/linux/workqueue.h
@@ -147,9 +147,9 @@ enum wq_affn_scope {
*/
struct workqueue_attrs {
/**
- * @nice: nice level
+ * @prio: priority encoded analoguous to task_struct->prio.
*/
- int nice;
+ int prio;
/**
* @cpumask: allowed CPUs
@@ -404,6 +404,7 @@ enum wq_flags {
*/
WQ_POWER_EFFICIENT = 1 << 7,
WQ_PERCPU = 1 << 8, /* bound to a specific cpu */
+ WQ_RT = 1 << 9, /* real-time priority, valid only with WQ_UNBOUND */
__WQ_DESTROYING = 1 << 15, /* internal: workqueue is destroying */
__WQ_DRAINING = 1 << 16, /* internal: workqueue is draining */
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index c83d68d7d0ee..31f3ea0dd62e 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -47,6 +47,7 @@
#include <linux/jhash.h>
#include <linux/hashtable.h>
#include <linux/rculist.h>
+#include <linux/sched/rt.h>
#include <linux/nodemask.h>
#include <linux/moduleparam.h>
#include <linux/uaccess.h>
@@ -126,7 +127,8 @@ enum wq_internal_consts {
* all cpus. Give MIN_NICE.
*/
RESCUER_NICE_LEVEL = MIN_NICE,
- HIGHPRI_NICE_LEVEL = MIN_NICE,
+ HIGHPRI_PRIORITY = NICE_TO_PRIO(MIN_NICE),
+ RT_PRIORITY = MAX_RT_PRIO - 1,
WQ_NAME_LEN = 32,
WORKER_ID_LEN = 10 + WQ_NAME_LEN, /* "kworker/R-" + WQ_NAME_LEN */
@@ -1275,7 +1277,7 @@ static bool assign_work(struct work_struct *work, struct worker *worker,
static struct irq_work *bh_pool_irq_work(struct worker_pool *pool)
{
- int high = pool->attrs->nice == HIGHPRI_NICE_LEVEL ? 1 : 0;
+ int high = pool->attrs->prio == HIGHPRI_PRIORITY ? 1 : 0;
return &per_cpu(bh_pool_irq_works, pool->cpu)[high];
}
@@ -1290,7 +1292,7 @@ static void kick_bh_pool(struct worker_pool *pool)
return;
}
#endif
- if (pool->attrs->nice == HIGHPRI_NICE_LEVEL)
+ if (pool->attrs->prio == HIGHPRI_PRIORITY)
raise_softirq_irqoff(HI_SOFTIRQ);
else
raise_softirq_irqoff(TASKLET_SOFTIRQ);
@@ -2959,7 +2961,8 @@ static int format_worker_id(char *buf, size_t size, struct worker *worker,
if (pool->cpu >= 0)
return scnprintf(buf, size, "kworker/%d:%d%s",
pool->cpu, worker->id,
- pool->attrs->nice < 0 ? "H" : "");
+ pool->attrs->prio < NICE_TO_PRIO(0) ?
+ "H" : "");
else
return scnprintf(buf, size, "kworker/u%d:%d",
pool->id, worker->id);
@@ -3018,7 +3021,12 @@ static struct worker *create_worker(struct worker_pool *pool)
goto fail;
}
- set_user_nice(worker->task, pool->attrs->nice);
+ if (rt_prio(pool->attrs->prio))
+ sched_set_fifo_low(worker->task);
+ else
+ set_user_nice(worker->task,
+ PRIO_TO_NICE(pool->attrs->prio));
+
kthread_bind_mask(worker->task, pool_allowed_cpus(pool));
}
@@ -3910,7 +3918,7 @@ static void bh_worker(struct worker *worker)
if (budget_exhausted)
trace_workqueue_bh_budget_yield(pool, restarts, timeout,
- pool->attrs->nice == HIGHPRI_NICE_LEVEL);
+ pool->attrs->prio == HIGHPRI_PRIORITY);
}
/*
@@ -3969,7 +3977,7 @@ static void drain_dead_softirq_workfn(struct work_struct *work)
* don't hog this CPU's BH.
*/
if (repeat) {
- if (pool->attrs->nice == HIGHPRI_NICE_LEVEL)
+ if (pool->attrs->prio == HIGHPRI_PRIORITY)
queue_work(system_bh_highpri_wq, work);
else
queue_work(system_bh_wq, work);
@@ -4001,7 +4009,7 @@ void workqueue_softirq_dead(unsigned int cpu)
dead_work.pool = pool;
init_completion(&dead_work.done);
- if (pool->attrs->nice == HIGHPRI_NICE_LEVEL)
+ if (pool->attrs->prio == HIGHPRI_PRIORITY)
queue_work(system_bh_highpri_wq, &dead_work.work);
else
queue_work(system_bh_wq, &dead_work.work);
@@ -5005,6 +5013,7 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(void)
goto fail;
cpumask_copy(attrs->cpumask, cpu_possible_mask);
+ attrs->prio = DEFAULT_PRIO;
attrs->affn_scope = WQ_AFFN_DFL;
return attrs;
fail:
@@ -5015,7 +5024,7 @@ struct workqueue_attrs *alloc_workqueue_attrs_noprof(void)
static void copy_workqueue_attrs(struct workqueue_attrs *to,
const struct workqueue_attrs *from)
{
- to->nice = from->nice;
+ to->prio = from->prio;
cpumask_copy(to->cpumask, from->cpumask);
cpumask_copy(to->__pod_cpumask, from->__pod_cpumask);
to->affn_strict = from->affn_strict;
@@ -5046,7 +5055,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs *attrs)
{
u32 hash = 0;
- hash = jhash_1word(attrs->nice, hash);
+ hash = jhash_1word(attrs->prio, hash);
hash = jhash_1word(attrs->affn_strict, hash);
hash = jhash(cpumask_bits(attrs->__pod_cpumask),
BITS_TO_LONGS(nr_cpumask_bits) * sizeof(long), hash);
@@ -5060,7 +5069,7 @@ static u32 wqattrs_hash(const struct workqueue_attrs *attrs)
static bool wqattrs_equal(const struct workqueue_attrs *a,
const struct workqueue_attrs *b)
{
- if (a->nice != b->nice)
+ if (a->prio != b->prio)
return false;
if (a->affn_strict != b->affn_strict)
return false;
@@ -5928,8 +5937,19 @@ static struct workqueue_attrs *alloc_wq_std_attrs(struct workqueue_struct *wq)
if (!attrs)
return NULL;
- if (wq->flags & WQ_HIGHPRI)
- attrs->nice = HIGHPRI_NICE_LEVEL;
+ if (wq->flags & WQ_RT) {
+ attrs->prio = RT_PRIORITY;
+ /*
+ * RT workqueues have strict CPU affinity for low
+ * latency execution.
+ */
+ attrs->affn_scope = WQ_AFFN_CPU;
+ attrs->affn_strict = true;
+ } else if (wq->flags & WQ_HIGHPRI) {
+ attrs->prio = HIGHPRI_PRIORITY;
+ } else {
+ attrs->prio = DEFAULT_PRIO;
+ }
if (wq->flags & __WQ_ORDERED)
attrs->ordered = true;
@@ -6115,6 +6135,12 @@ static struct workqueue_struct *__alloc_workqueue(const char *fmt,
return NULL;
}
+ if (flags & WQ_RT) {
+ if (WARN_ON_ONCE((flags & (WQ_HIGHPRI | WQ_UNBOUND)) !=
+ WQ_UNBOUND))
+ return NULL;
+ }
+
/* see the comment above the definition of WQ_POWER_EFFICIENT */
if ((flags & WQ_POWER_EFFICIENT) && wq_power_efficient)
flags = (flags & ~WQ_PERCPU) | WQ_UNBOUND;
@@ -6671,9 +6697,9 @@ static void pr_cont_pool_info(struct worker_pool *pool)
pr_cont(" flags=0x%x", pool->flags);
if (pool->flags & POOL_BH)
pr_cont(" bh%s",
- pool->attrs->nice == HIGHPRI_NICE_LEVEL ? "-hi" : "");
+ pool->attrs->prio == HIGHPRI_PRIORITY ? "-hi" : "");
else
- pr_cont(" nice=%d", pool->attrs->nice);
+ pr_cont(" nice=%d", PRIO_TO_NICE(pool->attrs->prio));
}
static void pr_cont_worker_id(struct worker *worker)
@@ -6682,7 +6708,7 @@ static void pr_cont_worker_id(struct worker *worker)
if (pool->flags & POOL_BH)
pr_cont("bh%s",
- pool->attrs->nice == HIGHPRI_NICE_LEVEL ? "-hi" : "");
+ pool->attrs->prio == HIGHPRI_PRIORITY ? "-hi" : "");
else
pr_cont("%d%s", task_pid_nr(worker->task),
worker->rescue_wq ? "(RESCUER)" : "");
@@ -7606,7 +7632,11 @@ static ssize_t nice_show(struct device *dev, struct device_attribute *attr,
int written;
mutex_lock(&wq->mutex);
- written = scnprintf(buf, PAGE_SIZE, "%d\n", wq->attrs->nice);
+ if (wq->attrs->prio == RT_PRIORITY)
+ written = scnprintf(buf, PAGE_SIZE, "rt\n");
+ else
+ written = scnprintf(buf, PAGE_SIZE, "%d\n",
+ PRIO_TO_NICE(wq->attrs->prio));
mutex_unlock(&wq->mutex);
return written;
@@ -7632,19 +7662,21 @@ static ssize_t nice_store(struct device *dev, struct device_attribute *attr,
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
- int ret = -ENOMEM;
+ int ret, nice = 0;
+
+ if (sscanf(buf, "%d", &nice) != 1 || nice < MIN_NICE || nice > MAX_NICE)
+ return -EINVAL;
mutex_lock(&wq_pool_mutex);
attrs = wq_sysfs_prep_attrs(wq);
- if (!attrs)
+ if (!attrs) {
+ ret = -ENOMEM;
goto out_unlock;
+ }
- if (sscanf(buf, "%d", &attrs->nice) == 1 &&
- attrs->nice >= MIN_NICE && attrs->nice <= MAX_NICE)
- ret = apply_workqueue_attrs_locked(wq, attrs);
- else
- ret = -EINVAL;
+ attrs->prio = NICE_TO_PRIO(nice);
+ ret = apply_workqueue_attrs_locked(wq, attrs);
out_unlock:
mutex_unlock(&wq_pool_mutex);
@@ -7673,6 +7705,10 @@ static ssize_t unbound_cpumask_store(struct device *dev,
struct workqueue_attrs *attrs;
int ret = -ENOMEM;
+ /* Do not allow cpumask changes for RT workers. */
+ if (wq->flags & WQ_RT)
+ return -EINVAL;
+
mutex_lock(&wq_pool_mutex);
attrs = wq_sysfs_prep_attrs(wq);
@@ -7716,6 +7752,10 @@ static ssize_t affinity_scope_store(struct device *dev,
struct workqueue_attrs *attrs;
int affn, ret = -ENOMEM;
+ /* Do not allow affinity changes for RT workers. */
+ if (wq->flags & WQ_RT)
+ return -EINVAL;
+
affn = parse_affn_scope(buf);
if (affn < 0)
return affn;
@@ -7748,6 +7788,10 @@ static ssize_t affinity_strict_store(struct device *dev,
struct workqueue_attrs *attrs;
int v, ret = -ENOMEM;
+ /* Do not allow affinity changes for RT workers. */
+ if (wq->flags & WQ_RT)
+ return -EINVAL;
+
if (sscanf(buf, "%d", &v) != 1)
return -EINVAL;
@@ -7786,6 +7830,10 @@ static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj,
if (!(wq->flags & WQ_UNBOUND))
return SYSFS_GROUP_INVISIBLE;
+ /* Do not allow priority changes for RT workers. */
+ if ((wq->flags & WQ_RT) && !strcmp(attr->name, "nice"))
+ return 0444;
+
return attr->mode;
}
@@ -8310,13 +8358,13 @@ static void __init restrict_unbound_cpumask(const char *name, const struct cpuma
cpumask_and(wq_unbound_cpumask, wq_unbound_cpumask, mask);
}
-static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int nice)
+static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int prio)
{
BUG_ON(init_worker_pool(pool));
pool->cpu = cpu;
cpumask_copy(pool->attrs->cpumask, cpumask_of(cpu));
cpumask_copy(pool->attrs->__pod_cpumask, cpumask_of(cpu));
- pool->attrs->nice = nice;
+ pool->attrs->prio = prio;
pool->attrs->affn_strict = true;
pool->node = cpu_to_node(cpu);
@@ -8339,7 +8387,7 @@ static void __init init_cpu_worker_pool(struct worker_pool *pool, int cpu, int n
void __init workqueue_init_early(void)
{
struct wq_pod_type *pt = &wq_pod_types[WQ_AFFN_SYSTEM];
- int std_nice[NR_STD_WORKER_POOLS] = { 0, HIGHPRI_NICE_LEVEL };
+ int std_prio[NR_STD_WORKER_POOLS] = { DEFAULT_PRIO, HIGHPRI_PRIORITY };
void (*irq_work_fns[NR_STD_WORKER_POOLS])(struct irq_work *) =
{ bh_pool_kick_normal, bh_pool_kick_highpri };
int i, cpu;
@@ -8391,7 +8439,7 @@ void __init workqueue_init_early(void)
i = 0;
for_each_bh_worker_pool(pool, cpu) {
- init_cpu_worker_pool(pool, cpu, std_nice[i]);
+ init_cpu_worker_pool(pool, cpu, std_prio[i]);
pool->flags |= POOL_BH;
init_irq_work(bh_pool_irq_work(pool), irq_work_fns[i]);
i++;
@@ -8399,7 +8447,7 @@ void __init workqueue_init_early(void)
i = 0;
for_each_cpu_worker_pool(pool, cpu)
- init_cpu_worker_pool(pool, cpu, std_nice[i++]);
+ init_cpu_worker_pool(pool, cpu, std_prio[i++]);
}
system_wq = alloc_workqueue("events", WQ_PERCPU | __WQ_DEPRECATED, 0);
diff --git a/tools/workqueue/wq_dump.py b/tools/workqueue/wq_dump.py
index 9313ebe0c525..371601b086ca 100644
--- a/tools/workqueue/wq_dump.py
+++ b/tools/workqueue/wq_dump.py
@@ -24,7 +24,7 @@ Worker Pools
Lists all worker pools indexed by their ID. For each pool:
ref number of pool_workqueue's associated with this pool
- nice nice value of the worker threads in the pool
+ prio priority of the worker threads in the pool
idle number of idle workers
workers number of all workers
cpu CPU the pool is associated with (per-cpu pool)
@@ -122,6 +122,8 @@ POOL_BH = prog['POOL_BH']
WQ_NAME_LEN = prog['WQ_NAME_LEN'].value_()
cpumask_str_len = len(cpumask_str(wq_unbound_cpumask))
+rt_prio = prog.constant('RT_PRIORITY', filename='kernel/workqueue.c')
+
print('Affinity Scopes')
print('===============')
@@ -163,7 +165,10 @@ for pi, pool in idr_for_each(worker_pool_idr):
for pi, pool in idr_for_each(worker_pool_idr):
pool = drgn.Object(prog, 'struct worker_pool', address=pool)
- print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} nice={pool.attrs.nice.value_():3} ', end='')
+ prio = pool.attrs.prio.value_()
+ if prio == rt_prio:
+ prio = 'rt'
+ print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} prio={prio:3} ', end='')
print(f'idle/workers={pool.nr_idle.value_():3}/{pool.nr_workers.value_():3} ', end='')
if pool.cpu >= 0:
print(f'cpu={pool.cpu.value_():3}', end='')
--
2.55.0
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration
2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
@ 2026-10-02 19:30 ` Tejun Heo
0 siblings, 0 replies; 8+ messages in thread
From: Tejun Heo @ 2026-10-02 19:30 UTC (permalink / raw)
To: Tvrtko Ursulin
Cc: linux-kernel, dri-devel, kernel-dev, Boris Brezillon,
Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost,
Steven Price, Lai Jiangshan, Breno Leitao
Hello,
The following is a Claude-generated review.
On Thu, Oct 01, 2026 at 05:07:09PM +0100, Tvrtko Ursulin wrote:
> +static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj,
> + struct attribute *attr, int n)
> +{
> + struct device *dev = kobj_to_dev(kobj);
> + struct workqueue_struct *wq = dev_to_wq(dev);
> +
> + if (!(wq->flags & WQ_UNBOUND))
> + return SYSFS_GROUP_INVISIBLE;
SYSFS_GROUP_INVISIBLE is documented for named groups only. For an unnamed
group, fs/sysfs/group.c masks the bit off and the leftover 0 is what hides
the attribute. Can you return 0 here like wq_sysfs_is_visible() does?
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [RFC v6.1 2/3] workqueue: Add support for real-time workers
2026-10-01 18:48 ` [RFC v6.1 " Tvrtko Ursulin
@ 2026-10-02 19:30 ` Tejun Heo
0 siblings, 0 replies; 8+ messages in thread
From: Tejun Heo @ 2026-10-02 19:30 UTC (permalink / raw)
To: Tvrtko Ursulin
Cc: linux-kernel, dri-devel, kernel-dev, Boris Brezillon,
Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost,
Steven Price, Lai Jiangshan, Breno Leitao
Hello,
The following is a Claude-generated review.
On Thu, Oct 01, 2026 at 07:48:48PM +0100, Tvrtko Ursulin wrote:
> For use cases such as the DRM scheduler submitting work to the GPU on
> behalf of low latency userspace applications, where latter have sufficient
> privileges to have had successfully obtained realtime Vulkan global
> priority, competing with random background CPU load can create large
> latency spikes which gets in the way of a smooth user experience.
Can you add the compositor / DRM master usage model from the v5 discussion
here? The cover letter also still says CAP_SYS_NICE is required while
group_priority_permit() accepts DRM master too.
> + HIGHPRI_PRIORITY = NICE_TO_PRIO(MIN_NICE),
> + RT_PRIORITY = MAX_RT_PRIO - 1,
MAX_RT_PRIO - 1 is what rt_priority 0 would map to, while
sched_set_fifo_low() gives the workers prio 98. Nothing depends on it today,
but maybe MAX_RT_PRIO - 2 with a comment tying it to sched_set_fifo_low(),
or soften the kerneldoc which says the encoding matches task_struct->prio?
Also, s/analoguous/analogous/ there.
> + if (wq->flags & WQ_RT) {
> + attrs->prio = RT_PRIORITY;
> + /*
> + * RT workqueues have strict CPU affinity for low
> + * latency execution.
> + */
> + attrs->affn_scope = WQ_AFFN_CPU;
> + attrs->affn_strict = true;
apply_workqueue_attrs() is unrestricted, so a caller can hand a WQ_RT wq the
attrs from alloc_workqueue_attrs() and turn it into a normal non-strict wq
while the flag stays set. Maybe apply_wqattrs_prepare() should force prio
and affinity for WQ_RT instead of restricting only the sysfs side?
> + if (flags & WQ_RT) {
> + if (WARN_ON_ONCE((flags & (WQ_HIGHPRI | WQ_UNBOUND)) !=
> + WQ_UNBOUND))
> + return NULL;
> + }
alloc_ordered_workqueue() with WQ_RT passes this and ends up with a single
pool spanning all CPUs because ordered wqs use dfl_pwq everywhere, while the
doc and sysfs say strict per-CPU. Should __WQ_ORDERED be rejected too? The
nested ifs can also be a single condition.
> - pr_cont(" nice=%d", pool->attrs->nice);
> + pr_cont(" nice=%d", PRIO_TO_NICE(pool->attrs->prio));
This prints nice=-21 for RT pools. Can you show "rt" here like nice_show()?
> + /* Do not allow cpumask changes for RT workers. */
> + if (wq->flags & WQ_RT)
> + return -EINVAL;
The three stores check WQ_RT and return -EINVAL while nice goes read-only
through is_visible below. Can all four go through
wq_sysfs_unbound_group_visible() returning 0444 for WQ_RT? That drops the
three store checks. Note that the global cpumask still applies to WQ_RT wqs
through workqueue_apply_unbound_cpumask(), so the comment overstates a bit.
The interface comment at the top of the sysfs section also still says nice
is RW int.
> + /* Do not allow priority changes for RT workers. */
> + if ((wq->flags & WQ_RT) && !strcmp(attr->name, "nice"))
> + return 0444;
attr == &dev_attr_nice.attr would avoid the strcmp.
> + prio = pool.attrs.prio.value_()
> + if prio == rt_prio:
> + prio = 'rt'
> + print(f'pool[{pi:0{max_pool_id_len}}] flags=0x{pool.flags.value_():02x} ref={pool.refcnt.value_():{max_ref_len}} prio={prio:3} ', end='')
Can we keep printing nice for non-RT pools? prio=120 is the internal
encoding, sysfs and the pool dumps print nice, and the example output in
workqueue.rst would go stale. prog['RT_PRIORITY'] would match the rest of
the file too.
One more thing which isn't in the diff. The rescuer of a WQ_RT |
WQ_MEM_RECLAIM wq, which is what panthor-drm-rt is, still runs at nice -20,
so under memory pressure the RT wq's work items run as CFS. Should
rescuer_thread() use sched_set_fifo_low() for WQ_RT?
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [RFC v6 3/3] drm/panthor: Create per queue priority workqueues
2026-10-01 16:07 ` [RFC v6 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin
@ 2026-10-02 19:30 ` Tejun Heo
0 siblings, 0 replies; 8+ messages in thread
From: Tejun Heo @ 2026-10-02 19:30 UTC (permalink / raw)
To: Tvrtko Ursulin
Cc: linux-kernel, dri-devel, kernel-dev, Boris Brezillon, Chia-I Wu,
Liviu Dudau, Matthew Brost, Steven Price
Hello,
The following is a Claude-generated review.
On Thu, Oct 01, 2026 at 05:07:11PM +0100, Tvrtko Ursulin wrote:
> Low and medium GPU priority are served by a normal workqueue,
> high is server by a WQ_HIGHPRI instance, while realtime GPU priority is
> using the newly added WQ_RTPRI flag for lowest possible latency.
s/server/served/ and the flag is WQ_RT now.
> + if (group->priority >= ARRAY_SIZE(group->ptdev->scheduler->submit_wq) ||
> + !group->ptdev->scheduler->submit_wq[group->priority]) {
> + ret = -EINVAL;
> + goto err_free_queue;
> + }
panthor_group_create() already rejects priorities >=
PANTHOR_CSG_PRIORITY_COUNT and all slots are populated once
panthor_sched_init() succeeded, so this can't fire.
> + sched_args.submit_wq = group->ptdev->scheduler->submit_wq[group->priority];
The .submit_wq = sched->wq initializer at the top of the function is still
there and gets overwritten here.
> + if (sched->submit_wq[PANTHOR_CSG_PRIORITY_RT])
> + destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]);
> +
> if (sched->wq)
> destroy_workqueue(sched->wq);
Work on sched->wq signals job fences whose callbacks queue onto the submit
wqs, so destroying sched->wq first would be the safer order.
> + sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] = alloc_workqueue("panthor-drm", WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
> + sched->submit_wq[PANTHOR_CSG_PRIORITY_LOW] = sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM];
> + sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] = alloc_workqueue("panthor-drm-high", WQ_HIGHPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
> + sched->submit_wq[PANTHOR_CSG_PRIORITY_RT] = alloc_workqueue("panthor-drm-rt", WQ_RT | WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
For an unbound wq max_active applies to the whole wq, so this is two
in-flight items per priority level across all the queues on the device,
where the shared wq before had no effective limit. queue_run_job() blocks
on sched->lock which tick_work() holds across FW round trips, so two
blocked run_jobs would stall every other queue's run and free work at that
level. What's the reason for 2? Also, these lines are over 100 columns.
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2026-10-02 19:30 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-01 16:07 [RFC v6 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin
2026-10-01 16:07 ` [RFC v6 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
2026-10-02 19:30 ` Tejun Heo
2026-10-01 16:07 ` [RFC v6 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin
2026-10-01 18:48 ` [RFC v6.1 " Tvrtko Ursulin
2026-10-02 19:30 ` Tejun Heo
2026-10-01 16:07 ` [RFC v6 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin
2026-10-02 19:30 ` Tejun Heo
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®