* [RFC v5 0/3] Realtime workqueues and panthor realtime submission
@ 2026-09-23 16:12 Tvrtko Ursulin
2026-09-23 16:12 ` [RFC v5 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
` (2 more replies)
0 siblings, 3 replies; 6+ messages in thread
From: Tvrtko Ursulin @ 2026-09-23 16:12 UTC (permalink / raw)
To: linux-kernel
Cc: kernel-dev, dri-devel, Tvrtko Ursulin, Boris Brezillon,
Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo
This is a continuation of the previous discussion which was here:
https://lore.kernel.org/dri-devel/20260702143745.79293-1-tvrtko.ursulin@igalia.com/
Work is now converted to a much simpler approach by adding real-time scheduling
workqueues based on Tejun's feedback
To re-cap, when an userspace thread submits GPU work, due how the DRM scheduler
uses workqueues to feed the GPU, and regardless of the GPU rendering context
priority, or the CPU scheduling priority of the userspace thread itself, the
use of workqueues can add significant latency to the submit path.
When CPU is busy with enough backround load this translates to severe latency
spikes measured as time between userspace submitting work and GPU actually being
given that work to execute.
With the panthor workqueue upgraded to use WQ_HIGHPRI and varying the CPU
priority of the submit thread, the test program from
https://gitlab.freedesktop.org/panfrost/linux/-/work_items/49 reproduces these
kind of latencies:
. N RT
M 27 28 32
95% 163 246 809
98% 924 991 1882
Legend:
M = Median submit latency in us
95% = Percentile latency in us
. = Userspace submit thread SCHED_OTHER
N = -||= nice -1
RT = -||- FIFO 1
Upgrading the panthor workqueues so that the realtime GPU priority queue uses
WQ_RTPRI, submit latency becomes completely controlled with the median of 14us
and 95 and 98-th percentiles at 23us and 25us respectively.
Important to note is that VK_QUEUE_GLOBAL_PRIORITY_REALTIME already required
the userspace to have CAP_SYS_NICE, meaning access to real-time workqueues is
effectively also guared behind this capability.
v2:
* See patch 1 changelog.
v3:
* See patch 1 changelog + apologies for v3 following so quickly after v2.
I have found a race condition as I expanded the testing to a second platform.
v4:
* See patch 1 changelog.
v5:
* Rebased on top of tj/for-7.4 branch.
* New patch in series simplifies unbound sysfs attribute handling.
* WQ_RTPRI reworked to fold RT marker into attr->nice.
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
Tvrtko Ursulin (3):
workqueue: Simplify unbound sysfs attribute registration
workqueue: Add support for real-time workers
drm/panthor: Create per queue priority workqueues
Documentation/core-api/workqueue.rst | 5 +
drivers/gpu/drm/panthor/panthor_sched.c | 37 ++++++-
include/linux/workqueue.h | 9 +-
kernel/workqueue.c | 131 +++++++++++++++---------
4 files changed, 125 insertions(+), 57 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC v5 1/3] workqueue: Simplify unbound sysfs attribute registration
2026-09-23 16:12 [RFC v5 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin
@ 2026-09-23 16:12 ` Tvrtko Ursulin
2026-09-23 17:12 ` Tvrtko Ursulin
2026-09-23 16:12 ` [RFC v5 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin
2026-09-23 16:12 ` [RFC v5 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin
2 siblings, 1 reply; 6+ messages in thread
From: Tvrtko Ursulin @ 2026-09-23 16:12 UTC (permalink / raw)
To: linux-kernel
Cc: kernel-dev, dri-devel, Tvrtko Ursulin, Boris Brezillon,
Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost,
Steven Price, Tejun Heo
Instead of manually registering each attribute we can put them in an
attribute group with a visibility check and device core will handle the
rest, which simplifies the registration and error unwind.
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Bradley Morgan <include@grrlz.net>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
---
kernel/workqueue.c | 98 ++++++++++++++++++++++++----------------------
1 file changed, 52 insertions(+), 46 deletions(-)
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index e618108c6127..e3a4ad56dae8 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -7599,8 +7599,8 @@ static const struct attribute_group wq_sysfs_group = {
};
__ATTRIBUTE_GROUPS(wq_sysfs);
-static ssize_t wq_nice_show(struct device *dev, struct device_attribute *attr,
- char *buf)
+static ssize_t nice_show(struct device *dev, struct device_attribute *attr,
+ char *buf)
{
struct workqueue_struct *wq = dev_to_wq(dev);
int written;
@@ -7627,8 +7627,8 @@ static struct workqueue_attrs *wq_sysfs_prep_attrs(struct workqueue_struct *wq)
return attrs;
}
-static ssize_t wq_nice_store(struct device *dev, struct device_attribute *attr,
- const char *buf, size_t count)
+static ssize_t nice_store(struct device *dev, struct device_attribute *attr,
+ const char *buf, size_t count)
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
@@ -7652,8 +7652,8 @@ static ssize_t wq_nice_store(struct device *dev, struct device_attribute *attr,
return ret ?: count;
}
-static ssize_t wq_cpumask_show(struct device *dev,
- struct device_attribute *attr, char *buf)
+static ssize_t unbound_cpumask_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
{
struct workqueue_struct *wq = dev_to_wq(dev);
int written;
@@ -7665,9 +7665,9 @@ static ssize_t wq_cpumask_show(struct device *dev,
return written;
}
-static ssize_t wq_cpumask_store(struct device *dev,
- struct device_attribute *attr,
- const char *buf, size_t count)
+static ssize_t unbound_cpumask_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t count)
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
@@ -7689,8 +7689,8 @@ static ssize_t wq_cpumask_store(struct device *dev,
return ret ?: count;
}
-static ssize_t wq_affn_scope_show(struct device *dev,
- struct device_attribute *attr, char *buf)
+static ssize_t affn_scope_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
{
struct workqueue_struct *wq = dev_to_wq(dev);
int written;
@@ -7708,9 +7708,9 @@ static ssize_t wq_affn_scope_show(struct device *dev,
return written;
}
-static ssize_t wq_affn_scope_store(struct device *dev,
- struct device_attribute *attr,
- const char *buf, size_t count)
+static ssize_t affn_scope_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t count)
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
@@ -7731,8 +7731,8 @@ static ssize_t wq_affn_scope_store(struct device *dev,
return ret ?: count;
}
-static ssize_t wq_affinity_strict_show(struct device *dev,
- struct device_attribute *attr, char *buf)
+static ssize_t affinity_strict_show(struct device *dev,
+ struct device_attribute *attr, char *buf)
{
struct workqueue_struct *wq = dev_to_wq(dev);
@@ -7740,9 +7740,9 @@ static ssize_t wq_affinity_strict_show(struct device *dev,
wq->attrs->affn_strict);
}
-static ssize_t wq_affinity_strict_store(struct device *dev,
- struct device_attribute *attr,
- const char *buf, size_t count)
+static ssize_t affinity_strict_store(struct device *dev,
+ struct device_attribute *attr,
+ const char *buf, size_t count)
{
struct workqueue_struct *wq = dev_to_wq(dev);
struct workqueue_attrs *attrs;
@@ -7762,14 +7762,40 @@ static ssize_t wq_affinity_strict_store(struct device *dev,
return ret ?: count;
}
-static struct device_attribute wq_sysfs_unbound_attrs[] = {
- __ATTR(nice, 0644, wq_nice_show, wq_nice_store),
- __ATTR(cpumask, 0644, wq_cpumask_show, wq_cpumask_store),
- __ATTR(affinity_scope, 0644, wq_affn_scope_show, wq_affn_scope_store),
- __ATTR(affinity_strict, 0644, wq_affinity_strict_show, wq_affinity_strict_store),
- __ATTR_NULL,
+static DEVICE_ATTR_RW(nice);
+static DEVICE_ATTR_RW(affn_scope);
+static DEVICE_ATTR_RW(affinity_strict);
+/* Avoid naming clash with the other cpumask */
+static struct device_attribute dev_attr_unbound_cpumask =
+ __ATTR(cpumask, 0644, unbound_cpumask_show, unbound_cpumask_store);
+
+static struct attribute *wq_sysfs_unbound_attrs[] = {
+ &dev_attr_nice.attr,
+ &dev_attr_unbound_cpumask.attr,
+ &dev_attr_affn_scope.attr,
+ &dev_attr_affinity_strict.attr,
+ NULL,
};
+static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj,
+ struct attribute *attr, int n)
+{
+ struct device *dev = kobj_to_dev(kobj);
+ struct workqueue_struct *wq = dev_to_wq(dev);
+
+ if (!(wq->flags & WQ_UNBOUND))
+ return SYSFS_GROUP_INVISIBLE;
+
+ return attr->mode;
+}
+
+static const struct attribute_group wq_sysfs_unbound_group = {
+ .is_visible = wq_sysfs_unbound_group_visible,
+ .attrs = wq_sysfs_unbound_attrs,
+};
+
+__ATTRIBUTE_GROUPS(wq_sysfs_unbound);
+
static const struct bus_type wq_subsys = {
.name = "workqueue",
.dev_groups = wq_sysfs_groups,
@@ -7907,14 +7933,9 @@ int workqueue_sysfs_register(struct workqueue_struct *wq)
wq_dev->wq = wq;
wq_dev->dev.bus = &wq_subsys;
wq_dev->dev.release = wq_device_release;
+ wq_dev->dev.groups = wq_sysfs_unbound_groups;
dev_set_name(&wq_dev->dev, "%s", wq->name);
- /*
- * attrs are created separately. Suppress uevent until
- * everything is ready.
- */
- dev_set_uevent_suppress(&wq_dev->dev, true);
-
ret = device_register(&wq_dev->dev);
if (ret) {
put_device(&wq_dev->dev);
@@ -7922,21 +7943,6 @@ int workqueue_sysfs_register(struct workqueue_struct *wq)
return ret;
}
- if (wq->flags & WQ_UNBOUND) {
- struct device_attribute *attr;
-
- for (attr = wq_sysfs_unbound_attrs; attr->attr.name; attr++) {
- ret = device_create_file(&wq_dev->dev, attr);
- if (ret) {
- device_unregister(&wq_dev->dev);
- wq->wq_dev = NULL;
- return ret;
- }
- }
- }
-
- dev_set_uevent_suppress(&wq_dev->dev, false);
- kobject_uevent(&wq_dev->dev.kobj, KOBJ_ADD);
return 0;
}
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC v5 2/3] workqueue: Add support for real-time workers
2026-09-23 16:12 [RFC v5 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin
2026-09-23 16:12 ` [RFC v5 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
@ 2026-09-23 16:12 ` Tvrtko Ursulin
2026-09-29 0:00 ` Tejun Heo
2026-09-23 16:12 ` [RFC v5 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin
2 siblings, 1 reply; 6+ messages in thread
From: Tvrtko Ursulin @ 2026-09-23 16:12 UTC (permalink / raw)
To: linux-kernel
Cc: kernel-dev, dri-devel, Tvrtko Ursulin, Boris Brezillon,
Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost,
Steven Price, Tejun Heo
For use cases such as the DRM scheduler submitting work to the GPU on
behalf of low latency userspace applications, where latter have sufficient
privileges to have had successfully obtained realtime Vulkan global
priority, competing with random background CPU load can create large
latency spikes which gets in the way of a smooth user experience.
For these situations the existing WQ_HIGHPRI does not bring a noticeable
improvement and a stronger hint is needed.
Lets add WQ_RTPRI which creates workers with a SCHED_FIFO scheduling class
to improve this.
We use a minimum priority level since we only care about winning the
contest against normal background CPU load.
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Bradley Morgan <include@grrlz.net>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
---
v2:
* Limit WQ_RTPRI to unbound workqueues and make it have strict CPU
affinitity. (Tejun)
* Fixed commit message typos. (AI)
* Fixed sysfs handling, max_active setting and user modified nice
application. (AI)
v3:
* Fix worker->pool null pointer dereference race by moving the
global decrement to detach_dying_workers().
* Rebase for upstream changes.
v4:
* Fixed onion unwind.
* Moved affinity setting to default attributes.
v5:
* Dropped global and local limits.
* Documented in workqueue.rst.
* Added NR_WQ_ATTRIBUTES.
* Reverted BH handling changes.
v6:
* Dropped separate attr->prio in favour of RTPRI_NICE_LEVEL checks. (Tejun)
* Reworked on top of tj/for-7.4.
---
Documentation/core-api/workqueue.rst | 5 +++++
include/linux/workqueue.h | 9 ++++----
kernel/workqueue.c | 33 +++++++++++++++++++++++++---
3 files changed, 40 insertions(+), 7 deletions(-)
diff --git a/Documentation/core-api/workqueue.rst b/Documentation/core-api/workqueue.rst
index bb770f556568..24df3d87e2dd 100644
--- a/Documentation/core-api/workqueue.rst
+++ b/Documentation/core-api/workqueue.rst
@@ -225,6 +225,11 @@ resources, scheduled and executed.
each other. Each maintains its separate pool of workers and
implements concurrency management among its workers.
+``WQ_RTPRI``
+ Real time priority workqueues must be created as unbound and have the strict
+ CPU affinity set. Their worker threads use the FIFO scheduling policy with
+ the lowest priority.
+
``WQ_CPU_INTENSIVE``
Work items of a CPU intensive wq do not contribute to the
concurrency level. In other words, runnable CPU intensive
diff --git a/include/linux/workqueue.h b/include/linux/workqueue.h
index a283766a192a..161f71f4c264 100644
--- a/include/linux/workqueue.h
+++ b/include/linux/workqueue.h
@@ -374,8 +374,9 @@ enum wq_flags {
WQ_FREEZABLE = 1 << 2, /* freeze during suspend */
WQ_MEM_RECLAIM = 1 << 3, /* may be used for memory reclaim */
WQ_HIGHPRI = 1 << 4, /* high priority */
- WQ_CPU_INTENSIVE = 1 << 5, /* cpu intensive workqueue */
- WQ_SYSFS = 1 << 6, /* visible in sysfs, see workqueue_sysfs_register() */
+ WQ_RTPRI = 1 << 5, /* real-time priority, valid only with WQ_UNBOUND */
+ WQ_CPU_INTENSIVE = 1 << 6, /* cpu intensive workqueue */
+ WQ_SYSFS = 1 << 7, /* visible in sysfs, see workqueue_sysfs_register() */
/*
* Per-cpu workqueues are generally preferred because they tend to
@@ -402,8 +403,8 @@ enum wq_flags {
*
* http://thread.gmane.org/gmane.linux.kernel/1480396
*/
- WQ_POWER_EFFICIENT = 1 << 7,
- WQ_PERCPU = 1 << 8, /* bound to a specific cpu */
+ WQ_POWER_EFFICIENT = 1 << 8,
+ WQ_PERCPU = 1 << 9, /* bound to a specific cpu */
__WQ_DESTROYING = 1 << 15, /* internal: workqueue is destroying */
__WQ_DRAINING = 1 << 16, /* internal: workqueue is draining */
diff --git a/kernel/workqueue.c b/kernel/workqueue.c
index e3a4ad56dae8..5cf486291af8 100644
--- a/kernel/workqueue.c
+++ b/kernel/workqueue.c
@@ -127,6 +127,7 @@ enum wq_internal_consts {
*/
RESCUER_NICE_LEVEL = MIN_NICE,
HIGHPRI_NICE_LEVEL = MIN_NICE,
+ RTPRI_NICE_LEVEL = MIN_NICE - 1,
WQ_NAME_LEN = 32,
WORKER_ID_LEN = 10 + WQ_NAME_LEN, /* "kworker/R-" + WQ_NAME_LEN */
@@ -3018,7 +3019,11 @@ static struct worker *create_worker(struct worker_pool *pool)
goto fail;
}
- set_user_nice(worker->task, pool->attrs->nice);
+ if (pool->attrs->nice == RTPRI_NICE_LEVEL)
+ sched_set_fifo_low(worker->task);
+ else
+ set_user_nice(worker->task, pool->attrs->nice);
+
kthread_bind_mask(worker->task, pool_allowed_cpus(pool));
}
@@ -5928,8 +5933,17 @@ static struct workqueue_attrs *alloc_wq_std_attrs(struct workqueue_struct *wq)
if (!attrs)
return NULL;
- if (wq->flags & WQ_HIGHPRI)
+ if (wq->flags & WQ_RTPRI) {
+ attrs->nice = RTPRI_NICE_LEVEL;
+ /*
+ * RT workqueues have strict CPU affinity for low
+ * latency execution.
+ */
+ attrs->affn_scope = WQ_AFFN_CPU;
+ attrs->affn_strict = true;
+ } else if (wq->flags & WQ_HIGHPRI) {
attrs->nice = HIGHPRI_NICE_LEVEL;
+ }
if (wq->flags & __WQ_ORDERED)
attrs->ordered = true;
@@ -6115,6 +6129,12 @@ static struct workqueue_struct *__alloc_workqueue(const char *fmt,
return NULL;
}
+ if (flags & WQ_RTPRI) {
+ if (WARN_ON_ONCE((flags & (WQ_HIGHPRI | WQ_UNBOUND)) !=
+ WQ_UNBOUND))
+ return NULL;
+ }
+
/* see the comment above the definition of WQ_POWER_EFFICIENT */
if ((flags & WQ_POWER_EFFICIENT) && wq_power_efficient)
flags = (flags & ~WQ_PERCPU) | WQ_UNBOUND;
@@ -7606,7 +7626,10 @@ static ssize_t nice_show(struct device *dev, struct device_attribute *attr,
int written;
mutex_lock(&wq->mutex);
- written = scnprintf(buf, PAGE_SIZE, "%d\n", wq->attrs->nice);
+ if (wq->attrs->nice == RTPRI_NICE_LEVEL)
+ written = scnprintf(buf, PAGE_SIZE, "rt\n");
+ else
+ written = scnprintf(buf, PAGE_SIZE, "%d\n", wq->attrs->nice);
mutex_unlock(&wq->mutex);
return written;
@@ -7786,6 +7809,10 @@ static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj,
if (!(wq->flags & WQ_UNBOUND))
return SYSFS_GROUP_INVISIBLE;
+ /* Do not allow priority changes for RT workers. */
+ if ((wq->flags & WQ_RTPRI) && !strcmp(attr->name, "nice"))
+ return 0444;
+
return attr->mode;
}
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* [RFC v5 3/3] drm/panthor: Create per queue priority workqueues
2026-09-23 16:12 [RFC v5 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin
2026-09-23 16:12 ` [RFC v5 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
2026-09-23 16:12 ` [RFC v5 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin
@ 2026-09-23 16:12 ` Tvrtko Ursulin
2 siblings, 0 replies; 6+ messages in thread
From: Tvrtko Ursulin @ 2026-09-23 16:12 UTC (permalink / raw)
To: linux-kernel
Cc: kernel-dev, dri-devel, Tvrtko Ursulin, Boris Brezillon,
Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo
Split the single workqueue shared between the driver internal logic and
DRM scheduler use into separate ones, where the DRM scheduler one is
created per GPU priority level using the appropriate mapping to
workqueue priorities.
Low and medium GPU priority are served by a normal workqueue,
high is server by a WQ_HIGHPRI instance, while realtime GPU priority is
using the newly added WQ_RTPRI flag for lowest possible latency.
These workqueues are device global and for all three we set the maximum
concurrency to two in order to keep the GPU optimally fed with work.
Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
Cc: Boris Brezillon <boris.brezillon@collabora.com>
Cc: Chia-I Wu <olv@google.com>
Cc: Liviu Dudau <liviu.dudau@arm.com>
Cc: Matthew Brost <matthew.brost@intel.com>
Cc: Steven Price <steven.price@arm.com>
Cc: Tejun Heo <tj@kernel.org>
---
drivers/gpu/drm/panthor/panthor_sched.c | 37 ++++++++++++++++++++++---
1 file changed, 33 insertions(+), 4 deletions(-)
diff --git a/drivers/gpu/drm/panthor/panthor_sched.c b/drivers/gpu/drm/panthor/panthor_sched.c
index 5b34032deff8..de32ad68230c 100644
--- a/drivers/gpu/drm/panthor/panthor_sched.c
+++ b/drivers/gpu/drm/panthor/panthor_sched.c
@@ -152,11 +152,18 @@ struct panthor_scheduler {
*
* Used for the scheduler tick, group update or other kind of FW
* event processing that can't be handled in the threaded interrupt
- * path. Also passed to the drm_gpu_scheduler instances embedded
- * in panthor_queue.
+ * path.
*/
struct workqueue_struct *wq;
+ /**
+ * @submit_wq: Per priority workqueues for the DRM scheduler
+ *
+ * Passed to the drm_gpu_scheduler instances embedded
+ * in panthor_queue based on the queue priority.
+ */
+ struct workqueue_struct *submit_wq[PANTHOR_CSG_PRIORITY_COUNT];
+
/**
* @heap_alloc_wq: Workqueue used to schedule tiler_oom works.
*
@@ -3582,8 +3589,14 @@ group_create_queue(struct panthor_group *group,
goto err_free_queue;
}
+ if (group->priority >= ARRAY_SIZE(group->ptdev->scheduler->submit_wq) ||
+ !group->ptdev->scheduler->submit_wq[group->priority]) {
+ ret = -EINVAL;
+ goto err_free_queue;
+ }
+
sched_args.name = queue->name;
-
+ sched_args.submit_wq = group->ptdev->scheduler->submit_wq[group->priority];
ret = drm_sched_init(&queue->scheduler, &sched_args);
if (ret)
goto err_free_queue;
@@ -4073,6 +4086,15 @@ static void panthor_sched_fini(struct drm_device *ddev, void *res)
if (!sched || !sched->csg_slot_count)
return;
+ if (sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM])
+ destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM]);
+
+ if (sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH])
+ destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH]);
+
+ if (sched->submit_wq[PANTHOR_CSG_PRIORITY_RT])
+ destroy_workqueue(sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]);
+
if (sched->wq)
destroy_workqueue(sched->wq);
@@ -4174,7 +4196,14 @@ int panthor_sched_init(struct panthor_device *ptdev)
*/
sched->heap_alloc_wq = alloc_workqueue("panthor-heap-alloc", WQ_UNBOUND, 0);
sched->wq = alloc_workqueue("panthor-csf-sched", WQ_MEM_RECLAIM | WQ_UNBOUND, 0);
- if (!sched->wq || !sched->heap_alloc_wq) {
+ sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] = alloc_workqueue("panthor-drm", WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
+ sched->submit_wq[PANTHOR_CSG_PRIORITY_LOW] = sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM];
+ sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] = alloc_workqueue("panthor-drm-high", WQ_HIGHPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
+ sched->submit_wq[PANTHOR_CSG_PRIORITY_RT] = alloc_workqueue("panthor-drm-rt", WQ_RTPRI | WQ_MEM_RECLAIM | WQ_UNBOUND, 2);
+ if (!sched->wq || !sched->heap_alloc_wq ||
+ !sched->submit_wq[PANTHOR_CSG_PRIORITY_MEDIUM] ||
+ !sched->submit_wq[PANTHOR_CSG_PRIORITY_HIGH] ||
+ !sched->submit_wq[PANTHOR_CSG_PRIORITY_RT]) {
panthor_sched_fini(&ptdev->base, sched);
drm_err(&ptdev->base, "Failed to allocate the workqueues");
return -ENOMEM;
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [RFC v5 1/3] workqueue: Simplify unbound sysfs attribute registration
2026-09-23 16:12 ` [RFC v5 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
@ 2026-09-23 17:12 ` Tvrtko Ursulin
0 siblings, 0 replies; 6+ messages in thread
From: Tvrtko Ursulin @ 2026-09-23 17:12 UTC (permalink / raw)
To: linux-kernel
Cc: kernel-dev, dri-devel, Boris Brezillon, Bradley Morgan,
Chia-I Wu, Liviu Dudau, Matthew Brost, Steven Price, Tejun Heo
On 23/09/2026 17:12, Tvrtko Ursulin wrote:
> Instead of manually registering each attribute we can put them in an
> attribute group with a visibility check and device core will handle the
> rest, which simplifies the registration and error unwind.
>
> Signed-off-by: Tvrtko Ursulin <tvrtko.ursulin@igalia.com>
> Cc: Boris Brezillon <boris.brezillon@collabora.com>
> Cc: Bradley Morgan <include@grrlz.net>
> Cc: Chia-I Wu <olv@google.com>
> Cc: Liviu Dudau <liviu.dudau@arm.com>
> Cc: Matthew Brost <matthew.brost@intel.com>
> Cc: Steven Price <steven.price@arm.com>
> Cc: Tejun Heo <tj@kernel.org>
> ---
> kernel/workqueue.c | 98 ++++++++++++++++++++++++----------------------
> 1 file changed, 52 insertions(+), 46 deletions(-)
>
> diff --git a/kernel/workqueue.c b/kernel/workqueue.c
> index e618108c6127..e3a4ad56dae8 100644
> --- a/kernel/workqueue.c
> +++ b/kernel/workqueue.c
> @@ -7599,8 +7599,8 @@ static const struct attribute_group wq_sysfs_group = {
> };
> __ATTRIBUTE_GROUPS(wq_sysfs);
>
> -static ssize_t wq_nice_show(struct device *dev, struct device_attribute *attr,
> - char *buf)
> +static ssize_t nice_show(struct device *dev, struct device_attribute *attr,
> + char *buf)
> {
> struct workqueue_struct *wq = dev_to_wq(dev);
> int written;
> @@ -7627,8 +7627,8 @@ static struct workqueue_attrs *wq_sysfs_prep_attrs(struct workqueue_struct *wq)
> return attrs;
> }
>
> -static ssize_t wq_nice_store(struct device *dev, struct device_attribute *attr,
> - const char *buf, size_t count)
> +static ssize_t nice_store(struct device *dev, struct device_attribute *attr,
> + const char *buf, size_t count)
> {
> struct workqueue_struct *wq = dev_to_wq(dev);
> struct workqueue_attrs *attrs;
> @@ -7652,8 +7652,8 @@ static ssize_t wq_nice_store(struct device *dev, struct device_attribute *attr,
> return ret ?: count;
> }
>
> -static ssize_t wq_cpumask_show(struct device *dev,
> - struct device_attribute *attr, char *buf)
> +static ssize_t unbound_cpumask_show(struct device *dev,
> + struct device_attribute *attr, char *buf)
> {
> struct workqueue_struct *wq = dev_to_wq(dev);
> int written;
> @@ -7665,9 +7665,9 @@ static ssize_t wq_cpumask_show(struct device *dev,
> return written;
> }
>
> -static ssize_t wq_cpumask_store(struct device *dev,
> - struct device_attribute *attr,
> - const char *buf, size_t count)
> +static ssize_t unbound_cpumask_store(struct device *dev,
> + struct device_attribute *attr,
> + const char *buf, size_t count)
> {
> struct workqueue_struct *wq = dev_to_wq(dev);
> struct workqueue_attrs *attrs;
> @@ -7689,8 +7689,8 @@ static ssize_t wq_cpumask_store(struct device *dev,
> return ret ?: count;
> }
>
> -static ssize_t wq_affn_scope_show(struct device *dev,
> - struct device_attribute *attr, char *buf)
> +static ssize_t affn_scope_show(struct device *dev,
> + struct device_attribute *attr, char *buf)
> {
> struct workqueue_struct *wq = dev_to_wq(dev);
> int written;
> @@ -7708,9 +7708,9 @@ static ssize_t wq_affn_scope_show(struct device *dev,
> return written;
> }
>
> -static ssize_t wq_affn_scope_store(struct device *dev,
> - struct device_attribute *attr,
> - const char *buf, size_t count)
> +static ssize_t affn_scope_store(struct device *dev,
> + struct device_attribute *attr,
> + const char *buf, size_t count)
> {
> struct workqueue_struct *wq = dev_to_wq(dev);
> struct workqueue_attrs *attrs;
> @@ -7731,8 +7731,8 @@ static ssize_t wq_affn_scope_store(struct device *dev,
> return ret ?: count;
> }
>
> -static ssize_t wq_affinity_strict_show(struct device *dev,
> - struct device_attribute *attr, char *buf)
> +static ssize_t affinity_strict_show(struct device *dev,
> + struct device_attribute *attr, char *buf)
> {
> struct workqueue_struct *wq = dev_to_wq(dev);
>
> @@ -7740,9 +7740,9 @@ static ssize_t wq_affinity_strict_show(struct device *dev,
> wq->attrs->affn_strict);
> }
>
> -static ssize_t wq_affinity_strict_store(struct device *dev,
> - struct device_attribute *attr,
> - const char *buf, size_t count)
> +static ssize_t affinity_strict_store(struct device *dev,
> + struct device_attribute *attr,
> + const char *buf, size_t count)
> {
> struct workqueue_struct *wq = dev_to_wq(dev);
> struct workqueue_attrs *attrs;
> @@ -7762,14 +7762,40 @@ static ssize_t wq_affinity_strict_store(struct device *dev,
> return ret ?: count;
> }
>
> -static struct device_attribute wq_sysfs_unbound_attrs[] = {
> - __ATTR(nice, 0644, wq_nice_show, wq_nice_store),
> - __ATTR(cpumask, 0644, wq_cpumask_show, wq_cpumask_store),
> - __ATTR(affinity_scope, 0644, wq_affn_scope_show, wq_affn_scope_store),
> - __ATTR(affinity_strict, 0644, wq_affinity_strict_show, wq_affinity_strict_store),
> - __ATTR_NULL,
> +static DEVICE_ATTR_RW(nice);
> +static DEVICE_ATTR_RW(affn_scope);
Sashiko on dri-devel pointed out I blundered with the accidental rename
here. But lets first see if people think this simplification is desired
to begin with. I think it is nicer than having to suppress and re-enable
uvents, and unwind on errors, plus, it's handy for the WQ_RTPRI patch to
restrict write access to the nice attribute.
Regards,
Tvrtko
> +static DEVICE_ATTR_RW(affinity_strict);
> +/* Avoid naming clash with the other cpumask */
> +static struct device_attribute dev_attr_unbound_cpumask =
> + __ATTR(cpumask, 0644, unbound_cpumask_show, unbound_cpumask_store);
> +
> +static struct attribute *wq_sysfs_unbound_attrs[] = {
> + &dev_attr_nice.attr,
> + &dev_attr_unbound_cpumask.attr,
> + &dev_attr_affn_scope.attr,
> + &dev_attr_affinity_strict.attr,
> + NULL,
> };
>
> +static umode_t wq_sysfs_unbound_group_visible(struct kobject *kobj,
> + struct attribute *attr, int n)
> +{
> + struct device *dev = kobj_to_dev(kobj);
> + struct workqueue_struct *wq = dev_to_wq(dev);
> +
> + if (!(wq->flags & WQ_UNBOUND))
> + return SYSFS_GROUP_INVISIBLE;
> +
> + return attr->mode;
> +}
> +
> +static const struct attribute_group wq_sysfs_unbound_group = {
> + .is_visible = wq_sysfs_unbound_group_visible,
> + .attrs = wq_sysfs_unbound_attrs,
> +};
> +
> +__ATTRIBUTE_GROUPS(wq_sysfs_unbound);
> +
> static const struct bus_type wq_subsys = {
> .name = "workqueue",
> .dev_groups = wq_sysfs_groups,
> @@ -7907,14 +7933,9 @@ int workqueue_sysfs_register(struct workqueue_struct *wq)
> wq_dev->wq = wq;
> wq_dev->dev.bus = &wq_subsys;
> wq_dev->dev.release = wq_device_release;
> + wq_dev->dev.groups = wq_sysfs_unbound_groups;
> dev_set_name(&wq_dev->dev, "%s", wq->name);
>
> - /*
> - * attrs are created separately. Suppress uevent until
> - * everything is ready.
> - */
> - dev_set_uevent_suppress(&wq_dev->dev, true);
> -
> ret = device_register(&wq_dev->dev);
> if (ret) {
> put_device(&wq_dev->dev);
> @@ -7922,21 +7943,6 @@ int workqueue_sysfs_register(struct workqueue_struct *wq)
> return ret;
> }
>
> - if (wq->flags & WQ_UNBOUND) {
> - struct device_attribute *attr;
> -
> - for (attr = wq_sysfs_unbound_attrs; attr->attr.name; attr++) {
> - ret = device_create_file(&wq_dev->dev, attr);
> - if (ret) {
> - device_unregister(&wq_dev->dev);
> - wq->wq_dev = NULL;
> - return ret;
> - }
> - }
> - }
> -
> - dev_set_uevent_suppress(&wq_dev->dev, false);
> - kobject_uevent(&wq_dev->dev.kobj, KOBJ_ADD);
> return 0;
> }
>
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [RFC v5 2/3] workqueue: Add support for real-time workers
2026-09-23 16:12 ` [RFC v5 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin
@ 2026-09-29 0:00 ` Tejun Heo
0 siblings, 0 replies; 6+ messages in thread
From: Tejun Heo @ 2026-09-29 0:00 UTC (permalink / raw)
To: Tvrtko Ursulin
Cc: linux-kernel, kernel-dev, dri-devel, Boris Brezillon,
Bradley Morgan, Chia-I Wu, Liviu Dudau, Matthew Brost,
Steven Price, Lai Jiangshan, Breno Leitao
Hello,
On Wed, Sep 23, 2026 at 05:12:50PM +0100, Tvrtko Ursulin wrote:
> For use cases such as the DRM scheduler submitting work to the GPU on
> behalf of low latency userspace applications, where latter have sufficient
> privileges to have had successfully obtained realtime Vulkan global
> priority, competing with random background CPU load can create large
> latency spikes which gets in the way of a smooth user experience.
panthor's group_priority_permit() also allows realtime groups for DRM
master without CAP_SYS_NICE. What's the usage model there? Should DRM
master be enough to get RT workers?
> @@ -374,8 +374,9 @@ enum wq_flags {
> WQ_FREEZABLE = 1 << 2, /* freeze during suspend */
> WQ_MEM_RECLAIM = 1 << 3, /* may be used for memory reclaim */
> WQ_HIGHPRI = 1 << 4, /* high priority */
> - WQ_CPU_INTENSIVE = 1 << 5, /* cpu intensive workqueue */
> - WQ_SYSFS = 1 << 6, /* visible in sysfs, see workqueue_sysfs_register() */
> + WQ_RTPRI = 1 << 5, /* real-time priority, valid only with WQ_UNBOUND */
> + WQ_CPU_INTENSIVE = 1 << 6, /* cpu intensive workqueue */
> + WQ_SYSFS = 1 << 7, /* visible in sysfs, see workqueue_sysfs_register() */
Can we name it just WQ_RT? Also, I think WQ_RT is closer to WQ_BH. We can
reorder the flags later if that helps but for now can you just put it in an
empty slot?
> @@ -127,6 +127,7 @@ enum wq_internal_consts {
> */
> RESCUER_NICE_LEVEL = MIN_NICE,
> HIGHPRI_NICE_LEVEL = MIN_NICE,
> + RTPRI_NICE_LEVEL = MIN_NICE - 1,
Can we match the scheduler's representation instead by replacing
attrs->nice with attrs->prio which uses the same encoding as p->prio?
Normal pools would be NICE_TO_PRIO(nice) and WQ_RT pools would be in the RT
range. create_worker() would then do:
if (rt_prio(pool->attrs->prio))
sched_set_fifo_low(worker->task);
else
set_user_nice(worker->task, PRIO_TO_NICE(pool->attrs->prio));
> @@ -7606,7 +7626,10 @@ static ssize_t nice_show(struct device *dev, struct device_attribute *attr,
> int written;
>
> mutex_lock(&wq->mutex);
> - written = scnprintf(buf, PAGE_SIZE, "%d\n", wq->attrs->nice);
> + if (wq->attrs->nice == RTPRI_NICE_LEVEL)
> + written = scnprintf(buf, PAGE_SIZE, "rt\n");
> + else
> + written = scnprintf(buf, PAGE_SIZE, "%d\n", wq->attrs->nice);
Can you also show rt in pr_cont_pool_info() and tools/workqueue/wq_dump.py?
Thanks.
--
tejun
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-09-29 0:00 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-23 16:12 [RFC v5 0/3] Realtime workqueues and panthor realtime submission Tvrtko Ursulin
2026-09-23 16:12 ` [RFC v5 1/3] workqueue: Simplify unbound sysfs attribute registration Tvrtko Ursulin
2026-09-23 17:12 ` Tvrtko Ursulin
2026-09-23 16:12 ` [RFC v5 2/3] workqueue: Add support for real-time workers Tvrtko Ursulin
2026-09-29 0:00 ` Tejun Heo
2026-09-23 16:12 ` [RFC v5 3/3] drm/panthor: Create per queue priority workqueues Tvrtko Ursulin
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®