From: Changwoo Min <changwoo@igalia.com>
To: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
Andrea Righi <arighi@nvidia.com>
Cc: sched-ext@lists.linux.dev, Emil Tsalapatis <emil@etsalapatis.com>,
linux-kernel@vger.kernel.org,
Cheng-Yang Chou <yphbchou0911@gmail.com>
Subject: Re: [PATCH 10/17] sched_ext: Add cmask, a base-windowed bitmap over cid space
Date: Wed, 29 Apr 2026 21:47:15 +0900 [thread overview]
Message-ID: <b8bc70c1-8914-4a39-825d-e94cfb983d60@igalia.com> (raw)
In-Reply-To: <20260428203545.181052-11-tj@kernel.org>
On 4/29/26 5:35 AM, Tejun Heo wrote:
> +/*
> + * x86 BPF JIT rejects BPF_OR | BPF_FETCH and BPF_AND | BPF_FETCH on arena
> + * pointers (see bpf_jit_supports_insn() in arch/x86/net/bpf_jit_comp.c). Only
> + * BPF_CMPXCHG / BPF_XCHG / BPF_ADD with FETCH are allowed. Implement
> + * test_and_{set,clear} and the atomic set/clear via a cmpxchg loop.
> + *
> + * CMASK_CAS_TRIES is far above what any non-pathological contention needs.
> + * Exhausting it means the bit update was lost, which corrupts the caller's view
> + * of the bitmap, so raise scx_bpf_error() to abort the scheduler.
> + */
> +#define CMASK_CAS_TRIES 1024
> +
> +static __always_inline void cmask_set(struct scx_cmask __arena *m, u32 cid)
> +{
> + u64 __arena *w;
> + u64 bit, old, new;
> + u32 i;
> +
> + if (!__cmask_contains(m, cid))
> + return;
> + w = __cmask_word(m, cid);
> + bit = BIT_U64(cid & 63);
> + bpf_for(i, 0, CMASK_CAS_TRIES) {
> + old = *w;
> + if (old & bit)
> + return;
> + new = old | bit;
> + if (__sync_val_compare_and_swap(w, old, new) == old)
> + return;
> + }
> + scx_bpf_error("cmask_set CAS exhausted at cid %u", cid);
> +}
> +
> +static __always_inline void cmask_clear(struct scx_cmask __arena *m, u32 cid)
> +{
> + u64 __arena *w;
> + u64 bit, old, new;
> + u32 i;
> +
> + if (!__cmask_contains(m, cid))
> + return;
> + w = __cmask_word(m, cid);
> + bit = BIT_U64(cid & 63);
> + bpf_for(i, 0, CMASK_CAS_TRIES) {
> + old = *w;
> + if (!(old & bit))
> + return;
> + new = old & ~bit;
> + if (__sync_val_compare_and_swap(w, old, new) == old)
> + return;
> + }
> + scx_bpf_error("cmask_clear CAS exhausted at cid %u", cid);
> +}
> +
> +static __always_inline bool cmask_test_and_set(struct scx_cmask __arena *m, u32 cid)
> +{
> + u64 __arena *w;
> + u64 bit, old, new;
> + u32 i;
> +
> + if (!__cmask_contains(m, cid))
> + return false;
> + w = __cmask_word(m, cid);
> + bit = BIT_U64(cid & 63);
> + bpf_for(i, 0, CMASK_CAS_TRIES) {
> + old = *w;
> + if (old & bit)
> + return true;
> + new = old | bit;
> + if (__sync_val_compare_and_swap(w, old, new) == old)
> + return false;
> + }
> + scx_bpf_error("cmask_test_and_set CAS exhausted at cid %u", cid);
> + return false;
> +}
> +
> +static __always_inline bool cmask_test_and_clear(struct scx_cmask __arena *m, u32 cid)
> +{
> + u64 __arena *w;
> + u64 bit, old, new;
> + u32 i;
> +
> + if (!__cmask_contains(m, cid))
> + return false;
> + w = __cmask_word(m, cid);
> + bit = BIT_U64(cid & 63);
> + bpf_for(i, 0, CMASK_CAS_TRIES) {
> + old = *w;
> + if (!(old & bit))
> + return false;
> + new = old & ~bit;
> + if (__sync_val_compare_and_swap(w, old, new) == old)
> + return true;
> + }
> + scx_bpf_error("cmask_test_and_clear CAS exhausted at cid %u", cid);
> + return false;
> +}
Exiting a BPF scheduler when CAS retries fail seems too brutal, while it
is extremely rare to happen. What about adding a kfunc for the slow path
that runs only when the CAS retry fails? For example,
scx_bpf_cmask_test_and_clear( ) does the same thing as
cmask_test_and_clear(), but it never fails. cmask_test_and_clear() calls
scx_bpf_cmask_test_and_clear( ) only when the CAS retry fails. In this
way, we can use the fast path in the BPF implementation while ensuring
the operation succeeds.
> +/**
> + * cmask_next_set - find the first set bit at or after @cid
> + * @m: cmask to search
> + * @cid: starting cid (clamped to @m->base if below)
> + *
> + * Returns the smallest set cid in [@cid, @m->base + @m->nr_bits), or
> + * @m->base + @m->nr_bits if none (the out-of-range sentinel matches the
> + * termination condition used by cmask_for_each()).
> + */
> +static __always_inline u32 cmask_next_set(const struct scx_cmask __arena *m, u32 cid)
> +{
> + u32 end = m->base + m->nr_bits;
> + u32 base = m->base / 64;
> + u32 last_wi = (end - 1) / 64 - base;
> + u32 start_wi, start_bit, i;
> +
> + if (cid < m->base)
> + cid = m->base;
> + if (cid >= end)
> + return end;
> +
> + start_wi = cid / 64 - base;
> + start_bit = cid & 63;
> +
> + bpf_for(i, 0, CMASK_MAX_WORDS) {
> + u32 wi = start_wi + i;
> + u64 word;
> + u32 found;
> +
> + if (wi > last_wi)
> + break;
> +
> + word = m->bits[wi];
> + if (i == 0)
> + word &= GENMASK_U64(63, start_bit);
> + if (!word)
> + continue;
> +
> + found = (base + wi) * 64 + __builtin_ctzll(word);
Some compiler versions (e.g., clang-18 or older) don’t support
__builtin_ctzll(). To handle this gracefully, there is already a
wrapper, ctzll(), in common.bpf.h. So, I suggest using ctzll()
for compatibility.
Reviewed-by: Changwoo Min <changwoo@igalia.com>
next prev parent reply other threads:[~2026-04-29 12:47 UTC|newest]
Thread overview: 33+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-04-28 20:35 [PATCHSET v3 sched_ext/for-7.2] sched_ext: Topological CPU IDs and cid-form struct_ops Tejun Heo
2026-04-28 20:35 ` [PATCH 01/17] sched_ext: Add ext_types.h for early subsystem-wide defs Tejun Heo
2026-04-28 20:35 ` [PATCH 02/17] sched_ext: Rename ops_cpu_valid() to scx_cpu_valid() and expose it Tejun Heo
2026-04-28 20:35 ` [PATCH 03/17] sched_ext: Move scx_exit(), scx_error() and friends to ext_internal.h Tejun Heo
2026-04-28 20:35 ` [PATCH 04/17] sched_ext: Shift scx_kick_cpu() validity check to scx_bpf_kick_cpu() Tejun Heo
2026-04-28 20:35 ` [PATCH 05/17] sched_ext: Relocate cpu_acquire/cpu_release to end of struct sched_ext_ops Tejun Heo
2026-04-28 20:35 ` [PATCH 06/17] sched_ext: Make scx_enable() take scx_enable_cmd Tejun Heo
2026-04-28 20:35 ` [PATCH 07/17] sched_ext: Add topological CPU IDs (cids) Tejun Heo
2026-04-28 20:35 ` [PATCH 08/17] sched_ext: Add scx_bpf_cid_override() kfunc Tejun Heo
2026-04-29 14:07 ` Andrea Righi
2026-04-29 17:06 ` Tejun Heo
2026-04-29 17:20 ` Andrea Righi
2026-04-28 20:35 ` [PATCH 09/17] tools/sched_ext: Add struct_size() helpers to common.bpf.h Tejun Heo
2026-04-28 20:35 ` [PATCH 10/17] sched_ext: Add cmask, a base-windowed bitmap over cid space Tejun Heo
2026-04-29 12:47 ` Changwoo Min [this message]
2026-04-29 17:16 ` Tejun Heo
2026-04-28 20:35 ` [PATCH 11/17] sched_ext: Add cid-form kfunc wrappers alongside cpu-form Tejun Heo
2026-04-28 20:35 ` [PATCH 12/17] sched_ext: Add bpf_sched_ext_ops_cid struct_ops type Tejun Heo
2026-04-28 20:35 ` [PATCH 13/17] sched_ext: Forbid cpu-form kfuncs from cid-form schedulers Tejun Heo
2026-04-28 20:35 ` [PATCH 14/17] tools/sched_ext: scx_qmap: Restart on hotplug instead of cpu_online/offline Tejun Heo
2026-04-28 20:35 ` [PATCH 15/17] tools/sched_ext: scx_qmap: Add cmask-based idle tracking and cid-based idle pick Tejun Heo
2026-04-28 20:35 ` [PATCH 16/17] tools/sched_ext: scx_qmap: Port to cid-form struct_ops Tejun Heo
2026-04-29 12:47 ` Changwoo Min
2026-04-29 13:53 ` Andrea Righi
2026-04-29 16:42 ` Tejun Heo
2026-04-28 20:35 ` [PATCH 17/17] sched_ext: Require cid-form struct_ops for sub-sched support Tejun Heo
2026-04-29 12:49 ` [PATCHSET v3 sched_ext/for-7.2] sched_ext: Topological CPU IDs and cid-form struct_ops Changwoo Min
2026-04-29 13:29 ` Andrea Righi
2026-04-29 14:11 ` Andrea Righi
2026-04-29 17:06 ` Tejun Heo
-- strict thread matches above, loose matches on Subject: below --
2026-04-29 18:21 [PATCHSET v4 " Tejun Heo
2026-04-29 18:21 ` [PATCH 10/17] sched_ext: Add cmask, a base-windowed bitmap over cid space Tejun Heo
2026-04-24 17:27 [PATCHSET v2 REPOST sched_ext/for-7.2] sched_ext: Topological CPU IDs and cid-form struct_ops Tejun Heo
2026-04-24 17:27 ` [PATCH 10/17] sched_ext: Add cmask, a base-windowed bitmap over cid space Tejun Heo
2026-04-24 1:32 Tejun Heo
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b8bc70c1-8914-4a39-825d-e94cfb983d60@igalia.com \
--to=changwoo@igalia.com \
--cc=arighi@nvidia.com \
--cc=emil@etsalapatis.com \
--cc=linux-kernel@vger.kernel.org \
--cc=sched-ext@lists.linux.dev \
--cc=tj@kernel.org \
--cc=void@manifault.com \
--cc=yphbchou0911@gmail.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome