mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Changwoo Min <changwoo@igalia.com>
To: Tejun Heo <tj@kernel.org>, David Vernet <void@manifault.com>,
	Andrea Righi <arighi@nvidia.com>
Cc: sched-ext@lists.linux.dev, Emil Tsalapatis <emil@etsalapatis.com>,
	linux-kernel@vger.kernel.org,
	Cheng-Yang Chou <yphbchou0911@gmail.com>
Subject: Re: [PATCH 10/17] sched_ext: Add cmask, a base-windowed bitmap over cid space
Date: Wed, 29 Apr 2026 21:47:15 +0900	[thread overview]
Message-ID: <b8bc70c1-8914-4a39-825d-e94cfb983d60@igalia.com> (raw)
In-Reply-To: <20260428203545.181052-11-tj@kernel.org>


On 4/29/26 5:35 AM, Tejun Heo wrote:

> +/*
> + * x86 BPF JIT rejects BPF_OR | BPF_FETCH and BPF_AND | BPF_FETCH on arena
> + * pointers (see bpf_jit_supports_insn() in arch/x86/net/bpf_jit_comp.c). Only
> + * BPF_CMPXCHG / BPF_XCHG / BPF_ADD with FETCH are allowed. Implement
> + * test_and_{set,clear} and the atomic set/clear via a cmpxchg loop.
> + *
> + * CMASK_CAS_TRIES is far above what any non-pathological contention needs.
> + * Exhausting it means the bit update was lost, which corrupts the caller's view
> + * of the bitmap, so raise scx_bpf_error() to abort the scheduler.
> + */
> +#define CMASK_CAS_TRIES		1024
> +
> +static __always_inline void cmask_set(struct scx_cmask __arena *m, u32 cid)
> +{
> +	u64 __arena *w;
> +	u64 bit, old, new;
> +	u32 i;
> +
> +	if (!__cmask_contains(m, cid))
> +		return;
> +	w = __cmask_word(m, cid);
> +	bit = BIT_U64(cid & 63);
> +	bpf_for(i, 0, CMASK_CAS_TRIES) {
> +		old = *w;
> +		if (old & bit)
> +			return;
> +		new = old | bit;
> +		if (__sync_val_compare_and_swap(w, old, new) == old)
> +			return;
> +	}
> +	scx_bpf_error("cmask_set CAS exhausted at cid %u", cid);
> +}
> +
> +static __always_inline void cmask_clear(struct scx_cmask __arena *m, u32 cid)
> +{
> +	u64 __arena *w;
> +	u64 bit, old, new;
> +	u32 i;
> +
> +	if (!__cmask_contains(m, cid))
> +		return;
> +	w = __cmask_word(m, cid);
> +	bit = BIT_U64(cid & 63);
> +	bpf_for(i, 0, CMASK_CAS_TRIES) {
> +		old = *w;
> +		if (!(old & bit))
> +			return;
> +		new = old & ~bit;
> +		if (__sync_val_compare_and_swap(w, old, new) == old)
> +			return;
> +	}
> +	scx_bpf_error("cmask_clear CAS exhausted at cid %u", cid);
> +}
> +
> +static __always_inline bool cmask_test_and_set(struct scx_cmask __arena *m, u32 cid)
> +{
> +	u64 __arena *w;
> +	u64 bit, old, new;
> +	u32 i;
> +
> +	if (!__cmask_contains(m, cid))
> +		return false;
> +	w = __cmask_word(m, cid);
> +	bit = BIT_U64(cid & 63);
> +	bpf_for(i, 0, CMASK_CAS_TRIES) {
> +		old = *w;
> +		if (old & bit)
> +			return true;
> +		new = old | bit;
> +		if (__sync_val_compare_and_swap(w, old, new) == old)
> +			return false;
> +	}
> +	scx_bpf_error("cmask_test_and_set CAS exhausted at cid %u", cid);
> +	return false;
> +}
> +
> +static __always_inline bool cmask_test_and_clear(struct scx_cmask __arena *m, u32 cid)
> +{
> +	u64 __arena *w;
> +	u64 bit, old, new;
> +	u32 i;
> +
> +	if (!__cmask_contains(m, cid))
> +		return false;
> +	w = __cmask_word(m, cid);
> +	bit = BIT_U64(cid & 63);
> +	bpf_for(i, 0, CMASK_CAS_TRIES) {
> +		old = *w;
> +		if (!(old & bit))
> +			return false;
> +		new = old & ~bit;
> +		if (__sync_val_compare_and_swap(w, old, new) == old)
> +			return true;
> +	}
> +	scx_bpf_error("cmask_test_and_clear CAS exhausted at cid %u", cid);
> +	return false;
> +}

Exiting a BPF scheduler when CAS retries fail seems too brutal, while it
is extremely rare to happen. What about adding a kfunc for the slow path
that runs only when the CAS retry fails? For example,
scx_bpf_cmask_test_and_clear( ) does the same thing as
cmask_test_and_clear(), but it never fails. cmask_test_and_clear() calls
scx_bpf_cmask_test_and_clear( ) only when the CAS retry fails. In this
way, we can use the fast path in the BPF implementation while ensuring
the operation succeeds.

> +/**
> + * cmask_next_set - find the first set bit at or after @cid
> + * @m: cmask to search
> + * @cid: starting cid (clamped to @m->base if below)
> + *
> + * Returns the smallest set cid in [@cid, @m->base + @m->nr_bits), or
> + * @m->base + @m->nr_bits if none (the out-of-range sentinel matches the
> + * termination condition used by cmask_for_each()).
> + */
> +static __always_inline u32 cmask_next_set(const struct scx_cmask __arena *m, u32 cid)
> +{
> +	u32 end = m->base + m->nr_bits;
> +	u32 base = m->base / 64;
> +	u32 last_wi = (end - 1) / 64 - base;
> +	u32 start_wi, start_bit, i;
> +
> +	if (cid < m->base)
> +		cid = m->base;
> +	if (cid >= end)
> +		return end;
> +
> +	start_wi = cid / 64 - base;
> +	start_bit = cid & 63;
> +
> +	bpf_for(i, 0, CMASK_MAX_WORDS) {
> +		u32 wi = start_wi + i;
> +		u64 word;
> +		u32 found;
> +
> +		if (wi > last_wi)
> +			break;
> +
> +		word = m->bits[wi];
> +		if (i == 0)
> +			word &= GENMASK_U64(63, start_bit);
> +		if (!word)
> +			continue;
> +
> +		found = (base + wi) * 64 + __builtin_ctzll(word);

Some compiler versions (e.g., clang-18 or older) don’t support
__builtin_ctzll(). To handle this gracefully, there is already a
wrapper, ctzll(), in common.bpf.h. So, I suggest using ctzll()
for compatibility.

Reviewed-by: Changwoo Min <changwoo@igalia.com>

  reply	other threads:[~2026-04-29 12:47 UTC|newest]

Thread overview: 33+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-04-28 20:35 [PATCHSET v3 sched_ext/for-7.2] sched_ext: Topological CPU IDs and cid-form struct_ops Tejun Heo
2026-04-28 20:35 ` [PATCH 01/17] sched_ext: Add ext_types.h for early subsystem-wide defs Tejun Heo
2026-04-28 20:35 ` [PATCH 02/17] sched_ext: Rename ops_cpu_valid() to scx_cpu_valid() and expose it Tejun Heo
2026-04-28 20:35 ` [PATCH 03/17] sched_ext: Move scx_exit(), scx_error() and friends to ext_internal.h Tejun Heo
2026-04-28 20:35 ` [PATCH 04/17] sched_ext: Shift scx_kick_cpu() validity check to scx_bpf_kick_cpu() Tejun Heo
2026-04-28 20:35 ` [PATCH 05/17] sched_ext: Relocate cpu_acquire/cpu_release to end of struct sched_ext_ops Tejun Heo
2026-04-28 20:35 ` [PATCH 06/17] sched_ext: Make scx_enable() take scx_enable_cmd Tejun Heo
2026-04-28 20:35 ` [PATCH 07/17] sched_ext: Add topological CPU IDs (cids) Tejun Heo
2026-04-28 20:35 ` [PATCH 08/17] sched_ext: Add scx_bpf_cid_override() kfunc Tejun Heo
2026-04-29 14:07   ` Andrea Righi
2026-04-29 17:06     ` Tejun Heo
2026-04-29 17:20       ` Andrea Righi
2026-04-28 20:35 ` [PATCH 09/17] tools/sched_ext: Add struct_size() helpers to common.bpf.h Tejun Heo
2026-04-28 20:35 ` [PATCH 10/17] sched_ext: Add cmask, a base-windowed bitmap over cid space Tejun Heo
2026-04-29 12:47   ` Changwoo Min [this message]
2026-04-29 17:16     ` Tejun Heo
2026-04-28 20:35 ` [PATCH 11/17] sched_ext: Add cid-form kfunc wrappers alongside cpu-form Tejun Heo
2026-04-28 20:35 ` [PATCH 12/17] sched_ext: Add bpf_sched_ext_ops_cid struct_ops type Tejun Heo
2026-04-28 20:35 ` [PATCH 13/17] sched_ext: Forbid cpu-form kfuncs from cid-form schedulers Tejun Heo
2026-04-28 20:35 ` [PATCH 14/17] tools/sched_ext: scx_qmap: Restart on hotplug instead of cpu_online/offline Tejun Heo
2026-04-28 20:35 ` [PATCH 15/17] tools/sched_ext: scx_qmap: Add cmask-based idle tracking and cid-based idle pick Tejun Heo
2026-04-28 20:35 ` [PATCH 16/17] tools/sched_ext: scx_qmap: Port to cid-form struct_ops Tejun Heo
2026-04-29 12:47   ` Changwoo Min
2026-04-29 13:53     ` Andrea Righi
2026-04-29 16:42       ` Tejun Heo
2026-04-28 20:35 ` [PATCH 17/17] sched_ext: Require cid-form struct_ops for sub-sched support Tejun Heo
2026-04-29 12:49 ` [PATCHSET v3 sched_ext/for-7.2] sched_ext: Topological CPU IDs and cid-form struct_ops Changwoo Min
2026-04-29 13:29 ` Andrea Righi
2026-04-29 14:11   ` Andrea Righi
2026-04-29 17:06   ` Tejun Heo
  -- strict thread matches above, loose matches on Subject: below --
2026-04-29 18:21 [PATCHSET v4 " Tejun Heo
2026-04-29 18:21 ` [PATCH 10/17] sched_ext: Add cmask, a base-windowed bitmap over cid space Tejun Heo
2026-04-24 17:27 [PATCHSET v2 REPOST sched_ext/for-7.2] sched_ext: Topological CPU IDs and cid-form struct_ops Tejun Heo
2026-04-24 17:27 ` [PATCH 10/17] sched_ext: Add cmask, a base-windowed bitmap over cid space Tejun Heo
2026-04-24  1:32 Tejun Heo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=b8bc70c1-8914-4a39-825d-e94cfb983d60@igalia.com \
    --to=changwoo@igalia.com \
    --cc=arighi@nvidia.com \
    --cc=emil@etsalapatis.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=tj@kernel.org \
    --cc=void@manifault.com \
    --cc=yphbchou0911@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome