mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Tejun Heo <tj@kernel.org>
To: Tao Cui <cui.tao@linux.dev>
Cc: Tao Cui <cuitao@kylinos.cn>,
	void@manifault.com, arighi@nvidia.com, changwoo@igalia.com,
	suzhidao@xiaomi.com, yphbchou0911@gmail.com,
	zhaomengmeng@kylinos.cn, sched-ext@lists.linux.dev,
	linux-kernel@vger.kernel.org, bpf@vger.kernel.org
Subject: Re: [PATCH] sched_ext/scx_flatcg: expire cached hweights on weight changes
Date: Fri, 14 Aug 2026 12:23:15 -1000	[thread overview]
Message-ID: <3432f79f1d8195a56a0b4fea53ceb099@kernel.org> (raw)
In-Reply-To: <80d78aac-7fc8-4c14-a2ae-419297d953f7@linux.dev>

Hello,

On Fri, Aug 14, 2026 at 10:48:35PM +0800, Tao Cui wrote:
> Is this the intended behavior of the budget clamping, or should the
> steady-state distribution converge to the compounded shares over time?
> The header comment's model doesn't seem to hold in this scenario.

The clamp isn't the culprit. I reproduced your setup (3 busy tasks per
leaf, 4 CPUs) and dug in.

The dominant factor is that D is runnable-task limited. At weight 800
its compounded share is 3.55 CPUs but it only has 3 tasks, so 75% is
the best it can do. Below that cap, the window granting loses more:
whenever all three of D's tasks are already running and another CPU
picks D, the pop from D's DSQ comes up empty, the cgv_node gets
stashed and that CPU grants a full cgrp_slice_ns window to another
cgroup. The loss scales with the window size. Measured D shares at
weight 800:

	3 tasks/leaf			55-59%
	3 tasks/leaf, 100ms slices	52%
	5 tasks/leaf			84-86%

Disabling the clamp entirely is the 59% above, so it barely matters.
With enough runnable tasks per cgroup, the distribution converges to
the documented compounding. The model holds, but only when no cgroup
is runnable-task limited, and the window granting degrades sooner than
per-task fair queueing would as that limit is approached. That's an
inherent simplification of this example scheduler.

Separately, while digging into this, I found that the true-up in
fcg_dispatch() is broken:

	__sync_fetch_and_add(&cgc->cvtime_delta,
			     (cpuc->cur_at + cgrp_slice_ns - now) *
			     FCG_HWEIGHT_ONE / (cgc->hweight ?: 1));

In the CNS_EXPIRE case, now is past cur_at + cgrp_slice_ns, so the u64
subexpression wraps. The multiplication preserves the two's complement
encoding but the unsigned division by hweight destroys it, adding about
2^64 / hweight per expiry instead of a small correction. The sign is
also inverted. The true-up should be actual minus charged, so the
expiry overrun should be added and the CNS_EMPTY unused portion
subtracted. Under saturation the budget clamp mostly masks the garbage,
which is why the numbers above barely move with it fixed (the 86% in
the 5 tasks/leaf row), but the accounting is broken all the same. The
following fixes it and tests fine (BPF division is unsigned, keep the
dividends positive):

	s64 delta = now - cpuc->cur_at - cgrp_slice_ns;

	if (delta >= 0)
		__sync_fetch_and_add(&cgc->cvtime_delta,
				     (u64)delta * FCG_HWEIGHT_ONE /
				     (cgc->hweight ?: 1));
	else
		__sync_fetch_and_sub(&cgc->cvtime_delta,
				     (u64)-delta * FCG_HWEIGHT_ONE /
				     (cgc->hweight ?: 1));

Care to send a patch?

Thanks.

--
tejun

  reply	other threads:[~2026-08-14 22:23 UTC|newest]

Thread overview: 6+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-08-14 14:41 Tao Cui
2026-08-14 14:48 ` Tao Cui
2026-08-14 22:23   ` Tejun Heo [this message]
2026-08-15 11:55     ` Tao Cui
2026-08-14 17:20 ` bot+bpf-ci
2026-08-14 20:56 ` Tejun Heo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=3432f79f1d8195a56a0b4fea53ceb099@kernel.org \
    --to=tj@kernel.org \
    --cc=arighi@nvidia.com \
    --cc=bpf@vger.kernel.org \
    --cc=changwoo@igalia.com \
    --cc=cui.tao@linux.dev \
    --cc=cuitao@kylinos.cn \
    --cc=linux-kernel@vger.kernel.org \
    --cc=sched-ext@lists.linux.dev \
    --cc=suzhidao@xiaomi.com \
    --cc=void@manifault.com \
    --cc=yphbchou0911@gmail.com \
    --cc=zhaomengmeng@kylinos.cn \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®