From: Sebastian Siewior <bigeasy@linutronix.de>
To: Anna-Maria Behnsen <anna-maria@linutronix.de>
Cc: linux-kernel@vger.kernel.org,
Peter Zijlstra <peterz@infradead.org>,
John Stultz <jstultz@google.com>,
Thomas Gleixner <tglx@linutronix.de>,
Eric Dumazet <edumazet@google.com>,
"Rafael J . Wysocki" <rafael.j.wysocki@intel.com>,
Arjan van de Ven <arjan@infradead.org>,
"Paul E . McKenney" <paulmck@kernel.org>,
Frederic Weisbecker <frederic@kernel.org>,
Rik van Riel <riel@surriel.com>,
Steven Rostedt <rostedt@goodmis.org>,
Giovanni Gherdovich <ggherdovich@suse.cz>,
Lukasz Luba <lukasz.luba@arm.com>,
"Gautham R . Shenoy" <gautham.shenoy@amd.com>,
Srinivas Pandruvada <srinivas.pandruvada@intel.com>,
K Prateek Nayak <kprateek.nayak@amd.com>
Subject: Re: [PATCH v9 30/32] timers: Implement the hierarchical pull model
Date: Wed, 6 Dec 2023 17:35:36 +0100 [thread overview]
Message-ID: <20231206163536.r9DcrsWQ@linutronix.de> (raw)
In-Reply-To: <20231201092654.34614-31-anna-maria@linutronix.de>
On 2023-12-01 10:26:52 [+0100], Anna-Maria Behnsen wrote:
…
> As long as a CPU is busy it expires both local and global timers. When a
> CPU goes idle it arms for the first expiring local timer. If the first
> expiring pinned (local) timer is before the first expiring movable timer,
> then no action is required because the CPU will wake up before the first
> movable timer expires. If the first expiring movable timer is before the
> first expiring pinned (local) timer, then this timer is queued into a idle
an
> timerqueue and eventually expired by some other active CPU.
s/some other/another ?
…
>
> Signed-off-by: Anna-Maria Behnsen <anna-maria@linutronix.de>
> ---
> diff --git a/kernel/time/timer.c b/kernel/time/timer.c
> index b6c9ac0c3712..ac3e888d053f 100644
> --- a/kernel/time/timer.c
> +++ b/kernel/time/timer.c
> @@ -2103,6 +2104,64 @@ void timer_lock_remote_bases(unsigned int cpu)
…
> +static void timer_use_tmigr(unsigned long basej, u64 basem,
> + unsigned long *nextevt, bool *tick_stop_path,
> + bool timer_base_idle, struct timer_events *tevt)
> +{
> + u64 next_tmigr;
> +
> + if (timer_base_idle)
> + next_tmigr = tmigr_cpu_new_timer(tevt->global);
> + else if (tick_stop_path)
> + next_tmigr = tmigr_cpu_deactivate(tevt->global);
> + else
> + next_tmigr = tmigr_quick_check();
> +
> + /*
> + * If the CPU is the last going idle in timer migration hierarchy, make
> + * sure the CPU will wake up in time to handle remote timers.
> + * next_tmigr == KTIME_MAX if other CPUs are still active.
> + */
> + if (next_tmigr < tevt->local) {
> + u64 tmp;
> +
> + /* If we missed a tick already, force 0 delta */
> + if (next_tmigr < basem)
> + next_tmigr = basem;
> +
> + tmp = div_u64(next_tmigr - basem, TICK_NSEC);
Is this considered a hot path? Asking because u64 divs are nice if can
be avoided ;)
I guess the original value is from fetch_next_timer_interrupt(). But
then you only need it if the caller (__get_next_timer_interrupt()) has
the `idle' value set. Otherwise the operation is pointless.
Would it somehow work to replace
base_local->is_idle = time_after(nextevt, basej + 1);
with maybe something like
base_local->is_idle = tevt.local > basem + TICK_NSEC
If so you could avoid the `nextevt' maneuver.
> + *nextevt = basej + (unsigned long)tmp;
> + tevt->local = next_tmigr;
> + }
> +}
> +# else
…
> @@ -2132,6 +2190,21 @@ static inline u64 __get_next_timer_interrupt(unsigned long basej, u64 basem,
> nextevt = fetch_next_timer_interrupt(basej, basem, base_local,
> base_global, &tevt);
>
> + /*
> + * When the when the next event is only one jiffie ahead there is no
If the next event is only one jiffy ahead then there is no
> + * need to call timer migration hierarchy related
> + * functions. @tevt->global will be KTIME_MAX, nevertheless if the next
> + * timer is a global timer. This is also true, when the timer base is
The second sentence is hard to parse.
> + * idle.
> + *
> + * The proper timer migration hierarchy function depends on the callsite
> + * and whether timer base is idle or not. @nextevt will be updated when
> + * this CPU needs to handle the first timer migration hierarchy event.
> + */
> + if (time_after(nextevt, basej + 1))
> + timer_use_tmigr(basej, basem, &nextevt, idle,
> + base_local->is_idle, &tevt);
> +
> /*
> * We have a fresh next event. Check whether we can forward the
> * base.
> diff --git a/kernel/time/timer_migration.c b/kernel/time/timer_migration.c
> new file mode 100644
> index 000000000000..05cd8f1bc45d
> --- /dev/null
> +++ b/kernel/time/timer_migration.c
> @@ -0,0 +1,1636 @@
…
> +/*
> + * The timer migration mechanism is built on a hierarchy of groups. The
> + * lowest level group contains CPUs, the next level groups of CPU groups
> + * and so forth. The CPU groups are kept per node so for the normal case
> + * lock contention won't happen across nodes. Depending on the number of
> + * CPUs per node even the next level might be kept as groups of CPU groups
> + * per node and only the levels above cross the node topology.
> + *
> + * Example topology for a two node system with 24 CPUs each.
> + *
> + * LVL 2 [GRP2:0]
> + * GRP1:0 = GRP1:M
> + *
> + * LVL 1 [GRP1:0] [GRP1:1]
> + * GRP0:0 - GRP0:2 GRP0:3 - GRP0:5
> + *
> + * LVL 0 [GRP0:0] [GRP0:1] [GRP0:2] [GRP0:3] [GRP0:4] [GRP0:5]
> + * CPUS 0-7 8-15 16-23 24-31 32-39 40-47
In the CPUS list between 24-31 and 32-39 is a tab while the other
separators are spaces. Could you please align it with spaces? Judging
form the top you have tabstop=8 but here tabstop=4 looks "nice".
> + *
> + * The groups hold a timer queue of events sorted by expiry time. These
> + * queues are updated when CPUs go in idle. When they come out of idle
> + * ignore flag of events is set.
> + *
Sebastian
next prev parent reply other threads:[~2023-12-06 16:35 UTC|newest]
Thread overview: 86+ messages / expand[flat|nested] mbox.gz Atom feed top
2023-12-01 9:26 [PATCH v9 00/32] timers: Move from a push remote at enqueue to a pull at expiry model Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 01/32] tick-sched: Fix function names in comments Anna-Maria Behnsen
2023-12-20 13:09 ` Frederic Weisbecker
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 02/32] tick/sched: Cleanup confusing variables Anna-Maria Behnsen
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 03/32] tick-sched: Warn when next tick seems to be in the past Anna-Maria Behnsen
2023-12-20 13:27 ` Frederic Weisbecker
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 04/32] tracing/timers: Enhance timer_start tracepoint Anna-Maria Behnsen
2023-12-20 13:35 ` Frederic Weisbecker
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 05/32] tracing/timers: Add tracepoint for tracking timer base is_idle flag Anna-Maria Behnsen
2023-12-20 13:43 ` Frederic Weisbecker
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 06/32] timers: Do not IPI for deferrable timers Anna-Maria Behnsen
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 07/32] timers: Move store of next event into __next_timer_interrupt() Anna-Maria Behnsen
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 08/32] timers: Clarify check in forward_timer_base() Anna-Maria Behnsen
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 09/32] timers: Split out forward timer base functionality Anna-Maria Behnsen
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 10/32] timers: Use already existing function for forwarding timer base Anna-Maria Behnsen
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 11/32] timers: Rework idle logic Anna-Maria Behnsen
2023-12-20 14:00 ` Frederic Weisbecker
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Thomas Gleixner
2023-12-01 9:26 ` [PATCH v9 12/32] timers: Fix nextevt calculation when no timers are pending Anna-Maria Behnsen
2023-12-04 16:03 ` Sebastian Siewior
2023-12-05 11:53 ` Anna-Maria Behnsen
2023-12-10 0:35 ` Frederic Weisbecker
2023-12-12 13:21 ` Anna-Maria Behnsen
2023-12-12 13:37 ` Frederic Weisbecker
2023-12-20 14:49 ` Frederic Weisbecker
2023-12-20 15:59 ` [tip: timers/core] " tip-bot2 for Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 13/32] timers: Restructure get_next_timer_interrupt() Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 14/32] timers: Split out get next timer interrupt Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 15/32] timers: Move marking timer bases idle into tick_nohz_stop_tick() Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 16/32] timers: Optimization for timer_base_try_to_set_idle() Anna-Maria Behnsen
2023-12-04 17:52 ` Sebastian Siewior
2023-12-05 12:05 ` Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 17/32] timers: Introduce add_timer() variants which modify timer flags Anna-Maria Behnsen
2023-12-05 18:28 ` Sebastian Siewior
2023-12-06 9:24 ` Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 18/32] workqueue: Use global variant for add_timer() Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 19/32] timers: add_timer_on(): Make sure TIMER_PINNED flag is set Anna-Maria Behnsen
2023-12-05 18:29 ` Sebastian Siewior
2023-12-06 9:57 ` Anna-Maria Behnsen
2023-12-06 10:26 ` Sebastian Siewior
2023-12-06 10:46 ` Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 20/32] timers: Ease code in run_local_timers() Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 21/32] timers: Split next timer interrupt logic Anna-Maria Behnsen
2023-12-05 18:29 ` Sebastian Siewior
2023-12-01 9:26 ` [PATCH v9 22/32] timers: Keep the pinned timers separate from the others Anna-Maria Behnsen
2023-12-05 21:11 ` Sebastian Siewior
2023-12-06 10:23 ` Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 23/32] timers: Retrieve next expiry of pinned/non-pinned timers separately Anna-Maria Behnsen
2023-12-06 9:47 ` Sebastian Siewior
2023-12-07 10:12 ` Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 24/32] timers: Split out "get next timer interrupt" functionality Anna-Maria Behnsen
2023-12-06 10:20 ` Sebastian Siewior
2023-12-01 9:26 ` [PATCH v9 25/32] timers: Add get next timer interrupt functionality for remote CPUs Anna-Maria Behnsen
2023-12-06 10:44 ` Sebastian Siewior
2023-12-07 10:27 ` Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 26/32] timers: Restructure internal locking Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 27/32] timers: Check if timers base is handled already Anna-Maria Behnsen
2023-12-06 10:58 ` Sebastian Siewior
2023-12-01 9:26 ` [PATCH v9 28/32] tick/sched: Split out jiffies update helper function Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 29/32] timers: Introduce function to check timer base is_idle flag Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 30/32] timers: Implement the hierarchical pull model Anna-Maria Behnsen
2023-12-06 16:35 ` Sebastian Siewior [this message]
2023-12-08 9:01 ` Anna-Maria Behnsen
2023-12-07 18:09 ` Sebastian Siewior
2023-12-08 10:31 ` Anna-Maria Behnsen
2023-12-08 18:18 ` Sebastian Siewior
2023-12-11 18:04 ` Sebastian Siewior
2023-12-12 11:31 ` Anna-Maria Behnsen
2023-12-12 11:43 ` Anna-Maria Behnsen
2023-12-12 15:59 ` Sebastian Siewior
2023-12-12 12:14 ` Sebastian Siewior
2023-12-12 14:52 ` Anna-Maria Behnsen
2023-12-12 17:08 ` Sebastian Siewior
2023-12-01 9:26 ` [PATCH v9 31/32] timer_migration: Add tracepoints Anna-Maria Behnsen
2023-12-01 9:26 ` [PATCH v9 32/32] timers: Always queue timers on the local CPU Anna-Maria Behnsen
2023-12-07 12:11 ` [PATCH v9 00/32] timers: Move from a push remote at enqueue to a pull at expiry model Anna-Maria Behnsen
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20231206163536.r9DcrsWQ@linutronix.de \
--to=bigeasy@linutronix.de \
--cc=anna-maria@linutronix.de \
--cc=arjan@infradead.org \
--cc=edumazet@google.com \
--cc=frederic@kernel.org \
--cc=gautham.shenoy@amd.com \
--cc=ggherdovich@suse.cz \
--cc=jstultz@google.com \
--cc=kprateek.nayak@amd.com \
--cc=linux-kernel@vger.kernel.org \
--cc=lukasz.luba@arm.com \
--cc=paulmck@kernel.org \
--cc=peterz@infradead.org \
--cc=rafael.j.wysocki@intel.com \
--cc=riel@surriel.com \
--cc=rostedt@goodmis.org \
--cc=srinivas.pandruvada@intel.com \
--cc=tglx@linutronix.de \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®