mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Julia Lawall <julia.lawall@inria.fr>
To: Peter Zijlstra <peterz@infradead.org>
Cc: Ingo Molnar <mingo@redhat.com>,
	 Vincent Guittot <vincent.guittot@linaro.org>,
	 Dietmar Eggemann <dietmar.eggemann@arm.com>,
	Mel Gorman <mgorman@suse.de>,
	 linux-kernel@vger.kernel.org
Subject: Re: EEVDF and NUMA balancing
Date: Mon, 18 Dec 2023 14:58:47 +0100 (CET)	[thread overview]
Message-ID: <b8ab29de-1775-46e-dd75-cdf98be8b0@inria.fr> (raw)
In-Reply-To: <20231009102949.GC14330@noisy.programming.kicks-ass.net>

Hello,

I have looked further into the NUMA balancing issue.

The context is that there are 2N threads running on 2N cores, one thread
gets NUMA balanced to the other socket, leaving N+1 threads on one socket
and N-1 threads on the other socket.  This condition typically persists
for one or more seconds.

Previously, I reported this on a 4-socket machine, but it can also occur
on a 2-socket machine, with other tests from the NAS benchmark suite
(sp.B, bt.B, etc).

Since there are N+1 threads on one of the sockets, it would seem that load
balancing would quickly kick in to bring some thread back to socket with
only N-1 threads.  This doesn't happen, though, because actually most of
the threads have some NUMA effects such that they have a preferred node.
So there is a high chance that an attempt to steal will fail, because both
threads have a preference for the socket.

At this point, the only hope is active balancing.  However, triggering
active balancing requires the success of the following condition in
imbalanced_active_balance:

        if ((env->migration_type == migrate_task) &&
            (sd->nr_balance_failed > sd->cache_nice_tries+2))

sd->nr_balance_failed does not increase because the core is idle.  When a
core is idle, it comes to the load_balance function from schedule() though
newidle_balance.  newidle_balance always sends in the flag CPU_NEWLY_IDLE,
even if the core has been idle for a long time.

Changing newidle_balance to use CPU_IDLE rather than CPU_NEWLY_IDLE when
the core was already idle before the call to schedule() is not enough
though, because there is also the constraint on the migration type.  That
turns out to be (mostly?) migrate_util.  Removing the following
code from find_busiest_queue:

                        /*
                         * Don't try to pull utilization from a CPU with one
                         * running task. Whatever its utilization, we will fail
                         * detach the task.
                         */
                        if (nr_running <= 1)
                                continue;

and changing the above test to:

        if ((env->migration_type == migrate_task || env->migration_type == migrate_util) &&
            (sd->nr_balance_failed > sd->cache_nice_tries+2))

seems to solve the problem.

I will test this on more applications.  But let me know if the above
solution seems completely inappropriate.  Maybe it violates some other
constraints.

I have no idea why this problem became more visible with EEVDF.  It seems
to have to do with the time slices all turning out to be the same.  I got
the same behavior in 6.5 by overwriting the timeslice calculation to
always return 1.  But I don't see the connection between the timeslice and
the behavior of the idle task.

thanks,
julia

  parent reply	other threads:[~2023-12-18 13:58 UTC|newest]

Thread overview: 46+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2023-10-03 20:25 Julia Lawall
2023-10-03 21:51 ` Peter Zijlstra
2023-10-04 12:01   ` Julia Lawall
2023-10-04 12:05     ` Peter Zijlstra
2023-10-04 16:24       ` Julia Lawall
2023-10-04 17:48         ` Peter Zijlstra
2023-10-04 18:04           ` Julia Lawall
2023-10-09 10:29             ` Peter Zijlstra
2023-10-09 14:07               ` Julia Lawall
2023-11-11 12:56               ` Julia Lawall
2023-12-18 13:58               ` Julia Lawall [this message]
2023-12-18 17:18                 ` Vincent Guittot
2023-12-18 22:31                   ` Julia Lawall
2023-12-19 17:38                     ` Vincent Guittot
2023-12-19 17:51                       ` Julia Lawall
2023-12-20 17:09                         ` Vincent Guittot
2023-12-21 18:20                           ` Julia Lawall
2023-12-22 14:55                             ` Vincent Guittot
2023-12-22 15:00                               ` Julia Lawall
2023-12-22 15:59                                 ` Vincent Guittot
2023-12-22 16:18                                   ` Julia Lawall
2023-12-22 16:29                                     ` Julia Lawall
2023-12-22 16:42                                       ` Vincent Guittot
2023-12-28 18:34                                         ` Julia Lawall
2023-12-29 15:18                                           ` Julia Lawall
2024-01-04 16:26                                             ` Vincent Guittot
2024-01-04 16:45                                               ` Julia Lawall
2024-01-05 14:51                                               ` Julia Lawall
2024-01-05 16:00                                                 ` Vincent Guittot
2024-01-05 16:39                                                   ` Julia Lawall
2024-01-05 17:27                                                     ` Julia Lawall
2024-01-18 16:35                                                       ` Vincent Guittot
2024-01-18 16:50                                                         ` Julia Lawall
2024-01-18 17:10                                                           ` Vincent Guittot
2024-01-18 17:43                                                             ` Julia Lawall
2024-01-18 22:13                                                         ` Julia Lawall
2024-01-19 11:26                                                           ` Vincent Guittot
2024-01-19 11:33                                                             ` Julia Lawall
2024-01-26 21:20                                                             ` Julia Lawall
2024-03-10  9:39                                                             ` Julia Lawall
2024-01-05 20:45                                                   ` Julia Lawall
2023-12-20 16:39                       ` Julia Lawall
2023-12-20 17:11                         ` Vincent Guittot
2023-10-04 18:15         ` Ingo Molnar
2023-10-04 18:20           ` Julia Lawall
2023-10-04 19:48           ` Julia Lawall

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=b8ab29de-1775-46e-dd75-cdf98be8b0@inria.fr \
    --to=julia.lawall@inria.fr \
    --cc=dietmar.eggemann@arm.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mgorman@suse.de \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=vincent.guittot@linaro.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®