mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Steven Sistare <steven.sistare@oracle.com>
To: Mike Galbraith <efault@gmx.de>,
	Rohit Jain <rohit.k.jain@oracle.com>,
	linux-kernel@vger.kernel.org
Cc: peterz@infradead.org, mingo@redhat.com, joelaf@google.com,
	jbacik@fb.com, riel@redhat.com, juri.lelli@redhat.com,
	dhaval.giani@oracle.com
Subject: Re: [RFC 1/2] sched: reduce migration cost between faster caches for idle_balance
Date: Thu, 15 Feb 2018 11:35:34 -0500	[thread overview]
Message-ID: <dd972024-b3d6-fa53-7cc5-9e71a01c1837@oracle.com> (raw)
In-Reply-To: <1518244651.10229.66.camel@gmx.de>

On 2/10/2018 1:37 AM, Mike Galbraith wrote:
> On Fri, 2018-02-09 at 11:08 -0500, Steven Sistare wrote:
>>>> @@ -8804,7 +8803,8 @@ static int idle_balance(struct rq *this_rq, struct rq_flags *rf)
>>>>  		if (!(sd->flags & SD_LOAD_BALANCE))
>>>>  			continue;
>>>>  
>>>> -		if (this_rq->avg_idle < curr_cost + sd->max_newidle_lb_cost) {
>>>> +		if (this_rq->avg_idle < curr_cost + sd->max_newidle_lb_cost +
>>>> +		    sd->sched_migration_cost) {
>>>>  			update_next_balance(sd, &next_balance);
>>>>  			break;
>>>>  		}
>>>
>>> Ditto.
>>
>> The old code did not migrate if the expected costs exceeded the expected idle
>> time.  The new code just adds the sd-specific penalty (essentially loss of cache 
>> footprint) to the costs.  The for_each_domain loop visit smallest to largest
>> sd's, hence visiting smallest to largest migration costs (though the tunables do 
>> not enforce an ordering), and bails at the first sd where the total cost is a lose.
> 
> Hrm..
> 
> You're now adding a hypothetical cost to the measured cost of running
> the LB machinery, which implies that the measurement is insufficient,
> but you still don't say why it is insufficient.  What happens if you
> don't do that?  I ask, because when I removed the...
> 
>    this_rq->avg_idle < sysctl_sched_migration_cost
> 
> ...bits to check removal effect for Peter, the original reason for it
> being added did not re-materialize, making me wonder why you need to
> make this cutoff more aggressive.

The current code with sysctl_sched_migration_cost discourages migration
too much, per our test results.  Deleting it entirely from idle_balance()
may be the right solution, or it may allow too much migration and
cause regressions due to loss of cache warmth on some workloads.
Rohit's patch deletes it and adds the sd->sched_migration_cost term
to allow a migration rate that is somewhere in the middle, and is
logically sound.  It discourages but does not prevent migration between
nodes, and encourages but does not always allow migration between cores.
By contrast, setting relax_domain_level to disable SD_BALANCE_NEWIDLE
at the SD_NUMA level is a big hammer.

I would be perfectly happy if deleting sysctl_sched_migration_cost from
idle_balance does the trick.  Last week in a different thread you mentioned
it did not hurt tbench:

>> Mike, do you remember what comes apart when we take
>> out the sysctl_sched_migration_cost test in idle_balance()?
>
> Used to be anything scheduling cross-core heftily suffered, ie pretty
> much any localhost communication heavy load.  I just tried disabling it
> in 4.13 though (pre pti cliff), tried tbench, and it made zip squat
> difference.  I presume that's due to the meanwhile added
> this_rq->rd->overload and/or curr_cost checks.

Can you provide more details on the sysbench oltp test that motivated you
to add sysctl_sched_migration_cost to idle_balance, so Rohit can re-test it?
   1b9508f6 sched: Rate-limit newidle
   Rate limit newidle to migration_cost. It's a win for all stages of
   sysbench oltp tests.

Rohit is running more tests with a patch that deletes
sysctl_sched_migration_cost from idle_balance, and for his patch but
with the 5000 usec mistake corrected back to 500 usec.  So far both
give improvements over the baseline, but for different cases, so we
need to try more workloads before we draw any conclusions.

Rohit, can you share your data so far?

- Steve

  reply	other threads:[~2018-02-15 16:36 UTC|newest]

Thread overview: 18+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2018-02-08 22:19 [RFC 0/2] sched: Make idle_balance smarter about topology Rohit Jain
2018-02-08 22:19 ` [RFC 1/2] sched: reduce migration cost between faster caches for idle_balance Rohit Jain
2018-02-09  3:42   ` Mike Galbraith
2018-02-09 16:08     ` Steven Sistare
2018-02-10  6:37       ` Mike Galbraith
2018-02-15 16:35         ` Steven Sistare [this message]
2018-02-15 18:07           ` Mike Galbraith
2018-02-15 18:21             ` Steven Sistare
2018-02-15 18:39               ` Mike Galbraith
2018-02-15 18:07           ` Rohit Jain
2018-02-16  4:53             ` Mike Galbraith
2018-02-08 22:19 ` [RFC 2/2] Introduce sysctl(s) for the migration costs Rohit Jain
2018-02-09  3:54   ` Mike Galbraith
2018-02-09 16:10     ` Steven Sistare
2018-02-09 17:08       ` Mike Galbraith
2018-02-09 17:33         ` Steven Sistare
2018-02-09 17:50           ` Mike Galbraith
2018-02-12 15:28   ` Peter Zijlstra

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=dd972024-b3d6-fa53-7cc5-9e71a01c1837@oracle.com \
    --to=steven.sistare@oracle.com \
    --cc=dhaval.giani@oracle.com \
    --cc=efault@gmx.de \
    --cc=jbacik@fb.com \
    --cc=joelaf@google.com \
    --cc=juri.lelli@redhat.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mingo@redhat.com \
    --cc=peterz@infradead.org \
    --cc=riel@redhat.com \
    --cc=rohit.k.jain@oracle.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®