From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1755283AbcI0Nsy (ORCPT ); Tue, 27 Sep 2016 09:48:54 -0400 Received: from foss.arm.com ([217.140.101.70]:47960 "EHLO foss.arm.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1752865AbcI0Nse (ORCPT ); Tue, 27 Sep 2016 09:48:34 -0400 Subject: Re: [PATCH] sched/fair: Do not decay new task load on first enqueue To: Vincent Guittot , Matt Fleming References: <20160923115808.2330-1-matt@codeblueprint.co.uk> Cc: Peter Zijlstra , Ingo Molnar , linux-kernel , Mike Galbraith , Yuyang Du From: Dietmar Eggemann Message-ID: <1774604b-1028-a4b1-a254-b2fb32cd91f1@arm.com> Date: Tue, 27 Sep 2016 14:48:31 +0100 User-Agent: Mozilla/5.0 (X11; Linux x86_64; rv:45.0) Gecko/20100101 Thunderbird/45.2.0 MIME-Version: 1.0 In-Reply-To: Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 7bit Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On 23/09/16 15:30, Vincent Guittot wrote: > Hi Matt, > > On 23 September 2016 at 13:58, Matt Fleming wrote: >> Since commit 7dc603c9028e ("sched/fair: Fix PELT integrity for new >> tasks") ::last_update_time will be set to a non-zero value in >> post_init_entity_util_avg(), which leads to p->se.avg.load_avg being >> decayed on enqueue before the task has even had a chance to run. >> >> For a NICE_0 task the sequence of events leading up to this with >> example load average changes might be, >> >> sched_fork() >> init_entity_runnable_average() >> p->se.avg.load_avg = scale_load_down(se->load.weight); // 1024 >> >> wake_up_new_task() >> post_init_entity_util_avg() >> attach_entity_load_avg() >> p->se.last_update_time = cfs_rq->avg.last_update_time; >> >> activate_task() >> enqueue_task() >> ... >> enqueue_entity_load_avg() >> migrated = !sa->last_update_time // false >> if (!migrated) >> __update_load_avg() >> p->se.avg.load_avg = 1002 > > Does it mean that you can see the perf drop that you mention below > because load is decayed to 1002 instead of staying to 1024 ? I think Matt is talking about the fact that the cfs->runnable_load_avg value is 0 once the hackbench task is initially dequeued. Without this patch the value of se->avg.load_avg (e.g. both times 1002) is exactly the same when we add it to cfs_rq->runnable_load_avg in enqueue_entity_load_avg() and when we subtract it in dequeue_entity_load_avg(). That's because the initial runtime is short (~250us on my hikey board). With this patch we add 1024 and subtract ~1002 which lets cfs_rq->runnable_load_avg still have a small positive value. This favours that for the next hackbench task another cpu will be chosen in (load-based) fork-balance. > > 1002 mainly comes from period_contrib being set to 1023 during > init_entity_runnable_average so any delay longer than 1us between > attach_entity_load_avg and enqueue_entity_load_avg will trig the decay > of the load from 1024 to 1002 > [...]