mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Vaidyanathan Srinivasan <svaidy@linux.vnet.ibm.com>
To: Len Brown <lenb@kernel.org>
Cc: Trinabh Gupta <trinabh@linux.vnet.ibm.com>,
	Arjan van de Ven <arjan@linux.intel.com>,
	peterz@infradead.org, suresh.b.siddha@intel.com,
	benh@kernel.crashing.org, venki@google.com,
	Andi Kleen <ak@linux.intel.com>,
	linux-kernel@vger.kernel.org
Subject: Re: [RFC PATCH V1 1/2] cpuidle: Data structure changes for global cpuidle device
Date: Fri, 25 Mar 2011 23:18:31 +0530	[thread overview]
Message-ID: <20110325174831.GB19214@dirshya.in.ibm.com> (raw)
In-Reply-To: <alpine.LFD.2.02.1103250330470.32565@x980>

* Len Brown <lenb@kernel.org> [2011-03-25 04:12:03]:

> I agree it is silly to allocate a cpuidle_device
> for every cpu in the system as we do today.
> 
> Yes, splitting the counters out of cpuidle_device
> is a necessary part of fixing that.
> 
> However, cpuidle_device.cpuidle_state[] is currently not per-driver,
> it is per-cpu, and it is writable.
> 
> In particular, the cpuidle_device->prepare() mechanism
> causes updates to the cpuidle_state[].flags,
> setting and clearing CPUIDLE_FLAG_IGNORE to
> tell the governor not to chose a state
> on a per-cpu basis at run-time.
> 
> I don't like that mechanism.
> I'd like to see it replaced, and when replaced,
> cpuidle_state[] can be per system-wide driver.

Thanks for the detailed review.  I agree that we should rework
handling of the cpuidle_state[].flags.  However, is the prepare()
mechanism used at all?  Can we remove the option completely?

> I think the real problem that prepare() was trying to solve
> is that the driver today does not have the ability to over-rule
> the choice made by the governor.  The driver may discover
> in the course of trying to satisfy the request of the governor
> that it needs to demote to a shallower state; or it may
> do its best to satisfy the governor's request, and the hardware
> may demote its request to a shallower state.
> 
> Unfortunately, when this happens, the driver dutifully
> returns the time spent in the state to cpuidle_idle_call(),
> who then updates the wrong last_residency, time, and usage counters.

I did not get this scenario.  Are you saying 

target_state->enter(dev, target_state) can enter a different state
than the one suggested by target_state?  

I understand the hardware demotion part, but can we really detect the
target 'demoted' state in that case?  I guess not.

> Sure is ironic for the driver to allocate the data structures and
> then hand the timer to the uppper layer, just to have the upper layer
> update the wrong data structures...
> 
> Surely the driver enter routine should update the counters
> that the driver was obligated to allocate, and it should return
> the state actually entered (for tracing), rather than the time spent
> there.

Can we do something like this:

last_state = target_state->enter(dev, target_state)

dev->last_state and dev->last_residency are updated inside
target_state->enter() 

The returned last_state is just for tracing, actual data is already
updated in the cpuidle_dev structure and used for sysfs display.

> The generic cpuidle code should simply handle where the counters live
> in the sysfs namespace, not updating the counters.
> This needs to be addressed before cpuidle_device.cpuidle_state[]
> can be made one/system.

Agreed.

Thanks again for the recommendations.

--Vaidy


  reply	other threads:[~2011-03-25 17:49 UTC|newest]

Thread overview: 8+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2011-03-22 12:47 [RFC PATCH V1 0/2] cpuidle: global registration of idle states with per-cpu statistics Trinabh Gupta
2011-03-22 12:48 ` [RFC PATCH V1 1/2] cpuidle: Data structure changes for global cpuidle device Trinabh Gupta
2011-03-25  8:12   ` Len Brown
2011-03-25 17:48     ` Vaidyanathan Srinivasan [this message]
2011-04-02  0:03       ` Len Brown
2011-03-22 12:48 ` [RFC PATCH V1 2/2] cpuidle: API changes in callers using new cpuidle_state_stats Trinabh Gupta
2011-03-25  7:28 ` [RFC PATCH V1 0/2] cpuidle: global registration of idle states with per-cpu statistics Len Brown
2011-03-25 17:15   ` Vaidyanathan Srinivasan

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20110325174831.GB19214@dirshya.in.ibm.com \
    --to=svaidy@linux.vnet.ibm.com \
    --cc=ak@linux.intel.com \
    --cc=arjan@linux.intel.com \
    --cc=benh@kernel.crashing.org \
    --cc=lenb@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=peterz@infradead.org \
    --cc=suresh.b.siddha@intel.com \
    --cc=trinabh@linux.vnet.ibm.com \
    --cc=venki@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®