mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Tejun Heo <tj@kernel.org>
To: Glauber Costa <glommer@openvz.org>
Cc: Peter Zijlstra <a.p.zijlstra@chello.nl>,
	Paul Turner <pjt@google.com>,
	linux-kernel@vger.kernel.org, cgroups@vger.kernel.org,
	Frederic Weisbecker <fweisbec@gmail.com>,
	devel@openvz.org
Subject: Re: [PATCH v7 06/11] sched: document the cpu cgroup.
Date: Thu, 6 Jun 2013 16:28:03 -0700	[thread overview]
Message-ID: <20130606232803.GP5045@htj.dyndns.org> (raw)
In-Reply-To: <1369825402-31046-7-git-send-email-glommer@openvz.org>

Hello, Glauber.

On Wed, May 29, 2013 at 03:03:17PM +0400, Glauber Costa wrote:
> The CPU cgroup is so far, undocumented. Although data exists in the
> Documentation directory about its functioning, it is usually spread,
> and/or presented in the context of something else. This file
> consolidates all cgroup-related information about it.
> 
> Signed-off-by: Glauber Costa <glommer@openvz.org>

Reviewed-by: Tejun Heo <tj@kernel.org>

Some minor points below.

> +Files
> +-----
> +
> +The CPU controller exposes the following files to the user:
> +
> + - cpu.shares: The weight of each group living in the same hierarchy, that
> + translates into the amount of CPU it is expected to get. Upon cgroup creation,
> + each group gets assigned a default of 1024. The percentage of CPU assigned to
> + the cgroup is the value of shares divided by the sum of all shares in all
> + cgroups in the same level.
> +
> + - cpu.cfs_period_us: The duration in microseconds of each scheduler period, for
> + bandwidth decisions. This defaults to 100000us or 100ms. Larger periods will
> + improve throughput at the expense of latency, since the scheduler will be able
> + to sustain a cpu-bound workload for longer. The opposite of true for smaller
                                                             ^
                                                             is?
> + periods. Note that this only affects non-RT tasks that are scheduled by the
> + CFS scheduler.
> +
> +- cpu.cfs_quota_us: The maximum time in microseconds during each cfs_period_us
> +  in for the current group will be allowed to run. For instance, if it is set to
    ^^^^^^^
    in for? doesn't parse for me.

> +  half of cpu_period_us, the cgroup will only be able to peak run for 50 % of
                                                          ^^^^^^^^^
							  to run at maximum?

> +  the time. One should note that this represents aggregate time over all CPUs
> +  in the system. Therefore, in order to allow full usage of two CPUs, for
> +  instance, one should set this value to twice the value of cfs_period_us.
> +
> +- cpu.stat: statistics about the bandwidth controls. No data will be presented
> +  if cpu.cfs_quota_us is not set. The file presents three

 Unnecessary line break?

> +  numbers:
> +	nr_periods: how many full periods have been elapsed.
> +	nr_throttled: number of times we exausted the full allowed bandwidth
> +	throttled_time: total time the tasks were not run due to being overquota
> +
> + - cpu.rt_runtime_us and cpu.rt_period_us: Those files are the RT-tasks
                                              ^^^^^
					      these

> +   analogous to the CFS files cfs_quota_us and cfs_period_us. One important
      ^^^^^^^^^^^^
      counterparts of?

> +   difference, though, is that while the cfs quotas are upper bounds that
> +   won't necessarily be met, the rt runtimes form a stricter guarantee.
                                       ^^^^^^^^^^^^^
				       runtimes are strict guarantees?

> +   Therefore, no overlap is allowed. Implications of that are that given a
                    ^^^^^^^
		 maybe overcommit is a better term?

> +   hierarchy with multiple children, the sum of all rt_runtime_us may not exceed
> +   the runtime of the parent. Also, a rt_runtime_us of 0, means that no rt tasks
                                                           ^
							   prolly unnecessary

> +   can ever be run in this cgroup. For more information about rt tasks runtime
> +   assignments, see scheduler/sched-rt-group.txt
      ^^^^^^^^^^^
      configuration?

> +
> + - cpuacct.usage: The aggregate CPU time, in nanoseconds, consumed by all tasks
> +   in this group.
> +
> + - cpuacct.usage_percpu: The CPU time, in nanoseconds, consumed by all tasks in
> +   this group, separated by CPU. The format is an space-separated array of time
> +   values, one for each present CPU.
> +
> + - cpuacct.stat: aggregate user and system time consumed by tasks in this group.
> +   The format is
> +	user: x
> +	system: y

Thanks.

-- 
tejun

  reply	other threads:[~2013-06-06 23:28 UTC|newest]

Thread overview: 21+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2013-05-29 11:03 [PATCH v7 00/11] per-cgroup cpu-stat Glauber Costa
2013-05-29 11:03 ` [PATCH v7 01/11] don't call cpuacct_charge in stop_task.c Glauber Costa
2013-05-29 11:03 ` [PATCH v7 02/11] cgroup: implement CFTYPE_NO_PREFIX Glauber Costa
2013-05-29 11:03 ` [PATCH v7 03/11] cgroup, sched: let cpu serve the same files as cpuacct Glauber Costa
2013-05-29 11:03 ` [PATCH v7 04/11] sched: adjust exec_clock to use it as cpu usage metric Glauber Costa
2013-06-06 23:00   ` Tejun Heo
2013-05-29 11:03 ` [PATCH v7 05/11] cpuacct: don't actually do anything Glauber Costa
2013-06-06 23:16   ` Tejun Heo
2013-05-29 11:03 ` [PATCH v7 06/11] sched: document the cpu cgroup Glauber Costa
2013-06-06 23:28   ` Tejun Heo [this message]
2013-05-29 11:03 ` [PATCH v7 07/11] sched: account guest time per-cgroup as well Glauber Costa
2013-06-06 23:48   ` Tejun Heo
2013-05-29 11:03 ` [PATCH v7 08/11] sched: Push put_prev_task() into pick_next_task() Glauber Costa
2013-06-06 23:56   ` Tejun Heo
2013-05-29 11:03 ` [PATCH v7 09/11] sched: record per-cgroup number of context switches Glauber Costa
2013-06-07  0:04   ` Tejun Heo
2013-05-29 11:03 ` [PATCH v7 10/11] sched: change nr_context_switches calculation Glauber Costa
2013-05-29 11:03 ` [PATCH v7 11/11] sched: introduce cgroup file stat_percpu Glauber Costa
2013-06-06  1:49 ` [PATCH v7 00/11] per-cgroup cpu-stat Tejun Heo
2013-06-06  7:58   ` Glauber Costa
2013-06-07  0:06 ` Tejun Heo

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20130606232803.GP5045@htj.dyndns.org \
    --to=tj@kernel.org \
    --cc=a.p.zijlstra@chello.nl \
    --cc=cgroups@vger.kernel.org \
    --cc=devel@openvz.org \
    --cc=fweisbec@gmail.com \
    --cc=glommer@openvz.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=pjt@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®