From: Michal Hocko <mhocko@kernel.org>
To: Kemi Wang <kemi.wang@intel.com>
Cc: "Luis R . Rodriguez" <mcgrof@kernel.org>,
Kees Cook <keescook@chromium.org>,
Andrew Morton <akpm@linux-foundation.org>,
Jonathan Corbet <corbet@lwn.net>,
Mel Gorman <mgorman@techsingularity.net>,
Johannes Weiner <hannes@cmpxchg.org>,
Christopher Lameter <cl@linux.com>,
Sebastian Andrzej Siewior <bigeasy@linutronix.de>,
Vlastimil Babka <vbabka@suse.cz>,
Dave <dave.hansen@linux.intel.com>,
Tim Chen <tim.c.chen@intel.com>,
Andi Kleen <andi.kleen@intel.com>,
Jesper Dangaard Brouer <brouer@redhat.com>,
Ying Huang <ying.huang@intel.com>, Aaron Lu <aaron.lu@intel.com>,
Proc sysctl <linux-fsdevel@vger.kernel.org>,
Linux MM <linux-mm@kvack.org>,
Linux Kernel <linux-kernel@vger.kernel.org>
Subject: Re: [PATCH v3] mm, sysctl: make NUMA stats configurable
Date: Tue, 3 Oct 2017 11:23:52 +0200 [thread overview]
Message-ID: <20171003092352.2wh2jbtt2dudfi5a@dhcp22.suse.cz> (raw)
In-Reply-To: <1506579101-5457-1-git-send-email-kemi.wang@intel.com>
On Thu 28-09-17 14:11:41, Kemi Wang wrote:
> This is the second step which introduces a tunable interface that allow
> numa stats configurable for optimizing zone_statistics(), as suggested by
> Dave Hansen and Ying Huang.
>
> =========================================================================
> When page allocation performance becomes a bottleneck and you can tolerate
> some possible tool breakage and decreased numa counter precision, you can
> do:
> echo [C|c]oarse > /proc/sys/vm/numa_stats_mode
> In this case, numa counter update is ignored. We can see about
> *4.8%*(185->176) drop of cpu cycles per single page allocation and reclaim
> on Jesper's page_bench01 (single thread) and *8.1%*(343->315) drop of cpu
> cycles per single page allocation and reclaim on Jesper's page_bench03 (88
> threads) running on a 2-Socket Broadwell-based server (88 threads, 126G
> memory).
>
> Benchmark link provided by Jesper D Brouer(increase loop times to
> 10000000):
> https://github.com/netoptimizer/prototype-kernel/tree/master/kernel/mm/
> bench
>
> =========================================================================
> When page allocation performance is not a bottleneck and you want all
> tooling to work, you can do:
> echo [S|s]trict > /proc/sys/vm/numa_stats_mode
>
> =========================================================================
> We recommend automatic detection of numa statistics by system, this is also
> system default configuration, you can do:
> echo [A|a]uto > /proc/sys/vm/numa_stats_mode
> In this case, numa counter update is skipped unless it has been read by
> users at least once, e.g. cat /proc/zoneinfo.
I am still not convinced the auto mode is worth all the additional code
and a safe default to use. The whole thing could have been 0/1 with a
simpler parsing and less code to catch readers.
E.g. why do we have to do static_branch_enable on any read or even
vmstat_stop? Wouldn't open be sufficient?
> @@ -153,6 +153,8 @@ static DEVICE_ATTR(meminfo, S_IRUGO, node_read_meminfo, NULL);
> static ssize_t node_read_numastat(struct device *dev,
> struct device_attribute *attr, char *buf)
> {
> + if (vm_numa_stats_mode == VM_NUMA_STAT_AUTO_MODE)
> + static_branch_enable(&vm_numa_stats_mode_key);
> return sprintf(buf,
> "numa_hit %lu\n"
> "numa_miss %lu\n"
> @@ -186,6 +188,8 @@ static ssize_t node_read_vmstat(struct device *dev,
> n += sprintf(buf+n, "%s %lu\n",
> vmstat_text[i + NR_VM_ZONE_STAT_ITEMS],
> sum_zone_numa_state(nid, i));
> + if (vm_numa_stats_mode == VM_NUMA_STAT_AUTO_MODE)
> + static_branch_enable(&vm_numa_stats_mode_key);
> #endif
>
> for (i = 0; i < NR_VM_NODE_STAT_ITEMS; i++)
[...]
> @@ -1582,6 +1703,10 @@ static int zoneinfo_show(struct seq_file *m, void *arg)
> {
> pg_data_t *pgdat = (pg_data_t *)arg;
> walk_zones_in_node(m, pgdat, false, false, zoneinfo_show_print);
> +#ifdef CONFIG_NUMA
> + if (vm_numa_stats_mode == VM_NUMA_STAT_AUTO_MODE)
> + static_branch_enable(&vm_numa_stats_mode_key);
> +#endif
> return 0;
> }
>
> @@ -1678,6 +1803,10 @@ static int vmstat_show(struct seq_file *m, void *arg)
>
> static void vmstat_stop(struct seq_file *m, void *arg)
> {
> +#ifdef CONFIG_NUMA
> + if (vm_numa_stats_mode == VM_NUMA_STAT_AUTO_MODE)
> + static_branch_enable(&vm_numa_stats_mode_key);
> +#endif
> kfree(m->private);
> m->private = NULL;
> }
> --
> 2.7.4
>
--
Michal Hocko
SUSE Labs
next prev parent reply other threads:[~2017-10-03 9:23 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2017-09-28 6:11 Kemi Wang
2017-09-28 21:29 ` Andrew Morton
2017-09-29 1:44 ` kemi
2017-09-29 7:09 ` Vlastimil Babka
2017-09-29 7:03 ` Vlastimil Babka
2017-10-09 2:20 ` kemi
2017-09-29 7:27 ` Vlastimil Babka
2017-10-03 9:23 ` Michal Hocko [this message]
2017-10-09 6:34 ` kemi
2017-10-09 7:55 ` Michal Hocko
2017-10-10 5:49 ` Michal Hocko
2017-10-10 5:54 ` kemi
2017-10-10 14:29 ` Dave Hansen
2017-10-10 14:31 ` Michal Hocko
2017-10-10 14:53 ` Dave Hansen
2017-10-10 14:57 ` Michal Hocko
2017-10-10 15:14 ` Christopher Lameter
2017-10-10 15:39 ` Dave Hansen
2017-10-10 17:51 ` Michal Hocko
2017-10-10 21:34 ` Andrew Morton
2017-10-11 6:16 ` Vlastimil Babka
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20171003092352.2wh2jbtt2dudfi5a@dhcp22.suse.cz \
--to=mhocko@kernel.org \
--cc=aaron.lu@intel.com \
--cc=akpm@linux-foundation.org \
--cc=andi.kleen@intel.com \
--cc=bigeasy@linutronix.de \
--cc=brouer@redhat.com \
--cc=cl@linux.com \
--cc=corbet@lwn.net \
--cc=dave.hansen@linux.intel.com \
--cc=hannes@cmpxchg.org \
--cc=keescook@chromium.org \
--cc=kemi.wang@intel.com \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-mm@kvack.org \
--cc=mcgrof@kernel.org \
--cc=mgorman@techsingularity.net \
--cc=tim.c.chen@intel.com \
--cc=vbabka@suse.cz \
--cc=ying.huang@intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®