* [PATCH 1/2] perf, tools: Document event specifications better
@ 2016-03-21 15:56 Andi Kleen
2016-03-21 15:56 ` [PATCH 2/2] perf, tools: Fix documentation of :ppp in perf list Andi Kleen
2016-03-22 7:57 ` [PATCH 1/2] perf, tools: Document event specifications better Jiri Olsa
0 siblings, 2 replies; 8+ messages in thread
From: Andi Kleen @ 2016-03-21 15:56 UTC (permalink / raw)
To: acme; +Cc: jolsa, mingo, peterz, linux-kernel, Andi Kleen
From: Andi Kleen <ak@linux.intel.com>
Document some undocumented features for specifying events in the perf
list manpage:
- Event groups
- Leader sampling
- How to specify raw PMU events in the new syntax
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/Documentation/perf-list.txt | 49 ++++++++++++++++++++++++++++++++++
1 file changed, 49 insertions(+)
diff --git a/tools/perf/Documentation/perf-list.txt b/tools/perf/Documentation/perf-list.txt
index 79483f4..240c8ff 100644
--- a/tools/perf/Documentation/perf-list.txt
+++ b/tools/perf/Documentation/perf-list.txt
@@ -91,6 +91,22 @@ raw encoding of 0x1A8 can be used:
You should refer to the processor specific documentation for getting these
details. Some of them are referenced in the SEE ALSO section below.
+ARBITRARY PMUS
+--------------
+
+perf also supports an extended syntax for specifying raw parameters
+to PMUs. Using this typically requires looking up the specific event
+in the CPU vendor specific documentation.
+
+The available PMUs and their raw parameters can be listed with
+
+ ls /sys/devices/*/format
+
+For example the raw event LSD.UOPS core pmu event above could
+be specified as
+
+ perf stat -e cpu/event=0xa8,umask=0x1,name=LSD.UOPS_CYCLES,cmask=1/ ...
+
PARAMETERIZED EVENTS
--------------------
@@ -104,6 +120,39 @@ also be supplied. For example:
perf stat -C 0 -e 'hv_gpci/dtbp_ptitc,phys_processor_idx=0x2/' ...
+EVENT GROUPS
+------------
+
+Perf supports time based multiplexing of events, when the number of events
+active exceeds the number of hardware performance counters. Multiplexing
+can cause measurement errors when the workload changes its execution
+profile.
+
+When metrics are computed using formulas from event counts, it is useful to
+ensure some events are always measured together as a group to minimize multiplexing
+errors. Event groups can be specified using { }.
+
+ perf stat -e '{instructions,cycles}' ...
+
+When too many events are specified in the group perf the event will not
+be measured. The number of available performance counters depend on the CPU.
+For example Intel Core CPUs have typically four generic performance counters
+for the core, plus a number of specialized fixed counters.
+
+Events from multiple different PMUs cannot be mixed in a group.
+
+LEADER SAMPLING
+---------------
+
+perf also supports group leader sampling using the :S specifier.
+
+ perf record -e '{cycles,instructions}:S' ...
+ perf report --group
+
+Normally all events in a event group sample, but with :S only
+the first event (the leader) samples, and it only reads the values of the
+other events in the group.
+
OPTIONS
-------
--
2.5.5
^ permalink raw reply [flat|nested] 8+ messages in thread* [PATCH 2/2] perf, tools: Fix documentation of :ppp in perf list
2016-03-21 15:56 [PATCH 1/2] perf, tools: Document event specifications better Andi Kleen
@ 2016-03-21 15:56 ` Andi Kleen
2016-03-22 7:58 ` Jiri Olsa
2016-03-24 7:38 ` [tip:perf/urgent] perf list: Fix documentation of :ppp tip-bot for Andi Kleen
2016-03-22 7:57 ` [PATCH 1/2] perf, tools: Document event specifications better Jiri Olsa
1 sibling, 2 replies; 8+ messages in thread
From: Andi Kleen @ 2016-03-21 15:56 UTC (permalink / raw)
To: acme; +Cc: jolsa, mingo, peterz, linux-kernel, Andi Kleen
From: Andi Kleen <ak@linux.intel.com>
Correctly document what is implemented for :ppp on Intel CPUs in
recent kernels.
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/Documentation/perf-list.txt | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/tools/perf/Documentation/perf-list.txt b/tools/perf/Documentation/perf-list.txt
index 240c8ff..fbe52ac 100644
--- a/tools/perf/Documentation/perf-list.txt
+++ b/tools/perf/Documentation/perf-list.txt
@@ -40,10 +40,12 @@ address should be. The 'p' modifier can be specified multiple times:
0 - SAMPLE_IP can have arbitrary skid
1 - SAMPLE_IP must have constant skid
2 - SAMPLE_IP requested to have 0 skid
- 3 - SAMPLE_IP must have 0 skid
+ 3 - SAMPLE_IP must have 0 skid, or uses randomization to avoid
+ sample shadowing effects.
For Intel systems precise event sampling is implemented with PEBS
-which supports up to precise-level 2.
+which supports up to precise-level 2, and precise level 3 for
+some special cases
On AMD systems it is implemented using IBS (up to precise-level 2).
The precise modifier works with event types 0x76 (cpu-cycles, CPU
--
2.5.5
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 2/2] perf, tools: Fix documentation of :ppp in perf list
2016-03-21 15:56 ` [PATCH 2/2] perf, tools: Fix documentation of :ppp in perf list Andi Kleen
@ 2016-03-22 7:58 ` Jiri Olsa
2016-03-24 7:38 ` [tip:perf/urgent] perf list: Fix documentation of :ppp tip-bot for Andi Kleen
1 sibling, 0 replies; 8+ messages in thread
From: Jiri Olsa @ 2016-03-22 7:58 UTC (permalink / raw)
To: Andi Kleen; +Cc: acme, jolsa, mingo, peterz, linux-kernel, Andi Kleen
On Mon, Mar 21, 2016 at 08:56:33AM -0700, Andi Kleen wrote:
> From: Andi Kleen <ak@linux.intel.com>
>
> Correctly document what is implemented for :ppp on Intel CPUs in
> recent kernels.
>
> Signed-off-by: Andi Kleen <ak@linux.intel.com>
Acked-by: Jiri Olsa <jolsa@kernel.org>
thanks,
jirka
^ permalink raw reply [flat|nested] 8+ messages in thread
* [tip:perf/urgent] perf list: Fix documentation of :ppp
2016-03-21 15:56 ` [PATCH 2/2] perf, tools: Fix documentation of :ppp in perf list Andi Kleen
2016-03-22 7:58 ` Jiri Olsa
@ 2016-03-24 7:38 ` tip-bot for Andi Kleen
1 sibling, 0 replies; 8+ messages in thread
From: tip-bot for Andi Kleen @ 2016-03-24 7:38 UTC (permalink / raw)
To: linux-tip-commits; +Cc: peterz, linux-kernel, acme, mingo, ak, tglx, hpa, jolsa
Commit-ID: 4ca0d8193f8b6b2c74d05d9bbcbaa99fdd553503
Gitweb: http://git.kernel.org/tip/4ca0d8193f8b6b2c74d05d9bbcbaa99fdd553503
Author: Andi Kleen <ak@linux.intel.com>
AuthorDate: Mon, 21 Mar 2016 08:56:33 -0700
Committer: Arnaldo Carvalho de Melo <acme@redhat.com>
CommitDate: Tue, 22 Mar 2016 10:01:45 -0300
perf list: Fix documentation of :ppp
Correctly document what is implemented for :ppp on Intel CPUs in recent
kernels.
Signed-off-by: Andi Kleen <ak@linux.intel.com>
Acked-by: Jiri Olsa <jolsa@kernel.org>
Cc: Peter Zijlstra <peterz@infradead.org>
Link: http://lkml.kernel.org/r/1458575793-12091-2-git-send-email-andi@firstfloor.org
Signed-off-by: Arnaldo Carvalho de Melo <acme@redhat.com>
---
tools/perf/Documentation/perf-list.txt | 6 ++++--
1 file changed, 4 insertions(+), 2 deletions(-)
diff --git a/tools/perf/Documentation/perf-list.txt b/tools/perf/Documentation/perf-list.txt
index 79483f4..ec723d0 100644
--- a/tools/perf/Documentation/perf-list.txt
+++ b/tools/perf/Documentation/perf-list.txt
@@ -40,10 +40,12 @@ address should be. The 'p' modifier can be specified multiple times:
0 - SAMPLE_IP can have arbitrary skid
1 - SAMPLE_IP must have constant skid
2 - SAMPLE_IP requested to have 0 skid
- 3 - SAMPLE_IP must have 0 skid
+ 3 - SAMPLE_IP must have 0 skid, or uses randomization to avoid
+ sample shadowing effects.
For Intel systems precise event sampling is implemented with PEBS
-which supports up to precise-level 2.
+which supports up to precise-level 2, and precise level 3 for
+some special cases
On AMD systems it is implemented using IBS (up to precise-level 2).
The precise modifier works with event types 0x76 (cpu-cycles, CPU
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: [PATCH 1/2] perf, tools: Document event specifications better
2016-03-21 15:56 [PATCH 1/2] perf, tools: Document event specifications better Andi Kleen
2016-03-21 15:56 ` [PATCH 2/2] perf, tools: Fix documentation of :ppp in perf list Andi Kleen
@ 2016-03-22 7:57 ` Jiri Olsa
2016-03-22 8:47 ` Peter Zijlstra
1 sibling, 1 reply; 8+ messages in thread
From: Jiri Olsa @ 2016-03-22 7:57 UTC (permalink / raw)
To: Andi Kleen; +Cc: acme, jolsa, mingo, peterz, linux-kernel, Andi Kleen
On Mon, Mar 21, 2016 at 08:56:32AM -0700, Andi Kleen wrote:
SNIP
> +EVENT GROUPS
> +------------
> +
> +Perf supports time based multiplexing of events, when the number of events
> +active exceeds the number of hardware performance counters. Multiplexing
> +can cause measurement errors when the workload changes its execution
> +profile.
> +
> +When metrics are computed using formulas from event counts, it is useful to
> +ensure some events are always measured together as a group to minimize multiplexing
> +errors. Event groups can be specified using { }.
> +
> + perf stat -e '{instructions,cycles}' ...
> +
> +When too many events are specified in the group perf the event will not
^^^^^^^^^^^^^^
maybe: 'none of them' will be meassured.
> +be measured. The number of available performance counters depend on the CPU.
> +For example Intel Core CPUs have typically four generic performance counters
> +for the core, plus a number of specialized fixed counters.
> +
> +Events from multiple different PMUs cannot be mixed in a group.
hum, but we allow that right?
only when there's mixture of SW and HW events we move
it all silently under HW event context
jirka
^ permalink raw reply [flat|nested] 8+ messages in thread* Re: [PATCH 1/2] perf, tools: Document event specifications better
2016-03-22 7:57 ` [PATCH 1/2] perf, tools: Document event specifications better Jiri Olsa
@ 2016-03-22 8:47 ` Peter Zijlstra
0 siblings, 0 replies; 8+ messages in thread
From: Peter Zijlstra @ 2016-03-22 8:47 UTC (permalink / raw)
To: Jiri Olsa; +Cc: Andi Kleen, acme, jolsa, mingo, linux-kernel, Andi Kleen
On Tue, Mar 22, 2016 at 08:57:35AM +0100, Jiri Olsa wrote:
> > +Events from multiple different PMUs cannot be mixed in a group.
>
> hum, but we allow that right?
>
> only when there's mixture of SW and HW events we move
> it all silently under HW event context
I forgot the exact details, but we put some tight restrictions on it.
SW events can indeed be added to regular HW events, and could with some
careful work be allowed on most other groups as well (but IIRC we don't
currently allow them onto things like uncore etc..).
^ permalink raw reply [flat|nested] 8+ messages in thread
* [PATCH 1/2] perf, tools: Document event specifications better
@ 2016-03-22 18:09 Andi Kleen
2016-03-27 11:37 ` Jiri Olsa
0 siblings, 1 reply; 8+ messages in thread
From: Andi Kleen @ 2016-03-22 18:09 UTC (permalink / raw)
To: acme; +Cc: jolsa, peterz, linux-kernel, Andi Kleen
From: Andi Kleen <ak@linux.intel.com>
Document some undocumented features for specifying events in the perf
list manpage:
- Event groups
- Leader sampling
- How to specify raw PMU events in the new syntax
- Global versus per process PMUs.
- Access restrictions
- Fix Intel SDM URL
v2: Lots of new content. address review feedback.
Signed-off-by: Andi Kleen <ak@linux.intel.com>
---
tools/perf/Documentation/perf-list.txt | 107 ++++++++++++++++++++++++++++++++-
1 file changed, 106 insertions(+), 1 deletion(-)
diff --git a/tools/perf/Documentation/perf-list.txt b/tools/perf/Documentation/perf-list.txt
index 79483f4..0059905 100644
--- a/tools/perf/Documentation/perf-list.txt
+++ b/tools/perf/Documentation/perf-list.txt
@@ -91,6 +91,67 @@ raw encoding of 0x1A8 can be used:
You should refer to the processor specific documentation for getting these
details. Some of them are referenced in the SEE ALSO section below.
+ARBITRARY PMUS
+--------------
+
+perf also supports an extended syntax for specifying raw parameters
+to PMUs. Using this typically requires looking up the specific event
+in the CPU vendor specific documentation.
+
+The available PMUs and their raw parameters can be listed with
+
+ ls /sys/devices/*/format
+
+For example the raw event LSD.UOPS core pmu event above could
+be specified as
+
+ perf stat -e cpu/event=0xa8,umask=0x1,name=LSD.UOPS_CYCLES,cmask=1/ ...
+
+PER SOCKET PMUS
+---------------
+
+Some PMUs are not associated with a core, but with a whole CPU socket.
+Events on these PMUs generally cannot be sampled, but only counted globally
+with perf stat -a. They can be bound to one logical CPU, but will measure
+all the CPUs in the same socket.
+
+This example measures memory bandwidth every second
+on the first memory controller on socket 0 of a Intel Xeon system
+
+ perf stat -C 0 -a uncore_imc_0/cas_count_read/,uncore_imc_0/cas_count_write/ -I 1000 ...
+
+Each memory controller has its own PMU. Measuring the complete system
+bandwidth would require specifying all imc PMUs (see perf list output),
+and adding the values together.
+
+This example measures the combined core power every second
+
+ perf stat -I 1000 -e power/energy-cores/ -a
+
+ACCESS RESTRICTIONS
+-------------------
+
+For non root users generally only context switched PMU events are available.
+This is normally only the events in the cpu PMU, the predefined events
+like cycles and instructions and some software events.
+
+Other PMUs and global measurements are normally root only.
+Some event qualifiers, such as any, are also root only.
+
+This can be overriden by setting the kernel.perf_event_paranoid
+sysctl to -1, which allows non root to use these events.
+
+For accessing trace point events perf needs to have read access to
+/sys/kernel/debug/tracing, even when perf_event_paranoid is in a relaxed
+setting.
+
+TRACING
+-------
+
+Some PMUs control advanced hardware tracing capabilities, such as Intel PT,
+that allows low overhead execution tracing. These are described in a separate
+intel-pt.txt document.
+
PARAMETERIZED EVENTS
--------------------
@@ -104,6 +165,50 @@ also be supplied. For example:
perf stat -C 0 -e 'hv_gpci/dtbp_ptitc,phys_processor_idx=0x2/' ...
+EVENT GROUPS
+------------
+
+Perf supports time based multiplexing of events, when the number of events
+active exceeds the number of hardware performance counters. Multiplexing
+can cause measurement errors when the workload changes its execution
+profile.
+
+When metrics are computed using formulas from event counts, it is useful to
+ensure some events are always measured together as a group to minimize multiplexing
+errors. Event groups can be specified using { }.
+
+ perf stat -e '{instructions,cycles}' ...
+
+The number of available performance counters depend on the CPU. A group
+cannot contain more events than available counters.
+For example Intel Core CPUs typically have four generic performance counters
+for the core, plus three fixed counters for instructions, cycles and
+ref-cycles. Some special events have restrictions on which counter they
+can schedule, and may not support multiple instances in a single group.
+When too many events are specified in the group none of them will not
+be measured.
+
+Globally pinned events can limit the number of counters available for
+other groups. On x86 systems, the NMI watchdog pins a counter by default.
+The nmi watchdog can be disabled as root with
+
+ echo 0 > /proc/sys/kernel/nmi_watchdog
+
+Events from multiple different PMUs cannot be mixed in a group, with
+some exceptions for software events.
+
+LEADER SAMPLING
+---------------
+
+perf also supports group leader sampling using the :S specifier.
+
+ perf record -e '{cycles,instructions}:S' ...
+ perf report --group
+
+Normally all events in a event group sample, but with :S only
+the first event (the leader) samples, and it only reads the values of the
+other events in the group.
+
OPTIONS
-------
@@ -141,5 +246,5 @@ SEE ALSO
--------
linkperf:perf-stat[1], linkperf:perf-top[1],
linkperf:perf-record[1],
-http://www.intel.com/Assets/PDF/manual/253669.pdf[Intel® 64 and IA-32 Architectures Software Developer's Manual Volume 3B: System Programming Guide],
+http://www.intel.com/sdm/[Intel® 64 and IA-32 Architectures Software Developer's Manual Volume 3B: System Programming Guide],
http://support.amd.com/us/Processor_TechDocs/24593_APM_v2.pdf[AMD64 Architecture Programmer’s Manual Volume 2: System Programming]
--
2.5.5
^ permalink raw reply [flat|nested] 8+ messages in thread* Re: [PATCH 1/2] perf, tools: Document event specifications better
2016-03-22 18:09 Andi Kleen
@ 2016-03-27 11:37 ` Jiri Olsa
0 siblings, 0 replies; 8+ messages in thread
From: Jiri Olsa @ 2016-03-27 11:37 UTC (permalink / raw)
To: Andi Kleen; +Cc: acme, jolsa, peterz, linux-kernel, Andi Kleen
On Tue, Mar 22, 2016 at 11:09:30AM -0700, Andi Kleen wrote:
> From: Andi Kleen <ak@linux.intel.com>
>
> Document some undocumented features for specifying events in the perf
> list manpage:
>
> - Event groups
> - Leader sampling
> - How to specify raw PMU events in the new syntax
> - Global versus per process PMUs.
> - Access restrictions
> - Fix Intel SDM URL
>
> v2: Lots of new content. address review feedback.
cool, thanks
Acked-by: Jiri Olsa <jolsa@kernel.org>
jirka
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2016-03-27 11:37 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2016-03-21 15:56 [PATCH 1/2] perf, tools: Document event specifications better Andi Kleen
2016-03-21 15:56 ` [PATCH 2/2] perf, tools: Fix documentation of :ppp in perf list Andi Kleen
2016-03-22 7:58 ` Jiri Olsa
2016-03-24 7:38 ` [tip:perf/urgent] perf list: Fix documentation of :ppp tip-bot for Andi Kleen
2016-03-22 7:57 ` [PATCH 1/2] perf, tools: Document event specifications better Jiri Olsa
2016-03-22 8:47 ` Peter Zijlstra
2016-03-22 18:09 Andi Kleen
2016-03-27 11:37 ` Jiri Olsa
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome