From: Stephane Eranian <eranian@google.com>
To: linux-kernel@vger.kernel.org
Cc: peterz@infradead.org, mingo@elte.hu, acme@redhat.com,
ming.m.lin@intel.com, andi@firstfloor.org,
robert.richter@amd.com, ravitillo@lbl.gov, will.deacon@arm.com,
paulus@samba.org, benh@kernel.crashing.org, rth@twiddle.net,
ralf@linux-mips.org, davem@davemloft.net, lethal@linux-sh.org
Subject: [PATCH 12/12] perf: add support for taken branch sampling to perf report (v2)
Date: Fri, 14 Oct 2011 14:37:13 +0200 [thread overview]
Message-ID: <1318595833-29984-13-git-send-email-eranian@google.com> (raw)
In-Reply-To: <1318595833-29984-1-git-send-email-eranian@google.com>
From: Roberto Agostino Vitillo <ravitillo@lbl.gov>
This patch adds support for taken branch sampling, i.e, the
PERF_SAMPLE_BRANCH_STACK feature to perf report. In other
words, to display histograms based on taken branches rather
than executed instructions addresses.
The new option is called -b and it takes no argument. To
generate meaningful output, the perf.data must have been
obtained using perf record -b xxx ... where xxx is a branch
filter option.
The output shows symbols, modules, sorted by 'who branches
where' the most often. The percentages reported in the first
column refer to the total number of branches captured and
not the usual number of samples.
Here is a quick example.
Here branchy is simple test program which looks as follows:
void f2(void)
{}
void f3(void)
{}
void f1(unsigned long n)
{
if (n & 1UL)
f2();
else
f3();
}
int main(void)
{
unsigned long i;
for (i=0; i < N; i++)
f1(i);
return 0;
}
Here is the output captured on Nehalem, if we are
only interested in user level function calls.
$ perf record -b any_call,u -e cycles:u branchy
$ perf report -b --sort=symbol
52.34% [.] main [.] f1
24.04% [.] f1 [.] f3
23.60% [.] f1 [.] f2
0.01% [k] _IO_new_file_xsputn [k] _IO_file_overflow
0.01% [k] _IO_vfprintf_internal [k] _IO_new_file_xsputn
0.01% [k] _IO_vfprintf_internal [k] strchrnul
0.01% [k] __printf [k] _IO_vfprintf_internal
0.01% [k] main [k] __printf
About half (52%) of the call branches captured are from main() -> f1().
The second half (24%+23%) is split in two equal shares between
f1() -> f2(), f1() ->f3(). The output is as expected given the code.
It should be noted, that using -b in perf record does not eliminate
information in the perf.data file. Consequently, a typical profile
can also be obtained by perf report by simply not using its -b option.
Signed-off-by: Roberto Agostino Vitillo <ravitillo@lbl.gov>
Signed-off-by: Stephane Eranian <eranian@google.com>
---
tools/perf/Documentation/perf-report.txt | 7 ++
tools/perf/builtin-report.c | 93 +++++++++++++++++++++++++++---
2 files changed, 91 insertions(+), 9 deletions(-)
diff --git a/tools/perf/Documentation/perf-report.txt b/tools/perf/Documentation/perf-report.txt
index 212f24d..3163be5 100644
--- a/tools/perf/Documentation/perf-report.txt
+++ b/tools/perf/Documentation/perf-report.txt
@@ -152,6 +152,13 @@ OPTIONS
information which may be very large and thus may clutter the display.
It currently includes: cpu and numa topology of the host system.
+-b::
+--branch-stack::
+ Use the addresses of sampled taken branches instead of the instruction
+ address to build the histograms. To generate meaningful output, the
+ perf.data file must have been obtained using perf record -b xxx where
+ xxx is a branch filter option.
+
SEE ALSO
--------
linkperf:perf-stat[1], linkperf:perf-annotate[1]
diff --git a/tools/perf/builtin-report.c b/tools/perf/builtin-report.c
index 4d7c834..f52f65c 100644
--- a/tools/perf/builtin-report.c
+++ b/tools/perf/builtin-report.c
@@ -55,6 +55,46 @@ static symbol_filter_t annotate_init;
static const char *cpu_list;
static DECLARE_BITMAP(cpu_bitmap, MAX_NR_CPUS);
+static int perf_session__add_branch_hist_entry(struct perf_session *session,
+ struct addr_location *al,
+ struct perf_sample *sample,
+ struct perf_evsel *evsel){
+ struct symbol *parent = NULL;
+ int err = 0;
+ unsigned i;
+ struct hist_entry *he;
+ struct branch_info *bi;
+
+ if ((sort__has_parent || symbol_conf.use_callchain) && sample->callchain) {
+ err = perf_session__resolve_callchain(session, al->thread,
+ sample->callchain, &parent);
+ if (err)
+ return err;
+ }
+
+ bi = perf_session__resolve_bstack(session, al->thread,
+ sample->branch_stack);
+ if (!bi)
+ return -ENOMEM;
+
+ for(i = 0; i < sample->branch_stack->nr; i++) {
+ if(hide_unresolved && !(bi[i].from.sym && bi[i].to.sym))
+ continue;
+ /*
+ * The report shows the percentage of total branches captured
+ * and not events sampled. Thus we use a pseudo period of 1.
+ */
+ he = __hists__add_branch_entry(&evsel->hists, al, parent,
+ &bi[i], 1);
+ if (he) {
+ evsel->hists.stats.total_period += 1;
+ hists__inc_nr_events(&evsel->hists, PERF_RECORD_SAMPLE);
+ } else
+ return -ENOMEM;
+ }
+ return err;
+}
+
static int perf_session__add_hist_entry(struct perf_session *session,
struct addr_location *al,
struct perf_sample *sample,
@@ -120,20 +160,28 @@ static int process_sample_event(union perf_event *event,
return -1;
}
- if (al.filtered || (hide_unresolved && al.sym == NULL))
- return 0;
-
if (cpu_list && !test_bit(sample->cpu, cpu_bitmap))
return 0;
- if (al.map != NULL)
- al.map->dso->hit = 1;
+ if (sort__branch_mode) {
+ if(perf_session__add_branch_hist_entry(session, &al, sample,
+ evsel)) {
+ pr_debug("problem adding lbr entry, skipping event\n");
+ return -1;
+ }
+ } else {
+ if (al.filtered || (hide_unresolved && al.sym == NULL))
+ return 0;
- if (perf_session__add_hist_entry(session, &al, sample, evsel)) {
- pr_debug("problem incrementing symbol period, skipping event\n");
- return -1;
- }
+ if (al.map != NULL)
+ al.map->dso->hit = 1;
+ if (perf_session__add_hist_entry(session, &al, sample, evsel)) {
+ pr_debug("problem incrementing symbol period, skipping"
+ " event\n");
+ return -1;
+ }
+ }
return 0;
}
@@ -183,6 +231,15 @@ static int perf_session__setup_sample_type(struct perf_session *self)
}
}
+ if(sort__branch_mode){
+ if(!(self->sample_type & PERF_SAMPLE_BRANCH_STACK)){
+ fprintf(stderr, "selected -b but no branch data."
+ " Did you call perf record without"
+ " -b?\n");
+ return -1;
+ }
+ }
+
return 0;
}
@@ -499,6 +556,8 @@ static const struct option options[] = {
"Specify disassembler style (e.g. -M intel for intel syntax)"),
OPT_BOOLEAN(0, "show-total-period", &symbol_conf.show_total_period,
"Show a column with the sum of periods"),
+ OPT_BOOLEAN('b', "branch-stack", &sort__branch_mode,
+ "use branch records for histogram filling"),
OPT_END()
};
@@ -514,6 +573,22 @@ int cmd_report(int argc, const char **argv, const char *prefix __used)
if (inverted_callchain)
callchain_param.order = ORDER_CALLER;
+ if (sort__branch_mode){
+ if(use_browser)
+ fprintf(stderr, "Warning: TUI interface not supported"
+ " in branch mode\n");
+ if(symbol_conf.dso_list_str != NULL)
+ fprintf(stderr, "Warning: dso filtering not supported"
+ " in branch mode\n");
+ if(symbol_conf.sym_list_str != NULL)
+ fprintf(stderr, "Warning: symbol filtering not supported"
+ " in branch mode\n");
+
+ use_browser = 0;
+ symbol_conf.dso_list_str = NULL;
+ symbol_conf.sym_list_str = NULL;
+ }
+
if (strcmp(input_name, "-") != 0)
setup_browser(true);
else
--
1.7.1
next prev parent reply other threads:[~2011-10-14 12:39 UTC|newest]
Thread overview: 36+ messages / expand[flat|nested] mbox.gz Atom feed top
2011-10-14 12:37 [PATCH 00/12] perf_events: add support for sampling taken branches (v2) Stephane Eranian
2011-10-14 12:37 ` [PATCH 01/12] perf_events: add generic taken branch sampling support (v2) Stephane Eranian
2011-12-05 21:06 ` Peter Zijlstra
2011-12-06 19:42 ` Stephane Eranian
2011-12-05 22:14 ` Peter Zijlstra
2011-12-06 19:27 ` Stephane Eranian
2011-10-14 12:37 ` [PATCH 02/12] perf_events: add Intel LBR MSR definitions (v2) Stephane Eranian
2011-10-14 12:37 ` [PATCH 03/12] perf_events: add Intel X86 LBR sharing logic (v2) Stephane Eranian
2011-10-14 12:37 ` [PATCH 04/12] perf_events: sync branch stack sampling with X86 precise_sampling (v2) Stephane Eranian
2011-10-14 12:37 ` [PATCH 05/12] perf_events: add LBR mappings for PERF_SAMPLE_BRANCH filters (v2) Stephane Eranian
2011-12-05 22:35 ` Peter Zijlstra
2011-12-07 4:22 ` Stephane Eranian
2011-10-14 12:37 ` [PATCH 06/12] perf_events: implement PERF_SAMPLE_BRANCH for Intel X86 (v2) Stephane Eranian
2011-10-14 12:37 ` [PATCH 07/12] perf_events: add LBR software filter support " Stephane Eranian
2011-12-05 22:29 ` Peter Zijlstra
2011-10-14 12:37 ` [PATCH 08/12] perf_events: disable PERF_SAMPLE_BRANCH_* when not supported (v2) Stephane Eranian
2011-10-14 12:37 ` [PATCH 09/12] perf_events: add hook to flush branch_stack on context switch (v2) Stephane Eranian
2011-12-05 21:10 ` Peter Zijlstra
2011-12-05 21:37 ` Peter Zijlstra
2011-12-07 18:25 ` Stephane Eranian
2011-12-08 10:49 ` Peter Zijlstra
2011-12-08 18:04 ` Stephane Eranian
2011-12-08 18:13 ` Peter Zijlstra
2011-12-08 22:06 ` Stephane Eranian
2011-12-09 9:00 ` Peter Zijlstra
2011-10-14 12:37 ` [PATCH 10/12] perf: add code to support PERF_SAMPLE_BRANCH_STACK (v2) Stephane Eranian
2011-10-14 12:37 ` [PATCH 11/12] perf: add support for sampling taken branch to perf record (v2) Stephane Eranian
2011-10-14 12:37 ` Stephane Eranian [this message]
2011-12-04 20:11 ` [PATCH 00/12] perf_events: add support for sampling taken branches (v2) Stephane Eranian
2011-12-05 15:27 ` Peter Zijlstra
2011-12-05 22:39 ` Peter Zijlstra
2011-12-06 9:49 ` Will Deacon
2011-12-06 11:03 ` Peter Zijlstra
2011-12-06 19:14 ` Stephane Eranian
2011-12-06 19:20 ` Peter Zijlstra
2011-12-06 19:22 ` Stephane Eranian
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1318595833-29984-13-git-send-email-eranian@google.com \
--to=eranian@google.com \
--cc=acme@redhat.com \
--cc=andi@firstfloor.org \
--cc=benh@kernel.crashing.org \
--cc=davem@davemloft.net \
--cc=lethal@linux-sh.org \
--cc=linux-kernel@vger.kernel.org \
--cc=ming.m.lin@intel.com \
--cc=mingo@elte.hu \
--cc=paulus@samba.org \
--cc=peterz@infradead.org \
--cc=ralf@linux-mips.org \
--cc=ravitillo@lbl.gov \
--cc=robert.richter@amd.com \
--cc=rth@twiddle.net \
--cc=will.deacon@arm.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®