mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Ingo Molnar <mingo@elte.hu>
To: Avi Kivity <avi@redhat.com>
Cc: Robert Richter <robert.richter@amd.com>,
	Benjamin Block <bebl@mageta.org>,
	Hans Rosenfeld <hans.rosenfeld@amd.com>,
	hpa@zytor.com, tglx@linutronix.de, suresh.b.siddha@intel.com,
	eranian@google.com, brgerst@gmail.com, Andreas.Herrmann3@amd.com,
	x86@kernel.org, linux-kernel@vger.kernel.org,
	Benjamin Block <benjamin.block@amd.com>
Subject: Re: [RFC 4/5] x86, perf: implements lwp-perf-integration (rc1)
Date: Tue, 20 Dec 2011 10:15:11 +0100	[thread overview]
Message-ID: <20111220091511.GB3091@elte.hu> (raw)
In-Reply-To: <4EEF26F0.1050709@redhat.com>

* Avi Kivity <avi@redhat.com> wrote:

> On 12/19/2011 01:40 PM, Ingo Molnar wrote:
> >
> > 2) the proper solution: creating a 'user-space vmalloc()' that 
> >    is per mm and that gets inherited transparently, across 
> >    fork() and exec(), and which lies outside the regular vma
> >    spaces. On 64-bit this should be straightforward.
> 
> That probably has uses outside perf too, but I can see mm nacks piling up.

This can be done in arch/x86/ code if it's too x86 specific - 
the platform controls the VM layout and can (and does) use 
special per CPU VM areas.

> >    These vmas are not actually 'known' to user-space 
> >    normally - the kernel PMU code knows about it and does 
> >    what we do with PEBS: flushes it when necessary and puts 
> >    it into the regular perf event channels.
> >
> >    This solves the inherited perf record workflow 
> >    immediately: the parent task just creates the buffer, 
> >    which gets inherited across exec() and fork(), into every 
> >    portion of the workload.
> 
> The buffer still needs to be managed. [...]

Of course, like we manage the DS buffer for PEBS.

> [...]  While you may be able to juggle different threads on 
> the same cpu using different events, threads on other cpus 
> need to use separate LWP contexts and buffers.

Yes, like different threads on different CPUs have different DS 
buffers, *here and today*.

Try this on (most) modern Intel CPUs:

  perf top -e cycles:pp

That will activate that exact mechanism.

The LWPCB and the LWP ring-buffer are really just an extension 
of that concept: per task buffers which are ring 3 visible. 

Note that user-space does not actually have to know about any of 
these LWP addresses (but can access them if it wants to - no 
strong feelings about that) - in the correctly implemented model 
it's fully kernel managed.

In fact the PEBS case had one more complication: there's the BTS 
branch-tracing feature which we support as well, and which 
overlaps PEBS use of the DS.

All these PMU hardware limitations can be supported, as long as 
the instrumentation *capability* adds value to the system in one 
way or another.

> >    System-wide profiling is a small additional variant of 
> >    this: creating such a user-vmalloc() area for all tasks 
> >    in the system so that the PMU code has them ready in the 
> >    context-switch code.
> 
> What about security?  Do we want to allow any userspace 
> process to mess up the buffers?  It can even reprogram the LWP 
> block, so you're counting different things, or at higher 
> frequencies, or into other processes ordinary vmas?

In most usecases it's the application messing up its own 
profiling - don't do that if it hurts.

I'd argue that future LWP versions should allow kernel-protected 
LWP pages, as long as the LWPCB is privileged as well as well. 
That would be useful for another purpose as well: LWP could be 
allowed to sample kernel-space execution as well, an obviously 
useful feature that was left out from LWP for barely explicable 
reasons.

Granted, LWP was mis-designed to quite a degree, those AMD chip 
engineers should have talked to people who understand how modern 
PMU abstractions are added to the OS kernel properly. But this 
mis-design does not keep us from utilizing this piece of 
hardware intelligently. PEBS/DS/BTS wasnt a beauty either.

> You could rebuild the LWP block on every context switch I 
> guess, but you need to prevent access to other cpus' LWP 
> blocks (since they may be running other processes).  I think 
> this calls for per-cpu cr3, even for threads in the same 
> process.

Why would we want to rebuild the LWPCB? Just keep one per task 
and do a lightweight switch to it during switch_to() - like we 
do it with the PEBS hardware-ring-buffer. It can be in the same 
single block of memory with the ring-buffer itself. (PEBS has 
similar characteristics)

Thanks,

	Ingo

  parent reply	other threads:[~2011-12-20  9:17 UTC|newest]

Thread overview: 54+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2011-11-29 12:41 [PATCH 0/9] rework of extended state handling, LWP support Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 1/9] x86, xsave: warn on #NM exceptions caused by the kernel Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 2/9] x86, xsave: cleanup fpu/xsave support Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 3/9] x86, xsave: cleanup fpu/xsave signal frame setup Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 4/9] x86, xsave: rework fpu/xsave support Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 5/9] x86, xsave: remove unused code Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 6/9] x86, xsave: more cleanups Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 7/9] x86, xsave: remove lazy allocation of xstate area Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 8/9] x86, xsave: add support for non-lazy xstates Hans Rosenfeld
2011-11-29 12:41 ` [PATCH 9/9] x86, xsave: add kernel support for AMDs Lightweight Profiling (LWP) Hans Rosenfeld
2011-11-29 21:31 ` [PATCH 0/9] rework of extended state handling, LWP support Andi Kleen
2011-11-30 17:37   ` Hans Rosenfeld
2011-11-30 21:52     ` Andi Kleen
2011-12-01 20:36       ` Hans Rosenfeld
2011-12-02  2:01         ` H. Peter Anvin
2011-12-02 11:20           ` Hans Rosenfeld
2011-12-07 19:57             ` Hans Rosenfeld
2011-12-07 20:00               ` [PATCH 7/8] x86, xsave: add support for non-lazy xstates Hans Rosenfeld
2011-12-07 20:00                 ` [PATCH 8/8] x86, xsave: add kernel support for AMDs Lightweight Profiling (LWP) Hans Rosenfeld
2011-12-05 10:22 ` [PATCH 0/9] rework of extended state handling, LWP support Ingo Molnar
2011-12-16 16:07   ` Hans Rosenfeld
2011-12-16 16:12     ` [RFC 1/5] x86, perf: Implement software-activation of lwp Hans Rosenfeld
2011-12-16 16:12       ` [RFC 2/5] perf: adds prototype for a new perf-context-type Hans Rosenfeld
2011-12-16 16:12       ` [RFC 3/5] perf: adds a new pmu-initialization-call Hans Rosenfeld
2011-12-16 16:12       ` [RFC 4/5] x86, perf: implements lwp-perf-integration (rc1) Hans Rosenfeld
2011-12-18  8:04         ` Ingo Molnar
2011-12-18 15:22           ` Benjamin Block
2011-12-18 23:43             ` Ingo Molnar
2011-12-19  9:09               ` Robert Richter
2011-12-19 10:54                 ` Ingo Molnar
2011-12-19 11:12                   ` Avi Kivity
2011-12-19 11:40                     ` Ingo Molnar
2011-12-19 11:58                       ` Avi Kivity
2011-12-19 18:13                         ` Benjamin
2011-12-20  8:56                           ` Ingo Molnar
2011-12-20  9:15                         ` Ingo Molnar [this message]
2011-12-20  9:47                           ` Avi Kivity
2011-12-20 10:09                             ` Ingo Molnar
2011-12-20 15:27                               ` Joerg Roedel
2011-12-20 18:40                                 ` Ingo Molnar
2011-12-21  0:07                                   ` Joerg Roedel
2011-12-21 12:34                                     ` Ingo Molnar
2011-12-21 12:44                                       ` Avi Kivity
2011-12-21 13:22                                         ` Ingo Molnar
2011-12-21 22:49                                           ` Joerg Roedel
2011-12-23 10:53                                             ` Ingo Molnar
2011-12-21 11:46                                   ` Gleb Natapov
2011-12-23 10:56                                     ` Ingo Molnar
2011-12-20 15:48                           ` Vince Weaver
2011-12-20 18:27                             ` Ingo Molnar
2011-12-20 22:47                               ` Vince Weaver
2011-12-21 12:00                                 ` Ingo Molnar
2011-12-21 13:55                                   ` Vince Weaver
2011-12-16 16:12       ` [RFC 5/5] x86, perf: adds support for the LWP threshold-int Hans Rosenfeld

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20111220091511.GB3091@elte.hu \
    --to=mingo@elte.hu \
    --cc=Andreas.Herrmann3@amd.com \
    --cc=avi@redhat.com \
    --cc=bebl@mageta.org \
    --cc=benjamin.block@amd.com \
    --cc=brgerst@gmail.com \
    --cc=eranian@google.com \
    --cc=hans.rosenfeld@amd.com \
    --cc=hpa@zytor.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=robert.richter@amd.com \
    --cc=suresh.b.siddha@intel.com \
    --cc=tglx@linutronix.de \
    --cc=x86@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®