mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* Re: [perfmon2] IV.3 - AMD IBS
@ 2009-06-25 11:28 stephane eranian
  0 siblings, 0 replies; 9+ messages in thread
From: stephane eranian @ 2009-06-25 11:28 UTC (permalink / raw)
  To: Ingo Molnar
  Cc: Drongowski, Paul, Peter Zijlstra, Rob Fowler, Philip Mucci, LKML,
	Andi Kleen, Paul Mackerras, Maynard Johnson, Andrew Morton,
	Thomas Gleixner, perfmon2-devel

Hi,

On Tue, Jun 23, 2009 at 4:55 PM, Ingo Molnar<mingo@elte.hu> wrote:
>
> The 20 bits delay is in cycles, right? So this in itself still lends
> itself to be transparently provided as a PERF_COUNT_HW_CPU_CYCLES
> counter.
>

I do not believe you can use  IBS as a better substitute for either CYCLES or
INSTRUCTIONS sampling. IBS simply does not operate in the same way.

But instead of me arguing with you guys for a long time, I have asked someone
at AMD who knows more than me about IBS. Paul posted his answer only on
the perfmon2 mailing list, I have forwarded it below.

You will also note that he is providing another example as to why support for
software sampling period randomization is useful.

I would like to thank Paul for spending time providing a lot of useful details
about IBS.

I am hoping this can clarify things.

On Wed, Jun 24, 2009 at 8:20 PM, Drongowski,
Paul<paul.drongowski@amd.com> wrote:
>
> Hi --
>
> I'm sorry to be joining this discussion so late. A few of my
> colleagues pointed me toward the current thread on IBS and I've tried
> to catch up by reading the archives. A short self-introduction: I'm a
> member of the AMD CodeAnalyst team, Ravi Bhargava and I wrote Appendix G
> (concerning IBS) of the AMD Software Optimization Guide for AMD
> Family 10h Processors and at one point in my life, I worked on DCPI
> (using ProfileMe).
>
> First off, Stephane and Rob have done a good job representing IBS and
> also ProfileMe. Thanks, guys!
>
> Rather than grossly disturb the current discussion, I'd like to offer
> a few points of clarification and maybe a little useful history.
>
> Peter's observation that IBS is a "mismatch with the traditional one
> value per counter thing" is quite apt. IBS has similarities to
> ProfileMe. Stephane's citation of the Itanium Data-EAR and
> Instruction-EAR are also very relevant as examples of profile data
> that do not fit with the "one value per counter thing."
>
> IBS Fetch.
>
>    IBS fetch sampling does not exactly sample x86 instructions. The
>    current fetch counter counts fetch operations where a fetch
> operation
>    may be a 32-byte fetch block (on AMD Family 10h) or it may be a
>    fetch operation initiated by a redirection such as a branch.
>    A fetch block is 32 bytes of instruction information which is
>    sent to the instruction decoder. The fetch address that is reported
>    may either be the start of a valid x86 instruction or the start of
>    a fetch block. In the second case, the address may be in the middle
> of
>    an x86 instruction.
>
>    IBS fetch sampling produces a number of event flags (e.g.,
> instruction
>    cache miss), but it also produces the latency (in cycles) of the
>    fetch operation. The latencies can be accumulated in either
>    descriptive statistics, or better, in a histogram since descriptive
>    statistics don't really show where an access is hitting in the
>    memory hierarchy. BTW, even though an IBS fetch sample may be
> reported,
>    the decoder may not use the instruction bytes due to a late arriving
>    redirection.
>
> IBS Op.
>
>    IBS op sampling does not sample x86 instructions. It samples the
>    ops which are issued from x86 instructions. Some x86 instructions
>    issue more than one op. Microcoded instructions are particularly
>    thorny as a single REP MOV may issue many ops, thereby affecting
>    the number of samples that fall on them (i.e., disproportionate to
> the
>    execution frequency of the surrounding basic block.) The number of
>    ops issued is data dependent and is unpredictable. Appendix C
>    of the Software Optimization Guide lists the number of ops issued
>    from x86 instructions (one, two or many).
>
>    Beginning with AMD Family 10h RevC, there are two op selection
>    (counting) modes for IBS: cycles-counting and dispatched op
> counting.
>
>    Cycles-counting is _not_ equivalent to CPU_CLK_UNHALTED -- it is
>    not a precise version of the performance monitoring counter (PMC)
>    event (event select 0x076). In cycles-mode, when the current count
>    reaches the max count, the next available dispatch group of ops is
>    selected and a secondary mechanism selects an op within the dispatch
>    group. The dispatch group may contain one, two or three ops. If you
>    smell a rat, you're right. The secondary scheme negatively affects
>    the desired pseudo-random selection scheme. Also, if a dispatch
>    group is not available, the sample is skipped and the counting
>    process is reset.
>
>    Further, cycles-mode selection is affected by pipeline stalls. This
>    affects the distribution of IBS op samples taken in cycles-mode.
>    With cycles-mode, one instruction may have more data cache miss
> events,
>    but the underlying sampling basis is so skewed that the comparison
> is
>    not meaningful. IBS op samples are generated only for ops that
> retire;
>    tagged ops on a "wrong path" are flushed without producing a sample.
>    Overall, I cannot personally say that IBS cycles-mode produces a
> precise
>    equivalent to CPU_CLK_UNHALTED. I cannot endorse or recommend
>    its use in this way.
>
>    Given these issues, dispatched op counting was added in RevC. This
> mode
>    is the _preferred_ mode. Ops are counted as they are dispatched and
> the
>    op that triggers the max count threshold is selected and tagged.
>    Dispatched op mode produces a distribution of op samples that
> reflects
>    the execution frequency of instructions/basic blocks. DirectPath
>    Double and VectorPath (microcoded) x86 instructions which issue more
> than
>    one op will still be oversampled, however. The distribution is
> important
>    because it allows meaningful comparison of event counts between
>    instructions.
>
>    Even though the distribution of samples in dispatched op mode
> reflects
>    execution frequency, it is not a substitute for RETIRED_INSTRUCTIONS
>    (event select 0x0c0). The number of IBS op samples in some
> workloads,
>    especially those with certain kinds of stack access and microcoded
>    instructions, diverges greatly from RETIRED_INSTRUCTIONS.
>
>    IBS is what it is.
>
> IBS derived events
>
>    Since ProfileMe and Data EAR didn't exactly take the world by storm,
>    (oh, yeah, I worked with HP Caliper on Itanium for a while, too ;-),
>    profiling infrastructures like OProfile and CodeAnalyst are largely
>    based on the PMC sampling model.
>
>    In order to get IBS into practice as quickly as possible, we defined
>    IBS derived events. This allowed us to implement basic support for
>    IBS in both OProfile and CodeAnalyst without major changes in
>    infrastructure. I should note that translation from raw IBS bits to
>    derived events is and was always intended to be performed by user
>    space tools. I personally believe that translation should not be
>    performed in the kernel -- kernel support should be simple and
>    lightweight.
>
>    An IBS op sample is a small "packet" of profile data:
>
>        A bunch of event flags (data cache miss, etc.)
>        Tag-to-retire time (cycles)
>        Completion-to-retire (cycles)
>        DC miss latency (cycles)
>        DC miss addresses (64-bit virtual and physical addresses)
>
>    These entities can be used to compute latency distributions,
>    memory access maps, etc. IBS enables new kinds of analysis such
>    as data-centric profiling that identifies hot data regions (that
>    could be used to tune data layout in NUMA environment).
>
>    Quite frankly, at this juncture, I find the derived event model to
> be
>    too limiting. DCPI had a much different way of organizing ProfileMe
>    data that allowed flexible formulation of queries during
> post-processing --
>    something that cannot be done with the derived event approach.
>
>    Further, the organization and use of DC miss addresses is open for
>    investigation. I would _love_ to encourage someone (anyone? anyone?)
>    to take up this investigation. There may also be unforeseen uses --
>    perhaps driving compile-time optimizations. The existing derived
> events
>    do not adequately support new applications of IBS data. Thus, I
> would
>    encourage kernel-level support that passes IBS data along without
>    modification.
>
> Filtering.
>
>    After our initial experience with IBS, we see the need for
> filtering.
>    One approach is to collect and report only those IBS register values
>    that are needed to support a certain kind of analysis. For example,
>    if the DC miss addresses are not needed, why collect them? Suravee
>    and Robert Richter (both terrific colleagues) have been
> investigating
>    this, so I will defer to their analysis and comments.
>
> Software randomization.
>
>    We've found that software randomization of the sampling period
> and/or
>    current count is needed to avoid certain situations where the
> pipeline
>    and the sampling process get into a periodic hard-loop that affects
>    the distribution of IBS op samples. BTW, forcing those low order
> four
>    bits to zero occasionally has a negative effect on op distribution.
>
> IBS future extensions
>
>    Of course, I can't discuss specific new features. However, here are
>    some possible variations:
>
>       * The current count and max count values may become longer.
>       * New event flags may be added.
>       * Existing event flags may be left out (i.e., not implemented
>         in a family or model)
>       * New ancillary data (like DC miss latency or DC miss address)
>         may be added.
>
>    It may be necessary to collect new 64-bit values that do not contain
>    event flags, for example.
>
> Thanks for enduring this long-winded message. I hope that I've
> communicated some information and requirements, and I'll be more than
> happy to answer questions about IBS (or get the answers).
>
> -- pj
>
> Dr. Paul Drongowski
> AMD CodeAnalyst team
> Boston Design Center
>
> -------------------------
> The information presented in this reply is for informational purposes
> only and may contain technical inaccuracies, omissions and
> typographical errors. Links to third party sites are for convenience
> only, and no endorsement is implied.
>
>
>
>
> ------------------------------------------------------------------------------
> _______________________________________________
> perfmon2-devel mailing list
> perfmon2-devel@lists.sourceforge.net
> https://lists.sourceforge.net/lists/listinfo/perfmon2-devel
>

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [perfmon2] IV.3 - AMD IBS
  2009-06-23 14:25           ` stephane eranian
@ 2009-06-23 14:55             ` Ingo Molnar
  0 siblings, 0 replies; 9+ messages in thread
From: Ingo Molnar @ 2009-06-23 14:55 UTC (permalink / raw)
  To: eranian
  Cc: Peter Zijlstra, Rob Fowler, Philip Mucci, LKML, Andi Kleen,
	Paul Mackerras, Maynard Johnson, Andrew Morton, Thomas Gleixner,
	perfmon2-devel


* stephane eranian <eranian@googlemail.com> wrote:

> On Tue, Jun 23, 2009 at 4:05 PM, Ingo Molnar<mingo@elte.hu> wrote:
> >
> > * stephane eranian <eranian@googlemail.com> wrote:
> >
> >> > The most natural way to support IBS would be to have a special
> >> > sampling cycle counter and use that as group lead and add non
> >> > sampling siblings to that group to get individual elements.
> >> >
> >> As discussed in my message, I think the way to support IBS is to
> >> create two pseudo-events (like your perf_hw_event_ids), one for
> >> fetch and one for op (because they could be measured
> >> simultaneously). The sample_period field would be used to express
> >> the IBS*CTL maxcnt, subject to the verification that the bottom 4
> >> bits must be 0. And then, you add two new sampling formats
> >> PERF_SAMPLE_IBSFETCH, PERF_SAMPLE_IBSOP. Those would only work
> >> with IBS pseudo events. Once you have the randomize option in
> >> perf_counter_attr, you could even enable IBSFETCH randomization.
> >
> > I'd suggest to start smaller, and first express the 'precise' 
> > nature of IBS transparently, by simply mapping it to one of the 
> > generic events. (cycles and instructions both appears to be 
> > possible)
>
> IBS is precise by nature.

(yes. Did you understand my comments above as saying the opposite?)

> [...] It does not work like PEBS. It tags an instruction and then 
> collects info about it. When it retires, IBS freezes and triggers 
> an interrupt like a regular counter interrupt. Except this time, 
> you don't care about the interrupted IP, you use the instruction 
> address in the IBS data register, it is guaranteed to correspond 
> to the tagged instruction.
> 
> The sampling period expresses the delay before picking the 
> instruction to tag. And as I said before, it is only 20 bits and 
> the bottom 4 bits must be zero (as they cannot be encoded).

The 20 bits delay is in cycles, right? So this in itself still lends 
itself to be transparently provided as a PERF_COUNT_HW_CPU_CYCLES 
counter.

	Ingo

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [perfmon2] IV.3 - AMD IBS
  2009-06-23  6:19     ` Peter Zijlstra
  2009-06-23  8:19       ` stephane eranian
@ 2009-06-23 14:40       ` Rob Fowler
  1 sibling, 0 replies; 9+ messages in thread
From: Rob Fowler @ 2009-06-23 14:40 UTC (permalink / raw)
  To: Peter Zijlstra
  Cc: Ingo Molnar, eranian, Philip Mucci, LKML, Andi Kleen,
	Paul Mackerras, Maynard Johnson, Andrew Morton, Thomas Gleixner,
	perfmon2-devel

I'm up to my neck in other stuff, so this will be short.

Yes, IBS is a very different model of performance measurement that
doesn't fit well with the traditional model.  It does do what the
HW engineers need for understanding multi unit out of order processors, though.

The separation of the fetch and op monitoring is an artifact of the separation
and de-coupling of the front and back end pipelines.  The front end IBS events
deal with stuff that happens in fetching:  TLB and cache misses, mis-predictions, etc.
The back end IBS events deal with computing and data fetch/store.

There are two "conventional" counters involved: the tag-to-retire and completion-to-retire
counts.  These can be accumulated, histogrammed, etc just like any conventional event, though
you need to add the counter contents to the accumulator rather than just increment.

The rest of the bits are predicates that can be used to filter the events into bins.
With n bits, you might need 2^n bins to accumulate all possibilities.
Sections 5 and 6 of the AMD software optimization guide provide some useful boolean expressions
for defining meaningful derived events.

In an old version of Rice HPCToolkit (now disappeared from the web) we
had a tool called xprof that processed DEC DCPI/ProfileMe binary files to
produce profiles with ~20 derived events that we thought would be useful.  The cost
of collecting all of this didn't vary by the amount we collected, so you would select
the ones you wanted to view at analysis time, not at execute time.  There was also a
mechanism for specifying other events. Nathan Tallent can provide details.

The Linear and Physical Address registers are an opportunity for someone to build data profiling
tools, or a combined instructions and data tool.

The critical thing is for the kernel, driver, and library builders to not do something that will
stand in the way of this.

Peter Zijlstra wrote:
> On Mon, 2009-06-22 at 10:08 -0400, Rob Fowler wrote:
>> Ingo Molnar wrote:
>>>> 3/ AMD IBS
>>>>
>>>> How is AMD IBS going to be implemented?
>>>>
>>>> IBS has two separate sets of registers. One to capture fetch
>>>> related data and another one to capture instruction execution
>>>> data. For each, there is one config register but multiple data
>>>> registers. In each mode, there is a specific sampling period and
>>>> IBS can interrupt.
>>>>
>>>> It looks like you could define two pseudo events or event types
>>>> and then define a new record_format and read_format. That formats
>>>> would only be valid for an IBS event.
>>>>
>>>> Is that how you intend to support IBS?
>>> That is indeed one of the ways we thought of, not really nice, but
>>> then, IBS is really weird, what were those AMD engineers thinking
>>> :-)
>> Actually, IBS has roots in DEC's "ProfileMe" for Alpha EV67 and later
>> processors.   Those of us who used it there found it to be an extremely
>> powerful, low-overhead mechanism for directly collecting information about
>> how well the micro-architecture is performing.  In particular, it can tell
>> you, not only which instructions take a long time to traverse the pipe, but
>> it also tells you which instructions delay other instructions and by how much.
>> This is extremely valuable if you are either working on instruction scheduling
>> in a compiler, or are modifying a program to give the compiler the opportunity
>> to do a good job.
>>
>> A core group of engineers who worked on Alpha went on to AMD.
>>
>> An unfortunate problem with IBS on AMD is that good support isn't common in the "mainstream"
>> open source community.
> 
> The 'problem' I have with IBS is that its basically a cycle counter
> coupled with a pretty arbitrary number of output dimensions separated
> into two groups, ops and fetches.
> 
> This is a very weird configuration in that it has a miss-match with the
> traditional one value per counter thing.
> 
> The most natural way to support IBS would be to have a special sampling
> cycle counter and use that as group lead and add non sampling siblings
> to that group to get individual elements.
> 
> This is however quite cumbersome.
> 
> One thing to consider when building an IBS interface is its future
> extensibility. In which fashion would IBS be extended?, additional
> output dimensions or something else all-together?

-- 
Robert J. Fowler
Chief Domain Scientist, HPC
Renaissance Computing Institute
The University of North Carolina at Chapel Hill
100 Europa Dr, Suite 540
Chapel Hill, NC 27517
V: 919.445.9670
F: 919 445.9669
rjf@renci.org

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [perfmon2] IV.3 - AMD IBS
  2009-06-23 14:05         ` Ingo Molnar
@ 2009-06-23 14:25           ` stephane eranian
  2009-06-23 14:55             ` Ingo Molnar
  0 siblings, 1 reply; 9+ messages in thread
From: stephane eranian @ 2009-06-23 14:25 UTC (permalink / raw)
  To: Ingo Molnar
  Cc: Peter Zijlstra, Rob Fowler, Philip Mucci, LKML, Andi Kleen,
	Paul Mackerras, Maynard Johnson, Andrew Morton, Thomas Gleixner,
	perfmon2-devel

On Tue, Jun 23, 2009 at 4:05 PM, Ingo Molnar<mingo@elte.hu> wrote:
>
> * stephane eranian <eranian@googlemail.com> wrote:
>
>> > The most natural way to support IBS would be to have a special
>> > sampling cycle counter and use that as group lead and add non
>> > sampling siblings to that group to get individual elements.
>> >
>> As discussed in my message, I think the way to support IBS is to
>> create two pseudo-events (like your perf_hw_event_ids), one for
>> fetch and one for op (because they could be measured
>> simultaneously). The sample_period field would be used to express
>> the IBS*CTL maxcnt, subject to the verification that the bottom 4
>> bits must be 0. And then, you add two new sampling formats
>> PERF_SAMPLE_IBSFETCH, PERF_SAMPLE_IBSOP. Those would only work
>> with IBS pseudo events. Once you have the randomize option in
>> perf_counter_attr, you could even enable IBSFETCH randomization.
>
> I'd suggest to start smaller, and first express the 'precise' nature
> of IBS transparently, by simply mapping it to one of the generic
> events. (cycles and instructions both appears to be possible)
>
IBS is precise by nature. It does not work like PEBS. It tags an instruction
and then collects info about it. When it retires, IBS freezes and triggers an
interrupt like a regular counter interrupt. Except this time, you don't care
about the interrupted IP, you use the instruction address in the IBS data
register, it is guaranteed to correspond to the tagged instruction.

The sampling period expresses the delay before picking the instruction
to tag. And as I said before, it is only 20 bits and the bottom 4 bits must
be zero (as they cannot be encoded).



> No extra sampling, no extra events - just a transparent side channel
> implementation for the specific case of PERF_COUNT_HW_CPU_CYCLES. (A
> bit like the fixed-purpose counters are done on the Intel side - a
> special-case - but none of the generic code knows about it.)
>

> This gives us immediate results with less code, and also gives us
> the platform to see how IBS is structured, what kind of general
> problems/quirks it has, and how popular its precision is, etc. We
> can always add extra sampling formats on top of that (i'm not
> opposed to that), to expose more and more of IBS.
>
> The same can be done on the PEBS side as well.
>
> Would you be interested in pursuing this?
>
>        Ingo
>

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [perfmon2] IV.3 - AMD IBS
  2009-06-23  8:19       ` stephane eranian
@ 2009-06-23 14:05         ` Ingo Molnar
  2009-06-23 14:25           ` stephane eranian
  0 siblings, 1 reply; 9+ messages in thread
From: Ingo Molnar @ 2009-06-23 14:05 UTC (permalink / raw)
  To: eranian
  Cc: Peter Zijlstra, Rob Fowler, Philip Mucci, LKML, Andi Kleen,
	Paul Mackerras, Maynard Johnson, Andrew Morton, Thomas Gleixner,
	perfmon2-devel


* stephane eranian <eranian@googlemail.com> wrote:

> > The most natural way to support IBS would be to have a special 
> > sampling cycle counter and use that as group lead and add non 
> > sampling siblings to that group to get individual elements.
> >
> As discussed in my message, I think the way to support IBS is to 
> create two pseudo-events (like your perf_hw_event_ids), one for 
> fetch and one for op (because they could be measured 
> simultaneously). The sample_period field would be used to express 
> the IBS*CTL maxcnt, subject to the verification that the bottom 4 
> bits must be 0. And then, you add two new sampling formats 
> PERF_SAMPLE_IBSFETCH, PERF_SAMPLE_IBSOP. Those would only work 
> with IBS pseudo events. Once you have the randomize option in 
> perf_counter_attr, you could even enable IBSFETCH randomization.

I'd suggest to start smaller, and first express the 'precise' nature 
of IBS transparently, by simply mapping it to one of the generic 
events. (cycles and instructions both appears to be possible)

No extra sampling, no extra events - just a transparent side channel 
implementation for the specific case of PERF_COUNT_HW_CPU_CYCLES. (A 
bit like the fixed-purpose counters are done on the Intel side - a 
special-case - but none of the generic code knows about it.)

This gives us immediate results with less code, and also gives us 
the platform to see how IBS is structured, what kind of general 
problems/quirks it has, and how popular its precision is, etc. We 
can always add extra sampling formats on top of that (i'm not 
opposed to that), to expose more and more of IBS.

The same can be done on the PEBS side as well.

Would you be interested in pursuing this?

	Ingo

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [perfmon2] IV.3 - AMD IBS
  2009-06-23  6:19     ` Peter Zijlstra
@ 2009-06-23  8:19       ` stephane eranian
  2009-06-23 14:05         ` Ingo Molnar
  2009-06-23 14:40       ` Rob Fowler
  1 sibling, 1 reply; 9+ messages in thread
From: stephane eranian @ 2009-06-23  8:19 UTC (permalink / raw)
  To: Peter Zijlstra
  Cc: Rob Fowler, Ingo Molnar, Philip Mucci, LKML, Andi Kleen,
	Paul Mackerras, Maynard Johnson, Andrew Morton, Thomas Gleixner,
	perfmon2-devel

On Tue, Jun 23, 2009 at 8:19 AM, Peter Zijlstra<a.p.zijlstra@chello.nl> wrote:
>
> The 'problem' I have with IBS is that its basically a cycle counter
> coupled with a pretty arbitrary number of output dimensions separated
> into two groups, ops and fetches.
>
Well, that's your view. Mine is different.

You have 2 independent cycle counters (one for fetch, one for op),
each is coupled with a value. it is just that the value does not fit
into 64 bits. the cycle count is not hosted in a generic counter but
in its own register. The captured data for fetch is constructed with
IBSFETCHCTL, IBSFETCHLINAD, IBSFETCHPHYSAD. The 3
registers are tied together. The values they contain represent the
same fetch event. Same thing with IBS op. I don't see the problem
because your API is able to cope with variable length output data.
The sampling buffer certainly can. Of course, internally you'd have
to manage it in a special way, but you already do this for fixed
counters, don't you?

> This is a very weird configuration in that it has a miss-match with the
> traditional one value per counter thing.
>
This is not the universal model. I can give lots of examples on Itanium
where you have one config register and multiple data registers
to capture the event: branch trace buffer (1 config, 33 data), Data EAR
(cache/TLB miss sampling, 1 config 3 data),Instruction-EAR
(cache/TLB miss sampling, 1 config, 2 data).

> The most natural way to support IBS would be to have a special sampling
> cycle counter and use that as group lead and add non sampling siblings
> to that group to get individual elements.
>
As discussed in my message, I think the way to support IBS is to create two
pseudo-events (like your perf_hw_event_ids), one for fetch and one for op
(because they could be measured simultaneously). The sample_period field
would be used to express the IBS*CTL maxcnt, subject to the verification
that the bottom 4 bits must be 0. And then, you add  two new sampling formats
PERF_SAMPLE_IBSFETCH, PERF_SAMPLE_IBSOP. Those would only work
with IBS pseudo events. Once you have the randomize option in perf_counter_attr,
you could even enable IBSFETCH randomization.

What is wrong with this approach?

Another question is: how do you present the values contained in the IBS data
registers:
   1 - leave it as raw (tool parses the raw register values)
   2 - decode it in the kernel and expose your own format

With 1/, you'd pick up automatically new fields if AMD adds some.
With 2/, you'd have to change your format if AMD change theirs.



> This is however quite cumbersome.
>
> One thing to consider when building an IBS interface is its future
> extensibility. In which fashion would IBS be extended?, additional
> output dimensions or something else all-together?
>
I don't know but it would be nice to provide better filtering capabilities
but for that, they can use some of the reserved bits they have, no
need to add more data registers.

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [perfmon2] IV.3 - AMD IBS
  2009-06-22 14:08   ` [perfmon2] " Rob Fowler
  2009-06-22 17:58     ` Maynard Johnson
@ 2009-06-23  6:19     ` Peter Zijlstra
  2009-06-23  8:19       ` stephane eranian
  2009-06-23 14:40       ` Rob Fowler
  1 sibling, 2 replies; 9+ messages in thread
From: Peter Zijlstra @ 2009-06-23  6:19 UTC (permalink / raw)
  To: Rob Fowler
  Cc: Ingo Molnar, eranian, Philip Mucci, LKML, Andi Kleen,
	Paul Mackerras, Maynard Johnson, Andrew Morton, Thomas Gleixner,
	perfmon2-devel

On Mon, 2009-06-22 at 10:08 -0400, Rob Fowler wrote:
> Ingo Molnar wrote:
> >> 3/ AMD IBS
> >>
> >> How is AMD IBS going to be implemented?
> >>
> >> IBS has two separate sets of registers. One to capture fetch
> >> related data and another one to capture instruction execution
> >> data. For each, there is one config register but multiple data
> >> registers. In each mode, there is a specific sampling period and
> >> IBS can interrupt.
> >>
> >> It looks like you could define two pseudo events or event types
> >> and then define a new record_format and read_format. That formats
> >> would only be valid for an IBS event.
> >>
> >> Is that how you intend to support IBS?
> > 
> > That is indeed one of the ways we thought of, not really nice, but
> > then, IBS is really weird, what were those AMD engineers thinking
> > :-)
> 
> Actually, IBS has roots in DEC's "ProfileMe" for Alpha EV67 and later
> processors.   Those of us who used it there found it to be an extremely
> powerful, low-overhead mechanism for directly collecting information about
> how well the micro-architecture is performing.  In particular, it can tell
> you, not only which instructions take a long time to traverse the pipe, but
> it also tells you which instructions delay other instructions and by how much.
> This is extremely valuable if you are either working on instruction scheduling
> in a compiler, or are modifying a program to give the compiler the opportunity
> to do a good job.
> 
> A core group of engineers who worked on Alpha went on to AMD.
> 
> An unfortunate problem with IBS on AMD is that good support isn't common in the "mainstream"
> open source community.

The 'problem' I have with IBS is that its basically a cycle counter
coupled with a pretty arbitrary number of output dimensions separated
into two groups, ops and fetches.

This is a very weird configuration in that it has a miss-match with the
traditional one value per counter thing.

The most natural way to support IBS would be to have a special sampling
cycle counter and use that as group lead and add non sampling siblings
to that group to get individual elements.

This is however quite cumbersome.

One thing to consider when building an IBS interface is its future
extensibility. In which fashion would IBS be extended?, additional
output dimensions or something else all-together?


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [perfmon2] IV.3 - AMD IBS
  2009-06-22 14:08   ` [perfmon2] " Rob Fowler
@ 2009-06-22 17:58     ` Maynard Johnson
  2009-06-23  6:19     ` Peter Zijlstra
  1 sibling, 0 replies; 9+ messages in thread
From: Maynard Johnson @ 2009-06-22 17:58 UTC (permalink / raw)
  To: Rob Fowler
  Cc: Andrew Morton, Andi Kleen, Peter Zijlstra, eranian, LKML,
	Ingo Molnar, Philip Mucci, Paul Mackerras, perfmon2-devel,
	Thomas Gleixner, Suravee.Suthikulpanit

Rob Fowler <rjf@renci.org> wrote on 06/22/2009 09:08:34 AM:

>
>
> Ingo Molnar wrote:
> >> 3/ AMD IBS
> >>
> >> How is AMD IBS going to be implemented?
> >>
> >> IBS has two separate sets of registers. One to capture fetch
> >> related data and another one to capture instruction execution
> >> data. For each, there is one config register but multiple data
> >> registers. In each mode, there is a specific sampling period and
> >> IBS can interrupt.
> >>
> >> It looks like you could define two pseudo events or event types
> >> and then define a new record_format and read_format. That formats
> >> would only be valid for an IBS event.
> >>
> >> Is that how you intend to support IBS?
> >
> > That is indeed one of the ways we thought of, not really nice, but
> > then, IBS is really weird, what were those AMD engineers thinking
> > :-)
>
> Actually, IBS has roots in DEC's "ProfileMe" for Alpha EV67 and later
> processors.   Those of us who used it there found it to be an extremely
> powerful, low-overhead mechanism for directly collecting information
about
> how well the micro-architecture is performing.  In particular, it can
tell
> you, not only which instructions take a long time to traverse the pipe,
but
> it also tells you which instructions delay other instructions and byhow
much.
> This is extremely valuable if you are either working on instruction
scheduling
> in a compiler, or are modifying a program to give the compiler the
opportunity
> to do a good job.
>
> A core group of engineers who worked on Alpha went on to AMD.
>
> An unfortunate problem with IBS on AMD is that good support isn't
> common in the "mainstream"
> open source community.
>
> Code Analyst from  AMD, also involving ex-DEC engineers, is
> the only place where it is supported at all decently throughout the
> tool chain.
> Last  time I looked, there was a tweaked version of oprofile underlying
CA.
> I haven't checked to see whether the tweaks have migrated back into
> the oprofile trunk.
Yes, IBS on AMD is now supported upstream in mainline, contributed by
Suravee Suthikulpanit.

-Maynard
>
>
> >
> > The point is - weird hardware gets expressed as a ... weird feature
> > under perfcounters too. Not all hardware weirdnesses can be
> > engineered away.
> >
> >
>
------------------------------------------------------------------------------

> > Are you an open source citizen? Join us for the Open Source Bridge
> conference!
> > Portland, OR, June 17-19. Two days of sessions, one day of
> unconference: $250.
> > Need another reason to go? 24-hour hacker lounge. Register today!
> > http://ad.doubleclick.net/clk;215844324;13503038;v?http:
> //opensourcebridge.org
> > _______________________________________________
> > perfmon2-devel mailing list
> > perfmon2-devel@lists.sourceforge.net
> > https://lists.sourceforge.net/lists/listinfo/perfmon2-devel
>
> --
> Robert J. Fowler
> Chief Domain Scientist, HPC
> Renaissance Computing Institute
> The University of North Carolina at Chapel Hill
> 100 Europa Dr, Suite 540
> Chapel Hill, NC 27517
> V: 919.445.9670
> F: 919 445.9669
> rjf@renci.org


^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: [perfmon2] IV.3 - AMD IBS
  2009-06-22 12:00 ` IV.3 - AMD IBS Ingo Molnar
@ 2009-06-22 14:08   ` Rob Fowler
  2009-06-22 17:58     ` Maynard Johnson
  2009-06-23  6:19     ` Peter Zijlstra
  0 siblings, 2 replies; 9+ messages in thread
From: Rob Fowler @ 2009-06-22 14:08 UTC (permalink / raw)
  To: Ingo Molnar
  Cc: eranian, Peter Zijlstra, Philip Mucci, LKML, Andi Kleen,
	Paul Mackerras, Maynard Johnson, Andrew Morton, Thomas Gleixner,
	perfmon2-devel



Ingo Molnar wrote:
>> 3/ AMD IBS
>>
>> How is AMD IBS going to be implemented?
>>
>> IBS has two separate sets of registers. One to capture fetch
>> related data and another one to capture instruction execution
>> data. For each, there is one config register but multiple data
>> registers. In each mode, there is a specific sampling period and
>> IBS can interrupt.
>>
>> It looks like you could define two pseudo events or event types
>> and then define a new record_format and read_format. That formats
>> would only be valid for an IBS event.
>>
>> Is that how you intend to support IBS?
> 
> That is indeed one of the ways we thought of, not really nice, but
> then, IBS is really weird, what were those AMD engineers thinking
> :-)

Actually, IBS has roots in DEC's "ProfileMe" for Alpha EV67 and later
processors.   Those of us who used it there found it to be an extremely
powerful, low-overhead mechanism for directly collecting information about
how well the micro-architecture is performing.  In particular, it can tell
you, not only which instructions take a long time to traverse the pipe, but
it also tells you which instructions delay other instructions and by how much.
This is extremely valuable if you are either working on instruction scheduling
in a compiler, or are modifying a program to give the compiler the opportunity
to do a good job.

A core group of engineers who worked on Alpha went on to AMD.

An unfortunate problem with IBS on AMD is that good support isn't common in the "mainstream"
open source community.

Code Analyst from  AMD, also involving ex-DEC engineers, is
the only place where it is supported at all decently throughout the tool chain.
Last  time I looked, there was a tweaked version of oprofile underlying CA.
I haven't checked to see whether the tweaks have migrated back into the oprofile
trunk.


> 
> The point is - weird hardware gets expressed as a ... weird feature
> under perfcounters too. Not all hardware weirdnesses can be
> engineered away.
> 
> ------------------------------------------------------------------------------
> Are you an open source citizen? Join us for the Open Source Bridge conference!
> Portland, OR, June 17-19. Two days of sessions, one day of unconference: $250.
> Need another reason to go? 24-hour hacker lounge. Register today!
> http://ad.doubleclick.net/clk;215844324;13503038;v?http://opensourcebridge.org
> _______________________________________________
> perfmon2-devel mailing list
> perfmon2-devel@lists.sourceforge.net
> https://lists.sourceforge.net/lists/listinfo/perfmon2-devel

-- 
Robert J. Fowler
Chief Domain Scientist, HPC
Renaissance Computing Institute
The University of North Carolina at Chapel Hill
100 Europa Dr, Suite 540
Chapel Hill, NC 27517
V: 919.445.9670
F: 919 445.9669
rjf@renci.org

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2009-06-25 11:29 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2009-06-25 11:28 [perfmon2] IV.3 - AMD IBS stephane eranian
  -- strict thread matches above, loose matches on Subject: below --
2009-06-16 17:42 v2 of comments on Performance Counters for Linux (PCL) stephane eranian
2009-06-22 12:00 ` IV.3 - AMD IBS Ingo Molnar
2009-06-22 14:08   ` [perfmon2] " Rob Fowler
2009-06-22 17:58     ` Maynard Johnson
2009-06-23  6:19     ` Peter Zijlstra
2009-06-23  8:19       ` stephane eranian
2009-06-23 14:05         ` Ingo Molnar
2009-06-23 14:25           ` stephane eranian
2009-06-23 14:55             ` Ingo Molnar
2009-06-23 14:40       ` Rob Fowler

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

Powered by JetHome