mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Alexei Starovoitov <alexei.starovoitov@gmail.com>
To: Wang Nan <wangnan0@huawei.com>
Cc: Alexei Starovoitov <ast@kernel.org>,
	Arnaldo Carvalho de Melo <acme@redhat.com>,
	Peter Zijlstra <peterz@infradead.org>,
	linux-kernel@vger.kernel.org,
	Brendan Gregg <brendan.d.gregg@gmail.com>,
	He Kuang <hekuang@huawei.com>, Jiri Olsa <jolsa@kernel.org>,
	Masami Hiramatsu <masami.hiramatsu.pt@hitachi.com>,
	Namhyung Kim <namhyung@kernel.org>,
	pi3orama@163.com, Zefan Li <lizefan@huawei.com>
Subject: Re: [PATCH 4/4] perf core: Add backward attribute to perf event
Date: Mon, 28 Mar 2016 17:28:56 -0700	[thread overview]
Message-ID: <20160329002855.GC31198@ast-mbp.thefacebook.com> (raw)
In-Reply-To: <1459147292-239310-5-git-send-email-wangnan0@huawei.com>

On Mon, Mar 28, 2016 at 06:41:32AM +0000, Wang Nan wrote:
> This patch introduces 'write_backward' bit to perf_event_attr, which
> controls the direction of a ring buffer. After set, the corresponding
> ring buffer is written from end to beginning. This feature is design to
> support reading from overwritable ring buffer.
> 
> Ring buffer can be created by mapping a perf event fd. Kernel puts event
> records into ring buffer, user programs like perf fetch them from
> address returned by mmap(). To prevent racing between kernel and perf,
> they communicate to each other through 'head' and 'tail' pointers.
> Kernel maintains 'head' pointer, points it to the next free area (tail
> of the last record). Perf maintains 'tail' pointer, points it to the
> tail of last consumed record (record has already been fetched). Kernel
> determines the available space in a ring buffer using these two
> pointers, prevents to overwrite unfetched records.
> 
> By mapping without 'PROT_WRITE', an overwritable ring buffer is created.
> Different from normal ring buffer, perf is unable to maintain 'tail'
> pointer because writing is forbidden. Therefore, for this type of ring
> buffers, kernel overwrite old records unconditionally, works like flight
> recorder. This feature would be useful if reading from overwritable ring
> buffer were as easy as reading from normal ring buffer. However,
> there's an obscure problem.
> 
> The following figure demonstrates the state of an overwritable ring
> buffer which is nearly full. In this figure, the 'head' pointer points
> to the end of last record, and a long record 'E' is pending. For a
> normal ring buffer, a 'tail' pointer would have pointed to position (X),
> so kernel knows there's no more space in the ring buffer. However, for
> an overwritable ring buffer, kernel doesn't care the 'tail' pointer.
> 
>    (X)                              head
>     .                                |
>     .                                V
>     +------+-------+----------+------+---+
>     |A....A|B.....B|C........C|D....D|   |
>     +------+-------+----------+------+---+
> 
> After writing record 'E', record 'A' is overwritten.
> 
>       head
>        |
>        V
>     +--+---+-------+----------+------+---+
>     |.E|..A|B.....B|C........C|D....D|E..|
>     +--+---+-------+----------+------+---+
> 
> Now perf decides to read from this ring buffer. However, none of the
> the two natural positions, 'head' and the start of this ring buffer,
> are pointing to the head of a record. Even perf can read the full ring
> buffer, it is unable to find the position to start decoding.
> 
> The first attempt tries to solve this problem AFAIK can be found from
> [1]. It makes kernel to maintain 'tail' pointer: updates it when ring
> buffer is half full. However, this approach introduces overhead to
> fast path. Test result shows a 1% overhead [2]. In addition, this method
> utilizes no more tham 50% records.
> 
> Another attempt can be found from [3], which allow putting the size of
> an event at the end of each record. This approach allows perf to find
> records in a backword manner from 'head' pointer by reading size of a
> record from its tail. However, because of alignment requirement, it
> needs 8 bytes to record the size of a record, which is a huge waste. Its
> performance is also not good, because more data need to be written.
> This approach also introduces some extra branch instructions to fast
> path.
> 
> 'write_backward' is a better solution to this problem.
> 
> Following figure demonstrates the state of the overwritable ring buffer
> when 'write_backward' is set before overwriting:
> 
>        head
>         |
>         V
>     +---+------+----------+-------+------+
>     |   |D....D|C........C|B.....B|A....A|
>     +---+------+----------+-------+------+
> 
> and after overwriting:
>                                      head
>                                       |
>                                       V
>     +---+------+----------+-------+---+--+
>     |..E|D....D|C........C|B.....B|A..|E.|
>     +---+------+----------+-------+---+--+
> 
> In each situation, 'head' points to the beginning of the newest record.
> From this record, perf can iterate over the full ring buffer, fetching
> as mush records as possible one by one.
> 
> The only limitation needs to consider is back-to-back reading. Due to
> the non-deterministic of user program, it is impossible to ensure the
> ring buffer keeps stable during reading. Consider an extreme situation:
> perf is scheduled out after reading record 'D', then a burst of events
> come, eat up the whole ring buffer (one or multiple rounds), but 'head'
> pointer happends to be at the same position when perf comes back.
> Continue reading after 'D' is incorrect now.
> 
> To prevent this problem, we need to find a way to ensure the ring buffer
> is stable during reading. ioctl(PERF_EVENT_IOC_PAUSE_OUTPUT) is
> suggested because its overhead is lower than
> ioctl(PERF_EVENT_IOC_ENABLE).
> 
> This patch utilizes event's default overflow_handler introduced
> previously. perf_event_output_backward() is created as the default
> overflow handler for backward ring buffers. To avoid extra overhead to
> fast path, original perf_event_output() becomes __perf_event_output()
> and marked '__always_inline'. In theory, there's no extra overhead
> introduced to fast path.
> 
> Performance result:
> 
> Calling 3000000 times of 'close(-1)', use gettimeofday() to check
> duration.  Use 'perf record -o /dev/null -e raw_syscalls:*' to capture
> system calls. In ns.
> 
> Testing environment:
> 
>  CPU    : Intel(R) Core(TM) i7-4790 CPU @ 3.60GHz
>  Kernel : v4.5.0
>                    MEAN         STDVAR
> BASE            800214.950    2853.083
> PRE1           2253846.700    9997.014
> PRE2           2257495.540    8516.293
> POST           2250896.100    8933.921
> 
> Where 'BASE' is pure performance without capturing. 'PRE1' is test
> result of pure 'v4.5.0' kernel. 'PRE2' is test result before this
> patch. 'POST' is test result after this patch. See [4] for detail
> experimental setup.
> 
> Considering the stdvar, this patch doesn't introduce performance
> overhead to fast path.
> 
> [1] http://lkml.iu.edu/hypermail/linux/kernel/1304.1/04584.html
> [2] http://lkml.iu.edu/hypermail/linux/kernel/1307.1/00535.html
> [3] http://lkml.iu.edu/hypermail/linux/kernel/1512.0/01265.html
> [4] http://lkml.kernel.org/g/56F89DCD.1040202@huawei.com
> 
> Signed-off-by: Wang Nan <wangnan0@huawei.com>
> Cc: He Kuang <hekuang@huawei.com>
> Cc: Alexei Starovoitov <ast@kernel.org>
> Cc: Arnaldo Carvalho de Melo <acme@redhat.com>
> Cc: Brendan Gregg <brendan.d.gregg@gmail.com>
> Cc: Jiri Olsa <jolsa@kernel.org>
> Cc: Masami Hiramatsu <masami.hiramatsu.pt@hitachi.com>
> Cc: Namhyung Kim <namhyung@kernel.org>
> Cc: Peter Zijlstra <peterz@infradead.org>
> Cc: Zefan Li <lizefan@huawei.com>
> Cc: pi3orama@163.com
> ---
>  include/linux/perf_event.h      | 28 +++++++++++++++++++++---
>  include/uapi/linux/perf_event.h |  3 ++-
>  kernel/events/core.c            | 48 ++++++++++++++++++++++++++++++++++++-----
>  kernel/events/ring_buffer.c     | 14 ++++++++++++
>  4 files changed, 84 insertions(+), 9 deletions(-)

Very useful feature. Looks good.
Acked-by: Alexei Starovoitov <ast@kernel.org>

  parent reply	other threads:[~2016-03-29  0:29 UTC|newest]

Thread overview: 32+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2016-03-28  6:41 [PATCH 0/4] perf core: Support reading from overwritable ring buffer Wang Nan
2016-03-28  6:41 ` [PATCH 1/4] perf core: Introduce new ioctl options to pause and resume " Wang Nan
2016-03-28 10:15   ` [PATCH][manpages 1/2] perf_event_open.2: Document PERF_EVENT_IOC_PAUSE_OUTPUT Wang Nan
2016-10-21  8:56     ` Michael Kerrisk (man-pages)
2016-10-21 14:37       ` Vince Weaver
2016-10-21 14:49         ` Michael Kerrisk (man-pages)
2016-03-29  0:27   ` [PATCH 1/4] perf core: Introduce new ioctl options to pause and resume ring buffer Alexei Starovoitov
2016-03-29  1:10     ` Wangnan (F)
2016-03-29  2:05     ` [PATCH 1/4 fix] " Wang Nan
2016-03-29  4:39       ` Alexei Starovoitov
2016-03-29 12:54   ` [PATCH 1/4] " Peter Zijlstra
2016-03-29 12:55     ` Peter Zijlstra
2016-03-30  1:57     ` Wangnan (F)
2016-03-30  6:46       ` Peter Zijlstra
2016-03-31  9:26   ` [tip:perf/core] perf/ring_buffer: Introduce new ioctl options to pause and resume the ring-buffer tip-bot for Wang Nan
2016-03-28  6:41 ` [PATCH 2/4] perf core: Set event's default overflow_handler Wang Nan
2016-03-31  9:26   ` [tip:perf/core] perf/core: Set event's default ::overflow_handler() tip-bot for Wang Nan
2016-03-28  6:41 ` [PATCH 3/4] perf core: Prepare writing into ring buffer from end Wang Nan
2016-03-29  0:25   ` Alexei Starovoitov
2016-03-31  9:26   ` [tip:perf/core] perf/ring_buffer: Prepare writing into the ring-buffer from the end tip-bot for Wang Nan
2016-03-28  6:41 ` [PATCH 4/4] perf core: Add backward attribute to perf event Wang Nan
2016-03-28 10:16   ` [PATCH][manpages 2/2] perf_event_open.2: Document write_backward Wang Nan
2016-10-21  8:57     ` Michael Kerrisk (man-pages)
2016-03-29  0:28   ` Alexei Starovoitov [this message]
2016-03-29  2:01   ` [PATCH 4/4] perf core: Add backward attribute to perf event Wangnan (F)
2016-03-29  4:59     ` Alexei Starovoitov
2016-03-29  5:59       ` Wangnan (F)
2016-03-29 14:04   ` Peter Zijlstra
2016-03-30  2:28     ` Wangnan (F)
2016-03-30  2:38       ` Wangnan (F)
2016-04-05 14:05         ` Wangnan (F)
2016-04-07  9:45     ` Wangnan (F)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20160329002855.GC31198@ast-mbp.thefacebook.com \
    --to=alexei.starovoitov@gmail.com \
    --cc=acme@redhat.com \
    --cc=ast@kernel.org \
    --cc=brendan.d.gregg@gmail.com \
    --cc=hekuang@huawei.com \
    --cc=jolsa@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=lizefan@huawei.com \
    --cc=masami.hiramatsu.pt@hitachi.com \
    --cc=namhyung@kernel.org \
    --cc=peterz@infradead.org \
    --cc=pi3orama@163.com \
    --cc=wangnan0@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®