From: Kent Overstreet <koverstreet@google.com>
To: Tejun Heo <tj@kernel.org>
Cc: Andrew Morton <akpm@linux-foundation.org>,
linux-kernel@vger.kernel.org, Oleg Nesterov <oleg@redhat.com>,
Christoph Lameter <cl@linux-foundation.org>,
Ingo Molnar <mingo@redhat.com>, Andi Kleen <andi@firstfloor.org>,
Jens Axboe <axboe@kernel.dk>,
"Nicholas A. Bellinger" <nab@linux-iscsi.org>,
Jeff Layton <jlayton@redhat.com>,
"J. Bruce Fields" <bfields@fieldses.org>
Subject: Re: [PATCH] Percpu tag allocator
Date: Thu, 13 Jun 2013 14:53:55 -0700 [thread overview]
Message-ID: <20130613215355.GA28664@moria.home.lan> (raw)
In-Reply-To: <20130613185318.GB12075@mtj.dyndns.org>
On Thu, Jun 13, 2013 at 11:53:18AM -0700, Tejun Heo wrote:
> Hello, Andrew, Kent.
>
> (cc'ing NFS folks for id[r|a] discussion)
>
> On Wed, Jun 12, 2013 at 08:03:11PM -0700, Andrew Morton wrote:
> > They all sound like pretty crappy reasons ;) If the idr/ida interface
> > is nasty then it can be wrapped to provide the same interface as the
> > percpu tag allocator.
> >
> > I could understand performance being an issue, but diligence demands
> > that we test that, or at least provide a convincing argument.
>
> The thing is that id[r|a] guarantee that the lowest available slot is
> allocated and this is important because it's used to name things which
> are visible to userland - things like block device minor number,
> device indicies and so on. That alone pretty much ensures that
> alloc/free paths can't be very scalable which usually is fine for most
> id[r|a] use cases as long as lookup is fast. I'm doubtful that it's a
> good idea to push per-cpu tag allocation into id[r|a]. The use cases
> are quite different.
>
> In fact, maybe what we can do is adding some features on top of the
> tag allocator and moving id[r|a] users which don't require strict
> in-order allocation to it. For example, NFS allocates an ID for each
> transaction it performs and uses it to index the associate command
> structure (Jeff, Bruce, please correct me if I'm getting it wrong).
> The only requirement on IDs is that they shouldn't be recycled too
> fast. Currently, idr implements cyclic mode for it but it can easily
> be replaced with per-cpu tag allocator like this one and it'd be a lot
> more scalable. There are a couple things to worry about tho - it
> probably should use the highbits as generation number as a tag is
> given out so that the actual ID doesn't get recycled quickly, and some
> form dynamic tag sizing would be nice too.
Yeah, that sounds like a perfect use.
Using the high bits as a gen number - that's something I've done before
in driver code, and that can be done completely outside the tag
allocator - no need for a cyclic mode.
For dynamic sizing, the issue is not so much dynamically sizing the tag
allocator's data structures - the tag allocator itself will use a
fraction of the memory of your tag structs - it's that you want to do
something slightly more intelligent than preallocating one giant array
of tag structs.
I already ran into this in the aio code - kiocbs are just big enough
that we don't want to preallocate them all when we allocate the kioctx.
I did the simplest thing I could think of for the aio code, but if other
users are going to be running into this too maybe it should be made
generic too.
Anyways, for aio I just use an array of pages for the kiocbs instead of
a flat array, and then the pages are allocated lazily.
http://evilpiepirate.org/git/linux-bcache.git/commit/?h=aio&id=999e7718f6b7ec99512fd576b166e5d63cd45ef2
Since the tag allocator uses stacks, it'll tend to give out ids that
were previously allocated and this should work pretty well in practice.
The one caveat right now is that if the workload is shifting across
cpus, tags being stranded on percpu freelists would cause us to allocate
pages sooner than we probably want to.
I don't think this is a big issue because the tag stealing is done based
on the worst case number of stranded tags - but I think I can improve it
with a bit of lazyness...
next prev parent reply other threads:[~2013-06-13 21:53 UTC|newest]
Thread overview: 29+ messages / expand[flat|nested] mbox.gz Atom feed top
2013-06-12 4:03 Kent Overstreet
2013-06-12 17:08 ` Oleg Nesterov
2013-06-12 17:59 ` Kent Overstreet
2013-06-12 19:14 ` Oleg Nesterov
2013-06-12 23:38 ` Andrew Morton
2013-06-13 2:05 ` Kent Overstreet
2013-06-13 3:03 ` Andrew Morton
2013-06-13 3:54 ` Kent Overstreet
2013-06-13 5:46 ` Andrew Morton
2013-06-13 18:53 ` Tejun Heo
2013-06-13 19:04 ` Andrew Morton
2013-06-13 19:15 ` Tejun Heo
2013-06-13 19:23 ` Andrew Morton
2013-06-13 19:35 ` Tejun Heo
2013-06-13 22:10 ` Andrew Morton
2013-06-13 22:30 ` Tejun Heo
2013-06-13 22:35 ` Andrew Morton
2013-06-13 23:13 ` Tejun Heo
2013-06-13 23:23 ` Tejun Heo
2013-06-19 1:32 ` Kent Overstreet
2013-06-13 19:08 ` J. Bruce Fields
2013-06-13 19:09 ` Jeff Layton
2013-06-13 21:53 ` Kent Overstreet [this message]
2013-06-13 19:06 ` Tejun Heo
2013-06-13 19:13 ` Andrew Morton
2013-06-13 19:21 ` Tejun Heo
2013-06-13 21:14 ` Kent Overstreet
2013-06-13 21:50 ` Tejun Heo
2013-06-13 21:58 ` Kent Overstreet
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20130613215355.GA28664@moria.home.lan \
--to=koverstreet@google.com \
--cc=akpm@linux-foundation.org \
--cc=andi@firstfloor.org \
--cc=axboe@kernel.dk \
--cc=bfields@fieldses.org \
--cc=cl@linux-foundation.org \
--cc=jlayton@redhat.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mingo@redhat.com \
--cc=nab@linux-iscsi.org \
--cc=oleg@redhat.com \
--cc=tj@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®