From: Hugh Dickins <hugh@veritas.com>
To: Christoph Lameter <clameter@sgi.com>
Cc: James Bottomley <James.Bottomley@HansenPartnership.com>,
Linus Torvalds <torvalds@linux-foundation.org>,
Andrew Morton <akpm@linux-foundation.org>,
FUJITA Tomonori <fujita.tomonori@lab.ntt.co.jp>,
Jens Axboe <jens.axboe@oracle.com>,
Pekka Enberg <penberg@cs.helsinki.fi>,
Peter Zijlstra <a.p.zijlstra@chello.nl>,
"Rafael J. Wysocki" <rjw@sisk.pl>,
linux-kernel@vger.kernel.org
Subject: Re: [PATCH] scsi: fix sense_slab/bio swapping livelock
Date: Mon, 7 Apr 2008 20:40:07 +0100 (BST) [thread overview]
Message-ID: <Pine.LNX.4.64.0804072000410.18591@blonde.site> (raw)
In-Reply-To: <Pine.LNX.4.64.0804062213230.18148@schroedinger.engr.sgi.com>
On Sun, 6 Apr 2008, Christoph Lameter wrote:
> On Sun, 6 Apr 2008, Hugh Dickins wrote:
> >
> > One very significant factor is SLUB, which
> > merges slab caches when it can, and on 64-bit happens to merge
> > both bio cache and sense_slab cache into kmalloc's 128-byte cache:
> > so that under this swapping load, bios above are liable to gobble
> > up all the slots needed for scsi_cmnd sense_buffers below.
>
> A reliance on free slots that the slab allocator may provide? That is a
> rather bad dependency since it is up to the slab allocator to implement
> the storage layout for the objects and thus the availability of slots may
> vary depending on the layout for the objects chosen by the allocator.
I'm not sure that I understand you. Yes, different slab allocators
may lay out slots differently. But a significant departure from
existing behaviour may be a bad idea in some circumstances.
(Hmm, maybe I've written a content-free sentence there!).
>
> Looking at mempool_alloc: Mempools may be used to do atomic allocations
> until they fail thereby exhausting reserves and available object in the
> partial lists of slab caches?
Mempools may be used for atomic allocations, but I think that's not
the case here. swap_writepage's get_swap_bio says GFP_NOIO, which
allows (indeed is) __GFP_WAIT, and does not give access to __GFP_HIGH
reserves.
Whereas at the __scsi_get_command end, there are GFP_ATOMIC sense_slab
allocations, which do give access to __GFP_HIGH reserves.
My supposition is that once a page has been allocated from __GFP_HIGH
reserves to a scsi sense_slab, swap_writepages are liable to gobble up
the rest of the page with bio allocations which they wouldn't have had
access to traditionally (i.e. under SLAB).
So an unexpected behaviour emerges from SLUB's slab merging.
Though of course the same might happen in other circumstances, even
without slab merging: if some kmem_cache allocations are made with
GFP_ATOMIC, those can give access to reserves to non-__GFP_HIGH
allocations from the same kmem_cache.
Maybe PF_MEMALLOC and __GFP_NOMEMALLOC complicate the situation:
I've given little thought to mempool_alloc's fiddling with the
gfp_mask (beyond repeatedly misreading it).
>
> In order to make this a significant factor we need to have already
> exhausted reserves right? Thus we are already operating at the boundary of
> what memory there is. Any non atomic alloc will then allocate a new page
> with N elements in order to get one object. The mempool_allocs from the
> atomic context will then gooble up the N-1 remaining objects? So the
> nonatomic alloc will then have to hit the page allocator again...
We need to have already exhausted reserves, yes: so this isn't an
issue hitting everyone all the time, and it may be nothing worse
than a surprising anomaly; but I'm pretty sure it's not how bio
and scsi command allocation is expected to interact.
What do you think a SLAB_NOMERGE flag? The last time I suggested
something like that (but I was thinking of debug), your comment
was "Ohh..", which left me in some doubt ;)
If we had a SLAB_NOMERGE flag, would we want to apply it to the
bio cache or to the scsi_sense_cache or to both? My difficulty
in answering that makes me wonder whether such a flag is right.
Hugh
next prev parent reply other threads:[~2008-04-07 19:36 UTC|newest]
Thread overview: 30+ messages / expand[flat|nested] mbox.gz Atom feed top
2008-04-06 22:56 Hugh Dickins
2008-04-06 23:35 ` James Bottomley
2008-04-07 1:01 ` Hugh Dickins
2008-04-07 17:51 ` Hugh Dickins
2008-04-07 18:04 ` James Bottomley
2008-04-07 18:26 ` Hugh Dickins
2008-04-07 2:48 ` FUJITA Tomonori
2008-04-07 18:07 ` Hugh Dickins
2008-04-08 14:04 ` FUJITA Tomonori
2008-04-07 5:26 ` Christoph Lameter
2008-04-07 19:40 ` Hugh Dickins [this message]
2008-04-07 19:55 ` Peter Zijlstra
2008-04-07 20:31 ` Hugh Dickins
2008-04-07 20:47 ` Peter Zijlstra
2008-04-07 21:00 ` Pekka Enberg
2008-04-07 21:05 ` Pekka Enberg
2008-04-07 21:15 ` Linus Torvalds
2008-04-07 21:34 ` Pekka Enberg
2008-04-07 21:39 ` Pekka Enberg
2008-04-07 22:05 ` Pekka J Enberg
2008-04-07 22:17 ` Linus Torvalds
2008-04-07 22:42 ` Pekka Enberg
2008-04-08 20:42 ` Pekka J Enberg
2008-04-08 20:44 ` Pekka Enberg
2008-04-08 20:45 ` Christoph Lameter
2008-04-08 21:11 ` Pekka Enberg
2008-04-08 21:40 ` Peter Zijlstra
2008-04-07 21:30 ` Hugh Dickins
2008-04-07 21:36 ` Pekka Enberg
2008-04-08 20:43 ` Christoph Lameter
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=Pine.LNX.4.64.0804072000410.18591@blonde.site \
--to=hugh@veritas.com \
--cc=James.Bottomley@HansenPartnership.com \
--cc=a.p.zijlstra@chello.nl \
--cc=akpm@linux-foundation.org \
--cc=clameter@sgi.com \
--cc=fujita.tomonori@lab.ntt.co.jp \
--cc=jens.axboe@oracle.com \
--cc=linux-kernel@vger.kernel.org \
--cc=penberg@cs.helsinki.fi \
--cc=rjw@sisk.pl \
--cc=torvalds@linux-foundation.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®