From: Linus Torvalds <torvalds@linux-foundation.org>
To: tytso@mit.edu
Cc: Jan Kara <jack@suse.cz>, Jens Axboe <jens.axboe@oracle.com>,
Linux Kernel <linux-kernel@vger.kernel.org>,
jengelh@medozas.de, stable@kernel.org, gregkh@suse.de
Subject: Re: [PATCH] writeback: Fix broken sync writeback
Date: Tue, 16 Feb 2010 21:16:46 -0800 (PST) [thread overview]
Message-ID: <alpine.LFD.2.00.1002162052230.4141@localhost.localdomain> (raw)
In-Reply-To: <20100217043009.GZ5337@thunk.org>
On Tue, 16 Feb 2010, tytso@mit.edu wrote:
>
> We've had this logic for a long time, and given the increase in disk
> density, and spindle speeds, the 4MB limit, which might have made
> sense 10 years ago, probably doesn't make sense now.
I still don't think that 4MB is enough on its own to suck quite that
much. Even a fast device should be perfectly happy with 4MB IOs, or it
must be sucking really badly.
In order to see the kinds of problems that got quoted in the original
thread, there must be something else going on too, methinks (disk light
was "blinking").
So I would guess that it's also getting stuck on that
inode_wait_for_writeback(inode);
inside that loop in wb_writeback().
In fact, I'm starting to wonder about that "Nothing written" case. The
code basically decides that "if I wrote zero pages, I didn't write
anything at all, so I must wait for the inode to complete old writes in
order to not busy-loop". Which sounds sensible on the face of it, but the
thing is, inodes can be dirty without actually having any dirty _pages_
associated with them.
Are we perhaps ending up in a situation where we essentially wait
synchronously on just the inode itself being written out? That would
explain the "40kB/s" kind of behavior.
If we were actually doing real 4MB chunks, that would _not_ explain 40kB/s
throughput.
But if we do a 4MB chunk (for the one file that had real dirty data in
it), and then do a few hundred trivial "write out the inode data
_synchronously_" (due to access time changes etc) in between until we hit
the file that has real dirty data again - now _that_ would explain 40kB/s
throughput. It's not just seeking around - it's not even trying to push
multiple IO's to get any elevator going or anything like that.
And then the patch that started this discussion makes sense: it improves
performance because in between those synchronous inode updates it now
writes big chunks. But again, it's mostly hiding us just doing insane
things.
I dunno. Just a theory. The more I look at that code, the uglier it looks.
And I do get the feeling that the "4MB chunking" is really just making the
more fundamental problems show up.
Linus
next prev parent reply other threads:[~2010-02-17 5:17 UTC|newest]
Thread overview: 39+ messages / expand[flat|nested] mbox.gz Atom feed top
2010-02-12 9:16 Jens Axboe
2010-02-12 15:45 ` Linus Torvalds
2010-02-13 12:58 ` Jan Engelhardt
2010-02-15 14:49 ` Jan Kara
2010-02-15 15:41 ` Jan Engelhardt
2010-02-15 15:58 ` Jan Kara
2010-06-27 16:44 ` Jan Engelhardt
2010-10-24 23:41 ` Sync writeback still broken Jan Engelhardt
2010-10-30 0:57 ` Linus Torvalds
2010-10-30 1:16 ` Linus Torvalds
2010-10-30 1:30 ` Linus Torvalds
2010-10-30 3:18 ` Andrew Morton
2010-10-30 13:15 ` Christoph Hellwig
2010-10-31 12:24 ` Jan Kara
2010-10-31 22:40 ` Jan Kara
2010-11-05 21:33 ` Jan Kara
2010-11-05 21:34 ` Jan Kara
2010-11-05 21:41 ` Linus Torvalds
2010-11-05 22:03 ` Jan Engelhardt
2010-11-07 12:57 ` Jan Kara
2011-01-20 22:50 ` Jan Engelhardt
2011-01-21 15:09 ` Jan Kara
2010-02-15 14:17 ` [PATCH] writeback: Fix broken sync writeback Jan Kara
2010-02-16 0:05 ` Linus Torvalds
2010-02-16 23:00 ` Jan Kara
2010-02-16 23:34 ` Linus Torvalds
2010-02-17 0:01 ` Linus Torvalds
2010-02-17 1:33 ` Jan Kara
2010-02-17 1:57 ` Dave Chinner
2010-02-17 3:35 ` Linus Torvalds
2010-02-17 4:30 ` tytso
2010-02-17 5:16 ` Linus Torvalds [this message]
2010-02-22 17:29 ` Jan Kara
2010-02-22 21:01 ` tytso
2010-02-22 22:26 ` Jan Kara
2010-02-23 2:53 ` Dave Chinner
2010-02-23 3:23 ` tytso
2010-02-23 5:53 ` Dave Chinner
2010-02-24 14:56 ` Jan Kara
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=alpine.LFD.2.00.1002162052230.4141@localhost.localdomain \
--to=torvalds@linux-foundation.org \
--cc=gregkh@suse.de \
--cc=jack@suse.cz \
--cc=jengelh@medozas.de \
--cc=jens.axboe@oracle.com \
--cc=linux-kernel@vger.kernel.org \
--cc=stable@kernel.org \
--cc=tytso@mit.edu \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome