From: "Stephen C. Tweedie" <sct@redhat.com>
To: Bartlomiej Zolnierkiewicz <B.Zolnierkiewicz@elka.pw.edu.pl>
Cc: Chris Mason <mason@suse.com>, Jens Axboe <axboe@suse.de>,
Jeff Garzik <jgarzik@pobox.com>,
Linux Kernel <linux-kernel@vger.kernel.org>,
Stephen Tweedie <sct@redhat.com>
Subject: Re: [PATCH] barrier patch set
Date: 30 Mar 2004 17:04:46 +0100 [thread overview]
Message-ID: <1080662685.1978.25.camel@sisko.scot.redhat.com> (raw)
In-Reply-To: <200403201805.26211.bzolnier@elka.pw.edu.pl>
Hi,
On Sat, 2004-03-20 at 17:05, Bartlomiej Zolnierkiewicz wrote:
> Jens, can you explain how this translates to the block layer?
> If "log blocks" is a separate request from "commit block",
> we can just do: log blocks, flush, commit block, flush cycle.
It's not so simple. There are two different operations here as far as
the fs is concerned --- ordering ("barrier"), and waiting ("flush").
For most journaling purposes, a filesystem really doesn't have to know
when data has hit disk --- it only needs to be assured that certain
ordering constraints will be obeyed. So in SCSI terms we can submit the
log blocks, then set an ORDERED queue tag on the commit block, and leave
it at that. We don't have to wait until that ORDERED tag actually
becomes persistent on disk --- we don't care, in most cases. This is
the "barrier", and in journaling, we basically need a barrier before and
after the commit block.
The tricky bit is that _sometimes_ we find, later on, that we need to
know when data has hit disk. If an application requests synchronised IO
(such as fsync, O_SYNC or O_DIRECT) then we cannot legally return until
the disk has told us that the data is persistent for certain. _That_ is
a full flush, and it's distinct from the simple ordering constraint that
we normally have for journal commit blocks.
The distinction becomes important if flushing and barriers have
significantly different performance characteristics. In the SCSI tagged
command case, we can insert an ORDERED queue tag into the pipeline
without having to wait for anything. But if we implement the barrier
and the flush by the same mechanism, then the fs ends up waiting; we get
write log blocks
wait for IO queue to empty
write commit block
wait for IO queue to empty
instead of
write log blocks
write commit block with ORDERED tag set
so the commits are miles slower.
Jens' patch happens to implement barriers via flushes on IDE, but at the
block layer blkdev_issue_flush() and set_buffer_ordered() are quite
distinct APIs.
--Stephen
next prev parent reply other threads:[~2004-03-30 16:07 UTC|newest]
Thread overview: 63+ messages / expand[flat|nested] mbox.gz Atom feed top
2004-03-19 15:35 Jens Axboe
2004-03-19 16:30 ` Mika Penttilä
2004-03-19 18:16 ` Jens Axboe
2004-03-19 18:44 ` Mika Penttilä
2004-03-20 9:55 ` Jens Axboe
2004-03-19 16:34 ` Jeff Garzik
2004-03-19 18:19 ` Jens Axboe
2004-03-19 23:01 ` Matthias Andree
2004-03-20 0:02 ` Bartlomiej Zolnierkiewicz
2004-03-20 1:48 ` Johannes Stezenbach
2004-03-20 2:13 ` Bartlomiej Zolnierkiewicz
2004-03-20 2:53 ` Johannes Stezenbach
2004-03-20 16:03 ` Bartlomiej Zolnierkiewicz
2004-03-20 11:36 ` Matthias Andree
2004-03-20 16:00 ` Bartlomiej Zolnierkiewicz
2004-03-20 23:36 ` Johannes Stezenbach
2004-03-21 1:33 ` Bartlomiej Zolnierkiewicz
2004-03-20 18:52 ` Helge Hafting
2004-03-22 11:15 ` Matthias Andree
2004-03-19 23:59 ` Bartlomiej Zolnierkiewicz
2004-03-20 0:14 ` Jeff Garzik
2004-03-20 0:40 ` Bartlomiej Zolnierkiewicz
2004-03-20 0:42 ` Jeff Garzik
2004-03-20 1:24 ` Bartlomiej Zolnierkiewicz
2004-03-20 9:58 ` Jens Axboe
2004-03-20 10:12 ` Jeff Garzik
2004-03-20 10:19 ` Jens Axboe
2004-03-20 10:37 ` Jeff Garzik
2004-03-20 16:30 ` Bartlomiej Zolnierkiewicz
2004-03-21 18:12 ` Jeff Garzik
2004-03-20 10:21 ` Jeff Garzik
2004-03-20 15:54 ` Bartlomiej Zolnierkiewicz
2004-03-20 0:17 ` Jeff Garzik
2004-03-20 9:53 ` Jens Axboe
2004-03-20 16:23 ` Bartlomiej Zolnierkiewicz
2004-03-20 16:27 ` Jens Axboe
2004-03-20 16:32 ` Chris Mason
2004-03-20 17:05 ` Bartlomiej Zolnierkiewicz
2004-03-20 17:10 ` Chris Mason
2004-03-20 20:16 ` Bartlomiej Zolnierkiewicz
2004-03-21 9:43 ` Jens Axboe
2004-03-30 16:04 ` Stephen C. Tweedie [this message]
2004-03-30 19:19 ` Chris Mason
2004-03-30 21:50 ` Stephen C. Tweedie
2004-03-30 22:13 ` Chris Mason
2004-03-31 14:03 ` Stephen C. Tweedie
2004-03-31 14:27 ` Chris Mason
2004-03-31 18:28 ` Ric Wheeler
2004-03-30 22:21 ` Jeff Garzik
2004-03-30 22:36 ` Chris Wedgwood
2004-03-30 22:39 ` Jeff Garzik
2004-03-30 22:41 ` Chris Wedgwood
2004-03-30 22:40 ` Bartlomiej Zolnierkiewicz
2004-03-30 22:38 ` Jeff Garzik
2004-03-31 14:08 ` Stephen C. Tweedie
2004-03-31 14:21 ` Chris Mason
2004-03-31 21:26 ` Jeff Garzik
2004-03-31 22:09 ` Chris Mason
2004-03-31 21:27 ` Jeff Garzik
2004-03-19 16:48 ` Marc-Christian Petersen
2004-03-19 18:19 ` Jens Axboe
2004-03-22 11:09 ` Andrew Morton
2004-03-22 11:10 ` Jens Axboe
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=1080662685.1978.25.camel@sisko.scot.redhat.com \
--to=sct@redhat.com \
--cc=B.Zolnierkiewicz@elka.pw.edu.pl \
--cc=axboe@suse.de \
--cc=jgarzik@pobox.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mason@suse.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®