mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Bodo Eggert <7eggert@gmx.de>
To: Theodore Tso <tytso@mit.edu>, Xu CanHao <xucanhao@gmail.com>,
	linux-kernel@vger.kernel.org
Subject: Re: Ext3 vs NTFS performance
Date: Sun, 06 May 2007 00:25:45 +0200	[thread overview]
Message-ID: <E1HkShF-0000zR-Ez@be1.lrz> (raw)
In-Reply-To: <8huGm-2W4-33@gated-at.bofh.it>

Theodore Tso <tytso@mit.edu> wrote:

> But as has already been discussed on this thread, in situations where
> the fileserver is under high memory pressure, any filesystem (XFS or
> ext4) would still end up allocating blocks out of order, resulting in
> fragmentation.  Explicit preallocation, as opposed to delayed
> allocation, is really the best long-term solution; and in order to do
> that, Samba needs to detect this scenario --- which as has been noted,
> there appears to be no good reason for the Windows CIFS client (or any
> other application)to be doing this, other than perhaps to deliberate
> trigger a worst case allocation pattern in ext3 --- and translate it
> into a explicit preallocation request.

There is an interface to tell the kernel about the way the file will be
accessed. IMO this interface should be used to do the preallocation, too.

The other question is: How to tell the poor-bill's preallocation from a
very clever application that communicates with another application and
which is supposed to zero out that exact byte from the data the other
application sent. I was tempted to say "just let samba cache these calls",
but it would be wrong. You'll need magic in the kernel to DTRT.

There are three correct ways of handling these one-zerobyte-writes after EOF:

1) Extend the file like truncate
2) Extend the file like write() (current behaviour)
3) Preallocate these blocks (to be implemented)
4) Write all zeroes (current behaviour for FAT)

(2) will cause bad allocations, it's obviously worse than (1). (3) would be
better than (1) and (2), but only xfs(?) and ext4 will support this in the
near future. (4) should double the write time, but give the best possible
read speed. According to [1], the expected read speed is about as high as (1)
gives, "playback performance improves to expected levels". If preallocation
does not seem to make a big difference, I don't think we should do (4) as
a replacement untill the filesystem does support real preallocations.


I suggest:

1) Make samba use fadvise(MIGHT_PREALLOCATE)
2) Make the kernel turn these 1-byte-writes-after-EOF into truncates
   on MIGHT_PREALLOCATE, and possibly turn off MIGHT_PREALLOCATE on
   other read/writes
3) Make the kernel fadvise(PREALLOCATE, $filesize)
   on MIGHT_PREALLOCATE + lseek(0), turning off the MIGHT_PREALLOCATE
   Possibly it might also turn on FADV_SEQUENTIAL.
4) Make the filesystems optionally preallocate the desired area, or
   ignore fadvise(PREALLOCATE, $filesize) instead.


[1] http://softwarecommunity.intel.com/articles/eng/1259.htm
-- 
It is still called paranoia when they really are out to get you.

Friß, Spammer: oA@cvb2dX.7eggert.dyndns.org
 CZCkzfiaNb@7eggert.dyndns.org nkp@7eggert.dyndns.org

       reply	other threads:[~2007-05-05 22:25 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
     [not found] <8hiYr-2fJ-1@gated-at.bofh.it>
     [not found] ` <8huGm-2W4-33@gated-at.bofh.it>
2007-05-05 22:25   ` Bodo Eggert [this message]
2007-05-06  5:04     ` Xu CanHao
     [not found] <8gShI-3hY-11@gated-at.bofh.it>
     [not found] ` <8h1bh-8sG-11@gated-at.bofh.it>
     [not found]   ` <8h2Al-280-1@gated-at.bofh.it>
     [not found]     ` <8hW9y-2Lp-3@gated-at.bofh.it>
2007-05-07 11:21       ` Bodo Eggert
2007-05-06  1:48 Albert Cahalan
  -- strict thread matches above, loose matches on Subject: below --
2007-05-05  3:13 Xu CanHao
2007-05-05 13:45 ` Theodore Tso
2007-05-03  3:51 Al Boldi
2007-05-01 20:43 Cabot, Mason B
2007-05-01 21:23 ` Andrew Morton
2007-05-02 12:21   ` Andi Kleen
2007-05-02 16:04     ` Theodore Tso
2007-05-02 18:40       ` Andi Kleen
2007-05-02 19:28         ` Theodore Tso
2007-05-02 16:16   ` Theodore Tso
2007-05-02 18:08     ` Jeremy Allison
2007-05-02 19:34       ` Theodore Tso
2007-05-02 20:38         ` Jeff Garzik
2007-05-02 22:01           ` Theodore Tso
2007-05-02  3:54 ` Gerhard Mack
2007-05-02 15:46   ` David Chinner
2007-05-02 15:44 ` David Chinner
2007-05-02 19:46   ` Chris Mason
2007-05-03  0:15     ` David Chinner
2007-05-03 12:57       ` Chris Mason
2007-05-03 21:14   ` Valerie Henson
2007-05-03 22:40     ` Bernd Eckenfels
2007-05-04  8:12       ` Anton Altaparmakov
2007-05-04  9:46         ` Christoph Hellwig
2007-05-04 14:47           ` Anton Altaparmakov
2007-05-04 15:49           ` Michael Tokarev
2007-05-04 18:41             ` Theodore Tso
2007-05-05  9:59             ` Christoph Hellwig
2007-05-06 20:59           ` Jörn Engel
2007-05-04 12:23     ` Theodore Tso
2007-05-04 19:40       ` Valerie Henson
2007-05-04 18:56 ` Phillip Susi
2007-05-04 19:52   ` Cabot, Mason B
2007-05-07 14:31     ` Phillip Susi

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=E1HkShF-0000zR-Ez@be1.lrz \
    --to=7eggert@gmx.de \
    --cc=linux-kernel@vger.kernel.org \
    --cc=tytso@mit.edu \
    --cc=xucanhao@gmail.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®