From: Charles Samuels <charles@cariden.com>
To: "linux-kernel@vger.kernel.org" <linux-kernel@vger.kernel.org>
Subject: Queuing of disk writes
Date: Fri, 1 Apr 2011 12:59:53 -0700 [thread overview]
Message-ID: <201104011259.53936.charles@cariden.com> (raw)
Kernel hackers,
I have an application that is writing large amounts of very fragmented data to
harddrives. That is, I could write megabytes of data in blocks of a few bytes
scattered around a multi-gigabyte file.
Obviously, doing this causes the harddrive to seek a lot and takes a while.
>From what I understand, if I allow linux to cache the writes, it will fill up
the kernel's write cache, and then consequently the disk drive's DMA queue. As
a result of that, the harddrive can pick the correct order to do these writes,
significantly reducing seek times.
However, there's a major cost in allowing the write cache to fill: fsync takes
*ages*. What's worse is that while fsync is proceeding, it seems *all* disk
operations in the OS are blocked. This is really terrible for performance of
my application: my application might want to do some reads (i.e. from another
thread) from the disk preempting the fsync temporarily. It's also really
terrible for me, because then my workstation becomes unresponsive for several
minutes.
My general question is how to mitigate this. Is it possible to get a signal
for when a file is out of the disk cache. Or can I ask linux approximately how
much data is in the write queue for that specific file, and just do a sleep()-
loop checking until it goes down to something managable at which point I do
the fsync? Or, does aio support this scenario well, and if so, from what
version of Linux? (I've determined that there are some scenarios in which it
does, but it still requires O_DIRECT, apparently, which is weird considering
how I've heard Linux kernel hackers feel about that particular flag).
And yes, I *know* fsync is a poor method to determine if data is actually
committed to something non-volatile. :)
Thanks for the help,
Charles
next reply other threads:[~2011-04-01 20:05 UTC|newest]
Thread overview: 8+ messages / expand[flat|nested] mbox.gz Atom feed top
2011-04-01 19:59 Charles Samuels [this message]
2011-04-01 20:10 ` Alan Cox
2011-04-01 20:34 ` Charles Samuels
2011-04-01 20:39 ` Alan Cox
2011-04-04 2:02 ` Ted Ts'o
2011-04-04 17:50 ` Charles Samuels
2011-04-04 17:54 ` david
2011-04-05 19:37 ` Ted Ts'o
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=201104011259.53936.charles@cariden.com \
--to=charles@cariden.com \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
Powered by JetHome