From: Andrew Morton <akpm@osdl.org>
To: Linus Torvalds <torvalds@osdl.org>
Cc: elenstev@mesatop.com, lm@bitmover.com, wli@holomorphy.com,
hugh@veritas.com, adi@bitmover.com, scole@lanl.gov,
support@bitmover.com, linux-kernel@vger.kernel.org
Subject: Re: 1352 NUL bytes at the end of a page? (was Re: Assertion `s && s->tree' failed: The saga continues.)
Date: Sun, 16 May 2004 23:11:20 -0700 [thread overview]
Message-ID: <20040516231120.405a0d14.akpm@osdl.org> (raw)
In-Reply-To: <Pine.LNX.4.58.0405162152290.25502@ppc970.osdl.org>
Linus Torvalds <torvalds@osdl.org> wrote:
>
> Andrew, the obvious culprit would be the memset() in fs/buffer.c
> (block_write_full_page() to be precise):
>
> memset(kaddr + offset, 0, PAGE_CACHE_SIZE - offset);
>
> imagine that the "write()" function updates i_size late - after having
> written out the new contents to the page, and _after_ havign unlocked the
> page, and now we get a writeback at the wrong time, and we decide to clear
> out the end of the page because we think it's past i_size.
>
> Andrew, what do you think?
Interesting. Playing with i_size like that in writepage() _is_ scary. My
immediate reaction is that if this race was real, it's so gross that we
would have spotted it before now in either 2.4 or 2.5->2.6.
Easy test: Steve, could you remove that memset from
block_write_full_page(), see if it changes anything?
It's not very important - it's there because if an application
(incorrectly) writes to mapped data outside EOF we're supposed to drop
their data and write zeroes instead.
> I think this race does exist, since generic_file_aio_write_nolock()
> literally _does_ update i_size only after it has written all the pages, so
> I don't see why a "block_write_full_page()" couldn't come in there between
> and zero them out again at the _old_ i_size boundary.
i_size is updated in generic_commit_write(), on a per-page basis, or I'm
missing something? I sure hope so.
Let's go through the scenarios.
On entry to block_write_full_page(), i_size is in the middle of this page
somewhere. We're worried that i_size can change, and that this will cause
block_write_full_page() to incorrectly zero out the tail of the page.
Well we can stop right there, because the only way someone can get some
more non-zero user data into this page before we memset and write it is by
locking the page beforehand, and block_write_full_page() has the page lock.
(Or they can write stuff into it via mmap, but writing to the page outside
i_size is an application bug).
Other ways in which i_size can change under block_write_full_page()'s feet
are:
- Someone did a truncate.
No problem - the page is about to be invalidated and chopped off the
file anyway.
- Someone did an extending truncate into another page.
OK, i_size will increase but we're still supposed to write zeroes into
the rest of this page outside the previous i_size.
- Someone extended the file into another page with lseek+write or pwrite.
Same argumentation as with extending truncate.
- Someone did an extending truncate to another i_size which lands
*within* this page.
Writing zeroes is still OK: nobody can get into this page to write new
user data anyway - it's locked.
Either all that, or I missed something ;) If Steve can try that test it
would be interesting. Even if removing the memset does make the corruption
go away, this might not be a kernel bug - it could be that the application
is incorrectly relying on mmapped writes outside i_size making it to disk.
As for O_DIRECT: I need to think about that a bit more. We hold i_sem and
have done an fdatasync prior to entering generic_file_aio_write_nolock() so
there should be no dirty pagecache at this stage anyway. The VM may decide
to dirty some pagecache and then block_write_full_page() could come in and
look at i_size and race against generic_file_aio_write_nolock()'s O_DIRECT
i_size_write(). But I doubt if bk is using direct-IO in combination with
MAP_SHARED...
next prev parent reply other threads:[~2004-05-17 6:12 UTC|newest]
Thread overview: 94+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <200405122234.06902.elenstev@mesatop.com>
[not found] ` <15594C37-A509-11D8-A7EA-000A95CC3A8A@lanl.gov>
[not found] ` <20040513183316.GE17965@bitmover.com>
[not found] ` <200405131723.15752.elenstev@mesatop.com>
[not found] ` <6616858C-A5AF-11D8-A7EA-000A95CC3A8A@lanl.gov>
[not found] ` <20040514144617.GE20197@work.bitmover.com>
2004-05-14 4:32 ` 1352 NUL bytes at the end of a page? Steven Cole
2004-05-14 16:53 ` 1352 NUL bytes at the end of a page? (was Re: Assertion `s && s->tree' failed: The saga continues.) Andy Isaacson
2004-05-14 17:23 ` Steven Cole
2004-05-15 0:54 ` Steven Cole
2004-05-15 1:55 ` 1352 NUL bytes at the end of a page? Wayne Scott
2004-05-15 3:15 ` 1352 NUL bytes at the end of a page? (was Re: Assertion `s && s->tree' failed: The saga continues.) Lincoln Dale
2004-05-15 3:41 ` Andrew Morton
2004-05-15 5:39 ` Steven Cole
2004-05-16 1:23 ` Steven Cole
2004-05-16 2:18 ` Linus Torvalds
2004-05-16 3:44 ` Linus Torvalds
2004-05-16 4:31 ` Steven Cole
2004-05-16 4:52 ` Linus Torvalds
2004-05-16 5:22 ` Andrea Arcangeli
2004-05-16 15:28 ` Steven Cole
2004-05-16 17:49 ` Rutger Nijlunsing
2004-05-16 20:38 ` Andrea Arcangeli
2004-05-16 21:19 ` Steven Cole
2004-05-16 21:29 ` Andrew Morton
2004-05-16 22:11 ` Steven Cole
2004-05-16 23:53 ` Andrea Arcangeli
2004-05-17 2:12 ` Steven Cole
2004-05-17 8:21 ` R. J. Wysocki
2004-05-16 5:54 ` Steven Cole
2004-05-16 6:09 ` Andrew Morton
2004-05-16 6:24 ` Andrew Morton
2004-05-16 10:01 ` Andrew Morton
2004-05-16 13:49 ` Steven Cole
2004-05-18 1:47 ` Benjamin Herrenschmidt
2004-05-16 3:20 ` Andrew Morton
2004-05-16 3:58 ` Linus Torvalds
2004-05-17 2:28 ` Larry McVoy
2004-05-17 2:42 ` Linus Torvalds
2004-05-17 3:36 ` Steven Cole
2004-05-17 5:17 ` Linus Torvalds
2004-05-17 6:11 ` Andrew Morton [this message]
2004-05-17 13:56 ` 1352 NUL bytes at the end of a page? Wayne Scott
2004-05-17 15:17 ` Theodore Ts'o
2004-05-17 15:20 ` Larry McVoy
2004-05-17 15:22 ` Linus Torvalds
2004-05-17 15:25 ` Larry McVoy
2004-05-17 15:37 ` viro
2004-05-17 17:30 ` Steven Cole
2004-05-17 17:40 ` viro
2004-05-17 17:39 ` Steven Cole
2004-05-17 19:06 ` viro
2004-05-17 15:40 ` Arjan van de Ven
2004-05-17 15:53 ` Steven Cole
2004-05-17 16:23 ` Davide Libenzi
2004-05-17 16:28 ` Davide Libenzi
2004-05-17 14:07 ` 1352 NUL bytes at the end of a page? (was Re: Assertion `s && s->tree' failed: The saga continues.) Larry McVoy
2004-05-17 14:12 ` Linus Torvalds
2004-05-17 7:25 ` Andrew Morton
2004-05-17 7:46 ` Andrew Morton
2004-05-17 8:39 ` Vladimir Saveliev
2004-05-17 8:44 ` Andrew Morton
2004-05-17 11:58 ` Steven Cole
2004-05-17 14:05 ` Larry McVoy
2004-05-17 14:14 ` Larry McVoy
2004-05-17 14:32 ` Linus Torvalds
2004-05-17 14:52 ` Larry McVoy
2004-05-17 15:02 ` Linus Torvalds
2004-05-17 15:05 ` Larry McVoy
2004-05-17 15:23 ` Chris Mason
2004-05-17 15:49 ` Steven Cole
2004-05-17 20:24 ` Chris Mason
2004-05-17 21:08 ` Steven Cole
2004-05-17 21:29 ` Andrew Morton
2004-05-17 22:15 ` Steven Cole
2004-05-17 23:52 ` Steven Cole
2004-05-18 0:03 ` Chris Mason
2004-05-18 0:15 ` Andrew Morton
2004-05-18 0:13 ` Andrew Morton
2004-05-18 0:45 ` Steven Cole
2004-05-18 1:34 ` Larry McVoy
2004-05-18 1:42 ` Andrew Morton
2004-05-18 1:56 ` Steven Cole
2004-05-17 14:11 ` Larry McVoy
[not found] ` <200405172142.52780.elenstev@mesatop.com>
[not found] ` <Pine.LNX.4.58.0405172056480.25502@ppc970.osdl.org>
[not found] ` <200405172319.38853.elenstev@mesatop.com>
2004-05-18 12:42 ` Chris Mason
2004-05-18 14:29 ` Steven Cole
2004-05-18 14:38 ` Linus Torvalds
2004-05-19 10:53 ` Steven Cole
2004-05-19 12:10 ` Chris Mason
2004-05-19 12:20 ` 1352 NUL bytes at the end of a page? Wayne Scott
2004-05-19 12:42 ` Nick Piggin
2004-05-19 13:28 ` Steven Cole
2004-05-19 13:36 ` Chris Mason
2004-05-19 13:59 ` Steven Cole
2004-05-19 14:03 ` Wayne Scott
2004-05-19 14:08 ` Chris Mason
2004-05-19 14:20 ` Steven Cole
2004-05-19 14:45 ` Steven Cole
2004-05-19 21:11 ` Non-regression for current kernel (was Re: 1352 NUL bytes at the end of a page?) Steven Cole
2004-05-17 15:35 1352 NUL bytes at the end of a page? (was Re: Assertion `s && s->tree' failed: The saga continues.) Albert Cahalan
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20040516231120.405a0d14.akpm@osdl.org \
--to=akpm@osdl.org \
--cc=adi@bitmover.com \
--cc=elenstev@mesatop.com \
--cc=hugh@veritas.com \
--cc=linux-kernel@vger.kernel.org \
--cc=lm@bitmover.com \
--cc=scole@lanl.gov \
--cc=support@bitmover.com \
--cc=torvalds@osdl.org \
--cc=wli@holomorphy.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®