From: Nick Piggin <nickpiggin@yahoo.com.au>
To: Federico Cuello <fedux@lugmen.org.ar>,
Ralf Hildebrandt <Ralf.Hildebrandt@charite.de>,
Artem Bityutskiy <Artem.Bityutskiy@nokia.com>
Cc: linux-kernel@vger.kernel.org
Subject: Re: sync-Regression in 2.6.28.2?
Date: Wed, 4 Feb 2009 17:17:26 +1100 [thread overview]
Message-ID: <200902041717.27320.nickpiggin@yahoo.com.au> (raw)
In-Reply-To: <4988A034.20406@lugmen.org.ar>
[-- Warning: decoded text below may be mangled, UTF-8 assumed --]
[-- Attachment #1: Type: text/plain; charset="utf-8", Size: 4237 bytes --]
On Wednesday 04 February 2009 06:51:16 Federico Cuello wrote:> Nick Piggin wrote:> > On Wednesday 28 January 2009 08:09:10 Federico Cuello wrote:> >> Ralf Hildebrandt escribió:> >>> I recently installed 2.6.28.2 on our postfix/dovecot-based> >>> mailboxserver. Previously, 2.6.28 and 2.6.28.1 have been running there> >>> without a hitch.> >>>> >>> Now with 2.6.28.2S I had two major lockups: All writes to the users'> >>> Maildirs (on ext4) would stall, the load would rise, "sync" would never> >>> return.> >>>> >>> I had to "reboot -f -n" to get the machine back. All hanging processes> >>> were unkillable, even with kill -9.> >>> [...]> >>> >> The same is happening to me, but I have some logs taken with sysrq.[...]> >> 0 2 99028 46536 26368 1569400 0 0 4 0 1407 387 0> >> 0 0 100> >>> >> Notice the 100% iowait.> >>> >> I also managed to reproduce it doing a rsync from one partition to a USB> >> drive. After the lockup I can't read any file from the source partition,> >> but the other partitions can be accessed normally.> >> > Hm, thanks for reporting, can you guys get a sysrq+W trace when the> > system reaches this state?>> Here it is, if you need something else please ask:
Thanks, could you reply-to-all when replying to retain ccs please?
Common theme is ext4, which uses no_nrwrite_index_update, and I introduceda bug in there which could possibly cause ext4 to go into a loop...
Would it be possible if you can test the following patch?
Thanks,Nick
---
Commit 05fe478dd04e02fa230c305ab9b5616669821dd3 introduced some@wbc->nr_to_write breakage. Here is the change from the commit:
>--- a/mm/page-writeback.c>+++ b/mm/page-writeback.c>@@ -963,8 +963,10 @@ retry:> }> }>> if (--nr_to_write <= 0)> done = 1;> if (wbc->sync_mode == WB_SYNC_NONE) {> if (--wbc->nr_to_write <= 0)> done = 1;> }> if (wbc->nonblocking && bdi_write_congested(bdi)) {> wbc->encountered_congestion = 1;> done = 1> }
It makes the following changes:1. Decrement wbc->nr_to_write instead of nr_to_write2. Decrement wbc->nr_to_write _only_ if wbc->sync_mode == WB_SYNC_NONE3. If synced nr_to_write pages, stop only if if wbc->sync_mode ==WB_SYNC_NONE, otherwise keep going.
However, according to the commit message, the intention was toonly make change 3. Change 1 is a bug. Change 2 does not seem to benecessary, and it breaks UBIFS expectations, so if needed, itshould be done separately later. And change 2 does not seem tobe documented in the commit message.
This patch does the following:1. Undo changes 1 and 22. Add a comment explaining change 3 (it very useful to have comments in_code_, not only in the commit).
Signed-off-by: Artem Bityutskiy <Artem.Bityutskiy@nokia.com>Acked-by: Nick Piggin <npiggin@suse.de>Cc: Andrew Morton <akpm@linux-foundation.org>--- mm/page-writeback.c | 21 +++++++++++++++------ 1 files changed, 15 insertions(+), 6 deletions(-)
diff --git a/mm/page-writeback.c b/mm/page-writeback.cindex b493db7..dc32dae 100644--- a/mm/page-writeback.c+++ b/mm/page-writeback.c@@ -1051,13 +1051,22 @@ continue_unlock: } } - if (wbc->sync_mode == WB_SYNC_NONE) {- wbc->nr_to_write--;- if (wbc->nr_to_write <= 0) {- done = 1;- break;- }+ if (nr_to_write > 0)+ nr_to_write--;+ else if (wbc->sync_mode == WB_SYNC_NONE) {+ /*+ * We stop writing back only if we are not+ * doing integrity sync. In case of integrity+ * sync we have to keep going because someone+ * may be concurrently dirtying pages, and we+ * might have synced a lot of newly appeared+ * dirty pages, but have not synced all of the+ * old dirty pages.+ */+ done = 1;+ break; }+ if (wbc->nonblocking && bdi_write_congested(bdi)) { wbc->encountered_congestion = 1; done = 1;\0ÿôèº{.nÇ+·®+%Ëÿ±éݶ\x17¥wÿº{.nÇ+·¥{±þG«éÿ{ayº\x1dÊÚë,j\a¢f£¢·hïêÿêçz_è®\x03(éÝ¢j"ú\x1a¶^[m§ÿÿ¾\a«þG«éÿ¢¸?¨èÚ&£ø§~á¶iOæ¬z·vØ^\x14\x04\x1a¶^[m§ÿÿÃ\fÿ¶ìÿ¢¸?I¥
next prev parent reply other threads:[~2009-02-04 6:18 UTC|newest]
Thread overview: 14+ messages / expand[flat|nested] mbox.gz Atom feed top
2009-01-27 9:35 Ralf Hildebrandt
2009-01-27 21:09 ` Federico Cuello
2009-02-03 1:09 ` Nick Piggin
2009-02-03 19:51 ` Federico Cuello
2009-02-04 6:17 ` Nick Piggin [this message]
2009-02-04 17:31 ` Federico Cuello
2009-02-05 3:25 ` Nick Piggin
2009-02-05 11:54 ` Federico Cuello
2009-02-09 13:45 ` Nick Piggin
2009-02-09 13:49 ` Ralf Hildebrandt
2009-02-15 13:42 ` Ralf Hildebrandt
2009-02-17 4:17 ` Nick Piggin
2009-02-17 15:15 ` Ralf Hildebrandt
2009-02-05 10:19 ` Ralf Hildebrandt
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=200902041717.27320.nickpiggin@yahoo.com.au \
--to=nickpiggin@yahoo.com.au \
--cc=Artem.Bityutskiy@nokia.com \
--cc=Ralf.Hildebrandt@charite.de \
--cc=fedux@lugmen.org.ar \
--cc=linux-kernel@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®