mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Hsiu-Hsien Lee <swinds24@gmail.com>
To: Theodore Ts'o <tytso@mit.edu>, linux-ext4@vger.kernel.org
Cc: Andreas Dilger <adilger.kernel@dilger.ca>,
	Jan Kara <jack@suse.cz>,
	Harshad Shirwadkar <harshadshirwadkar@gmail.com>,
	Baokun Li <libaokun@linux.alibaba.com>,
	Ojaswin Mujoo <ojaswin@linux.ibm.com>,
	Ritesh Harjani <ritesh.list@gmail.com>,
	Zhang Yi <yi.zhang@huawei.com>,
	linux-kernel@vger.kernel.org, Hsiu-Hsien Lee <swinds24@gmail.com>
Subject: [PATCH] ext4: don't fail journal recovery on an incomplete fast commit
Date: Tue,  6 Oct 2026 15:00:05 +0800	[thread overview]
Message-ID: <20261006070005.1209234-1-swinds24@gmail.com> (raw)

If the system crashes while the first fast commit after a full commit is
being written, the fast commit area can hold the head of that fast
commit without a valid tail.  ext4_fc_replay_scan() treats this as an
error as long as no valid tail has been seen: an invalid tag length or
an unknown tag returns -ECANCELED, a tail with a wrong tid or checksum
returns -EFSBADCRC.  jbd2_journal_recover() then fails, the file system
can be mounted neither read-write nor read-only, and the regular journal
transactions committed before the fast commit are not replayed either:

  JBD2: journal recovery failed
  EXT4-fs (sdb): error loading journal

The tail is written last and fsync() does not return before it has
completed, so an incomplete fast commit was never reported as durable.
Handle it like jbd2 handles a transaction without a valid commit block:
stop the scan and replay what is valid.  This is already what happens
when an earlier fast commit in the area has a valid tail.  Log a
warning so that the dropped fast commit is visible.

Reproducer, with the power cut emulated by copying the device while it
is mounted:

  mkfs.ext4 -O fast_commit /dev/sdb
  mount /dev/sdb /mnt; mkdir /mnt/d
  create 20 files in /mnt/d; sync
  write a fragmented file and modify the 20 files, fsync one file
      (one fast commit spanning several blocks)
  copy /dev/sdb while still mounted, then zero the block holding the
  fast commit tail in the copy
  mount the copy

Without this patch the mount fails as above.  With it, the mount
succeeds, the state after sync is recovered and e2fsck finds no errors.
The same holds when a full commit precedes the torn fast commit, in
which case the full commit is now replayed as well.

Fixes: 8016e29f4362 ("ext4: fast commit recovery path")
Assisted-by: Claude Opus 5.5
Signed-off-by: Hsiu-Hsien Lee <swinds24@gmail.com>
---
 fs/ext4/fast_commit.c | 36 +++++++++++++++++++++++++++++-------
 1 file changed, 29 insertions(+), 7 deletions(-)

diff --git a/fs/ext4/fast_commit.c b/fs/ext4/fast_commit.c
index 0cac890cf370..5649447b9524 100644
--- a/fs/ext4/fast_commit.c
+++ b/fs/ext4/fast_commit.c
@@ -2427,6 +2427,26 @@ static bool ext4_fc_value_len_isvalid(struct ext4_sb_info *sbi,
 	return false;
 }
 
+/*
+ * The fast commit area ends with an invalid tag or a tail that does not
+ * match.  This is how a fast commit that was only partially written before
+ * a crash looks like: its tail is written last and fsync() only returns
+ * after the tail write has completed, so nobody was told that this fast
+ * commit is durable and it can simply be dropped, the same way jbd2 drops
+ * a transaction without a valid commit block.  Everything up to the last
+ * valid tail is still replayed.
+ */
+static int ext4_fc_replay_scan_end(struct super_block *sb,
+				   struct ext4_fc_replay_state *state,
+				   int off, const char *reason)
+{
+	if (!state->fc_replay_num_tags)
+		ext4_msg(sb, KERN_WARNING,
+			 "ignoring incomplete fast commit at block %d: %s",
+			 off, reason);
+	return JBD2_FC_REPLAY_STOP;
+}
+
 /*
  * Recovery Scan phase handler
  *
@@ -2440,7 +2460,9 @@ static bool ext4_fc_value_len_isvalid(struct ext4_sb_info *sbi,
  * This function returns JBD2_FC_REPLAY_CONTINUE to indicate that SCAN is
  * incomplete and JBD2 should send more blocks. It returns JBD2_FC_REPLAY_STOP
  * to indicate that scan has finished and JBD2 can now start replay phase.
- * It returns a negative error to indicate that there was an error. At the end
+ * It returns a negative error to indicate that there was an error. An invalid
+ * or incomplete fast commit at the end of the area is not an error, see
+ * ext4_fc_replay_scan_end(). At the end
  * of a successful scan phase, sbi->s_fc_replay_state.fc_replay_num_tags is set
  * to indicate the number of tags that need to replayed during the replay phase.
  */
@@ -2489,8 +2511,8 @@ static int ext4_fc_replay_scan(journal_t *journal,
 		val = cur + EXT4_FC_TAG_BASE_LEN;
 		if (tl.fc_len > end - val ||
 		    !ext4_fc_value_len_isvalid(sbi, tl.fc_tag, tl.fc_len)) {
-			ret = state->fc_replay_num_tags ?
-				JBD2_FC_REPLAY_STOP : -ECANCELED;
+			ret = ext4_fc_replay_scan_end(sb, state, off,
+						      "invalid tag length");
 			goto out_err;
 		}
 		ext4_debug("Scan phase, tag:%s, blk %lld\n",
@@ -2530,8 +2552,8 @@ static int ext4_fc_replay_scan(journal_t *journal,
 				state->fc_regions_valid =
 					state->fc_regions_used;
 			} else {
-				ret = state->fc_replay_num_tags ?
-					JBD2_FC_REPLAY_STOP : -EFSBADCRC;
+				ret = ext4_fc_replay_scan_end(sb, state, off,
+							      "invalid tail");
 			}
 			state->fc_crc = 0;
 			break;
@@ -2551,8 +2573,8 @@ static int ext4_fc_replay_scan(journal_t *journal,
 				EXT4_FC_TAG_BASE_LEN + tl.fc_len);
 			break;
 		default:
-			ret = state->fc_replay_num_tags ?
-				JBD2_FC_REPLAY_STOP : -ECANCELED;
+			ret = ext4_fc_replay_scan_end(sb, state, off,
+						      "unknown tag");
 		}
 		if (ret < 0 || ret == JBD2_FC_REPLAY_STOP)
 			break;
-- 
2.43.0


             reply	other threads:[~2026-10-06  7:00 UTC|newest]

Thread overview: 2+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-06  7:00 Hsiu-Hsien Lee [this message]
2026-10-06 16:07 ` Jan Kara

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261006070005.1209234-1-swinds24@gmail.com \
    --to=swinds24@gmail.com \
    --cc=adilger.kernel@dilger.ca \
    --cc=harshadshirwadkar@gmail.com \
    --cc=jack@suse.cz \
    --cc=libaokun@linux.alibaba.com \
    --cc=linux-ext4@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=ojaswin@linux.ibm.com \
    --cc=ritesh.list@gmail.com \
    --cc=tytso@mit.edu \
    --cc=yi.zhang@huawei.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®