From: Siddh Raman Pant <siddh.raman.pant@oracle.com>
To: stable@vger.kernel.org, Greg KH <gregkh@linuxfoundation.org>,
"Darrick J . Wong" <djwong@kernel.org>,
Dave Chinner <dchinner@redhat.com>
Cc: linux-kernel@vger.kernel.org, Mark Tinguely <mark.tinguely@oracle.com>
Subject: [PATCH 5.10 17/21] xfs: catch stale AGF/AGF metadata
Date: Thu, 1 Oct 2026 20:21:30 +0530 [thread overview]
Message-ID: <b11672a1833e4454da1cc967efffca8564fe919c.1790866131.git.siddh.raman.pant@oracle.com> (raw)
In-Reply-To: <cover.1790866131.git.siddh.raman.pant@oracle.com>
From: Dave Chinner <dchinner@redhat.com>
There is a race condition that can trigger in dmflakey fstests that
can result in asserts in xfs_ialloc_read_agi() and
xfs_alloc_read_agf() firing. The asserts look like this:
XFS: Assertion failed: pag->pagf_freeblks == be32_to_cpu(agf->agf_freeblks), file: fs/xfs/libxfs/xfs_alloc.c, line: 3440
.....
Call Trace:
<TASK>
xfs_alloc_read_agf+0x2ad/0x3a0
xfs_alloc_fix_freelist+0x280/0x720
xfs_alloc_vextent_prepare_ag+0x42/0x120
xfs_alloc_vextent_iterate_ags+0x67/0x260
xfs_alloc_vextent_start_ag+0xe4/0x1c0
xfs_bmapi_allocate+0x6fe/0xc90
xfs_bmapi_convert_delalloc+0x338/0x560
xfs_map_blocks+0x354/0x580
iomap_writepages+0x52b/0xa70
xfs_vm_writepages+0xd7/0x100
do_writepages+0xe1/0x2c0
__writeback_single_inode+0x44/0x340
writeback_sb_inodes+0x2d0/0x570
__writeback_inodes_wb+0x9c/0xf0
wb_writeback+0x139/0x2d0
wb_workfn+0x23e/0x4c0
process_scheduled_works+0x1d4/0x400
worker_thread+0x234/0x2e0
kthread+0x147/0x170
ret_from_fork+0x3e/0x50
ret_from_fork_asm+0x1a/0x30
I've seen the AGI variant from scrub running on the filesysetm
after unmount failed due to systemd interference:
XFS: Assertion failed: pag->pagi_freecount == be32_to_cpu(agi->agi_freecount) || xfs_is_shutdown(pag->pag_mount), file: fs/xfs/libxfs/xfs_ialloc.c, line: 2804
.....
Call Trace:
<TASK>
xfs_ialloc_read_agi+0xee/0x150
xchk_perag_drain_and_lock+0x7d/0x240
xchk_ag_init+0x34/0x90
xchk_inode_xref+0x7b/0x220
xchk_inode+0x14d/0x180
xfs_scrub_metadata+0x2e2/0x510
xfs_ioc_scrub_metadata+0x62/0xb0
xfs_file_ioctl+0x446/0xbf0
__se_sys_ioctl+0x6f/0xc0
__x64_sys_ioctl+0x1d/0x30
x64_sys_call+0x1879/0x2ee0
do_syscall_64+0x68/0x130
? exc_page_fault+0x62/0xc0
entry_SYSCALL_64_after_hwframe+0x76/0x7e
Essentially, it is the same problem. When _flakey_drop_and_remount()
loads the drop-writes table, it makes all writes silently fail. Writes
are reported to the fs as completed successfully, but they are not
issued to the backing store. The filesystem sees the successful
write completion and marks the metadata buffer clean and removes it
from the AIL.
If this happens at the same time as memory pressure is occuring,
the now-clean AGF and/or AGI buffers can be reclaimed from memory.
Shortly afterwards, but before _flakey_drop_and_remount() runs
unmount, background writeback is kicked and it tries to allocate
blocks for the dirty pages in memory. This then tries to access the
AGF buffer we just turfed out of memory. It's not found, so it gets
read in from disk.
This is all fine, except for the fact that the last writeback of the
AGF did not actually reach disk. The AGF on disk is stale compared
to the in-memory state held by the perag, and so they don't match
and the assert fires.
Then other operations on that inode hang because the task was killed
whilst holding inode locks. e.g:
Workqueue: xfs-conv/dm-12 xfs_end_io
Call Trace:
<TASK>
__schedule+0x650/0xb10
schedule+0x6d/0xf0
schedule_preempt_disabled+0x15/0x30
rwsem_down_write_slowpath+0x31a/0x5f0
down_write+0x43/0x60
xfs_ilock+0x1a8/0x210
xfs_trans_alloc_inode+0x9c/0x240
xfs_iomap_write_unwritten+0xe3/0x300
xfs_end_ioend+0x90/0x130
xfs_end_io+0xce/0x100
process_scheduled_works+0x1d4/0x400
worker_thread+0x234/0x2e0
kthread+0x147/0x170
ret_from_fork+0x3e/0x50
ret_from_fork_asm+0x1a/0x30
</TASK>
and it's all down hill from there.
Memory pressure is one way to trigger this, another is to run "echo
3 > /proc/sys/vm/drop_caches" randomly while tests are running.
Regardless of how it is triggered, this effectively takes down the
system once umount hangs because it's holding a sb->s_umount lock
exclusive and now every sync(1) call gets stuck on it.
Fix this by replacing the asserts with a corruption detection check
and a shutdown.
Signed-off-by: Dave Chinner <dchinner@redhat.com>
Reviewed-by: Carlos Maiolino <cmaiolino@redhat.com>
Signed-off-by: Carlos Maiolino <cem@kernel.org>
(cherry picked from commit db6a2274162de615ff74b927d38942fe3134d298)
6.12:
Fixed conflict due to the following missing commit:
e9c4d8bfb26c ("xfs: factor out a generic xfs_group structure")
6.6:
Fixed conflict due to the following missing commit:
e45ea3645178 ("xfs: split the agf_roots and agf_levels arrays")
5.15:
Fixed conflict due to the following missing commits:
6f643c57d57c ("xfs: implement ->notify_failure() for XFS")
[Continuous range: a95fee40e3d4^..08d3e84feeb8]
08d3e84feeb8 ("xfs: pass perag to xfs_alloc_read_agf()")
76b47e528e3a ("xfs: kill xfs_alloc_pagf_init()")
99b13c7f0bd3 ("xfs: pass perag to xfs_ialloc_read_agi()")
a95fee40e3d4 ("xfs: kill xfs_ialloc_pagi_init()")
5.10:
Fixed conflicts due to the following missing commits:
fs/xfs/libxfs/xfs_alloc.c, fs/xfs/libxfs/xfs_ialloc.c:
75c8c50fa16a ("xfs: replace XFS_FORCED_SHUTDOWN with xfs_is_shutdown")
fs/xfs/libxfs/xfs_ialloc.c:
9bbafc71919a ("xfs: move xfs_perag_get/put to xfs_ag.[ch]")
Signed-off-by: Siddh Raman Pant <siddh.raman.pant@oracle.com>
---
fs/xfs/libxfs/xfs_alloc.c | 45 +++++++++++++++++++++++++++++---------
fs/xfs/libxfs/xfs_ialloc.c | 34 +++++++++++++++++++++++-----
2 files changed, 64 insertions(+), 15 deletions(-)
diff --git a/fs/xfs/libxfs/xfs_alloc.c b/fs/xfs/libxfs/xfs_alloc.c
index 7cb9f064ac64..896cf97fa4a0 100644
--- a/fs/xfs/libxfs/xfs_alloc.c
+++ b/fs/xfs/libxfs/xfs_alloc.c
@@ -26,6 +26,7 @@
#include "xfs_log.h"
#include "xfs_ag_resv.h"
#include "xfs_bmap.h"
+#include "xfs_health.h"
extern kmem_zone_t *xfs_bmap_free_item_zone;
@@ -3017,18 +3018,42 @@ xfs_alloc_read_agf(
pag->pagf_init = 1;
pag->pagf_agflreset = xfs_agfl_needs_reset(mp, agf);
}
+
#ifdef DEBUG
- else if (!XFS_FORCED_SHUTDOWN(mp)) {
- ASSERT(pag->pagf_freeblks == be32_to_cpu(agf->agf_freeblks));
- ASSERT(pag->pagf_btreeblks == be32_to_cpu(agf->agf_btreeblks));
- ASSERT(pag->pagf_flcount == be32_to_cpu(agf->agf_flcount));
- ASSERT(pag->pagf_longest == be32_to_cpu(agf->agf_longest));
- ASSERT(pag->pagf_levels[XFS_BTNUM_BNOi] ==
- be32_to_cpu(agf->agf_levels[XFS_BTNUM_BNOi]));
- ASSERT(pag->pagf_levels[XFS_BTNUM_CNTi] ==
- be32_to_cpu(agf->agf_levels[XFS_BTNUM_CNTi]));
+ /*
+ * It's possible for the AGF to be out of sync if the block device is
+ * silently dropping writes. This can happen in fstests with dmflakey
+ * enabled, which allows the buffer to be cleaned and reclaimed by
+ * memory pressure and then re-read from disk here. We will get a
+ * stale version of the AGF from disk, and nothing good can happen from
+ * here. Hence if we detect this situation, immediately shut down the
+ * filesystem.
+ *
+ * This can also happen if we are already in the middle of a forced
+ * shutdown, so don't bother checking if we are already shut down.
+ */
+ if (!XFS_FORCED_SHUTDOWN(mp)) {
+ bool ok = true;
+
+ ok &= pag->pagf_freeblks == be32_to_cpu(agf->agf_freeblks);
+ ok &= pag->pagf_freeblks == be32_to_cpu(agf->agf_freeblks);
+ ok &= pag->pagf_btreeblks == be32_to_cpu(agf->agf_btreeblks);
+ ok &= pag->pagf_flcount == be32_to_cpu(agf->agf_flcount);
+ ok &= pag->pagf_longest == be32_to_cpu(agf->agf_longest);
+ ok &= pag->pagf_levels[XFS_BTNUM_BNOi] ==
+ be32_to_cpu(agf->agf_levels[XFS_BTNUM_BNOi]);
+ ok &= pag->pagf_levels[XFS_BTNUM_CNTi] ==
+ be32_to_cpu(agf->agf_levels[XFS_BTNUM_CNTi]);
+
+ if (XFS_IS_CORRUPT(mp, !ok)) {
+ xfs_ag_mark_sick(pag, XFS_SICK_AG_AGF);
+ xfs_trans_brelse(tp, *bpp);
+ *bpp = NULL;
+ xfs_force_shutdown(mp, SHUTDOWN_CORRUPT_INCORE);
+ return -EFSCORRUPTED;
+ }
}
-#endif
+#endif /* DEBUG */
return 0;
}
diff --git a/fs/xfs/libxfs/xfs_ialloc.c b/fs/xfs/libxfs/xfs_ialloc.c
index 8fcb5c6810e9..3f0cfb50f61e 100644
--- a/fs/xfs/libxfs/xfs_ialloc.c
+++ b/fs/xfs/libxfs/xfs_ialloc.c
@@ -27,6 +27,7 @@
#include "xfs_trace.h"
#include "xfs_log.h"
#include "xfs_rmap.h"
+#include "xfs_health.h"
/*
* Lookup a record by ino in the btree given by cur.
@@ -2034,7 +2035,7 @@ xfs_difree_inobt(
goto error0;
}
- /*
+ /*
* Change the inode free counts and log the ag/sb changes.
*/
be32_add_cpu(&agi->agi_freecount, 1);
@@ -2657,12 +2658,35 @@ xfs_ialloc_read_agi(
pag->pagi_init = 1;
}
+#ifdef DEBUG
/*
- * It's possible for these to be out of sync if
- * we are in the middle of a forced shutdown.
+ * It's possible for the AGF to be out of sync if the block device is
+ * silently dropping writes. This can happen in fstests with dmflakey
+ * enabled, which allows the buffer to be cleaned and reclaimed by
+ * memory pressure and then re-read from disk here. We will get a
+ * stale version of the AGF from disk, and nothing good can happen from
+ * here. Hence if we detect this situation, immediately shut down the
+ * filesystem.
+ *
+ * This can also happen if we are already in the middle of a forced
+ * shutdown, so don't bother checking if we are already shut down.
*/
- ASSERT(pag->pagi_freecount == be32_to_cpu(agi->agi_freecount) ||
- XFS_FORCED_SHUTDOWN(mp));
+ if (!XFS_FORCED_SHUTDOWN(mp)) {
+ bool ok = true;
+
+ ok &= pag->pagi_freecount == be32_to_cpu(agi->agi_freecount);
+ ok &= pag->pagi_count == be32_to_cpu(agi->agi_count);
+
+ if (XFS_IS_CORRUPT(mp, !ok)) {
+ xfs_ag_mark_sick(pag, XFS_SICK_AG_AGI);
+ xfs_trans_brelse(tp, *bpp);
+ *bpp = NULL;
+ xfs_force_shutdown(mp, SHUTDOWN_CORRUPT_INCORE);
+ return -EFSCORRUPTED;
+ }
+ }
+#endif /* DEBUG */
+
return 0;
}
--
2.53.0
next prev parent reply other threads:[~2026-10-01 14:53 UTC|newest]
Thread overview: 56+ messages / expand[flat|nested] mbox.gz Atom feed top
[not found] <2026090954-revisit-dowry-37ee@gregkh>
2026-10-01 14:50 ` [PATCH 6.12 0/6] Backport of XFS umount hang fixes Siddh Raman Pant
2026-10-01 14:50 ` [PATCH 6.12 1/6] xfs: xfs_ifree_cluster vs xfs_iflush_shutdown_abort deadlock Siddh Raman Pant
2026-10-01 14:50 ` [PATCH 6.12 2/6] xfs: catch stale AGF/AGF metadata Siddh Raman Pant
2026-10-01 14:50 ` [PATCH 6.12 3/6] xfs: avoid dquot buffer pin deadlock Siddh Raman Pant
2026-10-01 14:50 ` [PATCH 6.12 4/6] xfs: rearrange code in xfs_buf_item.c Siddh Raman Pant
2026-10-01 14:50 ` [PATCH 6.12 5/6] xfs: factor out stale buffer item completion Siddh Raman Pant
2026-10-01 14:50 ` [PATCH 6.12 6/6] xfs: fix unmount hang with unflushable inodes stuck in the AIL Siddh Raman Pant
2026-10-02 14:19 ` [PATCH 6.12 0/6] Backport of XFS umount hang fixes Sasha Levin
2026-10-01 14:50 ` [PATCH 6.6 " Siddh Raman Pant
2026-10-01 14:50 ` [PATCH 6.6 1/6] xfs: xfs_ifree_cluster vs xfs_iflush_shutdown_abort deadlock Siddh Raman Pant
2026-10-01 14:50 ` [PATCH 6.6 2/6] xfs: catch stale AGF/AGF metadata Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.6 3/6] xfs: avoid dquot buffer pin deadlock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.6 4/6] xfs: rearrange code in xfs_buf_item.c Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.6 5/6] xfs: factor out stale buffer item completion Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.6 6/6] xfs: fix unmount hang with unflushable inodes stuck in the AIL Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 00/21] Backport of XFS umount hang fixes Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 01/21] xfs: hoist recovered bmap intent checks out of xfs_bui_item_recover Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 02/21] xfs: improve the code that checks recovered bmap intent items Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 03/21] xfs: hoist recovered rmap intent checks out of xfs_rui_item_recover Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 04/21] xfs: improve the code that checks recovered rmap intent items Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 05/21] xfs: hoist recovered refcount intent checks out of xfs_cui_item_recover Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 06/21] xfs: improve the code that checks recovered refcount intent items Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 07/21] xfs: don't nest icloglock inside ic_callback_lock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 08/21] xfs: convert XLOG_FORCED_SHUTDOWN() to xlog_is_shutdown() Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 09/21] xfs: log items should have a xlog pointer, not a mount Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 10/21] xfs: aborting inodes on shutdown may need buffer lock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 11/21] xfs: remove xfs_buf_t typedef Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 12/21] xfs: fix super block buf log item UAF during force shutdown Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 13/21] xfs: fix intermittent hang during quotacheck Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 14/21] xfs: dquot shrinker doesn't check for XFS_DQFLAG_FREEING Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 15/21] xfs: buffer pins need to hold a buffer reference Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 16/21] xfs: xfs_ifree_cluster vs xfs_iflush_shutdown_abort deadlock Siddh Raman Pant
2026-10-01 14:51 ` Siddh Raman Pant [this message]
2026-10-01 14:51 ` [PATCH 5.10 18/21] xfs: avoid dquot buffer pin deadlock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 19/21] xfs: rearrange code in xfs_buf_item.c Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 20/21] xfs: factor out stale buffer item completion Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.10 21/21] xfs: fix unmount hang with unflushable inodes stuck in the AIL Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 00/11] Backport of XFS umount hang fixes Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 01/11] xfs: log items should have a xlog pointer, not a mount Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 02/11] xfs: aborting inodes on shutdown may need buffer lock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 03/11] xfs: fix super block buf log item UAF during force shutdown Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 04/11] xfs: dquot shrinker doesn't check for XFS_DQFLAG_FREEING Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 05/11] xfs: buffer pins need to hold a buffer reference Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 06/11] xfs: xfs_ifree_cluster vs xfs_iflush_shutdown_abort deadlock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 07/11] xfs: catch stale AGF/AGF metadata Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 08/11] xfs: avoid dquot buffer pin deadlock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 09/11] xfs: rearrange code in xfs_buf_item.c Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 10/11] xfs: factor out stale buffer item completion Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 5.15 11/11] xfs: fix unmount hang with unflushable inodes stuck in the AIL Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.1 0/6] Backport of XFS umount hang fixes Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.1 1/6] xfs: xfs_ifree_cluster vs xfs_iflush_shutdown_abort deadlock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.1 2/6] xfs: catch stale AGF/AGF metadata Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.1 3/6] xfs: avoid dquot buffer pin deadlock Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.1 4/6] xfs: rearrange code in xfs_buf_item.c Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.1 5/6] xfs: factor out stale buffer item completion Siddh Raman Pant
2026-10-01 14:51 ` [PATCH 6.1 6/6] xfs: fix unmount hang with unflushable inodes stuck in the AIL Siddh Raman Pant
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=b11672a1833e4454da1cc967efffca8564fe919c.1790866131.git.siddh.raman.pant@oracle.com \
--to=siddh.raman.pant@oracle.com \
--cc=dchinner@redhat.com \
--cc=djwong@kernel.org \
--cc=gregkh@linuxfoundation.org \
--cc=linux-kernel@vger.kernel.org \
--cc=mark.tinguely@oracle.com \
--cc=stable@vger.kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®