* [PATCH v5 01/10] ceph: fix error handling when the OSD client is stopping in ceph_submit_write()
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 02/10] ceph: wait for pending fscache write in write_folio_nounlock() Tal Zussman
` (9 subsequent siblings)
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Sashiko, Tal Zussman
When ceph_inc_osd_stopping_blocker() fails, ceph_submit_write()
redirties and unlocks the folios remaining in ceph_wbc->fbatch.
However, none of those folios are locked at this point: locked folios
were moved to ceph_wbc->pages[] and removed from the fbatch by
ceph_process_folio_batch(). Unlocking them triggers
VM_BUG_ON_FOLIO(!folio_test_locked(folio)) in folio_unlock(), or
unlocks a folio locked by another task.
This error path also returns -EIO without fully cleaning up the folios in
ceph_wbc->pages[]: their references are never dropped, fscrypt bounce
pages are never freed, the writeback congestion count is never
decremented, and the pages array is leaked. ceph_wbc->locked_pages is
also left non-zero, so ceph_writepages_start() loops back around and
hits BUG_ON(ceph_wbc.locked_pages).
Leave the fbatch folios alone, as ceph_writepages_start() drops their
references when releasing the batch, and unwind ceph_wbc->pages[] the
same way writepages_finish() would have. Abort writeback in
ceph_writepages_start() instead of continuing, matching the handling
of ceph_inc_osd_stopping_blocker() failure on entry, as the OSD
client is being torn down and further writeback cannot make progress.
An LLM was used to verify that the cleanup was comprehensive, and
suggested aborting the writeback instead of continuing.
Fixes: 1551ec61dc55 ("ceph: introduce ceph_submit_write() method")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://sashiko.dev/#/patchset/20260804-remove-wait-on-page-writeback-v2-0-81f0ab065284%40columbia.edu?part=5
Assisted-by: Claude:fable-5
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 39 +++++++++++++++++++++++----------------
1 file changed, 23 insertions(+), 16 deletions(-)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index f8844390f88d..28fb76ded32c 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -1491,28 +1491,32 @@ int ceph_submit_write(struct address_space *mapping,
BUG_ON(len < ceph_fscrypt_page_offset(page) + thp_size(page) - offset);
if (!ceph_inc_osd_stopping_blocker(fsc->mdsc)) {
- for (i = 0; i < folio_batch_count(&ceph_wbc->fbatch); i++) {
- struct folio *folio = ceph_wbc->fbatch.folios[i];
-
- if (!folio)
- continue;
-
- page = &folio->page;
- redirty_page_for_writepage(wbc, page);
- unlock_page(page);
- }
-
for (i = 0; i < ceph_wbc->locked_pages; i++) {
- page = ceph_fscrypt_pagecache_page(ceph_wbc->pages[i]);
+ page = ceph_wbc->pages[i];
+ if (fscrypt_is_bounce_page(page)) {
+ page = fscrypt_pagecache_page(page);
+ fscrypt_free_bounce_page(ceph_wbc->pages[i]);
+ }
- if (!page)
- continue;
+ if (atomic_long_dec_return(&fsc->writeback_count) <
+ CONGESTION_OFF_THRESH(fsc->mount_options->congestion_kb))
+ fsc->write_congested = false;
ceph_undo_wrbuffer_claim(inode, page_folio(page));
redirty_page_for_writepage(wbc, page);
unlock_page(page);
+ put_page(page);
}
+ if (ceph_wbc->from_pool) {
+ mempool_free(ceph_wbc->pages, ceph_wb_pagevec_pool);
+ ceph_wbc->from_pool = false;
+ } else {
+ kfree(ceph_wbc->pages);
+ }
+ ceph_wbc->pages = NULL;
+ ceph_wbc->locked_pages = 0;
+
ceph_osdc_put_request(req);
return -EIO;
}
@@ -1745,8 +1749,11 @@ static int ceph_writepages_start(struct address_space *mapping,
}
rc = ceph_submit_write(mapping, wbc, &ceph_wbc);
- if (rc)
- goto release_folios;
+ if (rc) {
+ /* the OSD client is being torn down, don't retry */
+ folio_batch_release(&ceph_wbc.fbatch);
+ goto dec_osd_stopping_blocker;
+ }
ceph_wbc.locked_pages = 0;
ceph_wbc.strip_unit_end = 0;
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 02/10] ceph: wait for pending fscache write in write_folio_nounlock()
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
2026-09-02 16:11 ` [PATCH v5 01/10] ceph: fix error handling when the OSD client is stopping in ceph_submit_write() Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 03/10] ceph: add ceph_folio_snap_context() Tal Zussman
` (8 subsequent siblings)
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Sashiko, Tal Zussman
write_folio_nounlock() marks the folio with PG_private_2 via
ceph_set_page_fscache() and starts an async write to the cache via
ceph_fscache_write_to_cache(). PG_private_2 is only cleared once that
write completes.
If the OSD write fails, the error paths return without waiting for the
cache write. On retry, ceph_find_incompatible() would wait for
writeback, but not for PG_private_2, before calling
write_folio_nounlock() again, which would trip the VM_BUG_ON_FOLIO() in
folio_start_private_2().
Wait for any pending cache write before starting writeback on the
folio. Do so before bumping the writeback congestion count and
allocating the possibly mempool-backed OSD request, so that a thread
sleeping on the cache I/O does not hold either while it waits.
Add ceph_folio_wait_fscache(), a wrapper for folio_wait_private_2(), and
use it.
Fixes: 1702e7973410 ("ceph: add fscache writeback support")
Reported-by: Sashiko <sashiko-bot@kernel.org>
Link: https://sashiko.dev/#/patchset/20260804-remove-wait-on-page-writeback-v2-0-81f0ab065284%40columbia.edu?part=7
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 12 ++++++++++++
1 file changed, 12 insertions(+)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index 28fb76ded32c..5452e73ae11b 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -564,6 +564,11 @@ static void ceph_set_page_fscache(struct page *page)
folio_start_private_2(page_folio(page)); /* [DEPRECATED] */
}
+static void ceph_folio_wait_fscache(struct folio *folio)
+{
+ folio_wait_private_2(folio); /* [DEPRECATED] */
+}
+
static void ceph_fscache_write_terminated(void *priv, ssize_t error)
{
struct inode *inode = priv;
@@ -585,6 +590,10 @@ static inline void ceph_set_page_fscache(struct page *page)
{
}
+static inline void ceph_folio_wait_fscache(struct folio *folio)
+{
+}
+
static inline void ceph_fscache_write_to_cache(struct inode *inode, u64 off, u64 len, bool caching)
{
}
@@ -788,6 +797,9 @@ static int write_folio_nounlock(struct folio *folio,
ceph_vinop(inode), folio, folio->index, page_off, wlen, snapc,
snapc->seq);
+ /* wait for a cache write left pending by a previously failed attempt */
+ ceph_folio_wait_fscache(folio);
+
if (atomic_long_inc_return(&fsc->writeback_count) >
CONGESTION_ON_THRESH(fsc->mount_options->congestion_kb))
fsc->write_congested = true;
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 03/10] ceph: add ceph_folio_snap_context()
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
2026-09-02 16:11 ` [PATCH v5 01/10] ceph: fix error handling when the OSD client is stopping in ceph_submit_write() Tal Zussman
2026-09-02 16:11 ` [PATCH v5 02/10] ceph: wait for pending fscache write in write_folio_nounlock() Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 04/10] ceph: convert ceph_wait_until_current_writes_complete() to folios Tal Zussman
` (7 subsequent siblings)
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Tal Zussman
Add a folio-based counterpart to page_snap_context() that reads the
snap context from folio->private directly.
Convert the three writeback paths that passed &folio->page to
page_snap_context() to ceph_folio_snap_context(), removing three
open-coded folio-to-page conversions. No functional change.
Reviewed-by: Matthew Wilcox (Oracle) <willy@infradead.org>
Acked-by: Zi Yan <ziy@nvidia.com>
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 12 +++++++++---
1 file changed, 9 insertions(+), 3 deletions(-)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index 5452e73ae11b..2fd3d6d0ffa8 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -75,6 +75,12 @@ static inline struct ceph_snap_context *page_snap_context(struct page *page)
return NULL;
}
+static inline
+struct ceph_snap_context *ceph_folio_snap_context(const struct folio *folio)
+{
+ return folio->private;
+}
+
/*
* Dirty a page. Optimistically adjust accounting, on the assumption
* that we won't race with invalidate. If we do, readjust.
@@ -763,7 +769,7 @@ static int write_folio_nounlock(struct folio *folio,
return -EIO;
/* verify this is a writeable snap context */
- snapc = page_snap_context(&folio->page);
+ snapc = ceph_folio_snap_context(folio);
if (!snapc) {
doutc(cl, "%llx.%llx folio %p not dirty?\n", ceph_vinop(inode),
folio);
@@ -1192,7 +1198,7 @@ int ceph_check_page_before_write(struct address_space *mapping,
}
/* only if matching snap context */
- pgsnapc = page_snap_context(&folio->page);
+ pgsnapc = ceph_folio_snap_context(folio);
if (pgsnapc != ceph_wbc->snapc) {
doutc(cl, "folio snapc %p %lld != oldest %p %lld\n",
pgsnapc, pgsnapc->seq,
@@ -1863,7 +1869,7 @@ ceph_find_incompatible(struct folio *folio)
folio_wait_writeback(folio);
- snapc = page_snap_context(&folio->page);
+ snapc = ceph_folio_snap_context(folio);
if (!snapc || snapc == ci->i_head_snapc)
break;
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 04/10] ceph: convert ceph_wait_until_current_writes_complete() to folios
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
` (2 preceding siblings ...)
2026-09-02 16:11 ` [PATCH v5 03/10] ceph: add ceph_folio_snap_context() Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 05/10] ceph: convert get_writepages_data_length() " Tal Zussman
` (6 subsequent siblings)
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Tal Zussman
Iterate the writeback batch as folios rather than using
&folios[i]->page. Use ceph_folio_snap_context() for the snap context
check and folio_wait_writeback() instead of wait_on_page_writeback().
This removes one call to compound_head() and the last caller of
wait_on_page_writeback(). No functional change.
Reviewed-by: Matthew Wilcox (Oracle) <willy@infradead.org>
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 8 ++++----
1 file changed, 4 insertions(+), 4 deletions(-)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index 2fd3d6d0ffa8..5e4b9da5970b 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -1660,7 +1660,7 @@ void ceph_wait_until_current_writes_complete(struct address_space *mapping,
struct writeback_control *wbc,
struct ceph_writeback_ctl *ceph_wbc)
{
- struct page *page;
+ struct folio *folio;
unsigned i, nr;
if (wbc->sync_mode != WB_SYNC_NONE &&
@@ -1675,10 +1675,10 @@ void ceph_wait_until_current_writes_complete(struct address_space *mapping,
PAGECACHE_TAG_WRITEBACK,
&ceph_wbc->fbatch))) {
for (i = 0; i < nr; i++) {
- page = &ceph_wbc->fbatch.folios[i]->page;
- if (page_snap_context(page) != ceph_wbc->snapc)
+ folio = ceph_wbc->fbatch.folios[i];
+ if (ceph_folio_snap_context(folio) != ceph_wbc->snapc)
continue;
- wait_on_page_writeback(page);
+ folio_wait_writeback(folio);
}
folio_batch_release(&ceph_wbc->fbatch);
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 05/10] ceph: convert get_writepages_data_length() to folios
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
` (3 preceding siblings ...)
2026-09-02 16:11 ` [PATCH v5 04/10] ceph: convert ceph_wait_until_current_writes_complete() to folios Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 06/10] ceph: convert ceph_submit_write() " Tal Zussman
` (5 subsequent siblings)
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Tal Zussman
Introduce ceph_fscrypt_pagecache_folio() and ceph_fscrypt_folio_offset(),
folio equivalents of ceph_fscrypt_pagecache_page() and
ceph_fscrypt_page_offset(), and use them to convert
get_writepages_data_length() to folios.
This removes the last caller of page_snap_context(), so remove it as
well. This also removes a use of page->private and a call to
fscrypt_is_bounce_page().
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 25 +++++++++----------------
fs/ceph/crypto.h | 15 +++++++++++++++
2 files changed, 24 insertions(+), 16 deletions(-)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index 5e4b9da5970b..a72efe96e1f8 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -29,9 +29,9 @@
*
* There are a few funny things going on here.
*
- * The page->private field is used to reference a struct
- * ceph_snap_context for _every_ dirty page. This indicates which
- * snapshot the page was logically dirtied in, and thus which snap
+ * The folio->private field is used to reference a struct
+ * ceph_snap_context for _every_ dirty folio. This indicates which
+ * snapshot the folio was logically dirtied in, and thus which snap
* context needs to be associated with the osd write during writeback.
*
* Similarly, struct ceph_inode_info maintains a set of counters to
@@ -68,13 +68,6 @@
static int ceph_netfs_check_write_begin(struct file *file, loff_t pos, unsigned int len,
struct folio **foliop, void **_fsdata);
-static inline struct ceph_snap_context *page_snap_context(struct page *page)
-{
- if (PagePrivate(page))
- return (void *)page->private;
- return NULL;
-}
-
static inline
struct ceph_snap_context *ceph_folio_snap_context(const struct folio *folio)
{
@@ -706,7 +699,7 @@ get_oldest_context(struct inode *inode, struct ceph_writeback_ctl *ctl,
}
static u64 get_writepages_data_length(struct inode *inode,
- struct page *page, u64 start)
+ struct folio *folio, u64 start)
{
struct ceph_inode_info *ci = ceph_inode(inode);
struct ceph_snap_context *snapc;
@@ -714,7 +707,7 @@ static u64 get_writepages_data_length(struct inode *inode,
u64 end = i_size_read(inode);
u64 ret;
- snapc = page_snap_context(ceph_fscrypt_pagecache_page(page));
+ snapc = ceph_folio_snap_context(ceph_fscrypt_pagecache_folio(folio));
if (snapc != ci->i_head_snapc) {
bool found = false;
spin_lock(&ci->i_ceph_lock);
@@ -729,10 +722,10 @@ static u64 get_writepages_data_length(struct inode *inode,
spin_unlock(&ci->i_ceph_lock);
WARN_ON(!found);
}
- if (end > ceph_fscrypt_page_offset(page) + thp_size(page))
- end = ceph_fscrypt_page_offset(page) + thp_size(page);
+ if (end > ceph_fscrypt_folio_offset(folio) + folio_size(folio))
+ end = ceph_fscrypt_folio_offset(folio) + folio_size(folio);
ret = end > start ? end - start : 0;
- if (ret && fscrypt_is_bounce_page(page))
+ if (ret && fscrypt_is_bounce_folio(folio))
ret = round_up(ret, CEPH_FSCRYPT_BLOCK_SIZE);
return ret;
}
@@ -1601,7 +1594,7 @@ int ceph_submit_write(struct address_space *mapping,
* data length covers all locked pages */
u64 min_len = len + 1 - thp_size(page);
len = get_writepages_data_length(inode,
- ceph_wbc->pages[i - 1],
+ page_folio(ceph_wbc->pages[i - 1]),
offset);
len = max(len, min_len);
}
diff --git a/fs/ceph/crypto.h b/fs/ceph/crypto.h
index 79cb563fd887..948c8b5dca06 100644
--- a/fs/ceph/crypto.h
+++ b/fs/ceph/crypto.h
@@ -162,6 +162,11 @@ static inline struct page *ceph_fscrypt_pagecache_page(struct page *page)
return fscrypt_is_bounce_page(page) ? fscrypt_pagecache_page(page) : page;
}
+static inline struct folio *ceph_fscrypt_pagecache_folio(struct folio *folio)
+{
+ return fscrypt_is_bounce_folio(folio) ? fscrypt_pagecache_folio(folio) : folio;
+}
+
#else /* CONFIG_FS_ENCRYPTION */
static inline void ceph_fscrypt_set_ops(struct super_block *sb)
@@ -262,6 +267,11 @@ static inline struct page *ceph_fscrypt_pagecache_page(struct page *page)
{
return page;
}
+
+static inline struct folio *ceph_fscrypt_pagecache_folio(struct folio *folio)
+{
+ return folio;
+}
#endif /* CONFIG_FS_ENCRYPTION */
static inline loff_t ceph_fscrypt_page_offset(struct page *page)
@@ -269,4 +279,9 @@ static inline loff_t ceph_fscrypt_page_offset(struct page *page)
return page_offset(ceph_fscrypt_pagecache_page(page));
}
+static inline loff_t ceph_fscrypt_folio_offset(struct folio *folio)
+{
+ return folio_pos(ceph_fscrypt_pagecache_folio(folio));
+}
+
#endif /* _CEPH_CRYPTO_H */
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 06/10] ceph: convert ceph_submit_write() to folios
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
` (4 preceding siblings ...)
2026-09-02 16:11 ` [PATCH v5 05/10] ceph: convert get_writepages_data_length() " Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 07/10] ceph: remove page remnants from write_folio_nounlock() Tal Zussman
` (4 subsequent siblings)
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Tal Zussman
Convert the request assembly loop and error path in ceph_submit_write()
to folios. This drops ceph's uses of the set_page_writeback(),
redirty_page_for_writepage(), and unlock_page() compatibility wrappers
in the writeback submission path.
Add ceph_folio_start_fscache(), a folio counterpart of
ceph_set_page_fscache(). The remaining caller of the latter in
write_folio_nounlock() will be converted separately.
In total, this removes eight calls to compound_head() hidden in the
page-based APIs and one explicit page_folio() call in
ceph_undo_wrbuffer_claim()'s caller, while adding four explicit
page_folio() calls.
Note that get_writepages_data_length() must still be passed the
possibly-bounce folio, not the unwrapped pagecache folio, as it checks
fscrypt_is_bounce_folio() to round encrypted lengths up to the fscrypt
block size.
No functional change.
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 43 ++++++++++++++++++++++++++-----------------
1 file changed, 26 insertions(+), 17 deletions(-)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index a72efe96e1f8..b0deff481d95 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -558,6 +558,11 @@ const struct netfs_request_ops ceph_netfs_ops = {
};
#ifdef CONFIG_CEPH_FSCACHE
+static void ceph_folio_start_fscache(struct folio *folio)
+{
+ folio_start_private_2(folio); /* [DEPRECATED] */
+}
+
static void ceph_set_page_fscache(struct page *page)
{
folio_start_private_2(page_folio(page)); /* [DEPRECATED] */
@@ -585,6 +590,10 @@ static void ceph_fscache_write_to_cache(struct inode *inode, u64 off, u64 len, b
ceph_fscache_write_terminated, inode, true, caching);
}
#else
+static inline void ceph_folio_start_fscache(struct folio *folio)
+{
+}
+
static inline void ceph_set_page_fscache(struct page *page)
{
}
@@ -1467,14 +1476,14 @@ int ceph_submit_write(struct address_space *mapping,
struct ceph_client *cl = fsc->client;
struct ceph_vino vino = ceph_vino(inode);
struct ceph_osd_request *req = NULL;
- struct page *page = NULL;
+ struct folio *folio = NULL;
bool caching = ceph_is_cache_enabled(inode);
u64 offset;
u64 len;
unsigned i;
new_request:
- offset = ceph_fscrypt_page_offset(ceph_wbc->pages[0]);
+ offset = ceph_fscrypt_folio_offset(page_folio(ceph_wbc->pages[0]));
len = ceph_wbc->wsize;
req = ceph_osdc_new_request(&fsc->client->osdc,
@@ -1498,14 +1507,14 @@ int ceph_submit_write(struct address_space *mapping,
BUG_ON(IS_ERR(req));
}
- page = ceph_wbc->pages[ceph_wbc->locked_pages - 1];
- BUG_ON(len < ceph_fscrypt_page_offset(page) + thp_size(page) - offset);
+ folio = page_folio(ceph_wbc->pages[ceph_wbc->locked_pages - 1]);
+ BUG_ON(len < ceph_fscrypt_folio_offset(folio) + folio_size(folio) - offset);
if (!ceph_inc_osd_stopping_blocker(fsc->mdsc)) {
for (i = 0; i < ceph_wbc->locked_pages; i++) {
- page = ceph_wbc->pages[i];
- if (fscrypt_is_bounce_page(page)) {
- page = fscrypt_pagecache_page(page);
+ folio = page_folio(ceph_wbc->pages[i]);
+ if (fscrypt_is_bounce_folio(folio)) {
+ folio = fscrypt_pagecache_folio(folio);
fscrypt_free_bounce_page(ceph_wbc->pages[i]);
}
@@ -1513,10 +1522,10 @@ int ceph_submit_write(struct address_space *mapping,
CONGESTION_OFF_THRESH(fsc->mount_options->congestion_kb))
fsc->write_congested = false;
- ceph_undo_wrbuffer_claim(inode, page_folio(page));
- redirty_page_for_writepage(wbc, page);
- unlock_page(page);
- put_page(page);
+ ceph_undo_wrbuffer_claim(inode, folio);
+ folio_redirty_for_writepage(wbc, folio);
+ folio_unlock(folio);
+ folio_put(folio);
}
if (ceph_wbc->from_pool) {
@@ -1542,8 +1551,8 @@ int ceph_submit_write(struct address_space *mapping,
for (i = 0; i < ceph_wbc->locked_pages; i++) {
u64 cur_offset;
- page = ceph_fscrypt_pagecache_page(ceph_wbc->pages[i]);
- cur_offset = page_offset(page);
+ folio = ceph_fscrypt_pagecache_folio(page_folio(ceph_wbc->pages[i]));
+ cur_offset = folio_pos(folio);
/*
* Discontinuity in page range? Ceph can handle that by just passing
@@ -1576,12 +1585,12 @@ int ceph_submit_write(struct address_space *mapping,
ceph_wbc->op_idx++;
}
- set_page_writeback(page);
+ folio_start_writeback(folio);
if (caching)
- ceph_set_page_fscache(page);
+ ceph_folio_start_fscache(folio);
- len += thp_size(page);
+ len += folio_size(folio);
}
ceph_fscache_write_to_cache(inode, offset, len, caching);
@@ -1592,7 +1601,7 @@ int ceph_submit_write(struct address_space *mapping,
/* writepages_finish() clears writeback pages
* according to the data length, so make sure
* data length covers all locked pages */
- u64 min_len = len + 1 - thp_size(page);
+ u64 min_len = len + 1 - folio_size(folio);
len = get_writepages_data_length(inode,
page_folio(ceph_wbc->pages[i - 1]),
offset);
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 07/10] ceph: remove page remnants from write_folio_nounlock()
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
` (5 preceding siblings ...)
2026-09-02 16:11 ` [PATCH v5 06/10] ceph: convert ceph_submit_write() " Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 08/10] ceph: convert page cleanup loop in writepages_finish() to folios Tal Zussman
` (3 subsequent siblings)
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Tal Zussman
Use ceph_folio_start_fscache() and only use a struct page pointer at the
osd_req_op_extent_osd_data_pages() boundary, which requires a page
array. This removes a call to compound_head() in
ceph_set_page_fscache().
This was the last user of ceph_set_page_fscache(), so remove it.
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 18 ++++--------------
1 file changed, 4 insertions(+), 14 deletions(-)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index b0deff481d95..8e761144e8fb 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -563,11 +563,6 @@ static void ceph_folio_start_fscache(struct folio *folio)
folio_start_private_2(folio); /* [DEPRECATED] */
}
-static void ceph_set_page_fscache(struct page *page)
-{
- folio_start_private_2(page_folio(page)); /* [DEPRECATED] */
-}
-
static void ceph_folio_wait_fscache(struct folio *folio)
{
folio_wait_private_2(folio); /* [DEPRECATED] */
@@ -594,10 +589,6 @@ static inline void ceph_folio_start_fscache(struct folio *folio)
{
}
-static inline void ceph_set_page_fscache(struct page *page)
-{
-}
-
static inline void ceph_folio_wait_fscache(struct folio *folio)
{
}
@@ -748,7 +739,7 @@ static u64 get_writepages_data_length(struct inode *inode,
static int write_folio_nounlock(struct folio *folio,
struct writeback_control *wbc)
{
- struct page *page = &folio->page;
+ struct page *page;
struct inode *inode = folio->mapping->host;
struct ceph_inode_info *ci = ceph_inode(inode);
struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
@@ -830,7 +821,7 @@ static int write_folio_nounlock(struct folio *folio,
folio_start_writeback(folio);
if (caching)
- ceph_set_page_fscache(&folio->page);
+ ceph_folio_start_fscache(folio);
ceph_fscache_write_to_cache(inode, page_off, len, caching);
if (IS_ENCRYPTED(inode)) {
@@ -850,9 +841,8 @@ static int write_folio_nounlock(struct folio *folio,
/* it may be a short write due to an object boundary */
WARN_ON_ONCE(len > folio_size(folio));
- osd_req_op_extent_osd_data_pages(req, 0,
- bounce_page ? &bounce_page : &page, wlen, 0,
- false, false);
+ page = bounce_page ? bounce_page : &folio->page;
+ osd_req_op_extent_osd_data_pages(req, 0, &page, wlen, 0, false, false);
doutc(cl, "%llx.%llx %llu~%llu (%llu bytes, %sencrypted)\n",
ceph_vinop(inode), page_off, len, wlen,
IS_ENCRYPTED(inode) ? "" : "not ");
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 08/10] ceph: convert page cleanup loop in writepages_finish() to folios
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
` (6 preceding siblings ...)
2026-09-02 16:11 ` [PATCH v5 07/10] ceph: remove page remnants from write_folio_nounlock() Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 09/10] ceph: remove ceph_fscrypt_pagecache_page() and ceph_fscrypt_page_offset() Tal Zussman
` (2 subsequent siblings)
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Tal Zussman
Convert the page cleanup loop in writepages_finish() to work on folios.
Resolve the folio up front and unwrap any fscrypt bounce folio to its
page cache folio, storing the page cache page back into the array.
This removes a use of detach_page_private() and five calls to
compound_head() per page, while adding one back via page_folio().
While at it, remove the BUG_ON() in the num_pages loop, as it cannot be
triggered.
No functional change.
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 23 +++++++++++------------
1 file changed, 11 insertions(+), 12 deletions(-)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index 8e761144e8fb..f9741812e8a9 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -913,7 +913,6 @@ static void writepages_finish(struct ceph_osd_request *req)
struct ceph_inode_info *ci = ceph_inode(inode);
struct ceph_client *cl = ceph_inode_to_client(inode);
struct ceph_osd_data *osd_data;
- struct page *page;
int num_pages, total_pages = 0;
int i, j;
int rc = req->r_result;
@@ -960,35 +959,35 @@ static void writepages_finish(struct ceph_osd_request *req)
(u64)osd_data->length);
total_pages += num_pages;
for (j = 0; j < num_pages; j++) {
- page = osd_data->pages[j];
- if (fscrypt_is_bounce_page(page)) {
- page = fscrypt_pagecache_page(page);
+ struct folio *folio = page_folio(osd_data->pages[j]);
+
+ if (fscrypt_is_bounce_folio(folio)) {
+ folio = fscrypt_pagecache_folio(folio);
fscrypt_free_bounce_page(osd_data->pages[j]);
- osd_data->pages[j] = page;
+ osd_data->pages[j] = &folio->page;
}
- BUG_ON(!page);
- WARN_ON(!PageUptodate(page));
+ WARN_ON(!folio_test_uptodate(folio));
if (atomic_long_dec_return(&fsc->writeback_count) <
CONGESTION_OFF_THRESH(
fsc->mount_options->congestion_kb))
fsc->write_congested = false;
- ceph_put_snap_context(detach_page_private(page));
- end_page_writeback(page);
+ ceph_put_snap_context(folio_detach_private(folio));
+ folio_end_writeback(folio);
if (atomic64_dec_return(&mdsc->dirty_folios) <= 0) {
wake_up_all(&mdsc->flush_end_wq);
WARN_ON(atomic64_read(&mdsc->dirty_folios) < 0);
}
- doutc(cl, "unlocking %p\n", page);
+ doutc(cl, "unlocking %p\n", folio);
if (remove_page)
generic_error_remove_folio(inode->i_mapping,
- page_folio(page));
+ folio);
- unlock_page(page);
+ folio_unlock(folio);
}
doutc(cl, "%llx.%llx wrote %llu bytes cleaned %d pages\n",
ceph_vinop(inode), osd_data->length,
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 09/10] ceph: remove ceph_fscrypt_pagecache_page() and ceph_fscrypt_page_offset()
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
` (7 preceding siblings ...)
2026-09-02 16:11 ` [PATCH v5 08/10] ceph: convert page cleanup loop in writepages_finish() to folios Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-09-02 16:11 ` [PATCH v5 10/10] ceph: rename ceph_check_page_before_write() to ceph_check_folio_before_write() Tal Zussman
2026-10-08 16:00 ` [PATCH v5 00/10] ceph: convert writeback path to folios David Howells
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Tal Zussman
All callers have been converted to the folio equivalents, so remove
these unused helpers. This also removes a use of page_offset().
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/crypto.h | 15 ---------------
1 file changed, 15 deletions(-)
diff --git a/fs/ceph/crypto.h b/fs/ceph/crypto.h
index 948c8b5dca06..d199dc8410fe 100644
--- a/fs/ceph/crypto.h
+++ b/fs/ceph/crypto.h
@@ -157,11 +157,6 @@ int ceph_fscrypt_decrypt_extents(struct inode *inode, struct page **page,
int ceph_fscrypt_encrypt_pages(struct inode *inode, struct page **page, u64 off,
int len);
-static inline struct page *ceph_fscrypt_pagecache_page(struct page *page)
-{
- return fscrypt_is_bounce_page(page) ? fscrypt_pagecache_page(page) : page;
-}
-
static inline struct folio *ceph_fscrypt_pagecache_folio(struct folio *folio)
{
return fscrypt_is_bounce_folio(folio) ? fscrypt_pagecache_folio(folio) : folio;
@@ -263,22 +258,12 @@ static inline int ceph_fscrypt_encrypt_pages(struct inode *inode,
return 0;
}
-static inline struct page *ceph_fscrypt_pagecache_page(struct page *page)
-{
- return page;
-}
-
static inline struct folio *ceph_fscrypt_pagecache_folio(struct folio *folio)
{
return folio;
}
#endif /* CONFIG_FS_ENCRYPTION */
-static inline loff_t ceph_fscrypt_page_offset(struct page *page)
-{
- return page_offset(ceph_fscrypt_pagecache_page(page));
-}
-
static inline loff_t ceph_fscrypt_folio_offset(struct folio *folio)
{
return folio_pos(ceph_fscrypt_pagecache_folio(folio));
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* [PATCH v5 10/10] ceph: rename ceph_check_page_before_write() to ceph_check_folio_before_write()
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
` (8 preceding siblings ...)
2026-09-02 16:11 ` [PATCH v5 09/10] ceph: remove ceph_fscrypt_pagecache_page() and ceph_fscrypt_page_offset() Tal Zussman
@ 2026-09-02 16:11 ` Tal Zussman
2026-10-08 16:00 ` [PATCH v5 00/10] ceph: convert writeback path to folios David Howells
10 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-09-02 16:11 UTC (permalink / raw)
To: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, David Howells,
Matthew Wilcox (Oracle),
Zi Yan
Cc: ceph-devel, linux-kernel, Tal Zussman
This function only operates on folios, so rename it accordingly.
Signed-off-by: Tal Zussman <tz2294@columbia.edu>
---
fs/ceph/addr.c | 12 ++++++------
1 file changed, 6 insertions(+), 6 deletions(-)
diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c
index f9741812e8a9..8aa60e79a669 100644
--- a/fs/ceph/addr.c
+++ b/fs/ceph/addr.c
@@ -1172,10 +1172,10 @@ bool can_next_page_be_processed(struct ceph_writeback_ctl *ceph_wbc,
}
static
-int ceph_check_page_before_write(struct address_space *mapping,
- struct writeback_control *wbc,
- struct ceph_writeback_ctl *ceph_wbc,
- struct folio *folio)
+int ceph_check_folio_before_write(struct address_space *mapping,
+ struct writeback_control *wbc,
+ struct ceph_writeback_ctl *ceph_wbc,
+ struct folio *folio)
{
struct inode *inode = mapping->host;
struct ceph_fs_client *fsc = ceph_inode_to_fs_client(inode);
@@ -1359,8 +1359,8 @@ void ceph_process_folio_batch(struct address_space *mapping,
else if (!folio_trylock(folio))
break;
- rc = ceph_check_page_before_write(mapping, wbc,
- ceph_wbc, folio);
+ rc = ceph_check_folio_before_write(mapping, wbc,
+ ceph_wbc, folio);
if (rc == -ENODATA) {
folio_unlock(folio);
folio_put(folio);
--
2.39.5
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH v5 00/10] ceph: convert writeback path to folios
2026-09-02 16:11 [PATCH v5 00/10] ceph: convert writeback path to folios Tal Zussman
` (9 preceding siblings ...)
2026-09-02 16:11 ` [PATCH v5 10/10] ceph: rename ceph_check_page_before_write() to ceph_check_folio_before_write() Tal Zussman
@ 2026-10-08 16:00 ` David Howells
2026-10-10 0:16 ` Tal Zussman
10 siblings, 1 reply; 13+ messages in thread
From: David Howells @ 2026-10-08 16:00 UTC (permalink / raw)
To: Tal Zussman
Cc: dhowells, Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, Matthew Wilcox (Oracle),
Zi Yan, ceph-devel, linux-kernel
Hi Tal,
Apologies for not reaching out to you before, but I keep getting diverted on
to other thing - usually by AI patch reviews.
Just to note that I have patches in the works to convert ceph to use netfslib
and thus strip out all the dealing with pages/folios as much as possible -
even getting down into the protocol and replacing the five different page
carrying types with the same structure as is at the core of netfslib (bvecq
chains) and pass the bvecq chains all the way into sendmsg() for transmission.
This also lets me remove ceph's use of the PG_private_2 flag which the MM
people would really like back. That's then replaced - transparently inside of
netfslib - by setting a special ->private value on folios read from the
server, marking them dirty and then having writeback write them to the cache.
It's unfortunately stalled on getting some netfslib changes in, though there
may be hope for making some progress in this merge window.
If you go to my tree here:
https://git.kernel.org/pub/scm/linux/kernel/git/dhowells/linux-fs.git/
The netfs-next[12345] branches are the prerequisite patches I'm trying to get
in at the moment, divided into five batches.
On top of that is my netfs-crypt branch that has content crypto support
And then on top of that is my ceph-iter branch. Note that it's not fully
working as yet and I'm still testing and fixing bits of it.
If you want to push these patches first, I can always adapt my patches to take
account of them.
David
^ permalink raw reply [flat|nested] 13+ messages in thread* Re: [PATCH v5 00/10] ceph: convert writeback path to folios
2026-10-08 16:00 ` [PATCH v5 00/10] ceph: convert writeback path to folios David Howells
@ 2026-10-10 0:16 ` Tal Zussman
0 siblings, 0 replies; 13+ messages in thread
From: Tal Zussman @ 2026-10-10 0:16 UTC (permalink / raw)
To: David Howells
Cc: Ilya Dryomov, Alex Markuze, Viacheslav Dubeyko,
Christian Brauner, Jeff Layton, Matthew Wilcox (Oracle),
Zi Yan, ceph-devel, linux-kernel
H
On 10/8/26 12:00 PM, David Howells wrote:
> Hi Tal,
>
> Apologies for not reaching out to you before, but I keep getting diverted on
> to other thing - usually by AI patch reviews.
>
> Just to note that I have patches in the works to convert ceph to use netfslib
> and thus strip out all the dealing with pages/folios as much as possible -
> even getting down into the protocol and replacing the five different page
> carrying types with the same structure as is at the core of netfslib (bvecq
> chains) and pass the bvecq chains all the way into sendmsg() for transmission.
>
> This also lets me remove ceph's use of the PG_private_2 flag which the MM
> people would really like back. That's then replaced - transparently inside of
> netfslib - by setting a special ->private value on folios read from the
> server, marking them dirty and then having writeback write them to the cache.
>
> It's unfortunately stalled on getting some netfslib changes in, though there
> may be hope for making some progress in this merge window.
>
> If you go to my tree here:
>
> https://git.kernel.org/pub/scm/linux/kernel/git/dhowells/linux-fs.git/
>
> The netfs-next[12345] branches are the prerequisite patches I'm trying to get
> in at the moment, divided into five batches.
>
> On top of that is my netfs-crypt branch that has content crypto support
>
> And then on top of that is my ceph-iter branch. Note that it's not fully
> working as yet and I'm still testing and fixing bits of it.
>
That's great, I see it removes the ceph pagelist and with it the last couple
callers of __page_cache_alloc(). I wasn't planning on sending additional ceph
patches in the near future, as they would all require getting into the
protocol, so it's great to hear you're working on this!
> If you want to push these patches first, I can always adapt my patches to take
> account of them.
>
It'd be great to get this series in sooner rather than later. I think it's
ready to be merged and it would let us remove five page wrappers of folio
functions immediately. Hopefully since it's fairly small it won't cause you
too much trouble.
Thanks,
Tal
^ permalink raw reply [flat|nested] 13+ messages in thread