* [PATCH v12 01/10] Add a function to kmap one page of a multipage bio_vec
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 02/10] iov_iter: Add a segmented queue of bio_vec[] David Howells
` (9 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Jens Axboe, Christoph Hellwig, linux-block
Add a function to kmap one page of a multipage bio_vec by offset (which is
added to the offset in the bio_vec internally). The caller is responsible
for calculating how much of the page is then available.
Signed-off-by: David Howells <dhowells@redhat.com>
Reviewed-by: Paulo Alcantara (Red Hat) <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Jens Axboe <axboe@kernel.dk>
cc: Christoph Hellwig <hch@infradead.org>
cc: linux-block@vger.kernel.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
include/linux/bvec.h | 18 ++++++++++++++++++
1 file changed, 18 insertions(+)
diff --git a/include/linux/bvec.h b/include/linux/bvec.h
index fc566ee1c1ff..df9ce860e787 100644
--- a/include/linux/bvec.h
+++ b/include/linux/bvec.h
@@ -343,4 +343,22 @@ static inline phys_addr_t bvec_phys(const struct bio_vec *bvec)
return page_to_phys(bvec->bv_page) + bvec->bv_offset;
}
+/**
+ * bvec_kmap_partial - Map part of a bvec into the kernel virtual address space
+ * @bvec: bvec to map
+ * @offset: Offset into bvec
+ *
+ * Map the page containing the byte at @offset into the kernel virtual address
+ * space. The caller is responsible for making sure this doesn't overrun.
+ *
+ * Call kunmap_local on the returned address to unmap.
+ */
+static inline void *bvec_kmap_partial(struct bio_vec *bvec, size_t offset)
+{
+ offset += bvec->bv_offset;
+
+ return kmap_local_page(bvec->bv_page + (offset >> PAGE_SHIFT)) +
+ (offset & ~PAGE_MASK);
+}
+
#endif /* __LINUX_BVEC_H */
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 02/10] iov_iter: Add a segmented queue of bio_vec[]
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
2026-09-29 7:59 ` [PATCH v12 01/10] Add a function to kmap one page of a multipage bio_vec David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 03/10] netfs: Add some tools for managing bvecq chains David Howells
` (8 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Jens Axboe, Christoph Hellwig, linux-block
Add the concept of a segmented queue of bio_vec[] arrays. This allows an
indefinite quantity of elements to be handled and allows things like
network filesystems and crypto drivers to glue bits on the ends without
having to reallocate the array.
The bvecq struct that defines each segment also carries capacity/usage
information along with flags indicating whether the constituent memory
regions need freeing or unpinning. The bvecq structs are refcounted to
allow a queue to be extracted in batches and split between a number of
subrequests.
The bvecq can have the bio_vec[] it manages allocated in with it, but this
is not required. A flag is provided for if this is the case as comparing
->bv to ->__bv is not sufficient to detect this case.
Add an iterator type ITER_BVECQ for it. This is intended to replace
ITER_FOLIOQ (and ITER_XARRAY).
Note that the prev pointer is only really needed for iov_iter_revert() and
could be dispensed with if struct iov_iter contained the head information
as well as the current point.
Signed-off-by: David Howells <dhowells@redhat.com>
Reviewed-by: Paulo Alcantara (Red Hat) <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Jens Axboe <axboe@kernel.dk>
cc: Christoph Hellwig <hch@infradead.org>
cc: linux-block@vger.kernel.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
include/linux/bvecq.h | 50 +++++
include/linux/iov_iter.h | 74 ++++++-
include/linux/uio.h | 11 +
lib/iov_iter.c | 410 ++++++++++++++++++++++++++++++++++++-
lib/scatterlist.c | 67 +++++-
lib/tests/kunit_iov_iter.c | 260 +++++++++++++++++++++++
6 files changed, 866 insertions(+), 6 deletions(-)
create mode 100644 include/linux/bvecq.h
diff --git a/include/linux/bvecq.h b/include/linux/bvecq.h
new file mode 100644
index 000000000000..77fd07852c33
--- /dev/null
+++ b/include/linux/bvecq.h
@@ -0,0 +1,50 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/* Implementation of a segmented queue of bio_vec[].
+ *
+ * Copyright (C) 2026 Red Hat, Inc. All Rights Reserved.
+ * Written by David Howells (dhowells@redhat.com)
+ */
+
+#ifndef _LINUX_BVECQ_H
+#define _LINUX_BVECQ_H
+
+#include <linux/bvec.h>
+
+/*
+ * The type of memory retention used by the elements in bvecq->bv[] and how to
+ * clean it up.
+ */
+enum bvecq_mem {
+ BVECQ_MEM_EXTERNAL, /* Externally retained memory - no freeing */
+ BVECQ_MEM_PAGECACHE, /* Ref'd pagecache pages - must put */
+ BVECQ_MEM_GUP, /* Pinned memory from get_user_pages() - unpin */
+ BVECQ_MEM_ALLOCED, /* Memory alloc'd by bvecq - can be freed/pooled */
+} __mode(byte);
+
+/*
+ * Segmented bio_vec queue.
+ *
+ * These can be linked together to form messages of indefinite length and
+ * iterated over with an ITER_BVECQ iterator. The list is non-circular; next
+ * and prev are NULL at the ends.
+ *
+ * The bv pointer points to the bio_vec array; this may be __bv if allocated
+ * together. The caller is responsible for determining whether or not this is
+ * the case as the array pointed to by bv may be follow on directly from the
+ * bvecq by accident of allocation (ie. ->bv == ->__bv is *not* sufficient to
+ * determine this).
+ */
+struct bvecq {
+ struct bvecq *next; /* Next bvec in the list or NULL */
+ struct bvecq *prev; /* Prev bvec in the list or NULL */
+ refcount_t ref;
+ u32 priv; /* Private data */
+ u16 nr_slots; /* Number of elements in bv[] used */
+ u16 max_slots; /* Number of elements allocated in bv[] */
+ enum bvecq_mem mem_type:3; /* What sort of memory and how to free it */
+ bool inline_bv:1; /* T if __bv[] is being used */
+ struct bio_vec *bv; /* Pointer to array of page fragments */
+ struct bio_vec __bv[]; /* Default array (if ->inline_bv) */
+};
+
+#endif /* _LINUX_BVECQ_H */
diff --git a/include/linux/iov_iter.h b/include/linux/iov_iter.h
index f9a17fbbd398..ec0aa2893689 100644
--- a/include/linux/iov_iter.h
+++ b/include/linux/iov_iter.h
@@ -10,6 +10,7 @@
#include <linux/uio.h>
#include <linux/bvec.h>
+#include <linux/bvecq.h>
#include <linux/folio_queue.h>
typedef size_t (*iov_step_f)(void *iter_base, size_t progress, size_t len,
@@ -141,6 +142,71 @@ size_t iterate_bvec(struct iov_iter *iter, size_t len, void *priv, void *priv2,
return progress;
}
+/*
+ * Handle ITER_BVECQ.
+ */
+static __always_inline
+size_t iterate_bvecq(struct iov_iter *iter, size_t len, void *priv, void *priv2,
+ iov_step_f step)
+{
+ const struct bvecq *bq = iter->bvecq;
+ unsigned int slot = iter->bvecq_slot;
+ size_t progress = 0, skip = iter->iov_offset;
+
+ do {
+ const struct bio_vec *bvec;
+ struct page *page;
+ size_t poff, plen;
+ void *base;
+
+ if (slot >= bq->nr_slots) {
+ if (!bq->next)
+ break;
+ bq = bq->next;
+ slot = 0;
+ continue;
+ }
+
+ bvec = &bq->bv[slot];
+ /*
+ * The caller must ensure that a slot with bv_len>0 has a valid
+ * bv_page.
+ */
+ page = bvec->bv_page + (bvec->bv_offset + skip) / PAGE_SIZE;
+ poff = (bvec->bv_offset + skip) % PAGE_SIZE;
+ plen = min(bvec->bv_len - skip, len);
+
+ while (plen > 0) {
+ size_t part, remain, consumed;
+
+ part = min(plen, PAGE_SIZE - poff);
+ base = kmap_local_page(page) + poff;
+ remain = step(base, progress, part, priv, priv2);
+ kunmap_local(base);
+
+ consumed = part - remain;
+ progress += consumed;
+ skip += consumed;
+ len -= consumed;
+ if (!len || remain)
+ goto stop;
+ page++;
+ poff = 0;
+ plen -= consumed;
+ }
+
+ skip = 0;
+ slot++;
+ } while (len);
+
+stop:
+ iter->bvecq_slot = slot;
+ iter->bvecq = bq;
+ iter->iov_offset = skip;
+ iter->count -= progress;
+ return progress;
+}
+
/*
* Handle ITER_FOLIOQ.
*/
@@ -306,6 +372,8 @@ size_t iterate_and_advance2(struct iov_iter *iter, size_t len, void *priv,
return iterate_bvec(iter, len, priv, priv2, step);
if (iov_iter_is_kvec(iter))
return iterate_kvec(iter, len, priv, priv2, step);
+ if (iov_iter_is_bvecq(iter))
+ return iterate_bvecq(iter, len, priv, priv2, step);
if (iov_iter_is_folioq(iter))
return iterate_folioq(iter, len, priv, priv2, step);
if (iov_iter_is_xarray(iter))
@@ -342,8 +410,8 @@ size_t iterate_and_advance(struct iov_iter *iter, size_t len, void *priv,
* buffer is presented in segments, which for kernel iteration are broken up by
* physical pages and mapped, with the mapped address being presented.
*
- * [!] Note This will only handle BVEC, KVEC, FOLIOQ, XARRAY and DISCARD-type
- * iterators; it will not handle UBUF or IOVEC-type iterators.
+ * [!] Note This will only handle BVEC, KVEC, BVECQ, FOLIOQ, XARRAY and
+ * DISCARD-type iterators; it will not handle UBUF or IOVEC-type iterators.
*
* A step functions, @step, must be provided, one for handling mapped kernel
* addresses and the other is given user addresses which have the potential to
@@ -370,6 +438,8 @@ size_t iterate_and_advance_kernel(struct iov_iter *iter, size_t len, void *priv,
return iterate_bvec(iter, len, priv, priv2, step);
if (iov_iter_is_kvec(iter))
return iterate_kvec(iter, len, priv, priv2, step);
+ if (iov_iter_is_bvecq(iter))
+ return iterate_bvecq(iter, len, priv, priv2, step);
if (iov_iter_is_folioq(iter))
return iterate_folioq(iter, len, priv, priv2, step);
if (iov_iter_is_xarray(iter))
diff --git a/include/linux/uio.h b/include/linux/uio.h
index fe2e985d74d2..2d75a267da86 100644
--- a/include/linux/uio.h
+++ b/include/linux/uio.h
@@ -26,6 +26,7 @@ enum iter_type {
ITER_IOVEC,
ITER_BVEC,
ITER_KVEC,
+ ITER_BVECQ,
ITER_FOLIOQ,
ITER_XARRAY,
ITER_DISCARD,
@@ -68,6 +69,7 @@ struct iov_iter {
const struct iovec *__iov;
const struct kvec *kvec;
const struct bio_vec *bvec;
+ const struct bvecq *bvecq;
const struct folio_queue *folioq;
struct xarray *xarray;
void __user *ubuf;
@@ -77,6 +79,7 @@ struct iov_iter {
};
union {
unsigned long nr_segs;
+ u16 bvecq_slot;
u8 folioq_slot;
loff_t xarray_start;
};
@@ -145,6 +148,11 @@ static inline bool iov_iter_is_discard(const struct iov_iter *i)
return iov_iter_type(i) == ITER_DISCARD;
}
+static inline bool iov_iter_is_bvecq(const struct iov_iter *i)
+{
+ return iov_iter_type(i) == ITER_BVECQ;
+}
+
static inline bool iov_iter_is_folioq(const struct iov_iter *i)
{
return iov_iter_type(i) == ITER_FOLIOQ;
@@ -295,6 +303,9 @@ void iov_iter_kvec(struct iov_iter *i, unsigned int direction, const struct kvec
void iov_iter_bvec(struct iov_iter *i, unsigned int direction, const struct bio_vec *bvec,
unsigned long nr_segs, size_t count);
void iov_iter_discard(struct iov_iter *i, unsigned int direction, size_t count);
+void iov_iter_bvec_queue(struct iov_iter *i, unsigned int direction,
+ const struct bvecq *bvecq,
+ unsigned int first_slot, unsigned int offset, size_t count);
void iov_iter_folio_queue(struct iov_iter *i, unsigned int direction,
const struct folio_queue *folioq,
unsigned int first_slot, unsigned int offset, size_t count);
diff --git a/lib/iov_iter.c b/lib/iov_iter.c
index 2072c04e99d0..b76a297dc8d4 100644
--- a/lib/iov_iter.c
+++ b/lib/iov_iter.c
@@ -538,6 +538,40 @@ static void iov_iter_iovec_advance(struct iov_iter *i, size_t size)
i->__iov = iov;
}
+static void iov_iter_bvecq_advance(struct iov_iter *i, size_t by)
+{
+ const struct bvecq *bq = i->bvecq;
+ unsigned int slot = i->bvecq_slot;
+
+ if (!i->count)
+ return;
+ i->count -= by;
+
+ by += i->iov_offset; /* From beginning of current segment. */
+ do {
+ size_t len;
+
+ if (slot >= bq->nr_slots) {
+ if (!bq->next)
+ break;
+ bq = bq->next;
+ slot = 0;
+ continue;
+ }
+
+ len = bq->bv[slot].bv_len;
+
+ if (likely(by < len))
+ break;
+ by -= len;
+ slot++;
+ } while (by);
+
+ i->iov_offset = by;
+ i->bvecq_slot = slot;
+ i->bvecq = bq;
+}
+
static void iov_iter_folioq_advance(struct iov_iter *i, size_t size)
{
const struct folio_queue *folioq = i->folioq;
@@ -583,6 +617,8 @@ void iov_iter_advance(struct iov_iter *i, size_t size)
iov_iter_iovec_advance(i, size);
} else if (iov_iter_is_bvec(i)) {
iov_iter_bvec_advance(i, size);
+ } else if (iov_iter_is_bvecq(i)) {
+ iov_iter_bvecq_advance(i, size);
} else if (iov_iter_is_folioq(i)) {
iov_iter_folioq_advance(i, size);
} else if (iov_iter_is_discard(i)) {
@@ -591,6 +627,33 @@ void iov_iter_advance(struct iov_iter *i, size_t size)
}
EXPORT_SYMBOL(iov_iter_advance);
+static void iov_iter_bvecq_revert(struct iov_iter *i, size_t unroll)
+{
+ const struct bvecq *bq = i->bvecq;
+ unsigned int slot = i->bvecq_slot;
+
+ for (;;) {
+ size_t len;
+
+ if (slot == 0) {
+ bq = bq->prev;
+ slot = bq->nr_slots;
+ continue;
+ }
+ slot--;
+
+ len = bq->bv[slot].bv_len;
+ if (unroll <= len) {
+ i->iov_offset = len - unroll;
+ break;
+ }
+ unroll -= len;
+ }
+
+ i->bvecq_slot = slot;
+ i->bvecq = bq;
+}
+
static void iov_iter_folioq_revert(struct iov_iter *i, size_t unroll)
{
const struct folio_queue *folioq = i->folioq;
@@ -648,6 +711,9 @@ void iov_iter_revert(struct iov_iter *i, size_t unroll)
}
unroll -= n;
}
+ } else if (iov_iter_is_bvecq(i)) {
+ i->iov_offset = 0;
+ iov_iter_bvecq_revert(i, unroll);
} else if (iov_iter_is_folioq(i)) {
i->iov_offset = 0;
iov_iter_folioq_revert(i, unroll);
@@ -678,9 +744,30 @@ size_t iov_iter_single_seg_count(const struct iov_iter *i)
if (iov_iter_is_bvec(i))
return min(i->count, i->bvec->bv_len - i->iov_offset);
}
+ if (!i->count)
+ return 0;
+ if (unlikely(iov_iter_is_bvecq(i))) {
+ const struct bvecq *bq = i->bvecq;
+ unsigned int slot = i->bvecq_slot;
+ size_t offset = i->iov_offset;
+
+ for (;;) {
+ while (slot >= bq->nr_slots) {
+ bq = bq->next;
+ if (!bq)
+ return 0;
+ slot = 0;
+ offset = 0;
+ }
+ if (bq->bv[slot].bv_len > offset)
+ break;
+ slot++;
+ offset = 0;
+ }
+ return min(i->count, bq->bv[slot].bv_len - offset);
+ }
if (unlikely(iov_iter_is_folioq(i)))
- return !i->count ? 0 :
- umin(folioq_folio_size(i->folioq, i->folioq_slot), i->count);
+ return umin(folioq_folio_size(i->folioq, i->folioq_slot), i->count);
return i->count;
}
EXPORT_SYMBOL(iov_iter_single_seg_count);
@@ -717,6 +804,35 @@ void iov_iter_bvec(struct iov_iter *i, unsigned int direction,
}
EXPORT_SYMBOL(iov_iter_bvec);
+/**
+ * iov_iter_bvec_queue - Initialise an I/O iterator to use a segmented bvec queue
+ * @i: The iterator to initialise.
+ * @direction: The direction of the transfer.
+ * @bvecq: The starting point in the bvec queue.
+ * @first_slot: The first slot in the bvec queue to use
+ * @offset: The offset into the bvec in the first slot to start at
+ * @count: The size of the I/O buffer in bytes.
+ *
+ * Set up an I/O iterator to either draw data out of the buffers attached to an
+ * inode or to inject data into those buffers. The pages *must* be prevented
+ * from evaporation, either by the caller.
+ */
+void iov_iter_bvec_queue(struct iov_iter *i, unsigned int direction,
+ const struct bvecq *bvecq, unsigned int first_slot,
+ unsigned int offset, size_t count)
+{
+ WARN_ON(direction & ~(READ | WRITE));
+ *i = (struct iov_iter) {
+ .iter_type = ITER_BVECQ,
+ .data_source = direction,
+ .bvecq = bvecq,
+ .bvecq_slot = first_slot,
+ .count = count,
+ .iov_offset = offset,
+ };
+}
+EXPORT_SYMBOL(iov_iter_bvec_queue);
+
/**
* iov_iter_folio_queue - Initialise an I/O iterator to use the folios in a folio queue
* @i: The iterator to initialise.
@@ -839,6 +955,39 @@ static unsigned long iov_iter_alignment_bvec(const struct iov_iter *i)
return res;
}
+static unsigned long iov_iter_alignment_bvecq(const struct iov_iter *iter)
+{
+ const struct bvecq *bq;
+ unsigned long res = 0;
+ unsigned int slot = iter->bvecq_slot;
+ size_t skip = iter->iov_offset;
+ size_t size = iter->count;
+
+ if (!size)
+ return res;
+
+ for (bq = iter->bvecq; bq; bq = bq->next) {
+ for (; slot < bq->nr_slots; slot++) {
+ const struct bio_vec *bvec = &bq->bv[slot];
+ size_t part = min(bvec->bv_len - skip, size);
+
+ if (part) {
+ res |= bvec->bv_offset + skip;
+ res |= part;
+ }
+
+ size -= part;
+ if (size == 0)
+ return res;
+ skip = 0;
+ }
+
+ slot = 0;
+ }
+
+ return res;
+}
+
unsigned long iov_iter_alignment(const struct iov_iter *i)
{
if (likely(iter_is_ubuf(i))) {
@@ -854,6 +1003,8 @@ unsigned long iov_iter_alignment(const struct iov_iter *i)
if (iov_iter_is_bvec(i))
return iov_iter_alignment_bvec(i);
+ if (iov_iter_is_bvecq(i))
+ return iov_iter_alignment_bvecq(i);
/* With both xarray and folioq types, we're dealing with whole folios. */
if (iov_iter_is_folioq(i))
@@ -910,6 +1061,137 @@ static int want_pages_array(struct page ***res, size_t size,
return count;
}
+/*
+ * Count the number of virtually contiguous pages coming up next in an
+ * ITER_BVECQ iterator, up to the specified maxima.
+ */
+static unsigned int iter_count_bvecq_pages(const struct iov_iter *iter,
+ size_t maxsize,
+ unsigned int maxpages)
+{
+ const struct bvecq *bvecq = iter->bvecq;
+ unsigned int slot = iter->bvecq_slot;
+ ssize_t remain = umin(maxsize, iter->count);
+ size_t count = 0, offset = iter->iov_offset;
+
+ do {
+ const struct bio_vec *bv;
+ size_t boff, blen;
+
+ if (slot >= bvecq->nr_slots) {
+ if (!bvecq->next) {
+ WARN_ON_ONCE(remain > 0);
+ break;
+ }
+ bvecq = bvecq->next;
+ slot = 0;
+ offset = 0;
+ continue;
+ }
+
+ bv = &bvecq->bv[slot++];
+ boff = bv->bv_offset;
+ blen = bv->bv_len;
+
+ /* bv_page is not allowed to be NULL unless bv_len == 0. */
+ if (WARN_ON_ONCE(!bv->bv_page && blen > 0))
+ break;
+ if (!PAGE_ALIGNED(boff) && count > 0)
+ break;
+
+ boff += offset;
+ blen -= offset;
+ offset = 0;
+ if (!blen)
+ continue;
+
+ blen = umin(blen, remain);
+ remain -= blen;
+ blen += offset_in_page(boff);
+ count += DIV_ROUND_UP(blen, PAGE_SIZE);
+
+ if (!PAGE_ALIGNED(blen))
+ break;
+ } while (remain > 0 && count < maxpages);
+
+ return min(count, maxpages);
+}
+
+/*
+ * Get a list of virtually contiguous pages from an ITER_BVECQ iterator and pin
+ * them.
+ */
+static ssize_t iter_bvecq_get_pages(struct iov_iter *iter,
+ struct page ***ppages, size_t maxsize,
+ unsigned int maxpages, size_t *_start_offset)
+{
+ const struct bvecq *bvecq = iter->bvecq;
+ struct page **pages;
+ unsigned int slot = iter->bvecq_slot, nr = 0;
+ size_t extracted = 0, offset = iter->iov_offset;
+
+ /* Count the next run of virtually contiguous pages. */
+ maxpages = iter_count_bvecq_pages(iter, maxsize, maxpages);
+ if (!maxpages)
+ return 0;
+
+ maxpages = want_pages_array(ppages, maxsize, offset & ~PAGE_MASK, maxpages);
+ if (!maxpages)
+ return -ENOMEM;
+ pages = *ppages;
+
+ /* Now transcribe the page pointers. */
+ do {
+ const struct bio_vec *bv;
+ size_t boff, blen;
+
+ if (slot >= bvecq->nr_slots) {
+ if (!bvecq->next) {
+ WARN_ON_ONCE(extracted < iter->count);
+ break;
+ }
+ bvecq = bvecq->next;
+ slot = 0;
+ offset = 0;
+ continue;
+ }
+
+ bv = &bvecq->bv[slot];
+ boff = bv->bv_offset;
+ blen = bv->bv_len;
+
+ /* bv_page is not allowed to be NULL unless bv_len == 0. */
+
+ if (offset < blen) {
+ size_t poff = (boff + offset) % PAGE_SIZE;
+ size_t part = min(maxsize - extracted, blen - offset);
+ size_t pix = (boff + offset) / PAGE_SIZE;
+
+ if (poff + part > PAGE_SIZE)
+ part = PAGE_SIZE - poff;
+
+ if (!extracted)
+ *_start_offset = poff;
+
+ get_page(bv->bv_page + pix);
+ pages[nr++] = bv->bv_page + pix;
+ offset += part;
+ extracted += part;
+ }
+
+ if (offset >= blen) {
+ offset = 0;
+ slot++;
+ }
+ } while (nr < maxpages && extracted < maxsize);
+
+ iter->bvecq = bvecq;
+ iter->bvecq_slot = slot;
+ iter->iov_offset = offset;
+ iter->count -= extracted;
+ return extracted;
+}
+
static ssize_t iter_folioq_get_pages(struct iov_iter *iter,
struct page ***ppages, size_t maxsize,
unsigned maxpages, size_t *_start_offset)
@@ -1120,6 +1402,8 @@ static ssize_t __iov_iter_get_pages_alloc(struct iov_iter *i,
}
return maxsize;
}
+ if (iov_iter_is_bvecq(i))
+ return iter_bvecq_get_pages(i, pages, maxsize, maxpages, start);
if (iov_iter_is_folioq(i))
return iter_folioq_get_pages(i, pages, maxsize, maxpages, start);
if (iov_iter_is_xarray(i))
@@ -1192,6 +1476,38 @@ static int bvec_npages(const struct iov_iter *i, int maxpages)
return npages;
}
+static size_t iov_npages_bvecq(const struct iov_iter *iter, size_t maxpages)
+{
+ const struct bvecq *bq;
+ unsigned int slot = iter->bvecq_slot;
+ size_t npages = 0;
+ size_t skip = iter->iov_offset;
+ size_t size = iter->count;
+
+ for (bq = iter->bvecq; bq; bq = bq->next) {
+ for (; slot < bq->nr_slots; slot++) {
+ const struct bio_vec *bvec = &bq->bv[slot];
+ size_t offs = (bvec->bv_offset + skip) % PAGE_SIZE;
+ size_t part = min(bvec->bv_len - skip, size);
+
+ if (part) {
+ npages += DIV_ROUND_UP(offs + part, PAGE_SIZE);
+ if (npages >= maxpages)
+ goto out;
+ }
+
+ size -= part;
+ if (!size)
+ goto out;
+ skip = 0;
+ }
+
+ slot = 0;
+ }
+out:
+ return umin(npages, maxpages);
+}
+
int iov_iter_npages(const struct iov_iter *i, int maxpages)
{
if (unlikely(!i->count))
@@ -1206,6 +1522,8 @@ int iov_iter_npages(const struct iov_iter *i, int maxpages)
return iov_npages(i, maxpages);
if (iov_iter_is_bvec(i))
return bvec_npages(i, maxpages);
+ if (iov_iter_is_bvecq(i))
+ return iov_npages_bvecq(i, maxpages);
if (iov_iter_is_folioq(i)) {
unsigned offset = i->iov_offset % PAGE_SIZE;
int npages = DIV_ROUND_UP(offset + i->count, PAGE_SIZE);
@@ -1493,6 +1811,90 @@ void iov_iter_restore(struct iov_iter *i, struct iov_iter_state *state)
}
EXPORT_SYMBOL_FOR_MODULES(iov_iter_restore, "vmw_vsock_virtio_transport_common");
+/*
+ * Extract a list of virtually contiguous pages from an ITER_BVECQ iterator.
+ * This does not get references on the pages, nor does it get a pin on them.
+ */
+static ssize_t iov_iter_extract_bvecq_pages(struct iov_iter *iter,
+ struct page ***pages, size_t maxsize,
+ unsigned int maxpages,
+ iov_iter_extraction_t extraction_flags,
+ size_t *offset0)
+{
+ const struct bvecq *bvecq;
+ struct page **p;
+ unsigned int slot, nr = 0;
+ size_t extracted = 0, offset;
+
+ /* Count the next run of virtually contiguous pages. */
+ maxpages = iter_count_bvecq_pages(iter, maxsize, maxpages);
+ if (!maxpages)
+ return 0;
+
+ if (!*pages) {
+ *pages = kvmalloc_array(maxpages, sizeof(struct page *), GFP_KERNEL);
+ if (!*pages)
+ return -ENOMEM;
+ }
+
+ p = *pages;
+
+ /* Now transcribe the page pointers. */
+ extracted = 0;
+ bvecq = iter->bvecq;
+ offset = iter->iov_offset;
+ slot = iter->bvecq_slot;
+
+ do {
+ const struct bio_vec *bv;
+ size_t boff, blen;
+
+ if (slot >= bvecq->nr_slots) {
+ if (!bvecq->next) {
+ WARN_ON_ONCE(extracted < iter->count);
+ break;
+ }
+ bvecq = bvecq->next;
+ slot = 0;
+ offset = 0;
+ continue;
+ }
+
+ bv = &bvecq->bv[slot];
+ boff = bv->bv_offset;
+ blen = bv->bv_len;
+
+ /* bv_page is not allowed to be NULL unless bv_len == 0. */
+
+ if (offset < blen) {
+ size_t part = umin(maxsize - extracted, blen - offset);
+ size_t poff = (boff + offset) % PAGE_SIZE;
+ size_t pix = (boff + offset) / PAGE_SIZE;
+
+ if (poff + part > PAGE_SIZE)
+ part = PAGE_SIZE - poff;
+
+ if (!extracted)
+ *offset0 = poff;
+
+ p[nr++] = bv->bv_page + pix;
+ offset += part;
+ extracted += part;
+ }
+
+ if (offset >= blen) {
+ offset = 0;
+ slot++;
+ }
+ } while (nr < maxpages && extracted < maxsize);
+
+ iter->bvecq = bvecq;
+ iter->bvecq_slot = slot;
+ iter->iov_offset = offset;
+ iter->count -= extracted;
+ return extracted;
+}
+
/*
* Extract a list of contiguous pages from an ITER_FOLIOQ iterator. This does
* not get references on the pages, nor does it get a pin on them.
@@ -1853,6 +2255,10 @@ ssize_t iov_iter_extract_pages(struct iov_iter *i,
return iov_iter_extract_bvec_pages(i, pages, maxsize,
maxpages, extraction_flags,
offset0);
+ if (iov_iter_is_bvecq(i))
+ return iov_iter_extract_bvecq_pages(i, pages, maxsize,
+ maxpages, extraction_flags,
+ offset0);
if (iov_iter_is_folioq(i))
return iov_iter_extract_folioq_pages(i, pages, maxsize,
maxpages, extraction_flags,
diff --git a/lib/scatterlist.c b/lib/scatterlist.c
index 6ea40d2e6247..23e5a180103b 100644
--- a/lib/scatterlist.c
+++ b/lib/scatterlist.c
@@ -10,6 +10,7 @@
#include <linux/highmem.h>
#include <linux/kmemleak.h>
#include <linux/bvec.h>
+#include <linux/bvecq.h>
#include <linux/uio.h>
#include <linux/folio_queue.h>
@@ -1267,6 +1268,65 @@ static ssize_t extract_kvec_to_sg(struct iov_iter *iter,
return ret;
}
+/*
+ * Extract up to sg_max folios from an BVECQ-type iterator and add them to
+ * the scatterlist. The pages are not pinned.
+ */
+static ssize_t extract_bvecq_to_sg(struct iov_iter *iter,
+ ssize_t maxsize,
+ struct sg_table *sgtable,
+ unsigned int sg_max,
+ iov_iter_extraction_t extraction_flags)
+{
+ const struct bvecq *bvecq = iter->bvecq;
+ struct scatterlist *sg = sgtable->sgl + sgtable->nents;
+ unsigned int slot = iter->bvecq_slot;
+ ssize_t ret = 0;
+ size_t offset = iter->iov_offset;
+
+ maxsize = umin(maxsize, iov_iter_count(iter));
+
+ while (sg_max > 0 && ret < maxsize) {
+ const struct bio_vec *bv;
+ size_t blen, part;
+
+ if (slot >= bvecq->nr_slots) {
+ if (!bvecq->next) {
+ WARN_ON_ONCE(ret < iter->count);
+ break;
+ }
+ bvecq = bvecq->next;
+ slot = 0;
+ offset = 0;
+ continue;
+ }
+
+ bv = &bvecq->bv[slot];
+ blen = bv->bv_len;
+
+ if (offset >= blen) {
+ offset = 0;
+ slot++;
+ continue;
+ }
+
+ part = umin(maxsize - ret, blen - offset);
+
+ sg_set_page(sg, bv->bv_page, part, bv->bv_offset + offset);
+ sgtable->nents++;
+ sg++;
+ sg_max--;
+ offset += part;
+ ret += part;
+ }
+
+ iter->bvecq = bvecq;
+ iter->bvecq_slot = slot;
+ iter->iov_offset = offset;
+ iter->count -= ret;
+ return ret;
+}
+
/*
* Extract up to sg_max folios from an FOLIOQ-type iterator and add them to
* the scatterlist. The pages are not pinned.
@@ -1391,8 +1451,8 @@ static ssize_t extract_xarray_to_sg(struct iov_iter *iter,
* addition of @sg_max elements.
*
* The pages referred to by UBUF- and IOVEC-type iterators are extracted and
- * pinned; BVEC-, KVEC-, FOLIOQ- and XARRAY-type are extracted but aren't
- * pinned; DISCARD-type is not supported.
+ * pinned; BVEC-, BVECQ-, KVEC-, FOLIOQ- and XARRAY-type are extracted but
+ * aren't pinned; DISCARD-type is not supported.
*
* No end mark is placed on the scatterlist; that's left to the caller.
*
@@ -1424,6 +1484,9 @@ ssize_t extract_iter_to_sg(struct iov_iter *iter, size_t maxsize,
case ITER_KVEC:
return extract_kvec_to_sg(iter, maxsize, sgtable, sg_max,
extraction_flags);
+ case ITER_BVECQ:
+ return extract_bvecq_to_sg(iter, maxsize, sgtable, sg_max,
+ extraction_flags);
case ITER_FOLIOQ:
return extract_folioq_to_sg(iter, maxsize, sgtable, sg_max,
extraction_flags);
diff --git a/lib/tests/kunit_iov_iter.c b/lib/tests/kunit_iov_iter.c
index 32e42d8c7ca1..a3aeeca6ed58 100644
--- a/lib/tests/kunit_iov_iter.c
+++ b/lib/tests/kunit_iov_iter.c
@@ -12,6 +12,7 @@
#include <linux/mm.h>
#include <linux/uio.h>
#include <linux/bvec.h>
+#include <linux/bvecq.h>
#include <linux/folio_queue.h>
#include <linux/scatterlist.h>
#include <linux/minmax.h>
@@ -552,6 +553,183 @@ static void __init iov_kunit_copy_from_folioq(struct kunit *test)
KUNIT_SUCCEED(test);
}
+static void iov_kunit_destroy_bvecq(void *data)
+{
+ struct bvecq *bq, *next;
+
+ for (bq = data; bq; bq = next) {
+ next = bq->next;
+ /* The pages are freed by vmap with VM_MAP_PUT_PAGES. */
+ kfree(bq);
+ }
+}
+
+static struct bvecq *iov_kunit_alloc_bvecq(struct kunit *test, unsigned int max_slots)
+{
+ struct bvecq *bq;
+
+ bq = kzalloc(struct_size(bq, __bv, max_slots), GFP_KERNEL);
+ KUNIT_ASSERT_NOT_ERR_OR_NULL(test, bq);
+ bq->max_slots = max_slots;
+ bq->bv = bq->__bv;
+ bq->inline_bv = true;
+ return bq;
+}
+
+static struct bvecq *iov_kunit_create_bvecq(struct kunit *test, unsigned int max_slots)
+{
+ struct bvecq *bq;
+
+ bq = iov_kunit_alloc_bvecq(test, max_slots);
+ kunit_add_action_or_reset(test, iov_kunit_destroy_bvecq, bq);
+ return bq;
+}
+
+static void __init iov_kunit_load_bvecq(struct kunit *test,
+ struct iov_iter *iter, int dir,
+ struct bvecq *bq_head,
+ struct page **pages, size_t npages)
+{
+ struct bvecq *bq = bq_head;
+ size_t size = 0;
+
+ for (int i = 0; i < npages; i++) {
+ if (bq->nr_slots >= bq->max_slots) {
+ bq->next = iov_kunit_alloc_bvecq(test, 13);
+ bq->next->prev = bq;
+ bq = bq->next;
+ }
+ bvec_set_page(&bq->bv[bq->nr_slots], pages[i], PAGE_SIZE, 0);
+ bq->nr_slots++;
+ size += PAGE_SIZE;
+ }
+ iov_iter_bvec_queue(iter, dir, bq_head, 0, 0, size);
+}
+
+/*
+ * Test copying to a ITER_BVECQ-type iterator.
+ */
+static void __init iov_kunit_copy_to_bvecq(struct kunit *test)
+{
+ const struct kvec_test_range *pr;
+ struct iov_iter iter;
+ struct bvecq *bq;
+ struct page **spages, **bpages;
+ u8 *scratch, *buffer;
+ size_t bufsize, npages, size, copied;
+ int i, patt;
+
+ bufsize = 0x100000;
+ npages = bufsize / PAGE_SIZE;
+
+ bq = iov_kunit_create_bvecq(test, 13);
+
+ scratch = iov_kunit_create_buffer(test, &spages, npages);
+ for (i = 0; i < bufsize; i++)
+ scratch[i] = pattern(i);
+
+ buffer = iov_kunit_create_buffer(test, &bpages, npages);
+ memset(buffer, 0, bufsize);
+
+ iov_kunit_load_bvecq(test, &iter, READ, bq, bpages, npages);
+
+ i = 0;
+ for (pr = kvec_test_ranges; pr->from >= 0; pr++) {
+ size = pr->to - pr->from;
+ KUNIT_ASSERT_LE(test, pr->to, bufsize);
+
+ iov_iter_bvec_queue(&iter, READ, bq, 0, 0, pr->to);
+ iov_iter_advance(&iter, pr->from);
+ copied = copy_to_iter(scratch + i, size, &iter);
+
+ KUNIT_EXPECT_EQ(test, copied, size);
+ KUNIT_EXPECT_EQ(test, iter.count, 0);
+ i += size;
+ if (test->status == KUNIT_FAILURE)
+ goto stop;
+ }
+
+ /* Build the expected image in the scratch buffer. */
+ patt = 0;
+ memset(scratch, 0, bufsize);
+ for (pr = kvec_test_ranges; pr->from >= 0; pr++)
+ for (i = pr->from; i < pr->to; i++)
+ scratch[i] = pattern(patt++);
+
+ /* Compare the images */
+ for (i = 0; i < bufsize; i++) {
+ KUNIT_EXPECT_EQ_MSG(test, buffer[i], scratch[i], "at i=%x", i);
+ if (buffer[i] != scratch[i])
+ return;
+ }
+
+stop:
+ KUNIT_SUCCEED(test);
+}
+
+/*
+ * Test copying from a ITER_BVECQ-type iterator.
+ */
+static void __init iov_kunit_copy_from_bvecq(struct kunit *test)
+{
+ const struct kvec_test_range *pr;
+ struct iov_iter iter;
+ struct bvecq *bq;
+ struct page **spages, **bpages;
+ u8 *scratch, *buffer;
+ size_t bufsize, npages, size, copied;
+ int i, j;
+
+ bufsize = 0x100000;
+ npages = bufsize / PAGE_SIZE;
+
+ bq = iov_kunit_create_bvecq(test, 13);
+
+ buffer = iov_kunit_create_buffer(test, &bpages, npages);
+ for (i = 0; i < bufsize; i++)
+ buffer[i] = pattern(i);
+
+ scratch = iov_kunit_create_buffer(test, &spages, npages);
+ memset(scratch, 0, bufsize);
+
+ iov_kunit_load_bvecq(test, &iter, READ, bq, bpages, npages);
+
+ i = 0;
+ for (pr = kvec_test_ranges; pr->from >= 0; pr++) {
+ size = pr->to - pr->from;
+ KUNIT_ASSERT_LE(test, pr->to, bufsize);
+
+ iov_iter_bvec_queue(&iter, WRITE, bq, 0, 0, pr->to);
+ iov_iter_advance(&iter, pr->from);
+ copied = copy_from_iter(scratch + i, size, &iter);
+
+ KUNIT_EXPECT_EQ(test, copied, size);
+ KUNIT_EXPECT_EQ(test, iter.count, 0);
+ i += size;
+ }
+
+ /* Build the expected image in the main buffer. */
+ i = 0;
+ memset(buffer, 0, bufsize);
+ for (pr = kvec_test_ranges; pr->from >= 0; pr++) {
+ for (j = pr->from; j < pr->to; j++) {
+ buffer[i++] = pattern(j);
+ if (i >= bufsize)
+ goto stop;
+ }
+ }
+stop:
+
+ /* Compare the images */
+ for (i = 0; i < bufsize; i++) {
+ KUNIT_EXPECT_EQ_MSG(test, scratch[i], buffer[i], "at i=%x", i);
+ if (scratch[i] != buffer[i])
+ return;
+ }
+
+ KUNIT_SUCCEED(test);
+}
+
static void iov_kunit_destroy_xarray(void *data)
{
struct xarray *xarray = data;
@@ -867,6 +1045,85 @@ static void __init iov_kunit_extract_pages_bvec(struct kunit *test)
KUNIT_SUCCEED(test);
}
+/*
+ * Test the extraction of ITER_BVECQ-type iterators.
+ */
+static void __init iov_kunit_extract_pages_bvecq(struct kunit *test)
+{
+ const struct kvec_test_range *pr;
+ struct iov_iter iter;
+ struct bvecq *bq;
+ struct page **bpages, *pagelist[8], **pages = pagelist;
+ ssize_t len;
+ size_t bufsize, size = 0, npages;
+ int i, from;
+
+ bufsize = 0x100000;
+ npages = bufsize / PAGE_SIZE;
+
+ bq = iov_kunit_create_bvecq(test, 13);
+
+ iov_kunit_create_buffer(test, &bpages, npages);
+ iov_kunit_load_bvecq(test, &iter, READ, bq, bpages, npages);
+
+ for (pr = kvec_test_ranges; pr->from >= 0; pr++) {
+ from = pr->from;
+ size = pr->to - from;
+ KUNIT_ASSERT_LE(test, pr->to, bufsize);
+
+ iov_iter_bvec_queue(&iter, WRITE, bq, 0, 0, pr->to);
+ iov_iter_advance(&iter, from);
+
+ do {
+ size_t offset0 = LONG_MAX;
+
+ for (i = 0; i < ARRAY_SIZE(pagelist); i++)
+ pagelist[i] = (void *)(unsigned long)0xaa55aa55aa55aa55ULL;
+
+ len = iov_iter_extract_pages(&iter, &pages, 100 * 1024,
+ ARRAY_SIZE(pagelist), 0, &offset0);
+ KUNIT_EXPECT_GE(test, len, 0);
+ if (len < 0)
+ break;
+ KUNIT_EXPECT_LE(test, len, size);
+ KUNIT_EXPECT_EQ(test, iter.count, size - len);
+ if (len == 0)
+ break;
+ size -= len;
+ KUNIT_EXPECT_GE(test, (ssize_t)offset0, 0);
+ KUNIT_EXPECT_LT(test, offset0, PAGE_SIZE);
+
+ for (i = 0; i < ARRAY_SIZE(pagelist); i++) {
+ struct page *p;
+ ssize_t part = min_t(ssize_t, len, PAGE_SIZE - offset0);
+ int ix;
+
+ KUNIT_ASSERT_GE(test, part, 0);
+ ix = from / PAGE_SIZE;
+ KUNIT_ASSERT_LT(test, ix, npages);
+ p = bpages[ix];
+ KUNIT_EXPECT_PTR_EQ(test, pagelist[i], p);
+ KUNIT_EXPECT_EQ(test, offset0, from % PAGE_SIZE);
+ from += part;
+ len -= part;
+ KUNIT_ASSERT_GE(test, len, 0);
+ if (len == 0)
+ break;
+ offset0 = 0;
+ }
+
+ if (test->status == KUNIT_FAILURE)
+ goto stop;
+ } while (iov_iter_count(&iter) > 0);
+
+ KUNIT_EXPECT_EQ(test, size, 0);
+ KUNIT_EXPECT_EQ(test, iter.count, 0);
+ }
+
+stop:
+ KUNIT_SUCCEED(test);
+}
+
/*
* Test the extraction of ITER_FOLIOQ-type iterators.
*/
@@ -1226,12 +1483,15 @@ static struct kunit_case __refdata iov_kunit_cases[] = {
KUNIT_CASE(iov_kunit_copy_from_kvec),
KUNIT_CASE(iov_kunit_copy_to_bvec),
KUNIT_CASE(iov_kunit_copy_from_bvec),
+ KUNIT_CASE(iov_kunit_copy_to_bvecq),
+ KUNIT_CASE(iov_kunit_copy_from_bvecq),
KUNIT_CASE(iov_kunit_copy_to_folioq),
KUNIT_CASE(iov_kunit_copy_from_folioq),
KUNIT_CASE(iov_kunit_copy_to_xarray),
KUNIT_CASE(iov_kunit_copy_from_xarray),
KUNIT_CASE(iov_kunit_extract_pages_kvec),
KUNIT_CASE(iov_kunit_extract_pages_bvec),
+ KUNIT_CASE(iov_kunit_extract_pages_bvecq),
KUNIT_CASE(iov_kunit_extract_pages_folioq),
KUNIT_CASE(iov_kunit_extract_pages_xarray),
KUNIT_CASE(iov_kunit_iter_to_sg_kvec),
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 03/10] netfs: Add some tools for managing bvecq chains
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
2026-09-29 7:59 ` [PATCH v12 01/10] Add a function to kmap one page of a multipage bio_vec David Howells
2026-09-29 7:59 ` [PATCH v12 02/10] iov_iter: Add a segmented queue of bio_vec[] David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 04/10] afs: Use a bvecq to hold dir content rather than folioq David Howells
` (7 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Christoph Hellwig
Provide a selection of tools for managing bvec queue chains and a memory
pool for writeback to draw upon if needed. The tools include the
allocation, prepopulation and refcounting of bvecqs and bvecq chains. When
a chain is being cleaned up, the memory segments will be appropriately
disposed off according to the flags on each bvecq in that chain, allowing
for mixed memory cleanup types.
The main use for this is to replace the use of both folio_queues and
kvmalloc'd bio_vec arrays in netfslib and creating a structure that allows
network filesystems to easily build an entire message by gluing bits of
protocol on either side of the bvecq chains and also glue chains together
for sparse writes and compound operations. This allows TCP-based
transports to pass a message in a single sendmsg() call and avoid corking
as this makes things more efficient.
This will also be used to do things like creating an encryption buffer in
cifs or a directory content buffer in afs.
Signed-off-by: David Howells <dhowells@redhat.com>
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Christoph Hellwig <hch@infradead.org>
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
fs/netfs/Makefile | 1 +
fs/netfs/bvecq.c | 342 +++++++++++++++++++++++++++++++++++
fs/netfs/internal.h | 2 +
fs/netfs/main.c | 8 +
fs/netfs/stats.c | 4 +-
include/linux/bvecq.h | 116 ++++++++++++
include/linux/netfs.h | 1 +
include/trace/events/netfs.h | 24 +++
8 files changed, 497 insertions(+), 1 deletion(-)
create mode 100644 fs/netfs/bvecq.c
diff --git a/fs/netfs/Makefile b/fs/netfs/Makefile
index 54834cde7e56..b1ea4439c1bb 100644
--- a/fs/netfs/Makefile
+++ b/fs/netfs/Makefile
@@ -3,6 +3,7 @@
netfs-y := \
buffered_read.o \
buffered_write.o \
+ bvecq.o \
direct_read.o \
direct_write.o \
iterator.o \
diff --git a/fs/netfs/bvecq.c b/fs/netfs/bvecq.c
new file mode 100644
index 000000000000..56ec6d8953b4
--- /dev/null
+++ b/fs/netfs/bvecq.c
@@ -0,0 +1,342 @@
+// SPDX-License-Identifier: GPL-2.0-only
+/* Buffering helpers for bvec queues
+ *
+ * Copyright (C) 2026 Red Hat, Inc. All Rights Reserved.
+ * Written by David Howells (dhowells@redhat.com)
+ */
+
+#include <linux/bvecq.h>
+#include "internal.h"
+
+void bvecq_dump(const struct bvecq *bq)
+{
+ int b = 0;
+
+ for (; bq; bq = bvecq_next(bq), b++) {
+ int skipz = 0;
+
+ pr_notice("BQ[%u] %u/%u\n", b, bq->nr_slots, bq->max_slots);
+ for (int s = 0; s < bq->nr_slots; s++) {
+ const struct bio_vec *bv = &bq->bv[s];
+
+ if (!bv->bv_page && !bv->bv_len && skipz < 2) {
+ skipz = 1;
+ continue;
+ }
+ if (skipz == 1)
+ pr_notice("BQ[%u:00-%02u] ...\n", b, s - 1);
+ skipz = 2;
+ pr_notice("BQ[%u:%02u] %10lx %04x %04x %u\n",
+ b, s,
+ bv->bv_page ? page_to_pfn(bv->bv_page) : 0,
+ bv->bv_offset, bv->bv_len,
+ bv->bv_page ? page_count(bv->bv_page) : 0);
+ }
+ }
+}
+EXPORT_SYMBOL(bvecq_dump);
+
+/**
+ * bvecq_alloc_one - Allocate a single bvecq node with unpopulated slots
+ * @nr_slots: Number of slots to allocate
+ * @gfp: The allocation constraints.
+ * @for_writeback: True if allocating for writeback
+ *
+ * Allocate a single bvecq node and initialise the header. The allocation is
+ * rounded up to the size of the smallest slab granule that will accommodate it
+ * and all the remaining space is set as an inline slot array with max_slots
+ * set to the number of slots. The slot array is not initialised.
+ *
+ * If @for_writeback is set, then the emergency mempool may be used for
+ * allocation. If it does allocate from that pool, the maximum number of slots
+ * available will be BVECQ_POOL_SLOTS which may be less than requested. Also,
+ * if @for_writeback is set, the function will not fail - though it may have to
+ * wait.
+ *
+ * Return: The node pointer or NULL on allocation failure.
+ */
+struct bvecq *bvecq_alloc_one(size_t nr_slots, gfp_t gfp, bool for_writeback)
+{
+ struct bvecq *bq;
+ const size_t pool_size = struct_size_t(struct bvecq, __bv, BVECQ_POOL_SLOTS);
+ size_t size;
+ bool from_pool = false;
+
+ size = kmalloc_size_roundup(struct_size(bq, __bv, nr_slots));
+ gfp &= ~(GFP_ZONEMASK | __GFP_THISNODE);
+
+ if (for_writeback) {
+ if (size != pool_size) {
+ gfp_t gfp_temp = gfp;
+
+ gfp_temp |= __GFP_NOMEMALLOC | __GFP_NORETRY | __GFP_NOWARN;
+ gfp_temp &= ~(__GFP_DIRECT_RECLAIM | __GFP_IO);
+ bq = kmalloc(size, gfp_temp);
+ if (bq)
+ goto success;
+ }
+
+ bq = mempool_alloc_noreserve(&netfs_bvecq_pool, gfp);
+ if (!bq)
+ return bq;
+ from_pool = true;
+ size = pool_size;
+ } else {
+ bq = kmalloc(size, gfp);
+ if (!bq)
+ return bq;
+ }
+
+success:
+ *bq = (struct bvecq) {
+ .ref = REFCOUNT_INIT(1),
+ .bv = bq->__bv,
+ .inline_bv = true,
+ .max_slots = (size - sizeof(*bq)) / sizeof(bq->__bv[0]),
+ .from_pool = from_pool,
+ };
+ netfs_stat(&netfs_n_bvecq);
+ return bq;
+}
+EXPORT_SYMBOL(bvecq_alloc_one);
+
+/**
+ * bvecq_alloc_chain - Allocate an unpopulated bvecq chain
+ * @nr_slots: Number of slots to allocate
+ * @gfp: The allocation constraints.
+ * @for_writeback: True if allocating for writeback
+ *
+ * Allocate a chain of bvecq nodes providing at least the requested cumulative
+ * number of slots. Each node is a maximum of 4KiB in size.
+ *
+ * Return: The first node pointer or NULL on allocation failure.
+ */
+struct bvecq *bvecq_alloc_chain(size_t nr_slots, gfp_t gfp, bool for_writeback)
+{
+ struct bvecq *head = NULL, *tail = NULL;
+
+ _enter("%zu", nr_slots);
+
+ for (;;) {
+ struct bvecq *bq;
+
+ bq = bvecq_alloc_one(min(nr_slots, BVECQ_4KB_SLOTS), gfp, for_writeback);
+ if (!bq)
+ goto oom;
+
+ if (tail)
+ bvecq_append(tail, bq);
+ else
+ head = bq;
+ tail = bq;
+ if (tail->max_slots >= nr_slots)
+ break;
+ nr_slots -= tail->max_slots;
+ }
+
+ return head;
+oom:
+ bvecq_put(head);
+ return NULL;
+}
+EXPORT_SYMBOL(bvecq_alloc_chain);
+
+/**
+ * bvecq_alloc_buffer2 - Allocate a bvecq chain and populate with buffers
+ * @size: Target size of the buffer (can be 0 for an empty buffer)
+ * @pre_slots: Number of preamble slots to set aside
+ * @gfp: The allocation constraints.
+ * @for_writeback: True if allocating for writeback
+ *
+ * Allocate a chain of bvecq nodes and populate the slots with sufficient pages
+ * to provide at least the requested amount of space, leaving the first
+ * @pre_slots slots unset. The pre-slots must all fit into the the first
+ * bvecq.
+ *
+ * The pages allocated may be compound pages larger than PAGE_SIZE and thus
+ * occupy fewer slots. The pages have their refcounts set to 1 and can be
+ * passed to MSG_SPLICE_PAGES.
+ *
+ * Return: The first node pointer or NULL on allocation failure.
+ */
+struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp,
+ bool for_writeback)
+{
+ struct bvecq *head = NULL, *p = NULL;
+ size_t nr_per_bq = BVECQ_POOL_SLOTS;
+ size_t count = pre_slots + DIV_ROUND_UP(size, PAGE_SIZE);
+
+ _enter("%zx,%zx,%u", size, count, pre_slots);
+
+ if (WARN_ON_ONCE(pre_slots > nr_per_bq))
+ return NULL;
+
+ head = bvecq_alloc_chain(count, gfp, for_writeback);
+ if (!head)
+ return NULL;
+
+ p = head;
+ do {
+ struct page **pages;
+ size_t unused, want, got, slot;
+
+ if (!count)
+ break;
+ if (WARN_ON_ONCE(!p))
+ goto oom;
+
+ if (p->nr_slots == 0) {
+ /* Need to clear pre slots and pages[], so just clear all. */
+ memset(p->bv, 0, p->max_slots * sizeof(p->bv[0]));
+ p->mem_type = BVECQ_MEM_ALLOCED;
+ p->nr_slots = pre_slots;
+ count -= pre_slots;
+ pre_slots = 0;
+ if (!count)
+ break;
+ }
+
+ if (p->nr_slots >= p->max_slots) {
+ p = p->next;
+ continue;
+ }
+ unused = p->max_slots - p->nr_slots;
+
+ pages = (struct page **)&p->bv[p->max_slots];
+ pages -= unused;
+
+ want = min(count, unused);
+ got = alloc_pages_bulk(gfp, want, pages);
+ if (!got)
+ goto oom;
+
+ slot = p->nr_slots;
+ for (int i = 0; i < got; i++) {
+ set_page_count(pages[i], 1);
+ bvec_set_page(&p->bv[slot++], pages[i], PAGE_SIZE, 0);
+ }
+
+ bvecq_filled_to(p, slot);
+ count -= got;
+ } while (count > 0);
+
+ return head;
+oom:
+ bvecq_put(head);
+ return NULL;
+}
+EXPORT_SYMBOL(bvecq_alloc_buffer2);
+
+/*
+ * Free the page pointed to by a slot as necessary.
+ */
+static void bvecq_free_slot(struct bvecq *bq, unsigned int slot)
+{
+ struct page *page = bq->bv[slot].bv_page;
+
+ if (!page)
+ return;
+
+ switch (bq->mem_type) {
+ case BVECQ_MEM_EXTERNAL:
+ break;
+ case BVECQ_MEM_PAGECACHE:
+ put_page(page);
+ break;
+ case BVECQ_MEM_GUP:
+ unpin_user_page(page);
+ break;
+ case BVECQ_MEM_ALLOCED:
+ __free_pages(page, compound_order(page));
+ break;
+ default:
+ WARN_ON_ONCE(1);
+ break;
+ }
+}
+
+/**
+ * bvecq_put - Put a ref on a bvec queue
+ * @bq: The start of the folio queue to free
+ *
+ * Put the ref(s) on the nodes in a bvec queue, freeing up the node and the
+ * page fragments it points to as the refcounts become zero.
+ */
+void bvecq_put(struct bvecq *bq)
+{
+ struct bvecq *next;
+
+ for (; bq; bq = next) {
+ if (!refcount_dec_and_test(&bq->ref))
+ break;
+ for (int slot = 0; slot < bq->nr_slots; slot++)
+ bvecq_free_slot(bq, slot);
+ next = bq->next;
+ netfs_stat_d(&netfs_n_bvecq);
+ if (bq->from_pool)
+ mempool_free(bq, &netfs_bvecq_pool);
+ else
+ kfree(bq);
+ }
+}
+EXPORT_SYMBOL(bvecq_put);
+
+/**
+ * bvecq_expand_buffer - Allocate buffer space into a bvec queue
+ * @_buffer: Pointer to the bvecq chain to expand (may point to a NULL; updated).
+ * @_cur_size: Current size of the buffer (updated).
+ * @size: Target size of the buffer.
+ * @gfp: The allocation constraints.
+ *
+ * Append extra pages to a buffer to increase its capacity to the @size
+ * specified. If the current tail has space, but is not of the
+ * BVECQ_MEM_ALLOCED memory type, a separate bvecq will be allocated to hold
+ * the new memory.
+ */
+int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size, gfp_t gfp)
+{
+ struct bvecq *tail = *_buffer;
+
+ size = round_up(size, PAGE_SIZE);
+ if (tail)
+ while (tail->next)
+ tail = tail->next;
+
+ while (*_cur_size < size) {
+ struct page *page;
+ size_t need = size - *_cur_size;
+ int order = 0;
+
+ if (!tail || bvecq_is_full(tail) || tail->mem_type != BVECQ_MEM_ALLOCED) {
+ struct bvecq *p;
+
+ p = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, false);
+ if (!p)
+ return -ENOMEM;
+ if (tail)
+ bvecq_append(tail, p);
+ else
+ *_buffer = p;
+ tail = p;
+ p->mem_type = BVECQ_MEM_ALLOCED;
+ }
+
+ if (need > PAGE_SIZE)
+ order = umin(ilog2(need) - PAGE_SHIFT, MAX_PAGECACHE_ORDER);
+
+ page = alloc_pages(gfp | __GFP_COMP | __GFP_NORETRY | __GFP_NOWARN, order);
+ if (!page && order > 0) {
+ page = alloc_pages(gfp | __GFP_COMP, 0);
+ order = 0;
+ }
+ if (!page)
+ return -ENOMEM;
+
+ bvec_set_page(&tail->bv[tail->nr_slots], page, PAGE_SIZE << order, 0);
+ *_cur_size += PAGE_SIZE << order;
+ bvecq_filled_to(tail, tail->nr_slots + 1);
+ }
+
+ return 0;
+}
+EXPORT_SYMBOL(bvecq_expand_buffer);
diff --git a/fs/netfs/internal.h b/fs/netfs/internal.h
index b8591abc90a9..dd20a9f201be 100644
--- a/fs/netfs/internal.h
+++ b/fs/netfs/internal.h
@@ -45,6 +45,7 @@ extern struct list_head netfs_io_requests;
extern spinlock_t netfs_proc_lock;
extern mempool_t netfs_request_pool;
extern mempool_t netfs_subrequest_pool;
+extern mempool_t netfs_bvecq_pool;
extern mempool_t netfs_folioq_pool;
#ifdef CONFIG_PROC_FS
@@ -203,6 +204,7 @@ extern atomic_t netfs_n_wh_retry_write_subreq;
extern atomic_t netfs_n_wb_lock_skip;
extern atomic_t netfs_n_wb_lock_wait;
extern atomic_t netfs_n_folioq;
+extern atomic_t netfs_n_bvecq;
int netfs_stats_show(struct seq_file *m, void *v);
diff --git a/fs/netfs/main.c b/fs/netfs/main.c
index 609e22e8f76a..5c8517a66dde 100644
--- a/fs/netfs/main.c
+++ b/fs/netfs/main.c
@@ -28,6 +28,7 @@ static struct kmem_cache *netfs_request_slab;
static struct kmem_cache *netfs_subrequest_slab;
mempool_t netfs_request_pool;
mempool_t netfs_subrequest_pool;
+mempool_t netfs_bvecq_pool;
mempool_t netfs_folioq_pool;
#ifdef CONFIG_PROC_FS
@@ -111,6 +112,10 @@ static int __init netfs_init(void)
if (mempool_init_kmalloc_pool(&netfs_folioq_pool, 100, sizeof(struct folio_queue)) < 0)
goto error_folioq_pool;
+ if (mempool_init_kmalloc_pool(&netfs_bvecq_pool, 100,
+ struct_size_t(struct bvecq, __bv, BVECQ_POOL_SLOTS)) < 0)
+ goto error_bvecq_pool;
+
netfs_request_slab = kmem_cache_create("netfs_request",
sizeof(struct netfs_io_request), 0,
SLAB_HWCACHE_ALIGN | SLAB_ACCOUNT,
@@ -163,6 +168,8 @@ static int __init netfs_init(void)
error_reqpool:
kmem_cache_destroy(netfs_request_slab);
error_req:
+ mempool_exit(&netfs_bvecq_pool);
+error_bvecq_pool:
mempool_exit(&netfs_folioq_pool);
error_folioq_pool:
return ret;
@@ -177,6 +184,7 @@ static void __exit netfs_exit(void)
kmem_cache_destroy(netfs_subrequest_slab);
mempool_exit(&netfs_request_pool);
kmem_cache_destroy(netfs_request_slab);
+ mempool_exit(&netfs_bvecq_pool);
mempool_exit(&netfs_folioq_pool);
}
module_exit(netfs_exit);
diff --git a/fs/netfs/stats.c b/fs/netfs/stats.c
index 9a607c4e62dd..a10d34f88597 100644
--- a/fs/netfs/stats.c
+++ b/fs/netfs/stats.c
@@ -47,6 +47,7 @@ atomic_t netfs_n_wh_retry_write_subreq;
atomic_t netfs_n_wb_lock_skip;
atomic_t netfs_n_wb_lock_wait;
atomic_t netfs_n_folioq;
+atomic_t netfs_n_bvecq;
int netfs_stats_show(struct seq_file *m, void *v)
{
@@ -88,9 +89,10 @@ int netfs_stats_show(struct seq_file *m, void *v)
atomic_read(&netfs_n_rh_retry_read_subreq),
atomic_read(&netfs_n_wh_retry_write_req),
atomic_read(&netfs_n_wh_retry_write_subreq));
- seq_printf(m, "Objs : rr=%u sr=%u foq=%u wsc=%u\n",
+ seq_printf(m, "Objs : rr=%u sr=%u bq=%u foq=%u wsc=%u\n",
atomic_read(&netfs_n_rh_rreq),
atomic_read(&netfs_n_rh_sreq),
+ atomic_read(&netfs_n_bvecq),
atomic_read(&netfs_n_folioq),
atomic_read(&netfs_n_wh_wstream_conflict));
seq_printf(m, "WbLock : skip=%u wait=%u\n",
diff --git a/include/linux/bvecq.h b/include/linux/bvecq.h
index 77fd07852c33..b984aaa44908 100644
--- a/include/linux/bvecq.h
+++ b/include/linux/bvecq.h
@@ -43,8 +43,124 @@ struct bvecq {
u16 max_slots; /* Number of elements allocated in bv[] */
enum bvecq_mem mem_type:3; /* What sort of memory and how to free it */
bool inline_bv:1; /* T if __bv[] is being used */
+ bool from_pool:1; /* T if bvecq from mempool */
struct bio_vec *bv; /* Pointer to array of page fragments */
struct bio_vec __bv[]; /* Default array (if ->inline_bv) */
};
+/* Number of slots in a 512-byte mempool-backed bvecq. */
+#define BVECQ_POOL_SLOTS ((512 - sizeof(struct bvecq)) / sizeof(struct bio_vec))
+
+/* Number of slots in a 4K bvecq. */
+#define BVECQ_4KB_SLOTS ((4096 - sizeof(struct bvecq)) / sizeof(struct bio_vec))
+
+void bvecq_dump(const struct bvecq *bq);
+struct bvecq *bvecq_alloc_one(size_t nr_slots, gfp_t gfp, bool for_writeback);
+struct bvecq *bvecq_alloc_chain(size_t nr_slots, gfp_t gfp, bool for_writeback);
+struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp,
+ bool for_writeback);
+void bvecq_put(struct bvecq *bq);
+int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size, gfp_t gfp);
+
+/**
+ * bvecq_alloc_buffer - Allocate a bvecq chain and populate with buffers
+ * @size: Target size of the buffer (can be 0 for an empty buffer)
+ * @gfp: The allocation constraints.
+ * @for_writeback: True if allocating for writeback
+ *
+ * Wrapper around %bvecq_alloc_buffer2().
+ */
+static inline struct bvecq *bvecq_alloc_buffer(size_t size, gfp_t gfp, bool for_writeback)
+{
+ return bvecq_alloc_buffer2(size, 0, gfp, for_writeback);
+}
+
+/**
+ * bvecq_get - Get a ref on a bvecq
+ * @bq: The bvecq to get a ref on
+ */
+static inline struct bvecq *bvecq_get(struct bvecq *bq)
+{
+ refcount_inc(&bq->ref);
+ return bq;
+}
+
+/**
+ * bvecq_is_full - Determine if a bvecq is full
+ * @bvecq: The object to query
+ *
+ * Return: true if full; false if not.
+ */
+static inline bool bvecq_is_full(const struct bvecq *bvecq)
+{
+ return bvecq->nr_slots >= bvecq->max_slots;
+}
+
+/**
+ * bvecq_filled_to - Release filled slots with release barrier
+ * @bvecq: The object modified
+ * @to: The latest slot filled + 1
+ */
+static inline void bvecq_filled_to(struct bvecq *bvecq, unsigned int to)
+{
+ /* Set the slot counter after filling the slot */
+ smp_store_release(&bvecq->nr_slots, to);
+}
+
+/**
+ * bvecq_nr_slots_acquire - Get the number of filled slots with acquire barrier
+ * @bvecq: The object to query
+ *
+ * Return: The number of filled slots
+ */
+static inline unsigned int bvecq_nr_slots_acquire(const struct bvecq *bvecq)
+{
+ /* Read the slot counter before looking at the slot */
+ return smp_load_acquire(&bvecq->nr_slots);
+}
+
+/**
+ * bvecq_acquire_slot - Determine if a slot is valid with acquire barrier
+ * @bvecq: The object to query
+ * @slot: The next slot
+ *
+ * Return: true if valid; false if might not be valid
+ */
+static inline bool bvecq_acquire_slot(const struct bvecq *bvecq, unsigned int slot)
+{
+ /* Read the slot counter before looking at the slot */
+ return slot < bvecq_nr_slots_acquire(bvecq);
+}
+
+/**
+ * bvecq_append - Get the next bvecq with appropriate barrier
+ * @to: The bvecq to append to
+ * @add: The bvecq to append
+ *
+ * Attach a new bvecq to a chain using an appropriate barrier to protect the
+ * write.
+ *
+ * [!] Note that this function transfers the caller's ref to the chain.
+ */
+static inline void bvecq_append(struct bvecq *to, struct bvecq *add)
+{
+ add->prev = to;
+
+ /* Make sure the initialisation is stored before the next pointer. */
+ smp_store_release(&to->next, add);
+}
+
+/**
+ * bvecq_next - Get the next bvecq with appropriate barrier
+ * @bq: The bvecq to start from
+ *
+ * Return the next bvecq in a chain, using an appropriate barrier to protect
+ * the access.
+ */
+static inline struct bvecq *bvecq_next(const struct bvecq *bq)
+{
+ /* Read the contents of the next node after the pointer to it. */
+ return smp_load_acquire(&bq->next);
+}
+
#endif /* _LINUX_BVECQ_H */
diff --git a/include/linux/netfs.h b/include/linux/netfs.h
index 67e010b6994b..1723bcda86ba 100644
--- a/include/linux/netfs.h
+++ b/include/linux/netfs.h
@@ -17,6 +17,7 @@
#include <linux/workqueue.h>
#include <linux/fs.h>
#include <linux/pagemap.h>
+#include <linux/bvecq.h>
#include <linux/uio.h>
#include <linux/rolling_buffer.h>
diff --git a/include/trace/events/netfs.h b/include/trace/events/netfs.h
index bf1e1f185b05..3fe3d47ba55b 100644
--- a/include/trace/events/netfs.h
+++ b/include/trace/events/netfs.h
@@ -817,6 +817,30 @@ TRACE_EVENT(netfs_read_progress_at,
__entry->rreq, __entry->cleaned_to, __entry->progress_at)
);
+TRACE_EVENT(netfs_bv_slot,
+ TP_PROTO(const struct bvecq *bq, int slot),
+
+ TP_ARGS(bq, slot),
+
+ TP_STRUCT__entry(
+ __field(unsigned long, pfn)
+ __field(unsigned int, offset)
+ __field(unsigned int, len)
+ __field(unsigned int, slot)
+ ),
+
+ TP_fast_assign(
+ __entry->slot = slot;
+ __entry->pfn = page_to_pfn(bq->bv[slot].bv_page);
+ __entry->offset = bq->bv[slot].bv_offset;
+ __entry->len = bq->bv[slot].bv_len;
+ ),
+
+ TP_printk("bq[%x] p=%lx %x-%x",
+ __entry->slot,
+ __entry->pfn, __entry->offset, __entry->offset + __entry->len)
+ );
+
#undef EM
#undef E_
#endif /* _TRACE_NETFS_H */
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 04/10] afs: Use a bvecq to hold dir content rather than folioq
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
` (2 preceding siblings ...)
2026-09-29 7:59 ` [PATCH v12 03/10] netfs: Add some tools for managing bvecq chains David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 05/10] cifs: Use a bvecq for buffering instead of a folioq David Howells
` (6 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Christoph Hellwig
Use a bvecq to hold the contents of a directory rather than the folioq so
that the latter can be phased out.
Signed-off-by: David Howells <dhowells@redhat.com>
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Marc Dionne <marc.dionne@auristor.com>
cc: Christoph Hellwig <hch@infradead.org>
cc: linux-afs@lists.infradead.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
fs/afs/dir.c | 35 +++++----
fs/afs/dir_edit.c | 43 +++++------
fs/afs/dir_search.c | 33 ++++-----
fs/afs/inode.c | 2 +-
fs/afs/internal.h | 6 +-
fs/afs/symlink.c | 37 +++++-----
fs/netfs/write_issue.c | 164 ++++++-----------------------------------
7 files changed, 96 insertions(+), 224 deletions(-)
diff --git a/fs/afs/dir.c b/fs/afs/dir.c
index 2db534a2c7cc..9ea7930185a5 100644
--- a/fs/afs/dir.c
+++ b/fs/afs/dir.c
@@ -140,9 +140,9 @@ static void afs_dir_dump(struct afs_vnode *dvnode)
pr_warn("DIR %llx:%llx is=%llx\n",
dvnode->fid.vid, dvnode->fid.vnode, i_size);
- iov_iter_folio_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0, i_size);
- iterate_folioq(&iter, iov_iter_count(&iter), NULL, NULL,
- afs_dir_dump_step);
+ iov_iter_bvec_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0, i_size);
+ iterate_bvecq(&iter, iov_iter_count(&iter), NULL, NULL,
+ afs_dir_dump_step);
}
/*
@@ -203,9 +203,9 @@ static int afs_dir_check(struct afs_vnode *dvnode)
if (unlikely(!i_size))
return 0;
- iov_iter_folio_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0, i_size);
- checked = iterate_folioq(&iter, iov_iter_count(&iter), dvnode, NULL,
- afs_dir_check_step);
+ iov_iter_bvec_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0, i_size);
+ checked = iterate_bvecq(&iter, iov_iter_count(&iter), dvnode, NULL,
+ afs_dir_check_step);
if (checked != i_size) {
afs_dir_dump(dvnode);
return -EIO;
@@ -250,15 +250,14 @@ static ssize_t afs_do_read_single(struct afs_vnode *dvnode, struct file *file)
if (dvnode->directory_size < i_size) {
size_t cur_size = dvnode->directory_size;
- ret = netfs_alloc_folioq_buffer(NULL,
- &dvnode->directory, &cur_size, i_size,
- mapping_gfp_mask(dvnode->netfs.inode.i_mapping));
+ ret = bvecq_expand_buffer(&dvnode->directory, &cur_size,
+ round_up(i_size, PAGE_SIZE), GFP_KERNEL);
dvnode->directory_size = cur_size;
if (ret < 0)
return ret;
}
- iov_iter_folio_queue(&iter, ITER_DEST, dvnode->directory, 0, 0, dvnode->directory_size);
+ iov_iter_bvec_queue(&iter, ITER_DEST, dvnode->directory, 0, 0, dvnode->directory_size);
/* AFS requires us to perform the read of a directory synchronously as
* a single unit to avoid issues with the directory contents being
@@ -294,8 +293,8 @@ static ssize_t afs_read_single(struct afs_vnode *dvnode, struct file *file)
}
/*
- * Read the directory into a folio_queue buffer in one go, scrubbing the
- * previous contents. We return -ESTALE if the caller needs to call us again.
+ * Read the directory into the buffer in one go, scrubbing the previous
+ * contents. We return -ESTALE if the caller needs to call us again.
*/
ssize_t afs_read_dir(struct afs_vnode *dvnode, struct file *file)
__acquires(&dvnode->validate_lock)
@@ -483,7 +482,7 @@ static size_t afs_dir_iterate_step(void *iter_base, size_t progress, size_t len,
}
/*
- * Iterate through the directory folios.
+ * Iterate through the directory content.
*/
static int afs_dir_iterate_contents(struct inode *dir, struct dir_context *dir_ctx)
{
@@ -498,11 +497,11 @@ static int afs_dir_iterate_contents(struct inode *dir, struct dir_context *dir_c
if (i_size <= 0 || dir_ctx->pos >= i_size)
return 0;
- iov_iter_folio_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0, i_size);
+ iov_iter_bvec_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0, i_size);
iov_iter_advance(&iter, round_down(dir_ctx->pos, AFS_DIR_BLOCK_SIZE));
- iterate_folioq(&iter, iov_iter_count(&iter), dvnode, &ctx,
- afs_dir_iterate_step);
+ iterate_bvecq(&iter, iov_iter_count(&iter), dvnode, &ctx,
+ afs_dir_iterate_step);
if (ctx.error == -ESTALE)
afs_invalidate_dir(dvnode, afs_dir_invalid_iter_stale);
@@ -2228,8 +2227,8 @@ static int afs_dir_writepages(struct address_space *mapping,
}
if (test_bit(AFS_VNODE_DIR_VALID, &dvnode->flags)) {
- iov_iter_folio_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0,
- i_size_read(&dvnode->netfs.inode));
+ iov_iter_bvec_queue(&iter, ITER_SOURCE, dvnode->directory, 0, 0,
+ i_size_read(&dvnode->netfs.inode));
ret = netfs_writeback_single(mapping, wbc, &iter);
if (ret == 1)
ret = 0; /* Skipped write due to lock conflict. */
diff --git a/fs/afs/dir_edit.c b/fs/afs/dir_edit.c
index c31303059444..5f00c41403e5 100644
--- a/fs/afs/dir_edit.c
+++ b/fs/afs/dir_edit.c
@@ -10,7 +10,6 @@
#include <linux/namei.h>
#include <linux/pagemap.h>
#include <linux/iversion.h>
-#include <linux/folio_queue.h>
#include "internal.h"
#include "xdr_fs.h"
@@ -110,9 +109,8 @@ static void afs_clear_contig_bits(union afs_xdr_dir_block *block,
*/
static union afs_xdr_dir_block *afs_dir_get_block(struct afs_dir_iter *iter, size_t block)
{
- struct folio_queue *fq;
struct afs_vnode *dvnode = iter->dvnode;
- struct folio *folio;
+ struct bvecq *bq;
size_t blpos = block * AFS_DIR_BLOCK_SIZE;
size_t blend = (block + 1) * AFS_DIR_BLOCK_SIZE, fpos = iter->fpos;
int ret;
@@ -120,41 +118,38 @@ static union afs_xdr_dir_block *afs_dir_get_block(struct afs_dir_iter *iter, siz
if (dvnode->directory_size < blend) {
size_t cur_size = dvnode->directory_size;
- ret = netfs_alloc_folioq_buffer(
- NULL, &dvnode->directory, &cur_size, blend,
- mapping_gfp_mask(dvnode->netfs.inode.i_mapping));
+ ret = bvecq_expand_buffer(&dvnode->directory, &cur_size, blend,
+ GFP_KERNEL);
dvnode->directory_size = cur_size;
if (ret < 0)
goto fail;
}
- fq = iter->fq;
- if (!fq)
- fq = dvnode->directory;
+ bq = iter->bq;
+ if (!bq)
+ bq = dvnode->directory;
- /* Search the folio queue for the folio containing the block... */
- for (; fq; fq = fq->next) {
- for (int s = iter->fq_slot; s < folioq_count(fq); s++) {
- size_t fsize = folioq_folio_size(fq, s);
+ /* Search the contents for the region containing the block... */
+ for (; bq; bq = bq->next) {
+ for (int s = iter->bq_slot; s < bq->nr_slots; s++) {
+ struct bio_vec *bv = &bq->bv[s];
+ size_t bsize = bv->bv_len;
- if (blend <= fpos + fsize) {
+ if (blend <= fpos + bsize) {
/* ... and then return the mapped block. */
- folio = folioq_folio(fq, s);
- if (WARN_ON_ONCE(folio_pos(folio) != fpos))
- goto fail;
- iter->fq = fq;
- iter->fq_slot = s;
+ iter->bq = bq;
+ iter->bq_slot = s;
iter->fpos = fpos;
- return kmap_local_folio(folio, blpos - fpos);
+ return bvec_kmap_partial(bv, blpos - fpos);
}
- fpos += fsize;
+ fpos += bsize;
}
- iter->fq_slot = 0;
+ iter->bq_slot = 0;
}
fail:
- iter->fq = NULL;
- iter->fq_slot = 0;
+ iter->bq = NULL;
+ iter->bq_slot = 0;
afs_invalidate_dir(dvnode, afs_dir_invalid_edit_get_block);
return NULL;
}
diff --git a/fs/afs/dir_search.c b/fs/afs/dir_search.c
index 11ebdfffcb1d..e1f91b3ac6e2 100644
--- a/fs/afs/dir_search.c
+++ b/fs/afs/dir_search.c
@@ -66,12 +66,11 @@ bool afs_dir_init_iter(struct afs_dir_iter *iter, const struct qstr *name)
*/
union afs_xdr_dir_block *afs_dir_find_block(struct afs_dir_iter *iter, size_t block)
{
- struct folio_queue *fq = iter->fq;
struct afs_vnode *dvnode = iter->dvnode;
- struct folio *folio;
+ struct bvecq *bq = iter->bq;
size_t blpos = block * AFS_DIR_BLOCK_SIZE;
size_t blend = (block + 1) * AFS_DIR_BLOCK_SIZE, fpos = iter->fpos;
- int slot = iter->fq_slot;
+ int slot = iter->bq_slot;
_enter("%zx,%d", block, slot);
@@ -80,36 +79,34 @@ union afs_xdr_dir_block *afs_dir_find_block(struct afs_dir_iter *iter, size_t bl
if (dvnode->directory_size < blend)
goto fail;
- if (!fq || blpos < fpos) {
- fq = dvnode->directory;
+ if (!bq || blpos < fpos) {
+ bq = dvnode->directory;
slot = 0;
fpos = 0;
}
/* Search the folio queue for the folio containing the block... */
- for (; fq; fq = fq->next) {
- for (; slot < folioq_count(fq); slot++) {
- size_t fsize = folioq_folio_size(fq, slot);
+ for (; bq; bq = bq->next) {
+ for (; slot < bq->nr_slots; slot++) {
+ struct bio_vec *bv = &bq->bv[slot];
+ size_t bsize = bv->bv_len;
- if (blend <= fpos + fsize) {
+ if (blend <= fpos + bsize) {
/* ... and then return the mapped block. */
- folio = folioq_folio(fq, slot);
- if (WARN_ON_ONCE(folio_pos(folio) != fpos))
- goto fail;
- iter->fq = fq;
- iter->fq_slot = slot;
+ iter->bq = bq;
+ iter->bq_slot = slot;
iter->fpos = fpos;
- iter->block = kmap_local_folio(folio, blpos - fpos);
+ iter->block = bvec_kmap_partial(bv, blpos - fpos);
return iter->block;
}
- fpos += fsize;
+ fpos += bsize;
}
slot = 0;
}
fail:
- iter->fq = NULL;
- iter->fq_slot = 0;
+ iter->bq = NULL;
+ iter->bq_slot = 0;
afs_invalidate_dir(dvnode, afs_dir_invalid_edit_get_block);
return NULL;
}
diff --git a/fs/afs/inode.c b/fs/afs/inode.c
index 14f39a9bea6c..634fbf8eb212 100644
--- a/fs/afs/inode.c
+++ b/fs/afs/inode.c
@@ -683,7 +683,7 @@ void afs_evict_inode(struct inode *inode)
flush_delayed_work(&vnode->lock_work);
netfs_wait_for_outstanding_io(inode);
truncate_inode_pages_final(&inode->i_data);
- netfs_free_folioq_buffer(vnode->directory);
+ bvecq_put(vnode->directory);
if (vnode->symlink)
afs_evict_symlink(vnode);
diff --git a/fs/afs/internal.h b/fs/afs/internal.h
index d58273c00fbc..51078cc9b0e5 100644
--- a/fs/afs/internal.h
+++ b/fs/afs/internal.h
@@ -710,7 +710,7 @@ struct afs_vnode {
#define AFS_VNODE_MODIFYING 10 /* Set if we're performing a modification op */
#define AFS_VNODE_DIR_READ 11 /* Set if we've read a dir's contents */
- struct folio_queue *directory; /* Directory contents */
+ struct bvecq *directory; /* Directory contents */
struct afs_symlink __rcu *symlink; /* Symlink content */
struct list_head wb_keys; /* List of keys available for writeback */
struct list_head pending_locks; /* locks waiting to be granted */
@@ -991,9 +991,9 @@ static inline void afs_invalidate_cache(struct afs_vnode *vnode, unsigned int fl
struct afs_dir_iter {
struct afs_vnode *dvnode;
union afs_xdr_dir_block *block;
- struct folio_queue *fq;
+ struct bvecq *bq;
unsigned int fpos;
- int fq_slot;
+ int bq_slot;
unsigned int loop_check;
u8 nr_slots;
u8 bucket;
diff --git a/fs/afs/symlink.c b/fs/afs/symlink.c
index 6b8c122877ca..4e4da2b4ab0a 100644
--- a/fs/afs/symlink.c
+++ b/fs/afs/symlink.c
@@ -56,7 +56,6 @@ void afs_evict_symlink(struct afs_vnode *vnode)
void afs_init_new_symlink(struct afs_vnode *vnode, struct afs_operation *op)
{
struct afs_symlink *symlink = op->create.symlink;
- size_t dsize = 0;
size_t size = strlen(symlink->content) + 1;
char *p;
@@ -66,13 +65,18 @@ void afs_init_new_symlink(struct afs_vnode *vnode, struct afs_operation *op)
if (!fscache_cookie_enabled(netfs_i_cookie(&vnode->netfs)))
return;
- if (netfs_alloc_folioq_buffer(NULL, &vnode->directory, &dsize, size,
- mapping_gfp_mask(vnode->netfs.inode.i_mapping)) < 0)
+ vnode->directory =
+ bvecq_alloc_buffer(PAGE_SIZE,
+ mapping_gfp_mask(vnode->netfs.inode.i_mapping),
+ false);
+ if (!vnode->directory)
return;
- vnode->directory_size = dsize;
- p = kmap_local_folio(folioq_folio(vnode->directory, 0), 0);
+ vnode->directory_size = size;
+ p = bvec_kmap_partial(&vnode->directory->bv[0], 0);
memcpy(p, symlink->content, size);
+ if (size < PAGE_SIZE)
+ memset(p + size, 0, PAGE_SIZE - size);
kunmap_local(p);
netfs_single_mark_inode_dirty(&vnode->netfs.inode);
}
@@ -94,17 +98,12 @@ static ssize_t afs_do_read_symlink(struct afs_vnode *vnode)
}
if (!vnode->directory) {
- size_t cur_size = 0;
-
- ret = netfs_alloc_folioq_buffer(NULL,
- &vnode->directory, &cur_size, PAGE_SIZE,
- mapping_gfp_mask(vnode->netfs.inode.i_mapping));
- vnode->directory_size = PAGE_SIZE - 1;
- if (ret < 0)
- return ret;
+ vnode->directory = bvecq_alloc_buffer(PAGE_SIZE, GFP_KERNEL, false);
+ if (!vnode->directory)
+ return -ENOMEM;
}
- iov_iter_folio_queue(&iter, ITER_DEST, vnode->directory, 0, 0, PAGE_SIZE);
+ iov_iter_bvec_queue(&iter, ITER_DEST, vnode->directory, 0, 0, PAGE_SIZE);
/* AFS requires us to perform the read of a symlink as a single unit to
* avoid issues with the content being changed between reads.
@@ -126,7 +125,7 @@ static ssize_t afs_do_read_symlink(struct afs_vnode *vnode)
refcount_set(&symlink->ref, 1);
symlink->content[i_size] = 0;
- const char *s = kmap_local_folio(folioq_folio(vnode->directory, 0), 0);
+ const char *s = bvec_kmap_partial(&vnode->directory->bv[0], 0);
memcpy(symlink->content, s, i_size);
kunmap_local(s);
@@ -135,7 +134,7 @@ static ssize_t afs_do_read_symlink(struct afs_vnode *vnode)
}
if (!fscache_cookie_enabled(netfs_i_cookie(&vnode->netfs))) {
- netfs_free_folioq_buffer(vnode->directory);
+ bvecq_put(vnode->directory);
vnode->directory = NULL;
vnode->directory_size = 0;
}
@@ -248,14 +247,14 @@ int afs_symlink_writepages(struct address_space *mapping,
if (vnode->directory &&
atomic64_read(&vnode->cb_expires_at) != AFS_NO_CB_PROMISE) {
- iov_iter_folio_queue(&iter, ITER_SOURCE, vnode->directory, 0, 0,
- i_size_read(&vnode->netfs.inode));
+ iov_iter_bvec_queue(&iter, ITER_SOURCE, vnode->directory, 0, 0,
+ i_size_read(&vnode->netfs.inode));
ret = netfs_writeback_single(mapping, wbc, &iter);
}
if (ret == 0) {
netfs_wb_begin(&vnode->netfs, false);
- netfs_free_folioq_buffer(vnode->directory);
+ bvecq_put(vnode->directory);
vnode->directory = NULL;
vnode->directory_size = 0;
netfs_wb_end(&vnode->netfs);
diff --git a/fs/netfs/write_issue.c b/fs/netfs/write_issue.c
index 3989b4ec0c4b..c775c53834e3 100644
--- a/fs/netfs/write_issue.c
+++ b/fs/netfs/write_issue.c
@@ -624,129 +624,11 @@ int netfs_writepages(struct address_space *mapping,
}
EXPORT_SYMBOL(netfs_writepages);
-/*
- * Write some of a pending folio data back to the server and/or the cache.
- */
-static int netfs_write_folio_single(struct netfs_io_request *wreq,
- struct folio *folio)
-{
- struct netfs_io_stream *upload = &wreq->io_streams[0];
- struct netfs_io_stream *cache = &wreq->io_streams[1];
- struct netfs_io_stream *stream;
- size_t iter_off = 0;
- size_t fsize = folio_size(folio), flen;
- uoff_t fpos = folio_pos(folio);
- ssize_t ret;
- bool to_eof = false;
- bool no_debug = false;
-
- _enter("");
-
- flen = folio_size(folio);
- if (flen > wreq->i_size - fpos) {
- flen = wreq->i_size - fpos;
- folio_zero_segment(folio, flen, fsize);
- to_eof = true;
- } else if (flen == wreq->i_size - fpos) {
- to_eof = true;
- }
-
- _debug("folio %zx/%zx", flen, fsize);
-
- if (!upload->avail && !cache->avail) {
- trace_netfs_folio(folio, netfs_folio_trace_cancel_store);
- return 0;
- }
-
- if (!upload->construct)
- trace_netfs_folio(folio, netfs_folio_trace_store);
- else
- trace_netfs_folio(folio, netfs_folio_trace_store_plus);
-
- /* Attach the folio to the rolling buffer. */
- folio_get(folio);
- ret = rolling_buffer_append(&wreq->buffer, folio, NETFS_ROLLBUF_PUT_MARK, wreq->gfp);
- if (ret < 0) {
- folio_put(folio);
- return ret;
- }
-
- /* Move the submission point forward to allow for write-streaming data
- * not starting at the front of the page. We don't do write-streaming
- * with the cache as the cache requires DIO alignment.
- *
- * Also skip uploading for data that's been read and just needs copying
- * to the cache.
- */
- for (int s = 0; s < NR_IO_STREAMS; s++) {
- stream = &wreq->io_streams[s];
- stream->submit_off = 0;
- stream->submit_len = flen;
- if (!stream->avail) {
- stream->submit_off = UINT_MAX;
- stream->submit_len = 0;
- }
- }
-
- /* Attach the folio to one or more subrequests. For a big folio, we
- * could end up with thousands of subrequests if the wsize is small -
- * but we might need to wait during the creation of subrequests for
- * network resources (eg. SMB credits).
- */
- for (;;) {
- ssize_t part;
- size_t lowest_off = ULONG_MAX;
- int choose_s = -1;
-
- /* Always add to the lowest-submitted stream first. */
- for (int s = 0; s < NR_IO_STREAMS; s++) {
- stream = &wreq->io_streams[s];
- if (stream->submit_len > 0 &&
- stream->submit_off < lowest_off) {
- lowest_off = stream->submit_off;
- choose_s = s;
- }
- }
-
- if (choose_s < 0)
- break;
- stream = &wreq->io_streams[choose_s];
-
- /* Advance the iterator(s). */
- if (stream->submit_off > iter_off) {
- rolling_buffer_advance(&wreq->buffer, stream->submit_off - iter_off);
- iter_off = stream->submit_off;
- }
-
- atomic64_set(&wreq->issued_to, fpos + stream->submit_off);
- stream->submit_extendable_to = fsize - stream->submit_off;
- part = netfs_advance_write(wreq, stream, fpos + stream->submit_off,
- stream->submit_len, to_eof);
- stream->submit_off += part;
- if (part > stream->submit_len)
- stream->submit_len = 0;
- else
- stream->submit_len -= part;
- if (part > 0)
- no_debug = true;
- }
-
- wreq->buffer.iter.iov_offset = 0;
- if (fsize > iter_off)
- rolling_buffer_advance(&wreq->buffer, fsize - iter_off);
- atomic64_set(&wreq->issued_to, fpos + fsize);
-
- if (!no_debug)
- kdebug("R=%x: No submit", wreq->debug_id);
- _leave(" = 0");
- return 0;
-}
-
/**
* netfs_writeback_single - Write back a monolithic payload
* @mapping: The mapping to write from
* @wbc: Hints from the VM
- * @iter: Data to write, must be ITER_FOLIOQ.
+ * @iter: Data to write.
*
* Write a monolithic, non-pagecache object back to the server and/or
* the cache.
@@ -760,13 +642,8 @@ int netfs_writeback_single(struct address_space *mapping,
{
struct netfs_io_request *wreq;
struct netfs_inode *ictx = netfs_inode(mapping->host);
- struct folio_queue *fq;
- size_t size = iov_iter_count(iter);
int ret;
- if (WARN_ON_ONCE(!iov_iter_is_folioq(iter)))
- return -EIO;
-
if (!netfs_wb_begin(ictx, wbc->sync_mode == WB_SYNC_NONE)) {
/* The VFS will have undirtied the inode. */
netfs_single_mark_inode_dirty(&ictx->inode);
@@ -779,36 +656,41 @@ int netfs_writeback_single(struct address_space *mapping,
goto couldnt_start;
}
+ wreq->buffer.iter = *iter;
+ wreq->len = iov_iter_count(iter);
+
__set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, &wreq->flags);
trace_netfs_write(wreq, netfs_write_trace_writeback_single);
netfs_stat(&netfs_n_wh_writepages);
- if (__test_and_set_bit(NETFS_RREQ_UPLOAD_TO_SERVER, &wreq->flags))
+ if (test_bit(NETFS_RREQ_UPLOAD_TO_SERVER, &wreq->flags))
wreq->netfs_ops->begin_writeback(wreq);
- for (fq = (struct folio_queue *)iter->folioq; fq; fq = fq->next) {
- for (int slot = 0; slot < folioq_count(fq); slot++) {
- struct folio *folio = folioq_folio(fq, slot);
- size_t part = umin(folioq_folio_size(fq, slot), size);
+ for (int s = 0; s < NR_IO_STREAMS; s++) {
+ struct netfs_io_subrequest *subreq;
+ struct netfs_io_stream *stream = &wreq->io_streams[s];
- _debug("wbiter %lx %llx", folio->index, atomic64_read(&wreq->issued_to));
+ if (!stream->avail)
+ continue;
- ret = netfs_write_folio_single(wreq, folio);
- if (ret < 0)
- goto stop;
- size -= part;
- if (size <= 0)
- goto stop;
- }
+ netfs_prepare_write(wreq, stream, 0);
+
+ subreq = stream->construct;
+ subreq->len = wreq->len;
+ stream->submit_len = subreq->len;
+ stream->submit_extendable_to = round_up(wreq->len, PAGE_SIZE);
+
+ netfs_issue_write(wreq, stream);
}
-stop:
- for (int s = 0; s < NR_IO_STREAMS; s++)
- netfs_issue_write(wreq, &wreq->io_streams[s]);
netfs_all_subreqs_queued(wreq);
-
netfs_wake_collector(wreq);
+ /* TODO: Might want to be async here if WB_SYNC_NONE, but then need to
+ * wait before modifying.
+ */
+ ret = netfs_wait_for_write(wreq);
+
netfs_put_request(wreq, netfs_rreq_trace_put_return);
_leave(" = %d", ret);
return ret;
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 05/10] cifs: Use a bvecq for buffering instead of a folioq
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
` (3 preceding siblings ...)
2026-09-29 7:59 ` [PATCH v12 04/10] afs: Use a bvecq to hold dir content rather than folioq David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 06/10] smbdirect: Support ITER_BVECQ in smbdirect_map_sges_from_iter() David Howells
` (5 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Christoph Hellwig
Use a bvecq for internal buffering for crypto purposes instead of a folioq
so that the latter can be phased out.
Signed-off-by: David Howells <dhowells@redhat.com>
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Namjae Jeon <linkinjeon@kernel.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Christoph Hellwig <hch@infradead.org>
cc: linux-cifs@vger.kernel.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
fs/smb/client/cifsglob.h | 2 +-
fs/smb/client/smb2ops.c | 78 +++++++++++++++++++---------------------
2 files changed, 38 insertions(+), 42 deletions(-)
diff --git a/fs/smb/client/cifsglob.h b/fs/smb/client/cifsglob.h
index 79e4e84f8985..c8b1c23ef494 100644
--- a/fs/smb/client/cifsglob.h
+++ b/fs/smb/client/cifsglob.h
@@ -289,7 +289,7 @@ struct smb_rqst {
struct kvec *rq_iov; /* array of kvecs */
unsigned int rq_nvec; /* number of kvecs in array */
struct iov_iter rq_iter; /* Data iterator */
- struct folio_queue *rq_buffer; /* Buffer for encryption */
+ struct bvecq *rq_buffer; /* Buffer for encryption */
};
struct mid_q_entry;
diff --git a/fs/smb/client/smb2ops.c b/fs/smb/client/smb2ops.c
index aa142420dae2..317d10b22966 100644
--- a/fs/smb/client/smb2ops.c
+++ b/fs/smb/client/smb2ops.c
@@ -13,7 +13,7 @@
#include <linux/sort.h>
#include <crypto/aead.h>
#include <linux/fiemap.h>
-#include <linux/folio_queue.h>
+#include <linux/bvecq.h>
#include <uapi/linux/magic.h>
#include "cifsfs.h"
#include "cifsglob.h"
@@ -4830,19 +4830,18 @@ crypt_message(struct TCP_Server_Info *server, int num_rqst,
}
/*
- * Copy data from an iterator to the folios in a folio queue buffer.
+ * Copy data from an iterator to the pages in a bvec queue buffer.
*/
-static bool cifs_copy_iter_to_folioq(struct iov_iter *iter, size_t size,
- struct folio_queue *buffer)
+static bool cifs_copy_iter_to_bvecq(struct iov_iter *iter, size_t size,
+ struct bvecq *buffer)
{
for (; buffer; buffer = buffer->next) {
- for (int s = 0; s < folioq_count(buffer); s++) {
- struct folio *folio = folioq_folio(buffer, s);
- size_t part = folioq_folio_size(buffer, s);
+ for (int s = 0; s < bvecq_nr_slots_acquire(buffer); s++) {
+ struct bio_vec *bv = &buffer->bv[s];
+ size_t part = umin(bv->bv_len, size);
- part = umin(part, size);
-
- if (copy_folio_from_iter(folio, 0, part, iter) != part)
+ if (copy_page_from_iter(bv->bv_page, bv->bv_offset,
+ part, iter) != part)
return false;
size -= part;
}
@@ -4854,7 +4853,7 @@ void
smb3_free_compound_rqst(int num_rqst, struct smb_rqst *rqst)
{
for (int i = 0; i < num_rqst; i++)
- netfs_free_folioq_buffer(rqst[i].rq_buffer);
+ bvecq_put(rqst[i].rq_buffer);
}
/*
@@ -4881,7 +4880,7 @@ smb3_init_transform_rq(struct TCP_Server_Info *server, int num_rqst,
for (int i = 1; i < num_rqst; i++) {
struct smb_rqst *old = &old_rq[i - 1];
struct smb_rqst *new = &new_rq[i];
- struct folio_queue *buffer = NULL;
+ struct bvecq *buffer = NULL;
size_t size = iov_iter_count(&old->rq_iter);
orig_len += smb_rqst_len(server, old);
@@ -4889,17 +4888,16 @@ smb3_init_transform_rq(struct TCP_Server_Info *server, int num_rqst,
new->rq_nvec = old->rq_nvec;
if (size > 0) {
- size_t cur_size = 0;
- rc = netfs_alloc_folioq_buffer(NULL, &buffer, &cur_size,
- size, GFP_NOFS);
- new->rq_buffer = buffer;
- if (rc < 0)
+ rc = -ENOMEM;
+ buffer = bvecq_alloc_buffer(size, GFP_NOFS, true);
+ if (!buffer)
goto err_free;
- iov_iter_folio_queue(&new->rq_iter, ITER_SOURCE,
- buffer, 0, 0, size);
+ new->rq_buffer = buffer;
+ iov_iter_bvec_queue(&new->rq_iter, ITER_SOURCE,
+ buffer, 0, 0, size);
- if (!cifs_copy_iter_to_folioq(&old->rq_iter, size, buffer)) {
+ if (!cifs_copy_iter_to_bvecq(&old->rq_iter, size, buffer)) {
rc = smb_EIO1(smb_eio_trace_tx_copy_iter_to_buf, size);
goto err_free;
}
@@ -4989,22 +4987,20 @@ decrypt_raw_data(struct TCP_Server_Info *server, char *buf,
}
static int
-cifs_copy_folioq_to_iter(struct folio_queue *folioq, size_t data_size,
- size_t skip, struct iov_iter *iter)
+cifs_copy_bvecq_to_iter(struct bvecq *bq, size_t data_size,
+ size_t skip, struct iov_iter *iter)
{
- for (; folioq; folioq = folioq->next) {
- for (int s = 0; s < folioq_count(folioq); s++) {
- struct folio *folio;
- size_t fsize, n, len;
+ for (; bq; bq = bq->next) {
+ for (int s = 0; s < bvecq_nr_slots_acquire(bq); s++) {
+ struct bio_vec *bv = &bq->bv[s];
+ size_t n, len;
if (data_size == 0)
return 0;
- folio = folioq_folio(folioq, s);
- fsize = folio_size(folio);
- len = umin(fsize - skip, data_size);
+ len = umin(bv->bv_len - skip, data_size);
- n = copy_folio_to_iter(folio, skip, len, iter);
+ n = copy_page_to_iter(bv->bv_page, bv->bv_offset + skip, len, iter);
if (n != len) {
cifs_dbg(VFS, "%s: something went wrong\n", __func__);
return smb_EIO2(smb_eio_trace_rx_copy_to_iter,
@@ -5026,7 +5022,7 @@ cifs_copy_folioq_to_iter(struct folio_queue *folioq, size_t data_size,
static int
handle_read_data(struct TCP_Server_Info *server, struct mid_q_entry *mid,
- char *buf, unsigned int buf_len, struct folio_queue *buffer,
+ char *buf, unsigned int buf_len, struct bvecq *buffer,
unsigned int buffer_len, bool is_offloaded)
{
unsigned int data_offset;
@@ -5136,8 +5132,8 @@ handle_read_data(struct TCP_Server_Info *server, struct mid_q_entry *mid,
}
/* Copy the data to the output I/O iterator. */
- rdata->result = cifs_copy_folioq_to_iter(buffer, data_len,
- cur_off, &rdata->subreq.io_iter);
+ rdata->result = cifs_copy_bvecq_to_iter(buffer, data_len,
+ cur_off, &rdata->subreq.io_iter);
if (rdata->result != 0) {
if (is_offloaded)
mid->mid_state = MID_RESPONSE_MALFORMED;
@@ -5176,7 +5172,7 @@ handle_read_data(struct TCP_Server_Info *server, struct mid_q_entry *mid,
struct smb2_decrypt_work {
struct work_struct decrypt;
struct TCP_Server_Info *server;
- struct folio_queue *buffer;
+ struct bvecq *buffer;
char *buf;
unsigned int len;
};
@@ -5190,7 +5186,7 @@ static void smb2_decrypt_offload(struct work_struct *work)
struct mid_q_entry *mid;
struct iov_iter iter;
- iov_iter_folio_queue(&iter, ITER_DEST, dw->buffer, 0, 0, dw->len);
+ iov_iter_bvec_queue(&iter, ITER_DEST, dw->buffer, 0, 0, dw->len);
rc = decrypt_raw_data(dw->server, dw->buf, dw->server->vals->read_rsp_size,
&iter, true);
if (rc) {
@@ -5239,7 +5235,7 @@ static void smb2_decrypt_offload(struct work_struct *work)
}
free_pages:
- netfs_free_folioq_buffer(dw->buffer);
+ bvecq_put(dw->buffer);
cifs_small_buf_release(dw->buf);
kfree(dw);
}
@@ -5285,12 +5281,12 @@ receive_encrypted_read(struct TCP_Server_Info *server, struct mid_q_entry **mid,
dw->len = len;
len = round_up(dw->len, PAGE_SIZE);
- size_t cur_size = 0;
- rc = netfs_alloc_folioq_buffer(NULL, &dw->buffer, &cur_size, len, GFP_NOFS);
- if (rc < 0)
+ rc = -ENOMEM;
+ dw->buffer = bvecq_alloc_buffer(len, GFP_NOFS, false);
+ if (!dw->buffer)
goto discard_data;
- iov_iter_folio_queue(&iter, ITER_DEST, dw->buffer, 0, 0, len);
+ iov_iter_bvec_queue(&iter, ITER_DEST, dw->buffer, 0, 0, len);
/* Read the data into the buffer and clear excess bufferage. */
rc = cifs_read_iter_from_socket(server, &iter, dw->len);
@@ -5348,7 +5344,7 @@ receive_encrypted_read(struct TCP_Server_Info *server, struct mid_q_entry **mid,
}
free_pages:
- netfs_free_folioq_buffer(dw->buffer);
+ bvecq_put(dw->buffer);
free_dw:
kfree(dw);
return rc;
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 06/10] smbdirect: Support ITER_BVECQ in smbdirect_map_sges_from_iter()
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
` (4 preceding siblings ...)
2026-09-29 7:59 ` [PATCH v12 05/10] cifs: Use a bvecq for buffering instead of a folioq David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 07/10] netfs: Switch folioq to bvecq David Howells
` (4 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Shyam Prasad N, Tom Talpey, Christoph Hellwig
Add support for ITER_BVECQ to smbdirect_map_sges_from_iter().
Signed-off-by: David Howells <dhowells@redhat.com>
Acked-by: Stefan Metzmacher <metze@samba.org>
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Namjae Jeon <linkinjeon@kernel.org>
cc: Shyam Prasad N <sprasad@microsoft.com>
cc: Tom Talpey <tom@talpey.com>
cc: Christoph Hellwig <hch@infradead.org>
cc: linux-cifs@vger.kernel.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
fs/smb/smbdirect/connection.c | 67 +++++++++++++++++++++++++++++++++++
1 file changed, 67 insertions(+)
diff --git a/fs/smb/smbdirect/connection.c b/fs/smb/smbdirect/connection.c
index afd31fa12a36..7cf8e1f8556f 100644
--- a/fs/smb/smbdirect/connection.c
+++ b/fs/smb/smbdirect/connection.c
@@ -5,6 +5,7 @@
*/
#include "internal.h"
+#include <linux/bvecq.h>
#include <linux/folio_queue.h>
struct smbdirect_map_sges {
@@ -2015,6 +2016,69 @@ static ssize_t smbdirect_map_sges_from_bvec(struct iov_iter *iter,
return ret;
}
+/*
+ * Extract memory fragments from a BVECQ-class iterator and add them to an RDMA
+ * list. The fragments are not pinned.
+ */
+static ssize_t smbdirect_map_sges_from_bvecq(struct iov_iter *iter,
+ struct smbdirect_map_sges *state,
+ ssize_t maxsize)
+{
+ const struct bvecq *bq = iter->bvecq, *next;
+ unsigned int slot = iter->bvecq_slot;
+ ssize_t extracted = 0;
+ size_t offset = iter->iov_offset;
+
+ maxsize = umin(maxsize, iov_iter_count(iter));
+
+ do {
+ struct bio_vec *bv;
+ size_t bsize;
+
+ while (slot >= bq->nr_slots) {
+ next = bvecq_next(bq);
+ if (!next) {
+ if (WARN_ON_ONCE(maxsize > 0))
+ return -EIO;
+ goto out;
+ }
+ bq = next;
+ slot = 0;
+ }
+
+ bv = &bq->bv[slot];
+ bsize = bv->bv_len;
+
+ if (offset < bsize) {
+ size_t part = umin(maxsize, bsize - offset);
+ bool ok;
+
+ ok = smbdirect_map_sges_single_page(state,
+ bv->bv_page,
+ bv->bv_offset + offset,
+ part);
+ if (!ok)
+ return -EIO;
+
+ offset += part;
+ extracted += part;
+ maxsize -= part;
+ }
+
+ if (offset >= bsize) {
+ offset = 0;
+ slot++;
+ }
+ } while (state->num_sge < state->max_sge && maxsize > 0);
+
+out:
+ iter->bvecq = bq;
+ iter->bvecq_slot = slot;
+ iter->iov_offset = offset;
+ iter->count -= extracted;
+ return extracted;
+}
+
/*
* Extract fragments from a KVEC-class iterator and add them to an ib_sge list.
* This can deal with vmalloc'd buffers as well as kmalloc'd or static buffers.
@@ -2164,6 +2228,9 @@ static ssize_t smbdirect_map_sges_from_iter(struct iov_iter *iter, size_t len,
case ITER_BVEC:
ret = smbdirect_map_sges_from_bvec(iter, state, len);
break;
+ case ITER_BVECQ:
+ ret = smbdirect_map_sges_from_bvecq(iter, state, len);
+ break;
case ITER_KVEC:
ret = smbdirect_map_sges_from_kvec(iter, state, len);
break;
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 07/10] netfs: Switch folioq to bvecq
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
` (5 preceding siblings ...)
2026-09-29 7:59 ` [PATCH v12 06/10] smbdirect: Support ITER_BVECQ in smbdirect_map_sges_from_iter() David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 08/10] smbdirect: Remove support for ITER_FOLIOQ from smbdirect_map_sges_from_iter() David Howells
` (3 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Shyam Prasad N, Tom Talpey, Christoph Hellwig
In netfslib, perform more or less a straight switch from using folio_queue
to using bvecq to hold the lists of folios involved in various sorts of
buffered read and in buffered writeback.
It's not quite a straight swap, however, because:
(1) The bveq_alloc_*() routines want to know if the callers is doing
writeback for emergency pool use.
(2) folio_queue includes a fixed capacity folio_batch, but bvecq has a
variable size bio_vec array.
(3) folio_queue has a set of per-folio marks that are mostly unused; the
one exception is for netfs_prefetch_for_write() - but the marked folio
is then ignored because no_unlock_folio points to it (and the bvecq
mem cleanup type can handle that anyway). This will become an issue
if/when netfs_prefetch_for_write() expands the read to cope with large
cache granularity.
(4) The bvecq API doesn't have functions to get a folio's length or to
clear a folio pointer by slot, instead accessing the bio_vec array
directly. Also bvecq stores the folio length in bv_len instead of
storing the folio orders.
(5) rolling_buffer_bulk_load_from_ra() needs to work a bit differently as
the folio_queue contains a folio_batch and bvecq doesn't.
(6) rolling_buffer_delete_spent() has to clear the bvecq->next pointer
before putting the bvecq to avoid rolling up the entire list. On the
other hand, rolling_buffer_clear() just needs to put the tail.
Signed-off-by: David Howells <dhowells@redhat.com>
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Namjae Jeon <linkinjeon@kernel.org>
cc: Shyam Prasad N <sprasad@microsoft.com>
cc: Tom Talpey <tom@talpey.com>
cc: Christoph Hellwig <hch@infradead.org>
cc: linux-cifs@vger.kernel.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
fs/netfs/buffered_read.c | 34 ++++----
fs/netfs/iterator.c | 33 ++++----
fs/netfs/read_collect.c | 47 ++++++-----
fs/netfs/read_pgpriv2.c | 24 +++---
fs/netfs/read_retry.c | 28 +++----
fs/netfs/rolling_buffer.c | 143 ++++++++++++++-------------------
fs/netfs/write_collect.c | 26 +++---
fs/netfs/write_issue.c | 16 ++--
include/linux/rolling_buffer.h | 42 ++++------
include/trace/events/netfs.h | 2 +-
10 files changed, 180 insertions(+), 215 deletions(-)
diff --git a/fs/netfs/buffered_read.c b/fs/netfs/buffered_read.c
index e30bde80276a..aa1e4f8d46ab 100644
--- a/fs/netfs/buffered_read.c
+++ b/fs/netfs/buffered_read.c
@@ -215,7 +215,7 @@ static void netfs_issue_read(struct netfs_io_request *rreq,
* otherwise we set the deprecated PG_private_2.
*/
static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,
- struct folio_queue **fq,
+ struct bvecq **bq,
unsigned int *offset,
int *slot,
size_t len,
@@ -225,21 +225,21 @@ static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq,
struct folio *folio;
size_t fsize, overlap;
- if (!*fq)
+ if (!*bq)
break;
- if (*slot >= folioq_count(*fq)) {
- *fq = (*fq)->next;
+ if (!bvecq_acquire_slot(*bq, *slot)) {
+ *bq = bvecq_next(*bq);
*slot = 0;
*offset = 0;
continue;
}
/* Determine how much the subreq overlaps the folio, if at all. */
- fsize = folioq_folio_size(*fq, *slot);
+ fsize = (*bq)->bv[*slot].bv_len;
overlap = min(len, fsize - *offset);
if (overlap > 0 && copy) {
- folio = folioq_folio(*fq, *slot);
+ folio = bvec_folio(&(*bq)->bv[*slot]);
if (netfs_using_pgpriv2(rreq)) {
if (!folio_test_private_2(folio))
folio_start_private_2(folio);
@@ -275,7 +275,7 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)
.cached_to[1] = ULLONG_MAX,
};
struct fscache_occupancy *occ = &_occ;
- struct folio_queue *fq = rreq->buffer.tail;
+ struct bvecq *bq = rreq->buffer.tail;
unsigned int offset = 0;
ssize_t size = rreq->len;
uoff_t start = rreq->start;
@@ -408,10 +408,10 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq)
if (size <= 0)
netfs_all_subreqs_queued(rreq);
- if (fq) {
+ if (bq) {
/* See if the cache indicated this should be cached. */
copy = test_bit(NETFS_SREQ_COPY_TO_CACHE, &subreq->flags);
- netfs_mark_copy_to_cache(rreq, &fq, &slot, &offset, slice, copy);
+ netfs_mark_copy_to_cache(rreq, &bq, &slot, &offset, slice, copy);
}
trace_netfs_sreq(subreq, netfs_sreq_trace_submit);
@@ -479,8 +479,7 @@ void netfs_readahead(struct readahead_control *ractl)
* acquires a ref on each folio that we will need to release later -
* but we don't want to do that until after we've started the I/O.
*/
- added = rolling_buffer_bulk_load_from_ra(&rreq->buffer, ractl,
- rreq->debug_id, rreq->gfp);
+ added = rolling_buffer_bulk_load_from_ra(&rreq->buffer, ractl, rreq->gfp);
if (added < 0) {
ret = added;
goto cleanup_free;
@@ -503,15 +502,14 @@ EXPORT_SYMBOL(netfs_readahead);
/*
* Create a rolling buffer with a single occupying folio.
*/
-static int netfs_create_singular_buffer(struct netfs_io_request *rreq, struct folio *folio,
- unsigned int rollbuf_flags)
+static int netfs_create_singular_buffer(struct netfs_io_request *rreq, struct folio *folio)
{
ssize_t added;
- if (rolling_buffer_init(&rreq->buffer, rreq->debug_id, ITER_DEST, rreq->gfp) < 0)
+ if (rolling_buffer_init(&rreq->buffer, ITER_DEST, rreq->gfp, false) < 0)
return -ENOMEM;
- added = rolling_buffer_append(&rreq->buffer, folio, rollbuf_flags, rreq->gfp);
+ added = rolling_buffer_append(&rreq->buffer, folio, rreq->gfp);
if (added < 0)
return added;
rreq->submitted = rreq->start + added;
@@ -661,7 +659,7 @@ int netfs_read_folio(struct file *file, struct folio *folio)
trace_netfs_read(rreq, rreq->start, rreq->len, netfs_read_trace_readpage);
/* Set up the output buffer */
- ret = netfs_create_singular_buffer(rreq, folio, 0);
+ ret = netfs_create_singular_buffer(rreq, folio);
if (ret < 0)
goto discard;
@@ -818,7 +816,7 @@ int netfs_write_begin(struct netfs_inode *ctx,
trace_netfs_read(rreq, pos, len, netfs_read_trace_write_begin);
/* Set up the output buffer */
- ret = netfs_create_singular_buffer(rreq, folio, 0);
+ ret = netfs_create_singular_buffer(rreq, folio);
if (ret < 0)
goto error_put;
@@ -883,7 +881,7 @@ int netfs_prefetch_for_write(struct file *file, struct folio *folio,
trace_netfs_read(rreq, start, flen, netfs_read_trace_prefetch_for_write);
/* Set up the output buffer */
- ret = netfs_create_singular_buffer(rreq, folio, NETFS_ROLLBUF_PAGECACHE_MARK);
+ ret = netfs_create_singular_buffer(rreq, folio);
if (ret < 0)
goto error_put;
diff --git a/fs/netfs/iterator.c b/fs/netfs/iterator.c
index eb1efb17f53a..31748526d568 100644
--- a/fs/netfs/iterator.c
+++ b/fs/netfs/iterator.c
@@ -245,33 +245,34 @@ static size_t netfs_limit_xarray(const struct iov_iter *iter, size_t start_offse
}
/*
- * Select the span of a folio queue iterator we're going to use. Limit it by
- * both maximum size and maximum number of segments. Returns the size of the
- * span in bytes.
+ * Select the span of a bvecq iterator we're going to use. Limit it by both
+ * maximum size and maximum number of segments. Returns the size of the span
+ * in bytes.
*/
-static size_t netfs_limit_folioq(const struct iov_iter *iter, size_t start_offset,
- size_t max_size, size_t max_segs)
+static size_t netfs_limit_bvecq(const struct iov_iter *iter, size_t start_offset,
+ size_t max_size, size_t max_segs)
{
- const struct folio_queue *folioq = iter->folioq;
+ const struct bvecq *bq = iter->bvecq;
unsigned int nsegs = 0;
- unsigned int slot = iter->folioq_slot;
+ unsigned int slot = iter->bvecq_slot;
size_t span = 0, n = iter->count;
- if (WARN_ON(!iov_iter_is_folioq(iter)) ||
+ if (WARN_ON(!iov_iter_is_bvecq(iter)) ||
WARN_ON(start_offset > n) ||
n == 0)
return 0;
max_size = umin(max_size, n - start_offset);
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = folioq->next;
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = bvecq_next(bq);
slot = 0;
}
start_offset += iter->iov_offset;
do {
- size_t flen = folioq_folio_size(folioq, slot);
+ size_t flen;
+ flen = bq->bv[slot].bv_len;
if (start_offset < flen) {
span += flen - start_offset;
nsegs++;
@@ -283,11 +284,11 @@ static size_t netfs_limit_folioq(const struct iov_iter *iter, size_t start_offse
break;
slot++;
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = folioq->next;
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = bvecq_next(bq);
slot = 0;
}
- } while (folioq);
+ } while (bq);
return umin(span, max_size);
}
@@ -295,8 +296,8 @@ static size_t netfs_limit_folioq(const struct iov_iter *iter, size_t start_offse
size_t netfs_limit_iter(const struct iov_iter *iter, size_t start_offset,
size_t max_size, size_t max_segs)
{
- if (iov_iter_is_folioq(iter))
- return netfs_limit_folioq(iter, start_offset, max_size, max_segs);
+ if (iov_iter_is_bvecq(iter))
+ return netfs_limit_bvecq(iter, start_offset, max_size, max_segs);
if (iov_iter_is_bvec(iter))
return netfs_limit_bvec(iter, start_offset, max_size, max_segs);
if (iov_iter_is_xarray(iter))
diff --git a/fs/netfs/read_collect.c b/fs/netfs/read_collect.c
index 2625efd48a9b..6bbaaea69354 100644
--- a/fs/netfs/read_collect.c
+++ b/fs/netfs/read_collect.c
@@ -63,11 +63,11 @@ void netfs_cancel_copy_to_cache(struct netfs_io_request *rreq, struct folio *fol
* dirty and let writeback handle it.
*/
static void netfs_unlock_read_folio(struct netfs_io_request *rreq,
- struct folio_queue *folioq,
+ struct bvecq *bq,
int slot)
{
struct netfs_folio *finfo;
- struct folio *folio = folioq_folio(folioq, slot);
+ struct folio *folio = bvec_folio(&bq->bv[slot]);
if (unlikely(folio_pos(folio) < rreq->abandon_to)) {
trace_netfs_folio(folio, netfs_folio_trace_abandon);
@@ -98,7 +98,7 @@ static void netfs_unlock_read_folio(struct netfs_io_request *rreq,
trace_netfs_folio(folio, netfs_folio_trace_read_done);
}
- folioq_clear(folioq, slot);
+ bq->bv[slot].bv_page = NULL;
} else {
// TODO: Use of PG_private_2 is deprecated.
if (folio_test_private_2(folio))
@@ -114,7 +114,7 @@ static void netfs_unlock_read_folio(struct netfs_io_request *rreq,
folio_unlock(folio);
}
- folioq_clear(folioq, slot);
+ bq->bv[slot].bv_page = NULL;
}
/*
@@ -122,21 +122,21 @@ static void netfs_unlock_read_folio(struct netfs_io_request *rreq,
*/
void netfs_read_set_unlock_at(struct netfs_io_request *rreq)
{
- struct folio_queue *folioq = rreq->buffer.tail;
+ const struct bvecq *bq = rreq->buffer.tail;
unsigned int slot = rreq->buffer.first_tail_slot;
size_t cleaned_to = rreq->cleaned_to - rreq->start;
size_t progress_at = cleaned_to;
size_t minimum = 256 * 1024;
while (progress_at < rreq->len) {
- if (slot >= folioq_count(folioq)) {
- folioq = folioq->next;
- if (!folioq)
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = bvecq_next(bq);
+ if (!bq)
break;
slot = 0;
}
- progress_at += folioq_folio_size(folioq, slot);
+ progress_at += bq->bv[slot].bv_len;
if (progress_at - cleaned_to >= minimum)
break;
slot++;
@@ -152,7 +152,7 @@ void netfs_read_set_unlock_at(struct netfs_io_request *rreq)
static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
unsigned int *notes)
{
- struct folio_queue *folioq = rreq->buffer.tail;
+ struct bvecq *bq = rreq->buffer.tail;
unsigned int slot = rreq->buffer.first_tail_slot;
uoff_t collected_to = rreq->collected_to;
@@ -161,9 +161,9 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
// TODO: Begin decryption
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = rolling_buffer_delete_spent(&rreq->buffer);
- if (!folioq) {
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = rolling_buffer_delete_spent(&rreq->buffer);
+ if (!bq) {
WRITE_ONCE(rreq->progress_at, rreq->len);
return;
}
@@ -182,13 +182,13 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
uoff_t fpos, fend;
size_t fsize;
- folio = folioq_folio(folioq, slot);
+ folio = bvec_folio(&bq->bv[slot]);
if (WARN_ONCE(!folio_test_locked(folio),
"R=%08x: folio %lx is not locked\n",
rreq->debug_id, folio->index))
trace_netfs_folio(folio, netfs_folio_trace_not_locked);
- fsize = folioq_folio_size(folioq, slot);
+ fsize = bq->bv[slot].bv_len;
fpos = folio_pos(folio);
fend = fpos + fsize;
@@ -198,29 +198,28 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq,
if (collected_to < fend)
break;
- netfs_unlock_read_folio(rreq, folioq, slot);
+ netfs_unlock_read_folio(rreq, bq, slot);
WRITE_ONCE(rreq->cleaned_to, fpos + fsize);
*notes |= MADE_PROGRESS;
- /* Clean up the head folioq. If we clear an entire folioq, then
- * we can get rid of it provided it's not also the tail folioq
+ /* Clean up the head bq. If we clear an entire bq, then
+ * we can get rid of it provided it's not also the tail bq
* being filled by the issuer.
*/
- folioq_clear(folioq, slot);
+ bq->bv[slot].bv_page = NULL;
slot++;
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = rolling_buffer_delete_spent(&rreq->buffer);
- if (!folioq)
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = rolling_buffer_delete_spent(&rreq->buffer);
+ if (!bq)
goto done;
slot = 0;
- trace_netfs_folioq(folioq, netfs_trace_folioq_read_progress);
}
if (fpos + fsize >= collected_to)
break;
}
- rreq->buffer.tail = folioq;
+ rreq->buffer.tail = bq;
done:
rreq->buffer.first_tail_slot = slot;
diff --git a/fs/netfs/read_pgpriv2.c b/fs/netfs/read_pgpriv2.c
index 5280b606fda4..16d623ca9df3 100644
--- a/fs/netfs/read_pgpriv2.c
+++ b/fs/netfs/read_pgpriv2.c
@@ -53,7 +53,7 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio
trace_netfs_folio(folio, netfs_folio_trace_store_copy);
/* Attach the folio to the rolling buffer. */
- if (rolling_buffer_append(&creq->buffer, folio, 0, creq->gfp) < 0) {
+ if (rolling_buffer_append(&creq->buffer, folio, creq->gfp) < 0) {
set_bit(NETFS_RREQ_CANCEL_CACHING, &creq->flags);
folio_end_private_2(folio);
return;
@@ -173,13 +173,13 @@ void netfs_pgpriv2_end_copy_to_cache(struct netfs_io_request *rreq)
*/
bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)
{
- struct folio_queue *folioq = creq->buffer.tail;
+ struct bvecq *bq = creq->buffer.tail;
unsigned int slot = creq->buffer.first_tail_slot;
uoff_t collected_to = creq->collected_to;
bool made_progress = false;
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = rolling_buffer_delete_spent(&creq->buffer);
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = rolling_buffer_delete_spent(&creq->buffer);
slot = 0;
}
@@ -188,7 +188,7 @@ bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)
uoff_t fpos, fend;
size_t fsize, flen;
- folio = folioq_folio(folioq, slot);
+ folio = bvec_folio(&bq->bv[slot]);
if (WARN_ONCE(!folio_test_private_2(folio),
"R=%08x: folio %lx is not marked private_2\n",
creq->debug_id, folio->index))
@@ -211,15 +211,15 @@ bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)
creq->cleaned_to = fpos + fsize;
made_progress = true;
- /* Clean up the head folioq. If we clear an entire folioq, then
- * we can get rid of it provided it's not also the tail folioq
+ /* Clean up the head bq. If we clear an entire bq, then
+ * we can get rid of it provided it's not also the tail bq
* being filled by the issuer.
*/
- folioq_clear(folioq, slot);
+ bq->bv[slot].bv_page = NULL;
slot++;
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = rolling_buffer_delete_spent(&creq->buffer);
- if (!folioq)
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = rolling_buffer_delete_spent(&creq->buffer);
+ if (!bq)
goto done;
slot = 0;
}
@@ -228,7 +228,7 @@ bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq)
break;
}
- creq->buffer.tail = folioq;
+ creq->buffer.tail = bq;
done:
creq->buffer.first_tail_slot = slot;
return made_progress;
diff --git a/fs/netfs/read_retry.c b/fs/netfs/read_retry.c
index 5bd8dee5a834..e60024a75939 100644
--- a/fs/netfs/read_retry.c
+++ b/fs/netfs/read_retry.c
@@ -291,7 +291,7 @@ void netfs_retry_reads(struct netfs_io_request *rreq)
*/
void netfs_unlock_abandoned_read_pages(struct netfs_io_request *rreq)
{
- struct folio_queue *p;
+ struct bvecq *p;
/* We have to wait for readahead refs to have been released before we
* can unlock any folios as the ref-dropper walks i_pages and the only
@@ -301,23 +301,23 @@ void netfs_unlock_abandoned_read_pages(struct netfs_io_request *rreq)
netfs_wait_for_put_ra_refs(rreq);
for (p = rreq->buffer.tail; p; p = p->next) {
- for (int slot = 0; slot < folioq_count(p); slot++) {
- struct folio *folio = folioq_folio(p, slot);
+ for (int slot = rreq->buffer.first_tail_slot;
+ bvecq_acquire_slot(p, slot);
+ slot++) {
+ struct folio *folio;
- if (!folio)
+ if (!p->bv[slot].bv_page)
continue;
+
+ folio = bvec_folio(&p->bv[slot]);
netfs_cancel_copy_to_cache(rreq, folio);
- if (!folioq_is_marked2(p, slot)) {
- if (folio == rreq->no_unlock_folio &&
- test_bit(NETFS_RREQ_NO_UNLOCK_FOLIO,
- &rreq->flags)) {
- _debug("no unlock");
- } else {
- trace_netfs_folio(folio,
- netfs_folio_trace_abandon);
- folio_unlock(folio);
- }
+ if (folio == rreq->no_unlock_folio &&
+ test_bit(NETFS_RREQ_NO_UNLOCK_FOLIO, &rreq->flags)) {
+ _debug("no unlock");
+ } else {
+ trace_netfs_folio(folio, netfs_folio_trace_abandon);
+ folio_unlock(folio);
}
}
}
diff --git a/fs/netfs/rolling_buffer.c b/fs/netfs/rolling_buffer.c
index d30d5ef6d86e..76429fbb6920 100644
--- a/fs/netfs/rolling_buffer.c
+++ b/fs/netfs/rolling_buffer.c
@@ -63,45 +63,46 @@ EXPORT_SYMBOL(netfs_folioq_free);
* that the pointers can be independently driven by the producer and the
* consumer.
*/
-int rolling_buffer_init(struct rolling_buffer *roll, unsigned int rreq_id,
- unsigned int direction, gfp_t gfp)
+int rolling_buffer_init(struct rolling_buffer *roll, unsigned int direction,
+ gfp_t gfp, bool for_writeback)
{
- struct folio_queue *fq;
+ struct bvecq *bq;
+
+ roll->for_writeback = for_writeback;
- fq = netfs_folioq_alloc(rreq_id, gfp, netfs_trace_folioq_rollbuf_init);
- if (!fq)
+ bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, for_writeback);
+ if (!bq)
return -ENOMEM;
- roll->head = fq;
- roll->tail = fq;
- iov_iter_folio_queue(&roll->iter, direction, fq, 0, 0, 0);
+ roll->head = bq;
+ roll->tail = bq;
+ iov_iter_bvec_queue(&roll->iter, direction, bq, 0, 0, 0);
return 0;
}
/*
- * Add another folio_queue to a rolling buffer if there's no space left.
+ * Add another bvecq to a rolling buffer if there's no space left.
*/
int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp)
{
- struct folio_queue *fq, *head = roll->head;
+ struct bvecq *bq, *head = roll->head;
- if (!folioq_full(head))
+ if (!bvecq_is_full(head))
return 0;
- fq = netfs_folioq_alloc(head->rreq_id, gfp, netfs_trace_folioq_make_space);
- if (!fq)
+ bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, roll->for_writeback);
+ if (!bq)
return -ENOMEM;
- fq->prev = head;
- roll->head = fq;
- if (folioq_full(head)) {
+ roll->head = bq;
+ if (bvecq_is_full(head)) {
/* Make sure we don't leave the master iterator pointing to a
* block that might get immediately consumed.
*/
- if (roll->iter.folioq == head &&
- roll->iter.folioq_slot == folioq_nr_slots(head)) {
- roll->iter.folioq = fq;
- roll->iter.folioq_slot = 0;
+ if (roll->iter.bvecq == head &&
+ roll->iter.bvecq_slot == head->nr_slots) {
+ roll->iter.bvecq = bq;
+ roll->iter.bvecq_slot = 0;
}
}
@@ -110,7 +111,7 @@ int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp)
* [!] NOTE: After we set head->next, the consumer is at liberty to
* immediately delete the old head.
*/
- smp_store_release(&head->next, fq);
+ bvecq_append(head, bq);
return 0;
}
@@ -119,55 +120,59 @@ int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp)
*/
ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,
struct readahead_control *ractl,
- unsigned int rreq_id, gfp_t gfp)
+ gfp_t gfp)
{
- struct folio_queue *fq;
- ssize_t loaded = 0;
+ struct bvecq *bq;
+ size_t loaded = 0;
while (ractl->_nr_pages - ractl->_batch_count > 0) {
+ struct page **pages;
unsigned int nr;
- /* Allocate a folioq to put some folios into and attach it to
+ /* Allocate a bvecq to put some folios into and attach it to
* the rolling buffer.
*/
- fq = netfs_folioq_alloc(rreq_id, gfp,
- netfs_trace_folioq_make_space);
- if (!fq)
+ bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, false);
+ if (!bq)
goto nomem_unlock;
- fq->prev = roll->head;
+ bq->mem_type = BVECQ_MEM_EXTERNAL; /* Folio cleanup handled separately. */
+
if (!roll->tail)
- roll->tail = fq;
+ roll->tail = bq;
else
- roll->head->next = fq;
- roll->head = fq;
+ bvecq_append(roll->head, bq);
+ roll->head = bq;
- /* Get a batch of folios and note their orders. */
- nr = __readahead_batch(ractl, (struct page **)fq->vec.folios,
- folioq_nr_slots(fq));
+ /* Get a bunch of folios and note their sizes. */
+ pages = (struct page **)(bq->bv + bq->max_slots);
+ pages -= bq->max_slots;
+ nr = __readahead_batch(ractl, pages, bq->max_slots);
if (WARN_ON_ONCE(!nr))
break;
- fq->vec.nr = nr;
for (int slot = 0; slot < nr; slot++) {
- struct folio *folio = folioq_folio(fq, slot);
- unsigned int order;
+ struct folio *folio = page_folio(pages[slot]);
+ size_t len = folio_size(folio);
- order = folio_order(folio);
- fq->orders[slot] = order;
- loaded += PAGE_SIZE << order;
+ bvec_set_folio(&bq->bv[slot], folio, len, 0);
+ loaded += len;
trace_netfs_folio(folio, netfs_folio_trace_read);
}
+
+ bvecq_filled_to(bq, nr);
}
WRITE_ONCE(roll->iter.count, loaded);
- iov_iter_folio_queue(&roll->iter, ITER_DEST, roll->tail, 0, 0, loaded);
+ iov_iter_bvec_queue(&roll->iter, ITER_DEST, roll->tail, 0, 0, loaded);
return loaded;
nomem_unlock:
- for (fq = roll->tail; fq; fq = fq->next) {
- for (int slot = 0; slot < folioq_count(fq); slot++) {
- folio_unlock(fq->vec.folios[slot]);
- folioq_mark(fq, slot);
+ for (bq = roll->tail; bq; bq = bq->next) {
+ for (int slot = 0; slot < bq->nr_slots; slot++) {
+ struct folio *folio = bvec_folio(&bq->bv[slot]);
+
+ folio_unlock(folio);
+ folio_put(folio);
}
}
rolling_buffer_clear(roll);
@@ -180,7 +185,7 @@ ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,
* Append a folio to the rolling buffer.
*/
ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio,
- unsigned int flags, gfp_t gfp)
+ gfp_t gfp)
{
ssize_t size = folio_size(folio);
int slot;
@@ -188,16 +193,11 @@ ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio,
if (rolling_buffer_make_space(roll, gfp) < 0)
return -ENOMEM;
- slot = folioq_append(roll->head, folio);
- if (flags & ROLLBUF_MARK_1)
- folioq_mark(roll->head, slot);
- if (flags & ROLLBUF_MARK_2)
- folioq_mark2(roll->head, slot);
+ slot = roll->head->nr_slots;
+ bvec_set_folio(&roll->head->bv[slot], folio, size, 0);
+ bvecq_filled_to(roll->head, slot + 1);
WRITE_ONCE(roll->iter.count, roll->iter.count + size);
-
- /* Store the counter after setting the slot. */
- smp_store_release(&roll->next_head_slot, slot);
return size;
}
@@ -206,44 +206,23 @@ ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio,
* don't return the last buffer to keep the pointers independent, but return
* NULL instead.
*/
-struct folio_queue *rolling_buffer_delete_spent(struct rolling_buffer *roll)
+struct bvecq *rolling_buffer_delete_spent(struct rolling_buffer *roll)
{
- struct folio_queue *spent = roll->tail, *next = READ_ONCE(spent->next);
+ struct bvecq *spent = roll->tail, *next = bvecq_next(spent);
if (!next)
return NULL;
next->prev = NULL;
- netfs_folioq_free(spent, netfs_trace_folioq_delete);
roll->tail = next;
+ spent->next = NULL;
+ bvecq_put(spent);
return next;
}
/*
- * Clear out a rolling queue. Folios that have mark 1 set are put.
+ * Clear out a rolling queue.
*/
void rolling_buffer_clear(struct rolling_buffer *roll)
{
- struct folio_batch fbatch;
- struct folio_queue *p;
-
- folio_batch_init(&fbatch);
-
- while ((p = roll->tail)) {
- roll->tail = p->next;
- for (int slot = 0; slot < folioq_count(p); slot++) {
- struct folio *folio = folioq_folio(p, slot);
-
- if (!folio)
- continue;
- if (folioq_is_marked(p, slot)) {
- trace_netfs_folio(folio, netfs_folio_trace_put);
- if (!folio_batch_add(&fbatch, folio))
- folio_batch_release(&fbatch);
- }
- }
-
- netfs_folioq_free(p, netfs_trace_folioq_clear);
- }
-
- folio_batch_release(&fbatch);
+ bvecq_put(roll->tail);
}
diff --git a/fs/netfs/write_collect.c b/fs/netfs/write_collect.c
index 6e8ea534230d..80e6a3194a60 100644
--- a/fs/netfs/write_collect.c
+++ b/fs/netfs/write_collect.c
@@ -114,11 +114,11 @@ int netfs_folio_written_back(struct folio *folio)
static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
unsigned int *notes)
{
- struct folio_queue *folioq = wreq->buffer.tail;
+ struct bvecq *bq = wreq->buffer.tail;
unsigned int slot = wreq->buffer.first_tail_slot;
uoff_t collected_to = wreq->collected_to;
- if (WARN_ON_ONCE(!folioq)) {
+ if (WARN_ON_ONCE(!bq)) {
pr_err("[!] Writeback unlock found empty rolling buffer!\n");
netfs_dump_request(wreq);
return;
@@ -130,9 +130,9 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
return;
}
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = rolling_buffer_delete_spent(&wreq->buffer);
- if (!folioq)
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = rolling_buffer_delete_spent(&wreq->buffer);
+ if (!bq)
return;
slot = 0;
}
@@ -143,7 +143,7 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
uoff_t fpos, fend;
size_t fsize, flen;
- folio = folioq_folio(folioq, slot);
+ folio = bvec_folio(&bq->bv[slot]);
if (WARN_ONCE(!folio_test_writeback(folio),
"R=%08x: folio %lx is not under writeback\n",
wreq->debug_id, folio->index))
@@ -166,15 +166,15 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
wreq->cleaned_to = fpos + fsize;
*notes |= MADE_PROGRESS;
- /* Clean up the head folioq. If we clear an entire folioq, then
- * we can get rid of it provided it's not also the tail folioq
+ /* Clean up the head bq. If we clear an entire bq, then
+ * we can get rid of it provided it's not also the tail bq
* being filled by the issuer.
*/
- folioq_clear(folioq, slot);
+ bq->bv[slot].bv_page = NULL;
slot++;
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = rolling_buffer_delete_spent(&wreq->buffer);
- if (!folioq)
+ if (!bvecq_acquire_slot(bq, slot)) {
+ bq = rolling_buffer_delete_spent(&wreq->buffer);
+ if (!bq)
goto done;
slot = 0;
}
@@ -183,7 +183,7 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq,
break;
}
- wreq->buffer.tail = folioq;
+ wreq->buffer.tail = bq;
done:
wreq->buffer.first_tail_slot = slot;
}
diff --git a/fs/netfs/write_issue.c b/fs/netfs/write_issue.c
index c775c53834e3..5165f866332e 100644
--- a/fs/netfs/write_issue.c
+++ b/fs/netfs/write_issue.c
@@ -107,7 +107,9 @@ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping,
ictx = netfs_inode(wreq->inode);
if (is_cacheable)
fscache_begin_write_operation(&wreq->cache_resources, netfs_i_cookie(ictx));
- if (rolling_buffer_init(&wreq->buffer, wreq->debug_id, ITER_SOURCE, wreq->gfp) < 0)
+ if (rolling_buffer_init(&wreq->buffer, ITER_SOURCE, wreq->gfp,
+ (origin == NETFS_WRITEBACK ||
+ origin == NETFS_WRITEBACK_SINGLE)) < 0)
goto nomem;
wreq->cleaned_to = wreq->start;
@@ -162,12 +164,12 @@ void netfs_prepare_write(struct netfs_io_request *wreq,
struct netfs_io_subrequest *subreq;
struct iov_iter *wreq_iter = &wreq->buffer.iter;
- /* Make sure we don't point the iterator at a used-up folio_queue
- * struct being used as a placeholder to prevent the queue from
- * collapsing. In such a case, extend the queue.
+ /* Make sure we don't point the iterator at a used-up bvecq struct
+ * being used as a placeholder to prevent the queue from collapsing.
+ * In such a case, extend the queue.
*/
- if (iov_iter_is_folioq(wreq_iter) &&
- wreq_iter->folioq_slot >= folioq_nr_slots(wreq_iter->folioq))
+ if (iov_iter_is_bvecq(wreq_iter) &&
+ !bvecq_acquire_slot(wreq_iter->bvecq, wreq_iter->bvecq_slot))
rolling_buffer_make_space(&wreq->buffer, wreq->gfp);
subreq = netfs_alloc_subrequest(wreq, stream->source);
@@ -450,7 +452,7 @@ static int netfs_write_folio(struct netfs_io_request *wreq,
}
/* Attach the folio to the rolling buffer. */
- rolling_buffer_append(&wreq->buffer, folio, 0, wreq->gfp);
+ rolling_buffer_append(&wreq->buffer, folio, wreq->gfp);
/* Move the submission point forward to allow for write-streaming data
* not starting at the front of the page. We don't do write-streaming
diff --git a/include/linux/rolling_buffer.h b/include/linux/rolling_buffer.h
index a97f7cfaacaa..5c0bc4221f01 100644
--- a/include/linux/rolling_buffer.h
+++ b/include/linux/rolling_buffer.h
@@ -8,49 +8,35 @@
#ifndef _ROLLING_BUFFER_H
#define _ROLLING_BUFFER_H
-#include <linux/folio_queue.h>
+#include <linux/bvecq.h>
#include <linux/uio.h>
/*
- * Rolling buffer. Whilst the buffer is live and in use, folios and folio
- * queue segments can be added to one end by one thread and removed from the
- * other end by another thread. The buffer isn't allowed to be empty; it must
- * always have at least one folio_queue in it so that neither side has to
- * modify both queue pointers.
+ * Rolling buffer. Whilst the buffer is live and in use, folios and bvecq
+ * segments can be added to one end by one thread and removed from the other
+ * end by another thread. The buffer isn't allowed to be empty; it must always
+ * have at least one bvecq in it so that neither side has to modify both queue
+ * pointers.
*
* The iterator in the buffer is extended as buffers are inserted. It can be
* snapshotted to use a segment of the buffer.
*/
struct rolling_buffer {
- struct folio_queue *head; /* Producer's insertion point */
- struct folio_queue *tail; /* Consumer's removal point */
+ struct bvecq *head; /* Producer's insertion point */
+ struct bvecq *tail; /* Consumer's removal point */
struct iov_iter iter; /* Iterator tracking what's left in the buffer */
- u8 next_head_slot; /* Next slot in ->head */
u8 first_tail_slot; /* First slot in ->tail */
+ bool for_writeback; /* T if being used for writeback */
};
-/*
- * Snapshot of a rolling buffer.
- */
-struct rolling_buffer_snapshot {
- struct folio_queue *curr_folioq; /* Queue segment in which current folio resides */
- unsigned char curr_slot; /* Folio currently being read */
- unsigned char curr_order; /* Order of folio */
-};
-
-/* Marks to store per-folio in the internal folio_queue structs. */
-#define ROLLBUF_MARK_1 BIT(0)
-#define ROLLBUF_MARK_2 BIT(1)
-
-int rolling_buffer_init(struct rolling_buffer *roll, unsigned int rreq_id,
- unsigned int direction, gfp_t gfp);
+int rolling_buffer_init(struct rolling_buffer *roll, unsigned int direction,
+ gfp_t gfp, bool for_writeback);
int rolling_buffer_make_space(struct rolling_buffer *roll, gfp_t gfp);
ssize_t rolling_buffer_bulk_load_from_ra(struct rolling_buffer *roll,
struct readahead_control *ractl,
- unsigned int rreq_id, gfp_t gfp);
-ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio,
- unsigned int flags, gfp_t gfp);
-struct folio_queue *rolling_buffer_delete_spent(struct rolling_buffer *roll);
+ gfp_t gfp);
+ssize_t rolling_buffer_append(struct rolling_buffer *roll, struct folio *folio, gfp_t gfp);
+struct bvecq *rolling_buffer_delete_spent(struct rolling_buffer *roll);
void rolling_buffer_clear(struct rolling_buffer *roll);
static inline void rolling_buffer_advance(struct rolling_buffer *roll, size_t amount)
diff --git a/include/trace/events/netfs.h b/include/trace/events/netfs.h
index 3fe3d47ba55b..1fd62465e809 100644
--- a/include/trace/events/netfs.h
+++ b/include/trace/events/netfs.h
@@ -398,7 +398,7 @@ TRACE_EVENT(netfs_sreq,
__entry->len = sreq->len;
__entry->transferred = sreq->transferred;
__entry->start = sreq->start;
- __entry->slot = sreq->io_iter.folioq_slot;
+ __entry->slot = sreq->io_iter.bvecq_slot;
),
TP_printk("R=%08x[%x] %s %s f=%03x s=%llx %zx/%zx s=%u e=%d",
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 08/10] smbdirect: Remove support for ITER_FOLIOQ from smbdirect_map_sges_from_iter()
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
` (6 preceding siblings ...)
2026-09-29 7:59 ` [PATCH v12 07/10] netfs: Switch folioq to bvecq David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 09/10] iov_iter: Remove ITER_FOLIOQ David Howells
` (2 subsequent siblings)
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Shyam Prasad N, Tom Talpey, Christoph Hellwig
netfslib now only presents an bvecq queue and an associated ITER_BVECQ
iterator to the filesystem, so it isn't going to see the ITER_FOLIOQ
iterator. So remove that code.
Netfslib also won't supply ITER_BVEC/KVEC iterators, though smbdirect
might; further in future, it won't supply iterators at all, but rather a
bvecq slice (that can be used to construct an iterator).
Signed-off-by: David Howells <dhowells@redhat.com>
Acked-by: Stefan Metzmacher <metze@samba.org>
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Namjae Jeon <linkinjeon@kernel.org>
cc: Stefan Metzmacher <metze@samba.org>
cc: Shyam Prasad N <sprasad@microsoft.com>
cc: Tom Talpey <tom@talpey.com>
cc: Christoph Hellwig <hch@infradead.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: linux-cifs@vger.kernel.org
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
fs/smb/smbdirect/connection.c | 68 -----------------------------------
1 file changed, 68 deletions(-)
diff --git a/fs/smb/smbdirect/connection.c b/fs/smb/smbdirect/connection.c
index 7cf8e1f8556f..aa52ecdbef2b 100644
--- a/fs/smb/smbdirect/connection.c
+++ b/fs/smb/smbdirect/connection.c
@@ -6,7 +6,6 @@
#include "internal.h"
#include <linux/bvecq.h>
-#include <linux/folio_queue.h>
struct smbdirect_map_sges {
struct ib_sge *sge;
@@ -2140,70 +2139,6 @@ static ssize_t smbdirect_map_sges_from_kvec(struct iov_iter *iter,
return ret;
}
-/*
- * Extract folio fragments from a FOLIOQ-class iterator and add them to an
- * ib_sge list. The folios are not pinned.
- */
-static ssize_t smbdirect_map_sges_from_folioq(struct iov_iter *iter,
- struct smbdirect_map_sges *state,
- ssize_t maxsize)
-{
- const struct folio_queue *folioq = iter->folioq;
- unsigned int slot = iter->folioq_slot;
- ssize_t ret = 0;
- size_t offset = iter->iov_offset;
-
- if (WARN_ON_ONCE(!folioq))
- return -EIO;
-
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = folioq->next;
- if (WARN_ON_ONCE(!folioq))
- return -EIO;
- slot = 0;
- }
-
- do {
- struct folio *folio = folioq_folio(folioq, slot);
- size_t fsize = folioq_folio_size(folioq, slot);
-
- if (offset < fsize) {
- size_t part = umin(maxsize, fsize - offset);
- bool ok;
-
- ok = smbdirect_map_sges_single_page(state,
- folio_page(folio, 0),
- offset,
- part);
- if (!ok)
- return -EIO;
-
- offset += part;
- ret += part;
- maxsize -= part;
- }
-
- if (offset >= fsize) {
- offset = 0;
- slot++;
- if (slot >= folioq_nr_slots(folioq)) {
- if (!folioq->next) {
- WARN_ON_ONCE(ret < iter->count);
- break;
- }
- folioq = folioq->next;
- slot = 0;
- }
- }
- } while (state->num_sge < state->max_sge && maxsize > 0);
-
- iter->folioq = folioq;
- iter->folioq_slot = slot;
- iter->iov_offset = offset;
- iter->count -= ret;
- return ret;
-}
-
/*
* Extract page fragments from up to the given amount of the source iterator
* and build up an ib_sge list that refers to all of those bits. The ib_sge list
@@ -2234,9 +2169,6 @@ static ssize_t smbdirect_map_sges_from_iter(struct iov_iter *iter, size_t len,
case ITER_KVEC:
ret = smbdirect_map_sges_from_kvec(iter, state, len);
break;
- case ITER_FOLIOQ:
- ret = smbdirect_map_sges_from_folioq(iter, state, len);
- break;
default:
WARN_ONCE(1, "iov_iter_type[%u]\n", iov_iter_type(iter));
return -EIO;
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 09/10] iov_iter: Remove ITER_FOLIOQ
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
` (7 preceding siblings ...)
2026-09-29 7:59 ` [PATCH v12 08/10] smbdirect: Remove support for ITER_FOLIOQ from smbdirect_map_sges_from_iter() David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 7:59 ` [PATCH v12 10/10] netfs: Remove folio_queue David Howells
2026-09-29 8:21 ` [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Christoph Hellwig
Remove ITER_FOLIOQ as it's no longer used.
Signed-off-by: David Howells <dhowells@redhat.com>
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Christoph Hellwig <hch@infradead.org>
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
include/linux/iov_iter.h | 65 +--------
include/linux/uio.h | 12 --
lib/iov_iter.c | 236 +-------------------------------
lib/scatterlist.c | 69 +---------
lib/tests/kunit_iov_iter.c | 271 -------------------------------------
5 files changed, 7 insertions(+), 646 deletions(-)
diff --git a/include/linux/iov_iter.h b/include/linux/iov_iter.h
index ec0aa2893689..39ffd5878f50 100644
--- a/include/linux/iov_iter.h
+++ b/include/linux/iov_iter.h
@@ -11,7 +11,6 @@
#include <linux/uio.h>
#include <linux/bvec.h>
#include <linux/bvecq.h>
-#include <linux/folio_queue.h>
typedef size_t (*iov_step_f)(void *iter_base, size_t progress, size_t len,
void *priv, void *priv2);
@@ -207,62 +206,6 @@ size_t iterate_bvecq(struct iov_iter *iter, size_t len, void *priv, void *priv2,
return progress;
}
-/*
- * Handle ITER_FOLIOQ.
- */
-static __always_inline
-size_t iterate_folioq(struct iov_iter *iter, size_t len, void *priv, void *priv2,
- iov_step_f step)
-{
- const struct folio_queue *folioq = iter->folioq;
- unsigned int slot = iter->folioq_slot;
- size_t progress = 0, skip = iter->iov_offset;
-
- if (slot == folioq_nr_slots(folioq)) {
- /* The iterator may have been extended. */
- folioq = folioq->next;
- slot = 0;
- }
-
- do {
- struct folio *folio = folioq_folio(folioq, slot);
- size_t part, remain = 0, consumed;
- size_t fsize;
- void *base;
-
- if (!folio)
- break;
-
- fsize = folioq_folio_size(folioq, slot);
- if (skip < fsize) {
- base = kmap_local_folio(folio, skip);
- part = umin(len, PAGE_SIZE - skip % PAGE_SIZE);
- remain = step(base, progress, part, priv, priv2);
- kunmap_local(base);
- consumed = part - remain;
- len -= consumed;
- progress += consumed;
- skip += consumed;
- }
- if (skip >= fsize) {
- skip = 0;
- slot++;
- if (slot == folioq_nr_slots(folioq) && folioq->next) {
- folioq = folioq->next;
- slot = 0;
- }
- }
- if (remain)
- break;
- } while (len);
-
- iter->folioq_slot = slot;
- iter->folioq = folioq;
- iter->iov_offset = skip;
- iter->count -= progress;
- return progress;
-}
-
/*
* Handle ITER_XARRAY.
*/
@@ -374,8 +317,6 @@ size_t iterate_and_advance2(struct iov_iter *iter, size_t len, void *priv,
return iterate_kvec(iter, len, priv, priv2, step);
if (iov_iter_is_bvecq(iter))
return iterate_bvecq(iter, len, priv, priv2, step);
- if (iov_iter_is_folioq(iter))
- return iterate_folioq(iter, len, priv, priv2, step);
if (iov_iter_is_xarray(iter))
return iterate_xarray(iter, len, priv, priv2, step);
return iterate_discard(iter, len, priv, priv2, step);
@@ -410,8 +351,8 @@ size_t iterate_and_advance(struct iov_iter *iter, size_t len, void *priv,
* buffer is presented in segments, which for kernel iteration are broken up by
* physical pages and mapped, with the mapped address being presented.
*
- * [!] Note This will only handle BVEC, KVEC, BVECQ, FOLIOQ, XARRAY and
- * DISCARD-type iterators; it will not handle UBUF or IOVEC-type iterators.
+ * [!] Note This will only handle BVEC, KVEC, BVECQ, XARRAY and DISCARD-type
+ * iterators; it will not handle UBUF or IOVEC-type iterators.
*
* A step functions, @step, must be provided, one for handling mapped kernel
* addresses and the other is given user addresses which have the potential to
@@ -440,8 +381,6 @@ size_t iterate_and_advance_kernel(struct iov_iter *iter, size_t len, void *priv,
return iterate_kvec(iter, len, priv, priv2, step);
if (iov_iter_is_bvecq(iter))
return iterate_bvecq(iter, len, priv, priv2, step);
- if (iov_iter_is_folioq(iter))
- return iterate_folioq(iter, len, priv, priv2, step);
if (iov_iter_is_xarray(iter))
return iterate_xarray(iter, len, priv, priv2, step);
return iterate_discard(iter, len, priv, priv2, step);
diff --git a/include/linux/uio.h b/include/linux/uio.h
index 2d75a267da86..8d82a8258ef1 100644
--- a/include/linux/uio.h
+++ b/include/linux/uio.h
@@ -11,7 +11,6 @@
#include <uapi/linux/uio.h>
struct page;
-struct folio_queue;
typedef unsigned int __bitwise iov_iter_extraction_t;
@@ -27,7 +26,6 @@ enum iter_type {
ITER_BVEC,
ITER_KVEC,
ITER_BVECQ,
- ITER_FOLIOQ,
ITER_XARRAY,
ITER_DISCARD,
};
@@ -70,7 +68,6 @@ struct iov_iter {
const struct kvec *kvec;
const struct bio_vec *bvec;
const struct bvecq *bvecq;
- const struct folio_queue *folioq;
struct xarray *xarray;
void __user *ubuf;
};
@@ -80,7 +77,6 @@ struct iov_iter {
union {
unsigned long nr_segs;
u16 bvecq_slot;
- u8 folioq_slot;
loff_t xarray_start;
};
};
@@ -153,11 +149,6 @@ static inline bool iov_iter_is_bvecq(const struct iov_iter *i)
return iov_iter_type(i) == ITER_BVECQ;
}
-static inline bool iov_iter_is_folioq(const struct iov_iter *i)
-{
- return iov_iter_type(i) == ITER_FOLIOQ;
-}
-
static inline bool iov_iter_is_xarray(const struct iov_iter *i)
{
return iov_iter_type(i) == ITER_XARRAY;
@@ -306,9 +297,6 @@ void iov_iter_discard(struct iov_iter *i, unsigned int direction, size_t count);
void iov_iter_bvec_queue(struct iov_iter *i, unsigned int direction,
const struct bvecq *bvecq,
unsigned int first_slot, unsigned int offset, size_t count);
-void iov_iter_folio_queue(struct iov_iter *i, unsigned int direction,
- const struct folio_queue *folioq,
- unsigned int first_slot, unsigned int offset, size_t count);
void iov_iter_xarray(struct iov_iter *i, unsigned int direction, struct xarray *xarray,
loff_t start, size_t count);
ssize_t iov_iter_get_pages2(struct iov_iter *i, struct page **pages,
diff --git a/lib/iov_iter.c b/lib/iov_iter.c
index b76a297dc8d4..3c79f13fc1cf 100644
--- a/lib/iov_iter.c
+++ b/lib/iov_iter.c
@@ -572,39 +572,6 @@ static void iov_iter_bvecq_advance(struct iov_iter *i, size_t by)
i->bvecq = bq;
}
-static void iov_iter_folioq_advance(struct iov_iter *i, size_t size)
-{
- const struct folio_queue *folioq = i->folioq;
- unsigned int slot = i->folioq_slot;
-
- if (!i->count)
- return;
- i->count -= size;
-
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = folioq->next;
- slot = 0;
- }
-
- size += i->iov_offset; /* From beginning of current segment. */
- do {
- size_t fsize = folioq_folio_size(folioq, slot);
-
- if (likely(size < fsize))
- break;
- size -= fsize;
- slot++;
- if (slot >= folioq_nr_slots(folioq) && folioq->next) {
- folioq = folioq->next;
- slot = 0;
- }
- } while (size);
-
- i->iov_offset = size;
- i->folioq_slot = slot;
- i->folioq = folioq;
-}
-
void iov_iter_advance(struct iov_iter *i, size_t size)
{
if (unlikely(i->count < size))
@@ -619,8 +586,6 @@ void iov_iter_advance(struct iov_iter *i, size_t size)
iov_iter_bvec_advance(i, size);
} else if (iov_iter_is_bvecq(i)) {
iov_iter_bvecq_advance(i, size);
- } else if (iov_iter_is_folioq(i)) {
- iov_iter_folioq_advance(i, size);
} else if (iov_iter_is_discard(i)) {
i->count -= size;
}
@@ -654,32 +619,6 @@ static void iov_iter_bvecq_revert(struct iov_iter *i, size_t unroll)
i->bvecq = bq;
}
-static void iov_iter_folioq_revert(struct iov_iter *i, size_t unroll)
-{
- const struct folio_queue *folioq = i->folioq;
- unsigned int slot = i->folioq_slot;
-
- for (;;) {
- size_t fsize;
-
- if (slot == 0) {
- folioq = folioq->prev;
- slot = folioq_nr_slots(folioq);
- }
- slot--;
-
- fsize = folioq_folio_size(folioq, slot);
- if (unroll <= fsize) {
- i->iov_offset = fsize - unroll;
- break;
- }
- unroll -= fsize;
- }
-
- i->folioq_slot = slot;
- i->folioq = folioq;
-}
-
void iov_iter_revert(struct iov_iter *i, size_t unroll)
{
if (!unroll)
@@ -714,9 +653,6 @@ void iov_iter_revert(struct iov_iter *i, size_t unroll)
} else if (iov_iter_is_bvecq(i)) {
i->iov_offset = 0;
iov_iter_bvecq_revert(i, unroll);
- } else if (iov_iter_is_folioq(i)) {
- i->iov_offset = 0;
- iov_iter_folioq_revert(i, unroll);
} else { /* same logics for iovec and kvec */
const struct iovec *iov = iter_iov(i);
while (1) {
@@ -766,8 +702,6 @@ size_t iov_iter_single_seg_count(const struct iov_iter *i)
}
return min(i->count, bq->bv[slot].bv_len - offset);
}
- if (unlikely(iov_iter_is_folioq(i)))
- return umin(folioq_folio_size(i->folioq, i->folioq_slot), i->count);
return i->count;
}
EXPORT_SYMBOL(iov_iter_single_seg_count);
@@ -833,36 +767,6 @@ void iov_iter_bvec_queue(struct iov_iter *i, unsigned int direction,
}
EXPORT_SYMBOL(iov_iter_bvec_queue);
-/**
- * iov_iter_folio_queue - Initialise an I/O iterator to use the folios in a folio queue
- * @i: The iterator to initialise.
- * @direction: The direction of the transfer.
- * @folioq: The starting point in the folio queue.
- * @first_slot: The first slot in the folio queue to use
- * @offset: The offset into the folio in the first slot to start at
- * @count: The size of the I/O buffer in bytes.
- *
- * Set up an I/O iterator to either draw data out of the pages attached to an
- * inode or to inject data into those pages. The pages *must* be prevented
- * from evaporation, either by taking a ref on them or locking them by the
- * caller.
- */
-void iov_iter_folio_queue(struct iov_iter *i, unsigned int direction,
- const struct folio_queue *folioq, unsigned int first_slot,
- unsigned int offset, size_t count)
-{
- BUG_ON(direction & ~1);
- *i = (struct iov_iter) {
- .iter_type = ITER_FOLIOQ,
- .data_source = direction,
- .folioq = folioq,
- .folioq_slot = first_slot,
- .count = count,
- .iov_offset = offset,
- };
-}
-EXPORT_SYMBOL(iov_iter_folio_queue);
-
/**
* iov_iter_xarray - Initialise an I/O iterator to use the pages in an xarray
* @i: The iterator to initialise.
@@ -1006,9 +910,7 @@ unsigned long iov_iter_alignment(const struct iov_iter *i)
if (iov_iter_is_bvecq(i))
return iov_iter_alignment_bvecq(i);
- /* With both xarray and folioq types, we're dealing with whole folios. */
- if (iov_iter_is_folioq(i))
- return i->iov_offset | i->count;
+ /* With the xarray type, we're dealing with whole folios. */
if (iov_iter_is_xarray(i))
return (i->xarray_start + i->iov_offset) | i->count;
@@ -1192,65 +1094,6 @@ static ssize_t iter_bvecq_get_pages(struct iov_iter *iter,
return extracted;
}
-static ssize_t iter_folioq_get_pages(struct iov_iter *iter,
- struct page ***ppages, size_t maxsize,
- unsigned maxpages, size_t *_start_offset)
-{
- const struct folio_queue *folioq = iter->folioq;
- struct page **pages;
- unsigned int slot = iter->folioq_slot;
- size_t extracted = 0, count = iter->count, iov_offset = iter->iov_offset;
-
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = folioq->next;
- slot = 0;
- if (WARN_ON(iov_offset != 0))
- return -EIO;
- }
-
- maxpages = want_pages_array(ppages, maxsize, iov_offset & ~PAGE_MASK, maxpages);
- if (!maxpages)
- return -ENOMEM;
- *_start_offset = iov_offset & ~PAGE_MASK;
- pages = *ppages;
-
- for (;;) {
- struct folio *folio = folioq_folio(folioq, slot);
- size_t offset = iov_offset, fsize = folioq_folio_size(folioq, slot);
- size_t part = PAGE_SIZE - offset % PAGE_SIZE;
-
- if (offset < fsize) {
- part = umin(part, umin(maxsize - extracted, fsize - offset));
- count -= part;
- iov_offset += part;
- extracted += part;
-
- *pages = folio_page(folio, offset / PAGE_SIZE);
- get_page(*pages);
- pages++;
- maxpages--;
- }
-
- if (maxpages == 0 || extracted >= maxsize)
- break;
-
- if (iov_offset >= fsize) {
- iov_offset = 0;
- slot++;
- if (slot == folioq_nr_slots(folioq) && folioq->next) {
- folioq = folioq->next;
- slot = 0;
- }
- }
- }
-
- iter->count = count;
- iter->iov_offset = iov_offset;
- iter->folioq = folioq;
- iter->folioq_slot = slot;
- return extracted;
-}
-
static ssize_t iter_xarray_populate_pages(struct page **pages, struct xarray *xa,
pgoff_t index, unsigned int nr_pages)
{
@@ -1404,8 +1247,6 @@ static ssize_t __iov_iter_get_pages_alloc(struct iov_iter *i,
}
if (iov_iter_is_bvecq(i))
return iter_bvecq_get_pages(i, pages, maxsize, maxpages, start);
- if (iov_iter_is_folioq(i))
- return iter_folioq_get_pages(i, pages, maxsize, maxpages, start);
if (iov_iter_is_xarray(i))
return iter_xarray_get_pages(i, pages, maxsize, maxpages, start);
return -EFAULT;
@@ -1524,11 +1365,6 @@ int iov_iter_npages(const struct iov_iter *i, int maxpages)
return bvec_npages(i, maxpages);
if (iov_iter_is_bvecq(i))
return iov_npages_bvecq(i, maxpages);
- if (iov_iter_is_folioq(i)) {
- unsigned offset = i->iov_offset % PAGE_SIZE;
- int npages = DIV_ROUND_UP(offset + i->count, PAGE_SIZE);
- return min(npages, maxpages);
- }
if (iov_iter_is_xarray(i)) {
unsigned offset = (i->xarray_start + i->iov_offset) % PAGE_SIZE;
int npages = DIV_ROUND_UP(offset + i->count, PAGE_SIZE);
@@ -1895,68 +1731,6 @@ static ssize_t iov_iter_extract_bvecq_pages(struct iov_iter *iter,
return extracted;
}
-/*
- * Extract a list of contiguous pages from an ITER_FOLIOQ iterator. This does
- * not get references on the pages, nor does it get a pin on them.
- */
-static ssize_t iov_iter_extract_folioq_pages(struct iov_iter *i,
- struct page ***pages, size_t maxsize,
- unsigned int maxpages,
- iov_iter_extraction_t extraction_flags,
- size_t *offset0)
-{
- const struct folio_queue *folioq = i->folioq;
- struct page **p;
- unsigned int nr = 0;
- size_t extracted = 0, offset, slot = i->folioq_slot;
-
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = folioq->next;
- slot = 0;
- if (WARN_ON(i->iov_offset != 0))
- return -EIO;
- }
-
- offset = i->iov_offset & ~PAGE_MASK;
- *offset0 = offset;
-
- maxpages = want_pages_array(pages, maxsize, offset, maxpages);
- if (!maxpages)
- return -ENOMEM;
- p = *pages;
-
- for (;;) {
- struct folio *folio = folioq_folio(folioq, slot);
- size_t offset = i->iov_offset, fsize = folioq_folio_size(folioq, slot);
- size_t part = PAGE_SIZE - offset % PAGE_SIZE;
-
- if (offset < fsize) {
- part = umin(part, umin(maxsize - extracted, fsize - offset));
- i->count -= part;
- i->iov_offset += part;
- extracted += part;
-
- p[nr++] = folio_page(folio, offset / PAGE_SIZE);
- }
-
- if (nr >= maxpages || extracted >= maxsize)
- break;
-
- if (i->iov_offset >= fsize) {
- i->iov_offset = 0;
- slot++;
- if (slot == folioq_nr_slots(folioq) && folioq->next) {
- folioq = folioq->next;
- slot = 0;
- }
- }
- }
-
- i->folioq = folioq;
- i->folioq_slot = slot;
- return extracted;
-}
-
/*
* Extract a list of contiguous pages from an ITER_XARRAY iterator. This does not
* get references on the pages, nor does it get a pin on them.
@@ -2219,8 +1993,8 @@ static ssize_t iov_iter_extract_user_pages(struct iov_iter *i,
* added to the pages, but refs will not be taken.
* iov_iter_extract_will_pin() will return true.
*
- * (*) If the iterator is ITER_KVEC, ITER_BVEC, ITER_FOLIOQ or ITER_XARRAY, the
- * pages are merely listed; no extra refs or pins are obtained.
+ * (*) If the iterator is ITER_KVEC, ITER_BVEC, ITER_XARRAY, the pages are
+ * merely listed; no extra refs or pins are obtained.
* iov_iter_extract_will_pin() will return 0.
*
* Note also:
@@ -2259,10 +2033,6 @@ ssize_t iov_iter_extract_pages(struct iov_iter *i,
return iov_iter_extract_bvecq_pages(i, pages, maxsize,
maxpages, extraction_flags,
offset0);
- if (iov_iter_is_folioq(i))
- return iov_iter_extract_folioq_pages(i, pages, maxsize,
- maxpages, extraction_flags,
- offset0);
if (iov_iter_is_xarray(i))
return iov_iter_extract_xarray_pages(i, pages, maxsize,
maxpages, extraction_flags,
diff --git a/lib/scatterlist.c b/lib/scatterlist.c
index 23e5a180103b..b9a7298306d9 100644
--- a/lib/scatterlist.c
+++ b/lib/scatterlist.c
@@ -12,7 +12,6 @@
#include <linux/bvec.h>
#include <linux/bvecq.h>
#include <linux/uio.h>
-#include <linux/folio_queue.h>
/**
* sg_nents - return total count of entries in scatterlist
@@ -1327,67 +1326,6 @@ static ssize_t extract_bvecq_to_sg(struct iov_iter *iter,
return ret;
}
-/*
- * Extract up to sg_max folios from an FOLIOQ-type iterator and add them to
- * the scatterlist. The pages are not pinned.
- */
-static ssize_t extract_folioq_to_sg(struct iov_iter *iter,
- ssize_t maxsize,
- struct sg_table *sgtable,
- unsigned int sg_max,
- iov_iter_extraction_t extraction_flags)
-{
- const struct folio_queue *folioq = iter->folioq;
- struct scatterlist *sg = sgtable->sgl + sgtable->nents;
- unsigned int slot = iter->folioq_slot;
- ssize_t ret = 0;
- size_t offset = iter->iov_offset;
-
- BUG_ON(!folioq);
-
- if (slot >= folioq_nr_slots(folioq)) {
- folioq = folioq->next;
- if (WARN_ON_ONCE(!folioq))
- return 0;
- slot = 0;
- }
-
- do {
- struct folio *folio = folioq_folio(folioq, slot);
- size_t fsize = folioq_folio_size(folioq, slot);
-
- if (offset < fsize) {
- size_t part = umin(maxsize - ret, fsize - offset);
-
- sg_set_page(sg, folio_page(folio, 0), part, offset);
- sgtable->nents++;
- sg++;
- sg_max--;
- offset += part;
- ret += part;
- }
-
- if (offset >= fsize) {
- offset = 0;
- slot++;
- if (slot >= folioq_nr_slots(folioq)) {
- if (!folioq->next) {
- WARN_ON_ONCE(ret < iter->count);
- break;
- }
- folioq = folioq->next;
- slot = 0;
- }
- }
- } while (sg_max > 0 && ret < maxsize);
-
- iter->folioq = folioq;
- iter->folioq_slot = slot;
- iter->iov_offset = offset;
- iter->count -= ret;
- return ret;
-}
-
/*
* Extract up to sg_max folios from an XARRAY-type iterator and add them to
* the scatterlist. The pages are not pinned.
@@ -1451,8 +1389,8 @@ static ssize_t extract_xarray_to_sg(struct iov_iter *iter,
* addition of @sg_max elements.
*
* The pages referred to by UBUF- and IOVEC-type iterators are extracted and
- * pinned; BVEC-, BVECQ-, KVEC-, FOLIOQ- and XARRAY-type are extracted but
- * aren't pinned; DISCARD-type is not supported.
+ * pinned; BVEC-, BVECQ-, KVEC-, XARRAY-type are extracted but aren't pinned;
+ * DISCARD-type is not supported.
*
* No end mark is placed on the scatterlist; that's left to the caller.
*
@@ -1487,9 +1425,6 @@ ssize_t extract_iter_to_sg(struct iov_iter *iter, size_t maxsize,
case ITER_BVECQ:
return extract_bvecq_to_sg(iter, maxsize, sgtable, sg_max,
extraction_flags);
- case ITER_FOLIOQ:
- return extract_folioq_to_sg(iter, maxsize, sgtable, sg_max,
- extraction_flags);
case ITER_XARRAY:
return extract_xarray_to_sg(iter, maxsize, sgtable, sg_max,
extraction_flags);
diff --git a/lib/tests/kunit_iov_iter.c b/lib/tests/kunit_iov_iter.c
index a3aeeca6ed58..a2a83ba613fd 100644
--- a/lib/tests/kunit_iov_iter.c
+++ b/lib/tests/kunit_iov_iter.c
@@ -13,7 +13,6 @@
#include <linux/uio.h>
#include <linux/bvec.h>
#include <linux/bvecq.h>
-#include <linux/folio_queue.h>
#include <linux/scatterlist.h>
#include <linux/minmax.h>
#include <linux/mman.h>
@@ -383,176 +382,6 @@ static void __init iov_kunit_copy_from_bvec(struct kunit *test)
KUNIT_SUCCEED(test);
}
-static void iov_kunit_destroy_folioq(void *data)
-{
- struct folio_queue *folioq, *next;
-
- for (folioq = data; folioq; folioq = next) {
- next = folioq->next;
- kfree(folioq);
- }
-}
-
-static void __init iov_kunit_load_folioq(struct kunit *test,
- struct iov_iter *iter, int dir,
- struct folio_queue *folioq,
- struct page **pages, size_t npages)
-{
- struct folio_queue *p = folioq;
- size_t size = 0;
- int i;
-
- for (i = 0; i < npages; i++) {
- if (folioq_full(p)) {
- p->next = kzalloc_obj(struct folio_queue);
- KUNIT_ASSERT_NOT_ERR_OR_NULL(test, p->next);
- folioq_init(p->next, 0);
- p->next->prev = p;
- p = p->next;
- }
- folioq_append(p, page_folio(pages[i]));
- size += PAGE_SIZE;
- }
- iov_iter_folio_queue(iter, dir, folioq, 0, 0, size);
-}
-
-static struct folio_queue *iov_kunit_create_folioq(struct kunit *test)
-{
- struct folio_queue *folioq;
-
- folioq = kzalloc_obj(struct folio_queue);
- KUNIT_ASSERT_NOT_ERR_OR_NULL(test, folioq);
- kunit_add_action_or_reset(test, iov_kunit_destroy_folioq, folioq);
- folioq_init(folioq, 0);
- return folioq;
-}
-
-/*
- * Test copying to a ITER_FOLIOQ-type iterator.
- */
-static void __init iov_kunit_copy_to_folioq(struct kunit *test)
-{
- const struct kvec_test_range *pr;
- struct iov_iter iter;
- struct folio_queue *folioq;
- struct page **spages, **bpages;
- u8 *scratch, *buffer;
- size_t bufsize, npages, size, copied;
- int i, patt;
-
- bufsize = 0x100000;
- npages = bufsize / PAGE_SIZE;
-
- folioq = iov_kunit_create_folioq(test);
-
- scratch = iov_kunit_create_buffer(test, &spages, npages);
- for (i = 0; i < bufsize; i++)
- scratch[i] = pattern(i);
-
- buffer = iov_kunit_create_buffer(test, &bpages, npages);
- memset(buffer, 0, bufsize);
-
- iov_kunit_load_folioq(test, &iter, READ, folioq, bpages, npages);
-
- i = 0;
- for (pr = kvec_test_ranges; pr->from >= 0; pr++) {
- size = pr->to - pr->from;
- KUNIT_ASSERT_LE(test, pr->to, bufsize);
-
- iov_iter_folio_queue(&iter, READ, folioq, 0, 0, pr->to);
- iov_iter_advance(&iter, pr->from);
- copied = copy_to_iter(scratch + i, size, &iter);
-
- KUNIT_EXPECT_EQ(test, copied, size);
- KUNIT_EXPECT_EQ(test, iter.count, 0);
- KUNIT_EXPECT_EQ(test, iter.iov_offset, pr->to % PAGE_SIZE);
- i += size;
- if (test->status == KUNIT_FAILURE)
- goto stop;
- }
-
- /* Build the expected image in the scratch buffer. */
- patt = 0;
- memset(scratch, 0, bufsize);
- for (pr = kvec_test_ranges; pr->from >= 0; pr++)
- for (i = pr->from; i < pr->to; i++)
- scratch[i] = pattern(patt++);
-
- /* Compare the images */
- for (i = 0; i < bufsize; i++) {
- KUNIT_EXPECT_EQ_MSG(test, buffer[i], scratch[i], "at i=%x", i);
- if (buffer[i] != scratch[i])
- return;
- }
-
-stop:
- KUNIT_SUCCEED(test);
-}
-
-/*
- * Test copying from a ITER_FOLIOQ-type iterator.
- */
-static void __init iov_kunit_copy_from_folioq(struct kunit *test)
-{
- const struct kvec_test_range *pr;
- struct iov_iter iter;
- struct folio_queue *folioq;
- struct page **spages, **bpages;
- u8 *scratch, *buffer;
- size_t bufsize, npages, size, copied;
- int i, j;
-
- bufsize = 0x100000;
- npages = bufsize / PAGE_SIZE;
-
- folioq = iov_kunit_create_folioq(test);
-
- buffer = iov_kunit_create_buffer(test, &bpages, npages);
- for (i = 0; i < bufsize; i++)
- buffer[i] = pattern(i);
-
- scratch = iov_kunit_create_buffer(test, &spages, npages);
- memset(scratch, 0, bufsize);
-
- iov_kunit_load_folioq(test, &iter, READ, folioq, bpages, npages);
-
- i = 0;
- for (pr = kvec_test_ranges; pr->from >= 0; pr++) {
- size = pr->to - pr->from;
- KUNIT_ASSERT_LE(test, pr->to, bufsize);
-
- iov_iter_folio_queue(&iter, WRITE, folioq, 0, 0, pr->to);
- iov_iter_advance(&iter, pr->from);
- copied = copy_from_iter(scratch + i, size, &iter);
-
- KUNIT_EXPECT_EQ(test, copied, size);
- KUNIT_EXPECT_EQ(test, iter.count, 0);
- KUNIT_EXPECT_EQ(test, iter.iov_offset, pr->to % PAGE_SIZE);
- i += size;
- }
-
- /* Build the expected image in the main buffer. */
- i = 0;
- memset(buffer, 0, bufsize);
- for (pr = kvec_test_ranges; pr->from >= 0; pr++) {
- for (j = pr->from; j < pr->to; j++) {
- buffer[i++] = pattern(j);
- if (i >= bufsize)
- goto stop;
- }
- }
-stop:
-
- /* Compare the images */
- for (i = 0; i < bufsize; i++) {
- KUNIT_EXPECT_EQ_MSG(test, scratch[i], buffer[i], "at i=%x", i);
- if (scratch[i] != buffer[i])
- return;
- }
-
- KUNIT_SUCCEED(test);
-}
-
static void iov_kunit_destroy_bvecq(void *data)
{
struct bvecq *bq, *next;
@@ -1124,85 +953,6 @@ static void __init iov_kunit_extract_pages_bvecq(struct kunit *test)
KUNIT_SUCCEED(test);
}
-/*
- * Test the extraction of ITER_FOLIOQ-type iterators.
- */
-static void __init iov_kunit_extract_pages_folioq(struct kunit *test)
-{
- const struct kvec_test_range *pr;
- struct folio_queue *folioq;
- struct iov_iter iter;
- struct page **bpages, *pagelist[8], **pages = pagelist;
- ssize_t len;
- size_t bufsize, size = 0, npages;
- int i, from;
-
- bufsize = 0x100000;
- npages = bufsize / PAGE_SIZE;
-
- folioq = iov_kunit_create_folioq(test);
-
- iov_kunit_create_buffer(test, &bpages, npages);
- iov_kunit_load_folioq(test, &iter, READ, folioq, bpages, npages);
-
- for (pr = kvec_test_ranges; pr->from >= 0; pr++) {
- from = pr->from;
- size = pr->to - from;
- KUNIT_ASSERT_LE(test, pr->to, bufsize);
-
- iov_iter_folio_queue(&iter, WRITE, folioq, 0, 0, pr->to);
- iov_iter_advance(&iter, from);
-
- do {
- size_t offset0 = LONG_MAX;
-
- for (i = 0; i < ARRAY_SIZE(pagelist); i++)
- pagelist[i] = (void *)(unsigned long)0xaa55aa55aa55aa55ULL;
-
- len = iov_iter_extract_pages(&iter, &pages, 100 * 1024,
- ARRAY_SIZE(pagelist), 0, &offset0);
- KUNIT_EXPECT_GE(test, len, 0);
- if (len < 0)
- break;
- KUNIT_EXPECT_LE(test, len, size);
- KUNIT_EXPECT_EQ(test, iter.count, size - len);
- if (len == 0)
- break;
- size -= len;
- KUNIT_EXPECT_GE(test, (ssize_t)offset0, 0);
- KUNIT_EXPECT_LT(test, offset0, PAGE_SIZE);
-
- for (i = 0; i < ARRAY_SIZE(pagelist); i++) {
- struct page *p;
- ssize_t part = min_t(ssize_t, len, PAGE_SIZE - offset0);
- int ix;
-
- KUNIT_ASSERT_GE(test, part, 0);
- ix = from / PAGE_SIZE;
- KUNIT_ASSERT_LT(test, ix, npages);
- p = bpages[ix];
- KUNIT_EXPECT_PTR_EQ(test, pagelist[i], p);
- KUNIT_EXPECT_EQ(test, offset0, from % PAGE_SIZE);
- from += part;
- len -= part;
- KUNIT_ASSERT_GE(test, len, 0);
- if (len == 0)
- break;
- offset0 = 0;
- }
-
- if (test->status == KUNIT_FAILURE)
- goto stop;
- } while (iov_iter_count(&iter) > 0);
-
- KUNIT_EXPECT_EQ(test, size, 0);
- KUNIT_EXPECT_EQ(test, iter.count, 0);
- }
-
-stop:
- KUNIT_SUCCEED(test);
-}
-
/*
* Test the extraction of ITER_XARRAY-type iterators.
*/
@@ -1430,23 +1180,6 @@ static void __init iov_kunit_iter_to_sg_bvec(struct kunit *test)
iov_kunit_iter_to_sg_check(test, &iter, bufsize, &data);
}
-static void __init iov_kunit_iter_to_sg_folioq(struct kunit *test)
-{
- struct iov_kunit_iter_to_sg_data data;
- struct folio_queue *folioq;
- struct iov_iter iter;
- size_t bufsize;
-
- bufsize = 0x200000;
- iov_kunit_iter_to_sg_init(test, bufsize, false, &data);
-
- folioq = iov_kunit_create_folioq(test);
- iov_kunit_load_folioq(test, &iter, READ, folioq, data.pages,
- data.npages);
-
- iov_kunit_iter_to_sg_check(test, &iter, bufsize, &data);
-}
-
static void __init iov_kunit_iter_to_sg_xarray(struct kunit *test)
{
struct iov_kunit_iter_to_sg_data data;
@@ -1485,18 +1218,14 @@ static struct kunit_case __refdata iov_kunit_cases[] = {
KUNIT_CASE(iov_kunit_copy_from_bvec),
KUNIT_CASE(iov_kunit_copy_to_bvecq),
KUNIT_CASE(iov_kunit_copy_from_bvecq),
- KUNIT_CASE(iov_kunit_copy_to_folioq),
- KUNIT_CASE(iov_kunit_copy_from_folioq),
KUNIT_CASE(iov_kunit_copy_to_xarray),
KUNIT_CASE(iov_kunit_copy_from_xarray),
KUNIT_CASE(iov_kunit_extract_pages_kvec),
KUNIT_CASE(iov_kunit_extract_pages_bvec),
KUNIT_CASE(iov_kunit_extract_pages_bvecq),
- KUNIT_CASE(iov_kunit_extract_pages_folioq),
KUNIT_CASE(iov_kunit_extract_pages_xarray),
KUNIT_CASE(iov_kunit_iter_to_sg_kvec),
KUNIT_CASE(iov_kunit_iter_to_sg_bvec),
- KUNIT_CASE(iov_kunit_iter_to_sg_folioq),
KUNIT_CASE(iov_kunit_iter_to_sg_xarray),
KUNIT_CASE(iov_kunit_iter_to_sg_ubuf),
{}
^ permalink raw reply [flat|nested] 12+ messages in thread* [PATCH v12 10/10] netfs: Remove folio_queue
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
` (8 preceding siblings ...)
2026-09-29 7:59 ` [PATCH v12 09/10] iov_iter: Remove ITER_FOLIOQ David Howells
@ 2026-09-29 7:59 ` David Howells
2026-09-29 8:21 ` [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 7:59 UTC (permalink / raw)
To: Christian Brauner
Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel,
Christoph Hellwig
Remove folio_queue as it's no longer used.
Signed-off-by: David Howells <dhowells@redhat.com>
Reviewed-by: Paulo Alcantara <pc@manguebit.org>
cc: Matthew Wilcox <willy@infradead.org>
cc: Christoph Hellwig <hch@infradead.org>
cc: netfs@lists.linux.dev
cc: linux-fsdevel@vger.kernel.org
---
Documentation/core-api/folio_queue.rst | 209 ---------------
Documentation/core-api/index.rst | 1 -
Documentation/filesystems/netfs_library.rst | 2 +-
fs/netfs/internal.h | 5 -
fs/netfs/main.c | 7 -
fs/netfs/misc.c | 96 -------
fs/netfs/rolling_buffer.c | 46 ----
fs/netfs/stats.c | 4 +-
include/linux/folio_queue.h | 282 --------------------
include/linux/netfs.h | 13 -
include/trace/events/netfs.h | 23 --
kernel/bpf/btf.c | 2 -
12 files changed, 2 insertions(+), 688 deletions(-)
delete mode 100644 Documentation/core-api/folio_queue.rst
delete mode 100644 include/linux/folio_queue.h
diff --git a/Documentation/core-api/folio_queue.rst b/Documentation/core-api/folio_queue.rst
deleted file mode 100644
index b7628896d2b6..000000000000
--- a/Documentation/core-api/folio_queue.rst
+++ /dev/null
@@ -1,209 +0,0 @@
-.. SPDX-License-Identifier: GPL-2.0+
-
-===========
-Folio Queue
-===========
-
-:Author: David Howells <dhowells@redhat.com>
-
-.. Contents:
-
- * Overview
- * Initialisation
- * Adding and removing folios
- * Querying information about a folio
- * Querying information about a folio_queue
- * Folio queue iteration
- * Folio marks
- * Lockless simultaneous production/consumption issues
-
-
-Overview
-========
-
-The folio_queue struct forms a single segment in a segmented list of folios
-that can be used to form an I/O buffer. As such, the list can be iterated over
-using the ITER_FOLIOQ iov_iter type.
-
-The publicly accessible members of the structure are::
-
- struct folio_queue {
- struct folio_queue *next;
- struct folio_queue *prev;
- ...
- };
-
-A pair of pointers are provided, ``next`` and ``prev``, that point to the
-segments on either side of the segment being accessed. Whilst this is a
-doubly-linked list, it is intentionally not a circular list; the outward
-sibling pointers in terminal segments should be NULL.
-
-Each segment in the list also stores:
-
- * an ordered sequence of folio pointers,
- * the size of each folio and
- * three 1-bit marks per folio,
-
-but these should not be accessed directly as the underlying data structure may
-change, but rather the access functions outlined below should be used.
-
-The facility can be made accessible by::
-
- #include <linux/folio_queue.h>
-
-and to use the iterator::
-
- #include <linux/uio.h>
-
-
-Initialisation
-==============
-
-A segment should be initialised by calling::
-
- void folioq_init(struct folio_queue *folioq);
-
-with a pointer to the segment to be initialised. Note that this will not
-necessarily initialise all the folio pointers, so care must be taken to check
-the number of folios added.
-
-
-Adding and removing folios
-==========================
-
-Folios can be set in the next unused slot in a segment struct by calling one
-of::
-
- unsigned int folioq_append(struct folio_queue *folioq,
- struct folio *folio);
-
- unsigned int folioq_append_mark(struct folio_queue *folioq,
- struct folio *folio);
-
-Both functions update the stored folio count, store the folio and note its
-size. The second function also sets the first mark for the folio added. Both
-functions return the number of the slot used. [!] Note that no attempt is made
-to check that the capacity wasn't overrun and the list will not be extended
-automatically.
-
-A folio can be excised by calling::
-
- void folioq_clear(struct folio_queue *folioq, unsigned int slot);
-
-This clears the slot in the array and also clears all the marks for that folio,
-but doesn't change the folio count - so future accesses of that slot must check
-if the slot is occupied.
-
-
-Querying information about a folio
-==================================
-
-Information about the folio in a particular slot may be queried by the
-following function::
-
- struct folio *folioq_folio(const struct folio_queue *folioq,
- unsigned int slot);
-
-If a folio has not yet been set in that slot, this may yield an undefined
-pointer. The size of the folio in a slot may be queried with either of::
-
- unsigned int folioq_folio_order(const struct folio_queue *folioq,
- unsigned int slot);
-
- size_t folioq_folio_size(const struct folio_queue *folioq,
- unsigned int slot);
-
-The first function returns the size as an order and the second as a number of
-bytes.
-
-
-Querying information about a folio_queue
-========================================
-
-Information may be retrieved about a particular segment with the following
-functions::
-
- unsigned int folioq_nr_slots(const struct folio_queue *folioq);
-
- unsigned int folioq_count(struct folio_queue *folioq);
-
- bool folioq_full(struct folio_queue *folioq);
-
-The first function returns the maximum capacity of a segment. It must not be
-assumed that this won't vary between segments. The second returns the number
-of folios added to a segments and the third is a shorthand to indicate if the
-segment has been filled to capacity.
-
-Not that the count and fullness are not affected by clearing folios from the
-segment. These are more about indicating how many slots in the array have been
-initialised, and it assumed that slots won't get reused, but rather the segment
-will get discarded as the queue is consumed.
-
-
-Folio marks
-===========
-
-Folios within a queue can also have marks assigned to them. These marks can be
-used to note information such as if a folio needs folio_put() calling upon it.
-There are three marks available to be set for each folio.
-
-The marks can be set by::
-
- void folioq_mark(struct folio_queue *folioq, unsigned int slot);
- void folioq_mark2(struct folio_queue *folioq, unsigned int slot);
-
-Cleared by::
-
- void folioq_unmark(struct folio_queue *folioq, unsigned int slot);
- void folioq_unmark2(struct folio_queue *folioq, unsigned int slot);
-
-And the marks can be queried by::
-
- bool folioq_is_marked(const struct folio_queue *folioq, unsigned int slot);
- bool folioq_is_marked2(const struct folio_queue *folioq, unsigned int slot);
-
-The marks can be used for any purpose and are not interpreted by this API.
-
-
-Folio queue iteration
-=====================
-
-A list of segments may be iterated over using the I/O iterator facility using
-an ``iov_iter`` iterator of ``ITER_FOLIOQ`` type. The iterator may be
-initialised with::
-
- void iov_iter_folio_queue(struct iov_iter *i, unsigned int direction,
- const struct folio_queue *folioq,
- unsigned int first_slot, unsigned int offset,
- size_t count);
-
-This may be told to start at a particular segment, slot and offset within a
-queue. The iov iterator functions will follow the next pointers when advancing
-and prev pointers when reverting when needed.
-
-
-Lockless simultaneous production/consumption issues
-===================================================
-
-If properly managed, the list can be extended by the producer at the head end
-and shortened by the consumer at the tail end simultaneously without the need
-to take locks. The ITER_FOLIOQ iterator inserts appropriate barriers to aid
-with this.
-
-Care must be taken when simultaneously producing and consuming a list. If the
-last segment is reached and the folios it refers to are entirely consumed by
-the IOV iterators, an iov_iter struct will be left pointing to the last segment
-with a slot number equal to the capacity of that segment. The iterator will
-try to continue on from this if there's another segment available when it is
-used again, but care must be taken lest the segment got removed and freed by
-the consumer before the iterator was advanced.
-
-It is recommended that the queue always contain at least one segment, even if
-that segment has never been filled or is entirely spent. This prevents the
-head and tail pointers from collapsing.
-
-
-API Function Reference
-======================
-
-.. kernel-doc:: include/linux/folio_queue.h
diff --git a/Documentation/core-api/index.rst b/Documentation/core-api/index.rst
index 92f91c6a0d79..956f6803d2a2 100644
--- a/Documentation/core-api/index.rst
+++ b/Documentation/core-api/index.rst
@@ -39,7 +39,6 @@ Library functionality that is used throughout the kernel.
kref
cleanup
assoc_array
- folio_queue
xarray
maple_tree
idr
diff --git a/Documentation/filesystems/netfs_library.rst b/Documentation/filesystems/netfs_library.rst
index 0c9786ffe192..ebf07a37922e 100644
--- a/Documentation/filesystems/netfs_library.rst
+++ b/Documentation/filesystems/netfs_library.rst
@@ -452,7 +452,7 @@ be called from the writeback code to write the data to the cache, if there is
one.
The inode should be marked ``NETFS_ICTX_SINGLE_NO_UPLOAD`` if this API is to be
-used. The writeback function requires the buffer to be of ITER_FOLIOQ type.
+used.
High-Level VM API
==================
diff --git a/fs/netfs/internal.h b/fs/netfs/internal.h
index dd20a9f201be..583ffa967d5a 100644
--- a/fs/netfs/internal.h
+++ b/fs/netfs/internal.h
@@ -7,7 +7,6 @@
#include <linux/slab.h>
#include <linux/seq_file.h>
-#include <linux/folio_queue.h>
#include <linux/netfs.h>
#include <linux/fscache.h>
#include <linux/fscache-cache.h>
@@ -46,7 +45,6 @@ extern spinlock_t netfs_proc_lock;
extern mempool_t netfs_request_pool;
extern mempool_t netfs_subrequest_pool;
extern mempool_t netfs_bvecq_pool;
-extern mempool_t netfs_folioq_pool;
#ifdef CONFIG_PROC_FS
static inline void netfs_proc_add_rreq(struct netfs_io_request *rreq)
@@ -71,8 +69,6 @@ static inline void netfs_proc_del_rreq(struct netfs_io_request *rreq) {}
/*
* misc.c
*/
-struct folio_queue *netfs_buffer_make_space(struct netfs_io_request *rreq,
- enum netfs_folioq_trace trace);
void netfs_reset_iter(struct netfs_io_subrequest *subreq);
void netfs_wake_collector(struct netfs_io_request *rreq);
void netfs_subreq_clear_in_progress(struct netfs_io_subrequest *subreq);
@@ -203,7 +199,6 @@ extern atomic_t netfs_n_wh_retry_write_req;
extern atomic_t netfs_n_wh_retry_write_subreq;
extern atomic_t netfs_n_wb_lock_skip;
extern atomic_t netfs_n_wb_lock_wait;
-extern atomic_t netfs_n_folioq;
extern atomic_t netfs_n_bvecq;
int netfs_stats_show(struct seq_file *m, void *v);
diff --git a/fs/netfs/main.c b/fs/netfs/main.c
index 5c8517a66dde..36a1d98e593a 100644
--- a/fs/netfs/main.c
+++ b/fs/netfs/main.c
@@ -29,7 +29,6 @@ static struct kmem_cache *netfs_subrequest_slab;
mempool_t netfs_request_pool;
mempool_t netfs_subrequest_pool;
mempool_t netfs_bvecq_pool;
-mempool_t netfs_folioq_pool;
#ifdef CONFIG_PROC_FS
LIST_HEAD(netfs_io_requests);
@@ -109,9 +108,6 @@ static int __init netfs_init(void)
{
int ret = -ENOMEM;
- if (mempool_init_kmalloc_pool(&netfs_folioq_pool, 100, sizeof(struct folio_queue)) < 0)
- goto error_folioq_pool;
-
if (mempool_init_kmalloc_pool(&netfs_bvecq_pool, 100,
struct_size_t(struct bvecq, __bv, BVECQ_POOL_SLOTS)) < 0)
goto error_bvecq_pool;
@@ -170,8 +166,6 @@ static int __init netfs_init(void)
error_req:
mempool_exit(&netfs_bvecq_pool);
error_bvecq_pool:
- mempool_exit(&netfs_folioq_pool);
-error_folioq_pool:
return ret;
}
fs_initcall(netfs_init);
@@ -185,6 +179,5 @@ static void __exit netfs_exit(void)
mempool_exit(&netfs_request_pool);
kmem_cache_destroy(netfs_request_slab);
mempool_exit(&netfs_bvecq_pool);
- mempool_exit(&netfs_folioq_pool);
}
module_exit(netfs_exit);
diff --git a/fs/netfs/misc.c b/fs/netfs/misc.c
index a3cd76d584b8..40e0ff649132 100644
--- a/fs/netfs/misc.c
+++ b/fs/netfs/misc.c
@@ -8,102 +8,6 @@
#include <linux/swap.h>
#include "internal.h"
-/**
- * netfs_alloc_folioq_buffer - Allocate buffer space into a folio queue
- * @mapping: Address space to set on the folio (or NULL).
- * @_buffer: Pointer to the folio queue to add to (may point to a NULL; updated).
- * @_cur_size: Current size of the buffer (updated).
- * @size: Target size of the buffer.
- * @gfp: The allocation constraints.
- */
-int netfs_alloc_folioq_buffer(struct address_space *mapping,
- struct folio_queue **_buffer,
- size_t *_cur_size, ssize_t size, gfp_t gfp)
-{
- struct folio_queue *tail = *_buffer, *p;
-
- size = round_up(size, PAGE_SIZE);
- if (*_cur_size >= size)
- return 0;
-
- if (tail)
- while (tail->next)
- tail = tail->next;
-
- do {
- struct folio *folio;
- int order = 0, slot;
-
- if (!tail || folioq_full(tail)) {
- p = netfs_folioq_alloc(0, GFP_NOFS, netfs_trace_folioq_alloc_buffer);
- if (!p)
- return -ENOMEM;
- if (tail) {
- tail->next = p;
- p->prev = tail;
- } else {
- *_buffer = p;
- }
- tail = p;
- }
-
- if (size - *_cur_size > PAGE_SIZE)
- order = umin(ilog2(size - *_cur_size) - PAGE_SHIFT,
- MAX_PAGECACHE_ORDER);
-
- folio = folio_alloc(gfp, order);
- if (!folio && order > 0)
- folio = folio_alloc(gfp, 0);
- if (!folio)
- return -ENOMEM;
-
- folio->mapping = mapping;
- folio->index = *_cur_size / PAGE_SIZE;
- trace_netfs_folio(folio, netfs_folio_trace_alloc_buffer);
- slot = folioq_append_mark(tail, folio);
- *_cur_size += folioq_folio_size(tail, slot);
- } while (*_cur_size < size);
-
- return 0;
-}
-EXPORT_SYMBOL(netfs_alloc_folioq_buffer);
-
-/**
- * netfs_free_folioq_buffer - Free a folio queue.
- * @fq: The start of the folio queue to free
- *
- * Free up a chain of folio_queues and, if marked, the marked folios they point
- * to.
- */
-void netfs_free_folioq_buffer(struct folio_queue *fq)
-{
- struct folio_queue *next;
- struct folio_batch fbatch;
-
- folio_batch_init(&fbatch);
-
- for (; fq; fq = next) {
- for (int slot = 0; slot < folioq_count(fq); slot++) {
- struct folio *folio = folioq_folio(fq, slot);
-
- if (!folio ||
- !folioq_is_marked(fq, slot))
- continue;
-
- trace_netfs_folio(folio, netfs_folio_trace_put);
- if (folio_batch_add(&fbatch, folio))
- folio_batch_release(&fbatch);
- }
-
- netfs_stat_d(&netfs_n_folioq);
- next = fq->next;
- kfree(fq);
- }
-
- folio_batch_release(&fbatch);
-}
-EXPORT_SYMBOL(netfs_free_folioq_buffer);
-
/*
* Reset the subrequest iterator to refer just to the region remaining to be
* read. The iterator may or may not have been advanced by socket ops or
diff --git a/fs/netfs/rolling_buffer.c b/fs/netfs/rolling_buffer.c
index 76429fbb6920..66ce9add4012 100644
--- a/fs/netfs/rolling_buffer.c
+++ b/fs/netfs/rolling_buffer.c
@@ -12,52 +12,6 @@
#include <linux/slab.h>
#include "internal.h"
-static atomic_t debug_ids;
-
-/**
- * netfs_folioq_alloc - Allocate a folio_queue struct
- * @rreq_id: Associated debugging ID for tracing purposes
- * @gfp: Allocation constraints
- * @trace: Trace tag to indicate the purpose of the allocation
- *
- * Allocate, initialise and account the folio_queue struct and log a trace line
- * to mark the allocation.
- */
-struct folio_queue *netfs_folioq_alloc(unsigned int rreq_id, gfp_t gfp,
- unsigned int /*enum netfs_folioq_trace*/ trace)
-{
- struct folio_queue *fq;
-
- if (gfp == GFP_KERNEL)
- fq = mempool_alloc_noreserve(&netfs_folioq_pool, gfp);
- else
- fq = mempool_alloc(&netfs_folioq_pool, gfp);
- if (fq) {
- netfs_stat(&netfs_n_folioq);
- folioq_init(fq, rreq_id);
- fq->debug_id = atomic_inc_return(&debug_ids);
- trace_netfs_folioq(fq, trace);
- }
- return fq;
-}
-EXPORT_SYMBOL(netfs_folioq_alloc);
-
-/**
- * netfs_folioq_free - Free a folio_queue struct
- * @folioq: The object to free
- * @trace: Trace tag to indicate which free
- *
- * Free and unaccount the folio_queue struct.
- */
-void netfs_folioq_free(struct folio_queue *folioq,
- unsigned int /*enum netfs_trace_folioq*/ trace)
-{
- trace_netfs_folioq(folioq, trace);
- netfs_stat_d(&netfs_n_folioq);
- mempool_free(folioq, &netfs_folioq_pool);
-}
-EXPORT_SYMBOL(netfs_folioq_free);
-
/*
* Initialise a rolling buffer. We allocate an empty folio queue struct to so
* that the pointers can be independently driven by the producer and the
diff --git a/fs/netfs/stats.c b/fs/netfs/stats.c
index a10d34f88597..0ba6ce9295b2 100644
--- a/fs/netfs/stats.c
+++ b/fs/netfs/stats.c
@@ -46,7 +46,6 @@ atomic_t netfs_n_wh_retry_write_req;
atomic_t netfs_n_wh_retry_write_subreq;
atomic_t netfs_n_wb_lock_skip;
atomic_t netfs_n_wb_lock_wait;
-atomic_t netfs_n_folioq;
atomic_t netfs_n_bvecq;
int netfs_stats_show(struct seq_file *m, void *v)
@@ -89,11 +88,10 @@ int netfs_stats_show(struct seq_file *m, void *v)
atomic_read(&netfs_n_rh_retry_read_subreq),
atomic_read(&netfs_n_wh_retry_write_req),
atomic_read(&netfs_n_wh_retry_write_subreq));
- seq_printf(m, "Objs : rr=%u sr=%u bq=%u foq=%u wsc=%u\n",
+ seq_printf(m, "Objs : rr=%u sr=%u bq=%u wsc=%u\n",
atomic_read(&netfs_n_rh_rreq),
atomic_read(&netfs_n_rh_sreq),
atomic_read(&netfs_n_bvecq),
- atomic_read(&netfs_n_folioq),
atomic_read(&netfs_n_wh_wstream_conflict));
seq_printf(m, "WbLock : skip=%u wait=%u\n",
atomic_read(&netfs_n_wb_lock_skip),
diff --git a/include/linux/folio_queue.h b/include/linux/folio_queue.h
deleted file mode 100644
index f6d5f1f127c9..000000000000
--- a/include/linux/folio_queue.h
+++ /dev/null
@@ -1,282 +0,0 @@
-/* SPDX-License-Identifier: GPL-2.0-or-later */
-/* Queue of folios definitions
- *
- * Copyright (C) 2024 Red Hat, Inc. All Rights Reserved.
- * Written by David Howells (dhowells@redhat.com)
- *
- * See:
- *
- * Documentation/core-api/folio_queue.rst
- *
- * for a description of the API.
- */
-
-#ifndef _LINUX_FOLIO_QUEUE_H
-#define _LINUX_FOLIO_QUEUE_H
-
-#include <linux/folio_batch.h>
-#include <linux/mm.h>
-
-/*
- * Segment in a queue of running buffers. Each segment can hold a number of
- * folios and a portion of the queue can be referenced with the ITER_FOLIOQ
- * iterator. The possibility exists of inserting non-folio elements into the
- * queue (such as gaps).
- *
- * Explicit prev and next pointers are used instead of a list_head to make it
- * easier to add segments to tail and remove them from the head without the
- * need for a lock.
- */
-struct folio_queue {
- struct folio_batch vec; /* Folios in the queue segment */
- u8 orders[FOLIO_BATCH_SIZE]; /* Order of each folio */
- struct folio_queue *next; /* Next queue segment or NULL */
- struct folio_queue *prev; /* Previous queue segment of NULL */
- unsigned long marks; /* 1-bit mark per folio */
- unsigned long marks2; /* Second 1-bit mark per folio */
-#if FOLIO_BATCH_SIZE > BITS_PER_LONG
-#error marks is not big enough
-#endif
- unsigned int rreq_id;
- unsigned int debug_id;
-};
-
-/**
- * folioq_init - Initialise a folio queue segment
- * @folioq: The segment to initialise
- * @rreq_id: The request identifier to use in tracelines.
- *
- * Initialise a folio queue segment and set an identifier to be used in traces.
- *
- * Note that the folio pointers are left uninitialised.
- */
-static inline void folioq_init(struct folio_queue *folioq, unsigned int rreq_id)
-{
- folio_batch_init(&folioq->vec);
- folioq->next = NULL;
- folioq->prev = NULL;
- folioq->marks = 0;
- folioq->marks2 = 0;
- folioq->rreq_id = rreq_id;
- folioq->debug_id = 0;
-}
-
-/**
- * folioq_nr_slots: Query the capacity of a folio queue segment
- * @folioq: The segment to query
- *
- * Query the number of folios that a particular folio queue segment might hold.
- * [!] NOTE: This must not be assumed to be the same for every segment!
- */
-static inline unsigned int folioq_nr_slots(const struct folio_queue *folioq)
-{
- return FOLIO_BATCH_SIZE;
-}
-
-/**
- * folioq_count: Query the occupancy of a folio queue segment
- * @folioq: The segment to query
- *
- * Query the number of folios that have been added to a folio queue segment.
- * Note that this is not decreased as folios are removed from a segment.
- */
-static inline unsigned int folioq_count(struct folio_queue *folioq)
-{
- return folio_batch_count(&folioq->vec);
-}
-
-/**
- * folioq_full: Query if a folio queue segment is full
- * @folioq: The segment to query
- *
- * Query if a folio queue segment is fully occupied. Note that this does not
- * change if folios are removed from a segment.
- */
-static inline bool folioq_full(struct folio_queue *folioq)
-{
- //return !folio_batch_space(&folioq->vec);
- return folioq_count(folioq) >= folioq_nr_slots(folioq);
-}
-
-/**
- * folioq_is_marked: Check first folio mark in a folio queue segment
- * @folioq: The segment to query
- * @slot: The slot number of the folio to query
- *
- * Determine if the first mark is set for the folio in the specified slot in a
- * folio queue segment.
- */
-static inline bool folioq_is_marked(const struct folio_queue *folioq, unsigned int slot)
-{
- return test_bit(slot, &folioq->marks);
-}
-
-/**
- * folioq_mark: Set the first mark on a folio in a folio queue segment
- * @folioq: The segment to modify
- * @slot: The slot number of the folio to modify
- *
- * Set the first mark for the folio in the specified slot in a folio queue
- * segment.
- */
-static inline void folioq_mark(struct folio_queue *folioq, unsigned int slot)
-{
- set_bit(slot, &folioq->marks);
-}
-
-/**
- * folioq_unmark: Clear the first mark on a folio in a folio queue segment
- * @folioq: The segment to modify
- * @slot: The slot number of the folio to modify
- *
- * Clear the first mark for the folio in the specified slot in a folio queue
- * segment.
- */
-static inline void folioq_unmark(struct folio_queue *folioq, unsigned int slot)
-{
- clear_bit(slot, &folioq->marks);
-}
-
-/**
- * folioq_is_marked2: Check second folio mark in a folio queue segment
- * @folioq: The segment to query
- * @slot: The slot number of the folio to query
- *
- * Determine if the second mark is set for the folio in the specified slot in a
- * folio queue segment.
- */
-static inline bool folioq_is_marked2(const struct folio_queue *folioq, unsigned int slot)
-{
- return test_bit(slot, &folioq->marks2);
-}
-
-/**
- * folioq_mark2: Set the second mark on a folio in a folio queue segment
- * @folioq: The segment to modify
- * @slot: The slot number of the folio to modify
- *
- * Set the second mark for the folio in the specified slot in a folio queue
- * segment.
- */
-static inline void folioq_mark2(struct folio_queue *folioq, unsigned int slot)
-{
- set_bit(slot, &folioq->marks2);
-}
-
-/**
- * folioq_unmark2: Clear the second mark on a folio in a folio queue segment
- * @folioq: The segment to modify
- * @slot: The slot number of the folio to modify
- *
- * Clear the second mark for the folio in the specified slot in a folio queue
- * segment.
- */
-static inline void folioq_unmark2(struct folio_queue *folioq, unsigned int slot)
-{
- clear_bit(slot, &folioq->marks2);
-}
-
-/**
- * folioq_append: Add a folio to a folio queue segment
- * @folioq: The segment to add to
- * @folio: The folio to add
- *
- * Add a folio to the tail of the sequence in a folio queue segment, increasing
- * the occupancy count and returning the slot number for the folio just added.
- * The folio size is extracted and stored in the queue and the marks are left
- * unmodified.
- *
- * Note that it's left up to the caller to check that the segment capacity will
- * not be exceeded and to extend the queue.
- */
-static inline unsigned int folioq_append(struct folio_queue *folioq, struct folio *folio)
-{
- unsigned int slot = folioq->vec.nr++;
-
- folioq->vec.folios[slot] = folio;
- folioq->orders[slot] = folio_order(folio);
- return slot;
-}
-
-/**
- * folioq_append_mark: Add a folio to a folio queue segment
- * @folioq: The segment to add to
- * @folio: The folio to add
- *
- * Add a folio to the tail of the sequence in a folio queue segment, increasing
- * the occupancy count and returning the slot number for the folio just added.
- * The folio size is extracted and stored in the queue, the first mark is set
- * and and the second and third marks are left unmodified.
- *
- * Note that it's left up to the caller to check that the segment capacity will
- * not be exceeded and to extend the queue.
- */
-static inline unsigned int folioq_append_mark(struct folio_queue *folioq, struct folio *folio)
-{
- unsigned int slot = folioq->vec.nr++;
-
- folioq->vec.folios[slot] = folio;
- folioq->orders[slot] = folio_order(folio);
- folioq_mark(folioq, slot);
- return slot;
-}
-
-/**
- * folioq_folio: Get a folio from a folio queue segment
- * @folioq: The segment to access
- * @slot: The folio slot to access
- *
- * Retrieve the folio in the specified slot from a folio queue segment. Note
- * that no bounds check is made and if the slot hasn't been added into yet, the
- * pointer will be undefined. If the slot has been cleared, NULL will be
- * returned.
- */
-static inline struct folio *folioq_folio(const struct folio_queue *folioq, unsigned int slot)
-{
- return folioq->vec.folios[slot];
-}
-
-/**
- * folioq_folio_order: Get the order of a folio from a folio queue segment
- * @folioq: The segment to access
- * @slot: The folio slot to access
- *
- * Retrieve the order of the folio in the specified slot from a folio queue
- * segment. Note that no bounds check is made and if the slot hasn't been
- * added into yet, the order returned will be 0.
- */
-static inline unsigned int folioq_folio_order(const struct folio_queue *folioq, unsigned int slot)
-{
- return folioq->orders[slot];
-}
-
-/**
- * folioq_folio_size: Get the size of a folio from a folio queue segment
- * @folioq: The segment to access
- * @slot: The folio slot to access
- *
- * Retrieve the size of the folio in the specified slot from a folio queue
- * segment. Note that no bounds check is made and if the slot hasn't been
- * added into yet, the size returned will be PAGE_SIZE.
- */
-static inline size_t folioq_folio_size(const struct folio_queue *folioq, unsigned int slot)
-{
- return PAGE_SIZE << folioq_folio_order(folioq, slot);
-}
-
-/**
- * folioq_clear: Clear a folio from a folio queue segment
- * @folioq: The segment to clear
- * @slot: The folio slot to clear
- *
- * Clear a folio from a sequence in a folio queue segment and clear its marks.
- * The occupancy count is left unchanged.
- */
-static inline void folioq_clear(struct folio_queue *folioq, unsigned int slot)
-{
- folioq->vec.folios[slot] = NULL;
- folioq_unmark(folioq, slot);
- folioq_unmark2(folioq, slot);
-}
-
-#endif /* _LINUX_FOLIO_QUEUE_H */
diff --git a/include/linux/netfs.h b/include/linux/netfs.h
index 1723bcda86ba..31a80f4377d7 100644
--- a/include/linux/netfs.h
+++ b/include/linux/netfs.h
@@ -24,7 +24,6 @@
enum netfs_sreq_ref_trace;
typedef struct mempool mempool_t;
struct fscache_occupancy;
-struct folio_queue;
/**
* folio_start_private_2 - Start an fscache write on a folio. [DEPRECATED]
@@ -470,18 +469,6 @@ void netfs_end_io_write(struct inode *inode);
int netfs_start_io_direct(struct inode *inode);
void netfs_end_io_direct(struct inode *inode);
-/* Miscellaneous APIs. */
-struct folio_queue *netfs_folioq_alloc(unsigned int rreq_id, gfp_t gfp,
- unsigned int trace /*enum netfs_folioq_trace*/);
-void netfs_folioq_free(struct folio_queue *folioq,
- unsigned int trace /*enum netfs_trace_folioq*/);
-
-/* Buffer wrangling helpers API. */
-int netfs_alloc_folioq_buffer(struct address_space *mapping,
- struct folio_queue **_buffer,
- size_t *_cur_size, ssize_t size, gfp_t gfp);
-void netfs_free_folioq_buffer(struct folio_queue *fq);
-
/* Writeback exclusion API. */
bool netfs_wb_begin(struct netfs_inode *ictx, bool nowait);
void netfs_wb_end(struct netfs_inode *ictx);
diff --git a/include/trace/events/netfs.h b/include/trace/events/netfs.h
index 1fd62465e809..c6ba891e0413 100644
--- a/include/trace/events/netfs.h
+++ b/include/trace/events/netfs.h
@@ -773,29 +773,6 @@ TRACE_EVENT(netfs_collect_stream,
__entry->collected_to, __entry->issued_to)
);
-TRACE_EVENT(netfs_folioq,
- TP_PROTO(const struct folio_queue *fq,
- enum netfs_folioq_trace trace),
-
- TP_ARGS(fq, trace),
-
- TP_STRUCT__entry(
- __field(unsigned int, rreq)
- __field(unsigned int, id)
- __field(enum netfs_folioq_trace, trace)
- ),
-
- TP_fast_assign(
- __entry->rreq = fq ? fq->rreq_id : 0;
- __entry->id = fq ? fq->debug_id : 0;
- __entry->trace = trace;
- ),
-
- TP_printk("R=%08x fq=%x %s",
- __entry->rreq, __entry->id,
- __print_symbolic(__entry->trace, netfs_folioq_traces))
- );
-
TRACE_EVENT(netfs_read_progress_at,
TP_PROTO(const struct netfs_io_request *rreq),
diff --git a/kernel/bpf/btf.c b/kernel/bpf/btf.c
index d870bc5e50bc..63fd6156ffa7 100644
--- a/kernel/bpf/btf.c
+++ b/kernel/bpf/btf.c
@@ -6803,8 +6803,6 @@ static const struct bpf_raw_tp_null_args raw_tp_null_args[] = {
/* amdgpu */
{ "amdgpu_vm_bo_map", 0x1 },
{ "amdgpu_vm_bo_unmap", 0x1 },
- /* netfs */
- { "netfs_folioq", 0x1 },
/* xfs from xfs_defer_pending_class */
{ "xfs_defer_create_intent", 0x1 },
{ "xfs_defer_cancel_list", 0x1 },
^ permalink raw reply [flat|nested] 12+ messages in thread* Re: [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related
2026-09-29 7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
` (9 preceding siblings ...)
2026-09-29 7:59 ` [PATCH v12 10/10] netfs: Remove folio_queue David Howells
@ 2026-09-29 8:21 ` David Howells
10 siblings, 0 replies; 12+ messages in thread
From: David Howells @ 2026-09-29 8:21 UTC (permalink / raw)
To: Christian Brauner
Cc: dhowells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel
I forgot to change the cover subject line. It should be something like:
netfs, iov_iter: Use a chain of bio_vec arrays instead of folio_queue
David
^ permalink raw reply [flat|nested] 12+ messages in thread