mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related
@ 2026-09-29  7:59 David Howells
  2026-09-29  7:59 ` [PATCH v12 01/10] Add a function to kmap one page of a multipage bio_vec David Howells
                   ` (10 more replies)
  0 siblings, 11 replies; 12+ messages in thread
From: David Howells @ 2026-09-29  7:59 UTC (permalink / raw)
  To: Christian Brauner
  Cc: David Howells, Paulo Alcantara, Matthew Wilcox, Namjae Jeon,
	Marc Dionne, Stefan Metzmacher, Eric Van Hensbergen,
	Dominique Martinet, Ilya Dryomov, netfs, linux-afs, linux-cifs,
	linux-nfs, ceph-devel, v9fs, linux-fsdevel, linux-kernel

Hi Christian,

Could you pull these patches please into your vfs-7.4.netfs branch?  This
is the third of four (maybe five) batches, in this case replacing the
folio_queue struct with a bvecq struct, and is based upon the
aforementioned branch.  This was split from v11 of a larger netfslib
series[1].

This series adds a new type, struct bvecq.  This points to a bounded array
of bio_vecs with information about how to clean up the memory they point
to.  A number of bvecq structs can be linked together to form a chain and
an iterator type, ITER_BVECQ, is provided that can be pointed at the chain
and can walk it.  Netfslib is then switched to use the new type and
folio_queue is removed along with ITER_FOLIOQ.

This is intended as a more flexible replacement for struct folio_queue as
the latter can only point to whole folios and cannot point to partial pages
that might be extracted from an ITER_IOVEC/ITER_UBUF for direct or
unbuffered I/O purposes.

These patches make a more or less straight swap between struct folio_queue
and struct bvecq.  The change is not quite trivial as the folio_queue
struct contains a folio_batch struct and is basically a list of whole
folios only, whereas the new bvecq struct is a chain of bio_vec arrays.

A further patchset will switch the netfslib unbuffered I/O code to using
bvecqs and then the bvecqs will be passed down into the filesystem rather
than passing an iterator.

The primary reason behind making this change is so that bvecq chains can be
used in the assembly of network filesystem messages.  Netfslib can furnish
the filesystem with bvecq chains representing the data buffers that need to
be read or written and then, for example, for a write RPC, the filesystem
can add protocol headers and trailers onto the chain without the need to
copy it and can also glue multiple buffers together to form compound
operations and/or perform sparse operations.

One option here is to add two more fields to the bvecq struct, one to
increase the offset of the first bio_vec and one to decrease the length of
the last, thereby allowing a bvecq struct to point at a slice of another
bio_vec array without having to copy the array - but at the cost of adding
one or two extra conditional ops when getting the offset or length of a
segment.

I have in-progress patch sets to rewrite the cifs transport[2] and the ceph
and rbd transport[3] to make use of bvecq chains.  With the cifs transport,
the idea is to extend the SMB message concept further up the stack and have
the PDU creation routines attach individual RPC requests blobs that can
then be automatically chained by the transport when a compound is being
assembled for transmission.

With the ceph transport, the idea is to convert all the different data
containers it has into just passing around bvecqs.  These can then be
chained together in order to transmit them.

This improves efficiency in the TCP stack as we no longer need to cork the
TCP socket, call sendmsg() multiple times and then uncork; rather we can
preassemble the message in a bvecq chain and just send the entire messsage
in one shot with a single sendmsg() and reduce the number of places doing
loops.  This got an improvement in nfsd performance[4].

I also have some changes on the TCP receive side for the cifs transport
(which is also likely applicable to the ceph TCP transport) whereby the
receive buffers are 'spliced' out into a bvecq in the cifs I/O thread
rather than being copied.  Using a bvecq chain here is advantageous as we
don't know in advance how many segments we're going to have.  This allows
us firstly to avoid copying data with the socket lock held (thus holding up
sendmsg) and secondly to avoid copying data in the I/O thread (copying can
be offloaded to the app thread).  The last time I benchmarked this, it
appeared to get fio reading tests on cifs a 5% speedup.

Unfortunately, this doesn't help AFS much as that uses a UDP transport, but
it might also help 9P, at least with its TCP transport.

The patches can also be found here:

	https://git.kernel.org/pub/scm/linux/kernel/git/dhowells/linux-fs.git/log/?h=netfs-next-3

Thanks,
David

Changes
=======

ver #12)
- Split from v11 of "netfs: Keep track of folios in a segmented bio_vec[]
  chain"[1]

[1] https://lore.kernel.org/r/20260902173350.3468672-1-dhowells@redhat.com/
[2] https://git.kernel.org/pub/scm/linux/kernel/git/dhowells/linux-fs.git/log/?h=cifs-experimental
[3] https://git.kernel.org/pub/scm/linux/kernel/git/dhowells/linux-fs.git/log/?h=ceph-iter
[4] https://lore.kernel.org/netdev/168979108540.1905271.9720708849149797793.stgit@morisot.1015granger.net/

David Howells (10):
  Add a function to kmap one page of a multipage bio_vec
  iov_iter: Add a segmented queue of bio_vec[]
  netfs: Add some tools for managing bvecq chains
  afs: Use a bvecq to hold dir content rather than folioq
  cifs: Use a bvecq for buffering instead of a folioq
  smbdirect: Support ITER_BVECQ in smbdirect_map_sges_from_iter()
  netfs: Switch folioq to bvecq
  smbdirect: Remove support for ITER_FOLIOQ from
    smbdirect_map_sges_from_iter()
  iov_iter: Remove ITER_FOLIOQ
  netfs: Remove folio_queue

 Documentation/core-api/folio_queue.rst      | 209 ---------
 Documentation/core-api/index.rst            |   1 -
 Documentation/filesystems/netfs_library.rst |   2 +-
 fs/afs/dir.c                                |  35 +-
 fs/afs/dir_edit.c                           |  43 +-
 fs/afs/dir_search.c                         |  33 +-
 fs/afs/inode.c                              |   2 +-
 fs/afs/internal.h                           |   6 +-
 fs/afs/symlink.c                            |  37 +-
 fs/netfs/Makefile                           |   1 +
 fs/netfs/buffered_read.c                    |  34 +-
 fs/netfs/bvecq.c                            | 342 ++++++++++++++
 fs/netfs/internal.h                         |   7 +-
 fs/netfs/iterator.c                         |  33 +-
 fs/netfs/main.c                             |  13 +-
 fs/netfs/misc.c                             |  96 ----
 fs/netfs/read_collect.c                     |  47 +-
 fs/netfs/read_pgpriv2.c                     |  24 +-
 fs/netfs/read_retry.c                       |  28 +-
 fs/netfs/rolling_buffer.c                   | 189 +++-----
 fs/netfs/stats.c                            |   6 +-
 fs/netfs/write_collect.c                    |  26 +-
 fs/netfs/write_issue.c                      | 180 ++-----
 fs/smb/client/cifsglob.h                    |   2 +-
 fs/smb/client/smb2ops.c                     |  78 ++--
 fs/smb/smbdirect/connection.c               | 135 +++---
 include/linux/bvec.h                        |  18 +
 include/linux/bvecq.h                       | 166 +++++++
 include/linux/folio_queue.h                 | 282 -----------
 include/linux/iov_iter.h                    |  87 ++--
 include/linux/netfs.h                       |  14 +-
 include/linux/rolling_buffer.h              |  42 +-
 include/linux/uio.h                         |  17 +-
 include/trace/events/netfs.h                |  49 +-
 kernel/bpf/btf.c                            |   2 -
 lib/iov_iter.c                              | 494 +++++++++++++-------
 lib/scatterlist.c                           |  82 ++--
 lib/tests/kunit_iov_iter.c                  | 131 +++---
 38 files changed, 1437 insertions(+), 1556 deletions(-)
 delete mode 100644 Documentation/core-api/folio_queue.rst
 create mode 100644 fs/netfs/bvecq.c
 create mode 100644 include/linux/bvecq.h
 delete mode 100644 include/linux/folio_queue.h


^ permalink raw reply	[flat|nested] 12+ messages in thread

end of thread, other threads:[~2026-09-29  8:22 UTC | newest]

Thread overview: 12+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29  7:59 [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells
2026-09-29  7:59 ` [PATCH v12 01/10] Add a function to kmap one page of a multipage bio_vec David Howells
2026-09-29  7:59 ` [PATCH v12 02/10] iov_iter: Add a segmented queue of bio_vec[] David Howells
2026-09-29  7:59 ` [PATCH v12 03/10] netfs: Add some tools for managing bvecq chains David Howells
2026-09-29  7:59 ` [PATCH v12 04/10] afs: Use a bvecq to hold dir content rather than folioq David Howells
2026-09-29  7:59 ` [PATCH v12 05/10] cifs: Use a bvecq for buffering instead of a folioq David Howells
2026-09-29  7:59 ` [PATCH v12 06/10] smbdirect: Support ITER_BVECQ in smbdirect_map_sges_from_iter() David Howells
2026-09-29  7:59 ` [PATCH v12 07/10] netfs: Switch folioq to bvecq David Howells
2026-09-29  7:59 ` [PATCH v12 08/10] smbdirect: Remove support for ITER_FOLIOQ from smbdirect_map_sges_from_iter() David Howells
2026-09-29  7:59 ` [PATCH v12 09/10] iov_iter: Remove ITER_FOLIOQ David Howells
2026-09-29  7:59 ` [PATCH v12 10/10] netfs: Remove folio_queue David Howells
2026-09-29  8:21 ` [PATCH v12 00/10] netfs, cachefiles: Changes for next, primarily occupancy tracking-related David Howells

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®