From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1DABB3783D5 for ; Mon, 5 Oct 2026 07:13:24 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791184407; cv=none; b=KOyIR9yaTY1HKs454v0KQK2tI6wwuNOnvVNWchCU+5lx4C0rKRe68lh4epCTbYKFFhJ6+dJQV5kuwyG2QIQ6KjiiHC62tWaSYy8XPcjJKLBwSoIN6cPzJIPypcwIBPQ4D4ItnGJHamPm6M6pvccvR8Ya16kA97khB2+SWBxXncU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1791184407; c=relaxed/simple; bh=sKuRWcgfdX8V7JdPvPOaDytTIG1TX8c2snVDaw3p44U=; h=From:To:Cc:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version; b=uTPx3W31QyGieXvb+28iepc43z/ptgESzgnS75UZ7lrHtWTsTlJh0CIfxdwx0F0IeFNfpg6qDU4ihJvEAyFoYhLjsGfIh/tHxMggLc3SozdecqDZcjftvcJy/eo4m/LCkkl2erh1D0UnedwfFUHQBPGlSFwBlfr1DwOHmrXEfWE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=iBC4NjhK; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="iBC4NjhK" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1791184404; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=TPYXwowLJtPyU6pZVRwxzMoXAk/gZZDH9gxVcE6MDi0=; b=iBC4NjhK/LT5PC1+LOGMYHBgoJPv9WxSNWPxgJWD34xOryqmQyoYx3ChtN/tLJomIyNE4v OizAiOJiTRU3Dia2BRnvJE/9jzl9fj45dnvmmcQyHWLljA9Z5Qz9zotwacZf6M+tmkvynZ +9Ja4FFXCrf9tpWsF6EG6ycdOg4NIoE= Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-651-B2XHdbWDOr2kSIP3jeM8yw-1; Mon, 05 Oct 2026 03:13:17 -0400 X-MC-Unique: B2XHdbWDOr2kSIP3jeM8yw-1 X-Mimecast-MFC-AGG-ID: B2XHdbWDOr2kSIP3jeM8yw_1791184395 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id BFE46195FD31; Mon, 5 Oct 2026 07:13:14 +0000 (UTC) Received: from warthog.com (unknown [10.44.32.90]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id D967F19373D8; Mon, 5 Oct 2026 07:13:09 +0000 (UTC) From: David Howells To: Christian Brauner Cc: David Howells , Paulo Alcantara , Matthew Wilcox , Namjae Jeon , Marc Dionne , Stefan Metzmacher , Eric Van Hensbergen , Dominique Martinet , Ilya Dryomov , netfs@lists.linux.dev, linux-afs@lists.infradead.org, linux-cifs@vger.kernel.org, linux-nfs@vger.kernel.org, ceph-devel@vger.kernel.org, v9fs@lists.linux.dev, linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org Subject: [PATCH v12 4/8] netfs: Add some tools for managing a position in a bvecq chain Date: Mon, 5 Oct 2026 08:12:21 +0100 Message-ID: <20261005071227.147182-5-dhowells@redhat.com> In-Reply-To: <20261005071227.147182-1-dhowells@redhat.com> References: <20261005071227.147182-1-dhowells@redhat.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Provide a selection of tools for holding and managing a position within a bvec queue chain. Signed-off-by: David Howells Reviewed-by: Paulo Alcantara cc: Matthew Wilcox cc: netfs@lists.linux.dev cc: linux-fsdevel@vger.kernel.org --- fs/netfs/bvecq.c | 250 ++++++++++++++++++++++++++++++++++++++++++ include/linux/bvecq.h | 183 +++++++++++++++++++++++++++++++ 2 files changed, 433 insertions(+) diff --git a/fs/netfs/bvecq.c b/fs/netfs/bvecq.c index 5b747f6b5938..1493e028bb17 100644 --- a/fs/netfs/bvecq.c +++ b/fs/netfs/bvecq.c @@ -342,3 +342,253 @@ int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size, return 0; } EXPORT_SYMBOL(bvecq_expand_buffer); + +/** + * bvecq_buffer_init - Initialise a buffer and set position + * @pos: The position to point at the new buffer. + * @gfp: The allocation constraints. + * @for_writeback: True if allocating for writeback + * + * Initialise a rolling buffer. We allocate an unpopulated bvecq node to so + * that the pointers can be independently driven by the producer and the + * consumer. + * + * Return 0 if successful; -ENOMEM on allocation failure. + */ +int bvecq_buffer_init(struct bvecq_pos *pos, gfp_t gfp, bool for_writeback) +{ + struct bvecq *bq; + + bq = bvecq_alloc_one(BVECQ_POOL_SLOTS, gfp, for_writeback); + if (!bq) + return -ENOMEM; + + pos->bvecq = bq; /* Comes with a ref. */ + pos->slot = 0; + pos->offset = 0; + return 0; +} + +/** + * bvecq_buffer_append - Append a new bvecq node to a buffer + * @pos: The position of the last node. + * @bq: The buffer to add. + * + * Add a new node on to the buffer chain at the specified position, either + * because the previous one is full or because we have a discontiguity to + * contend with, and update @pos to point to it. + */ +void bvecq_buffer_append(struct bvecq_pos *pos, struct bvecq *bq) +{ + struct bvecq *head = pos->bvecq; + + pos->bvecq = bvecq_get(bq); + pos->slot = 0; + pos->offset = 0; + + /* [!] NOTE: After we set head->next, the consumer is at liberty to + * immediately delete the old head. + */ + bvecq_append(head, bq); + bvecq_put(head); +} + +/** + * bvecq_pos_advance - Advance a bvecq position + * @pos: The position to advance. + * @amount: The amount of bytes to advance by. + * + * Advance the specified bvecq position by @amount bytes. @pos is updated and + * bvecq ref counts may have been manipulated. If the position hits the end of + * the queue, then it is left pointing beyond the last slot of the last bvecq + * so that it doesn't break the chain. + */ +void bvecq_pos_advance(struct bvecq_pos *pos, size_t amount) +{ + struct bvecq *bq = pos->bvecq, *next; + unsigned int slot = pos->slot; + size_t offset = pos->offset; + + while (amount) { + size_t part; + + if (!bvecq_acquire_slot(bq, slot)) { + next = bvecq_next(bq); + if (!next) { + WARN_ON_ONCE(amount > 0); + break; + } + if (bvecq_acquire_slot(bq, slot)) + continue; /* More slots got added. */ + bq = next; + slot = 0; + offset = 0; + continue; + } + + part = bq->bv[slot].bv_len - offset; + + if (part > amount) { + offset += amount; + break; + } + amount -= part; + offset = 0; + slot++; + } + + pos->slot = slot; + pos->offset = offset; + bvecq_pos_move(pos, bq); +} + +/* + * Clear part of the memory pointed to by a bio_vec. + */ +static void bvec_zero(const struct bio_vec *bv, size_t offset, size_t len) +{ + struct page *page = bv->bv_page; + + offset += bv->bv_offset; + + page += offset / PAGE_SIZE; + offset = offset % PAGE_SIZE; + + while (len) { + size_t part = min(len, PAGE_SIZE - offset); + char *p = kmap_local_page(page); + + memset(p + offset, 0, part); + kunmap_local(p); + + len -= part; + offset = 0; + page++; + } +} + +/** + * bvecq_zero - Clear memory starting at the bvecq position. + * @pos: The position in the bvecq chain to start clearing. + * @amount: The number of bytes to clear. + * + * Clear memory fragments pointed to by a bvec queue. @pos is updated and + * bvecq ref counts may have been manipulated. If the position hits the end of + * the queue, then it is left pointing beyond the last slot of the last bvecq + * so that it doesn't break the chain. + * + * Return: The number of bytes cleared. + */ +ssize_t bvecq_zero(struct bvecq_pos *pos, size_t amount) +{ + struct bvecq *bq = pos->bvecq, *next; + unsigned int slot = pos->slot; + ssize_t cleared = 0; + size_t offset = pos->offset; + + while (amount) { + const struct bio_vec *bv; + size_t part; + + if (!bvecq_acquire_slot(bq, slot)) { + next = bvecq_next(bq); + if (!next) { + WARN_ON_ONCE(amount > 0); + break; + } + if (bvecq_acquire_slot(bq, slot)) + continue; /* More slots got added. */ + bq = next; + slot = 0; + offset = 0; + continue; + } + + bv = &bq->bv[slot]; + if (offset >= bv->bv_len) { + slot++; + offset = 0; + continue; + } + + part = min(bv->bv_len - offset, amount); + bvec_zero(bv, offset, part); + cleared += part; + offset += part; + amount -= part; + } + + pos->slot = slot; + pos->offset = offset; + bvecq_pos_move(pos, bq); + return cleared; +} + +/** + * bvecq_slice - Find a slice of a bvecq queue + * @pos: The position to start at. + * @max_size: The maximum size of the slice (or ULONG_MAX). + * @max_slots: The maximum number of slots in the slice (or INT_MAX). + * @_nr_slots: Where to put the number of slots (updated). + * + * Determine the size and number of slots that can be obtained the next slice + * of bvec queue up to the maximum size and slot count specified. + * + * @pos is updated to the end of the slice. If the position hits the end of + * the queue, then it is left pointing beyond the last slot of the last bvecq + * so that it doesn't break the chain. + * + * Return: The number of bytes in the slice. + */ +size_t bvecq_slice(struct bvecq_pos *pos, size_t max_size, + unsigned int max_slots, unsigned int *_nr_slots) +{ + struct bvecq *bq, *next; + unsigned int slot = pos->slot, nslots = 0; + size_t size = 0, offset = pos->offset; + + bq = pos->bvecq; + for (;;) { + for (; slot < bvecq_nr_slots_acquire(bq); slot++) { + const struct bio_vec *bvec = &bq->bv[slot]; + + if (offset < bvec->bv_len && bvec->bv_page) { + size_t part = min(bvec->bv_len - offset, max_size); + + size += part; + offset += part; + max_size -= part; + nslots++; + if (!max_size || nslots >= max_slots) + goto out; + } + offset = 0; + } + + /* pos->bvecq isn't allowed to go NULL as the queue may get + * extended and we would lose our place. + */ + next = bvecq_next(bq); + if (!next) + break; + if (bvecq_acquire_slot(bq, slot)) + continue; /* More slots got added. */ + slot = 0; + bq = next; + } + +out: + *_nr_slots = nslots; + if (slot == bvecq_nr_slots_acquire(bq)) { + next = bvecq_next(bq); + if (next) { + bq = next; + slot = 0; + offset = 0; + } + } + bvecq_pos_move(pos, bq); + pos->slot = slot; + pos->offset = offset; + return size; +} diff --git a/include/linux/bvecq.h b/include/linux/bvecq.h index b984aaa44908..f1513f507cfb 100644 --- a/include/linux/bvecq.h +++ b/include/linux/bvecq.h @@ -54,6 +54,16 @@ struct bvecq { /* Number of slots in a 4K bvecq. */ #define BVECQ_4KB_SLOTS ((4096 - sizeof(struct bvecq)) / sizeof(struct bio_vec)) +/* + * Position in a bio_vec queue. The bvecq holds a ref on the queue segment it + * points to. + */ +struct bvecq_pos { + struct bvecq *bvecq; /* The first bvecq */ + unsigned int offset; /* The offset within the starting slot */ + u16 slot; /* The starting slot */ +}; + void bvecq_dump(const struct bvecq *bq); struct bvecq *bvecq_alloc_one(size_t nr_slots, gfp_t gfp, bool for_writeback); struct bvecq *bvecq_alloc_chain(size_t nr_slots, gfp_t gfp, bool for_writeback); @@ -61,6 +71,12 @@ struct bvecq *bvecq_alloc_buffer2(size_t size, unsigned int pre_slots, gfp_t gfp bool for_writeback); void bvecq_put(struct bvecq *bq); int bvecq_expand_buffer(struct bvecq **_buffer, size_t *_cur_size, size_t size, gfp_t gfp); +int bvecq_buffer_init(struct bvecq_pos *pos, gfp_t gfp, bool for_writeback); +void bvecq_buffer_append(struct bvecq_pos *pos, struct bvecq *bq); +void bvecq_pos_advance(struct bvecq_pos *pos, size_t amount); +ssize_t bvecq_zero(struct bvecq_pos *pos, size_t amount); +size_t bvecq_slice(struct bvecq_pos *pos, size_t max_size, + unsigned int max_slots, unsigned int *_nr_slots); /** * bvecq_alloc_buffer - Allocate a bvecq chain and populate with buffers @@ -163,4 +179,171 @@ static inline struct bvecq *bvecq_next(const struct bvecq *bq) return smp_load_acquire(&bq->next); } +/** + * bvecq_pos_set - Set one position to be the same as another + * @pos: The position object to set + * @at: The source position. + * + * Set @pos to have the same position as @at. This may take a ref on the + * bvecq pointed to. + */ +static inline void bvecq_pos_set(struct bvecq_pos *pos, const struct bvecq_pos *at) +{ + *pos = *at; + bvecq_get(pos->bvecq); +} + +/** + * bvecq_pos_unset - Unset a position + * @pos: The position object to unset + * + * Unset @pos. This does any needed ref cleanup. + */ +static inline void bvecq_pos_unset(struct bvecq_pos *pos) +{ + bvecq_put(pos->bvecq); + pos->bvecq = NULL; + pos->slot = 0; + pos->offset = 0; +} + +/** + * bvecq_pos_transfer - Transfer one position to another, clearing the first + * @pos: The position object to set + * @from: The source position to clear. + * + * Set @pos to have the same position as @from and then clear @from. This may + * transfer a ref on the bvecq pointed to. + */ +static inline void bvecq_pos_transfer(struct bvecq_pos *pos, struct bvecq_pos *from) +{ + *pos = *from; + from->bvecq = NULL; + from->slot = 0; + from->offset = 0; +} + +/** + * bvecq_pos_move - Update a position to a new bvecq + * @pos: The position object to update. + * @to: The new bvecq to point at. + * + * Update @pos to point to @to if it doesn't already do so. This may + * manipulate refs on the bvecqs pointed to. + */ +static inline void bvecq_pos_move(struct bvecq_pos *pos, struct bvecq *to) +{ + struct bvecq *old = pos->bvecq; + + if (old != to) { + pos->bvecq = bvecq_get(to); + bvecq_put(old); + } +} + +/** + * bvecq_pos_nudge - Nudge a position onto the next segment if current used up + * @pos: The position object to nudge. + * + * Update @pos to point to the next segment in the chain if we've used up the + * current segment. This may manipulate refs on the bvecqs pointed to. + * + * Return: true if found a new segment, false if hit the end. + */ +static inline bool bvecq_pos_nudge(struct bvecq_pos *pos) +{ + struct bvecq *bq = pos->bvecq; + + for (;;) { + if (!bvecq_acquire_slot(bq, pos->slot)) { + bq = bvecq_next(bq); + if (!bq) + return false; + if (bvecq_acquire_slot(bq, pos->slot)) + continue; /* More slots got added. */ + bvecq_pos_move(pos, bq); + pos->slot = 0; + pos->offset = 0; + continue; + } + if (pos->offset >= bq->bv[pos->slot].bv_len) { + pos->slot++; + pos->offset = 0; + continue; + } + return true; + } +} + +/** + * bvecq_pos_step - Step a position to the next slot if possible + * @pos: The position object to step. + * + * Update @pos to point to the next slot in the queue if not at the end. This + * may manipulate refs on the bvecqs pointed to. + * + * Return: true if successful, false if was at the end. + */ +static inline bool bvecq_pos_step(struct bvecq_pos *pos) +{ + struct bvecq *bq = pos->bvecq, *next; + + pos->slot++; + pos->offset = 0; + if (bvecq_acquire_slot(bq, pos->slot)) + return true; + next = bvecq_next(bq); + if (!next) + return false; + if (bvecq_acquire_slot(bq, pos->slot)) + return true; + bvecq_pos_move(pos, next); + pos->slot = 0; + return true; +} + +/** + * bvecq_delete_spent - Delete the bvecq at the front if possible + * @pos: The position object to update. + * + * Delete the used up bvecq at the front of the queue that @pos points to if it + * is not the last node in the queue; if it is the last node in the queue, it + * is kept so that the queue doesn't become detached from the other end. This + * may manipulate refs on the bvecqs pointed to. It is also possible that the + * producer will fill more slots in the current bvecq. + * + * Also, we have to be very careful: the consumer can catch the producer, which + * could lead to us having nothing left in the queue, causing the front and + * back pointers to end up on different tracks. To avoid this, we must always + * keep at least one segment in the queue. + * + * The caller must reload from @pos after calling this. + * + * Return: true if there's more available; false if not. + */ +static inline bool bvecq_delete_spent(struct bvecq_pos *pos) +{ + struct bvecq *spent = pos->bvecq; + struct bvecq *next; + unsigned int slot = pos->slot; + +again: + /* Read the contents of the queue node after the pointer to it. */ + next = bvecq_next(spent); + if (!next) + return false; /* Nothing more to consume at the moment. */ + if (slot < bvecq_nr_slots_acquire(spent)) + return true; /* The producer added more. */ + next->prev = NULL; + bvecq_pos_move(pos, next); + pos->slot = 0; + pos->offset = 0; + if (!bvecq_acquire_slot(next, 0)) { + spent = next; + slot = 0; + goto again; + } + return true; +} + #endif /* _LINUX_BVECQ_H */