From: Andrea Righi <arighi@nvidia.com>
To: Alexei Starovoitov <ast@kernel.org>,
Daniel Borkmann <daniel@iogearbox.net>,
Andrii Nakryiko <andrii@kernel.org>,
Eduard Zingerman <eddyz87@gmail.com>,
Kumar Kartikeya Dwivedi <memxor@gmail.com>
Cc: Martin KaFai Lau <martin.lau@linux.dev>,
Song Liu <song@kernel.org>,
Yonghong Song <yonghong.song@linux.dev>,
Jiri Olsa <jolsa@kernel.org>,
Emil Tsalapatis <emil@etsalapatis.com>,
Ihor Solodrai <ihor.solodrai@linux.dev>,
John Fastabend <john.fastabend@gmail.com>,
Josh Don <joshdon@google.com>,
Vineeth Pillai <vineethrp@google.com>,
Balbir Singh <balbirs@nvidia.com>,
Shameer Kolothum <skolothumtho@nvidia.com>,
Himadri Chhaya-Shailesh <himadrics@protonmail.com>,
NchangRoy <royfru44@gmail.com>,
bpf@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: [PATCH 2/5] bpf: Expose pinned arenas as sized bpffs files
Date: Sun, 11 Oct 2026 00:45:57 +0200 [thread overview]
Message-ID: <20261010224902.1232178-3-arighi@nvidia.com> (raw)
In-Reply-To: <20261010224902.1232178-1-arighi@nvidia.com>
Allow userspace consumers to access pinned BPF arena memory through
ordinary file-backed mappings. This enables shared-memory communication
with consumers that require a sized file, such as QEMU guest RAM
backends.
Add BPF_F_ARENA_EXPORT to expose pinned arenas as sized bpffs files that
userspace can open and mmap. Keep the map FD as the canonical BPF
mapping, while exported mappings can use their own virtual addresses and
bounded file offsets and lengths. Resolve faults by file offset and zap
only the overlap with each exported VMA, allowing consumers to map
disjoint slices of one arena.
Require full-capacity canonical mappings for exported arenas so the
initialized address range matches the size reported by the pinned file.
Keep shorter canonical mappings available for arenas without export.
Validate exported mappings against the map capacity and canonical range
before resolving faults. Keep the pin's size fixed while allowing
permission and ownership updates.
Allow read-only shared mappings through O_RDONLY pin FDs by restoring
VM_SHARED when generic mmap retains VM_MAYSHARE, without granting
VM_MAYWRITE. Keep private mappings rejected and preserve existing BPF
map and LSM checks, allowing pin permissions to distinguish readers from
writers.
Require BPF_F_ARENA_NO_FREE with BPF_F_ARENA_EXPORT, so BPF programs
cannot free pages while external consumers may still use them.
Retention-only arenas keep their ordinary pin behavior. File mappings
and map references keep the arena and its pages alive until the last
reference is released.
Initialize the pin's capacity before publishing its inode, and support
bounded SEEK_SET, SEEK_CUR, and SEEK_END operations. Use the canonical
range end to distinguish initialization from a valid mapping at address
zero.
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
include/uapi/linux/bpf.h | 3 ++
kernel/bpf/arena.c | 68 +++++++++++++++++++-----
kernel/bpf/inode.c | 95 +++++++++++++++++++++++++++++++---
tools/include/uapi/linux/bpf.h | 3 ++
4 files changed, 150 insertions(+), 19 deletions(-)
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 75101d602bd32..40a5489ac9fea 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1503,6 +1503,9 @@ enum {
/* Keep arena pages allocated until the map is destroyed. */
BPF_F_ARENA_NO_FREE = (1U << 20),
+
+ /* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */
+ BPF_F_ARENA_EXPORT = (1U << 21),
};
/* Flags for BPF_PROG_QUERY. */
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 343e84139ec29..98f6dffe0138f 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -284,7 +284,12 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
!(attr->map_flags & BPF_F_MMAPABLE) ||
/* No unsupported flags present */
(attr->map_flags & ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE |
- BPF_F_NO_USER_CONV | BPF_F_ARENA_NO_FREE)))
+ BPF_F_NO_USER_CONV | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT)))
+ return ERR_PTR(-EINVAL);
+
+ /* Exported memory must remain allocated while consumers use it. */
+ if ((attr->map_flags & BPF_F_ARENA_EXPORT) && !(attr->map_flags & BPF_F_ARENA_NO_FREE))
return ERR_PTR(-EINVAL);
if (attr->map_extra & ~PAGE_MASK)
@@ -496,7 +501,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
int ret;
kbase = bpf_arena_get_kern_vm_start(arena);
- kaddr = kbase + (u32)(vmf->address);
+ /* vmf->pgoff includes the file offset of a bpffs-backed slice. */
+ kaddr = kbase + (u32)(arena->user_vm_start + ((u64)vmf->pgoff << PAGE_SHIFT));
page = vmalloc_to_page((void *)kaddr);
if (!page && !(arena->map.map_flags & BPF_F_SEGV_ON_FAULT)) {
@@ -613,13 +619,23 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
long ret;
+ if (filp->f_op != &bpf_map_fops) {
+ if (!len || pgoff >= map->max_entries ||
+ len > ((u64)map->max_entries - pgoff) << PAGE_SHIFT)
+ return -EINVAL;
+ return mm_get_unmapped_area(filp, addr, len, pgoff, flags);
+ }
+
if (pgoff)
return -EINVAL;
if (len > SZ_4G)
return -E2BIG;
+ if ((map->map_flags & BPF_F_ARENA_EXPORT) &&
+ len != (u64)map->max_entries << PAGE_SHIFT)
+ return -EINVAL;
- /* if user_vm_start was specified at arena creation time */
- if (arena->user_vm_start) {
+ /* Once established, the canonical range cannot change. */
+ if (arena->user_vm_end) {
if (len > arena->user_vm_end - arena->user_vm_start)
return -E2BIG;
if (len != arena->user_vm_end - arena->user_vm_start)
@@ -633,7 +649,7 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
return ret;
if ((ret >> 32) == ((ret + len - 1) >> 32))
return ret;
- if (WARN_ON_ONCE(arena->user_vm_start))
+ if (WARN_ON_ONCE(arena->user_vm_end))
/* checks at map creation time should prevent this */
return -EFAULT;
return round_up(ret, SZ_4G);
@@ -642,9 +658,13 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
{
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
+ bool exported = vma->vm_file->f_op != &bpf_map_fops;
guard(mutex)(&arena->lock);
- if (arena->user_vm_start && arena->user_vm_start != vma->vm_start)
+ if (!exported && (map->map_flags & BPF_F_ARENA_EXPORT) &&
+ vma->vm_end - vma->vm_start != (u64)map->max_entries << PAGE_SHIFT)
+ return -EINVAL;
+ if (!exported && arena->user_vm_end && arena->user_vm_start != vma->vm_start)
/*
* If map_extra was not specified at arena creation time then
* 1st user process can do mmap(NULL, ...) to pick user_vm_start
@@ -655,19 +675,31 @@ static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
*/
return -EBUSY;
- if (arena->user_vm_end && arena->user_vm_end != vma->vm_end)
+ if (exported && !arena->user_vm_end)
+ return -EINVAL;
+ if (!exported && arena->user_vm_end && arena->user_vm_end != vma->vm_end)
/* all user processes must have the same size of mmap-ed region */
return -EBUSY;
- /* Earlier checks should prevent this */
- if (WARN_ON_ONCE(vma->vm_end - vma->vm_start > SZ_4G || vma->vm_pgoff))
+ if (exported) {
+ u64 page_cnt = (arena->user_vm_end - arena->user_vm_start) >> PAGE_SHIFT;
+
+ if (vma->vm_pgoff >= page_cnt ||
+ (vma->vm_end - vma->vm_start) >> PAGE_SHIFT >
+ page_cnt - vma->vm_pgoff)
+ return -EINVAL;
+ } else if (WARN_ON_ONCE(vma->vm_end - vma->vm_start > SZ_4G || vma->vm_pgoff)) {
+ /* Earlier checks should prevent this for map FD mappings. */
return -EFAULT;
+ }
if (remember_vma(arena, vma))
return -ENOMEM;
- arena->user_vm_start = vma->vm_start;
- arena->user_vm_end = vma->vm_end;
+ if (!exported) {
+ arena->user_vm_start = vma->vm_start;
+ arena->user_vm_end = vma->vm_end;
+ }
/*
* bpf_map_mmap() checks that it's being mmaped as VM_SHARED and
* clears VM_MAYEXEC. Set VM_DONTEXPAND to avoid potential change
@@ -841,8 +873,12 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
struct mm_struct *mm;
struct vma_list *vml;
unsigned long vm_start;
+ u64 start, end, vma_start, vma_end;
u64 my_gen;
+ start = uaddr - arena->user_vm_start;
+ end = start + size;
+
/*
* Taking mmap_read_lock() under arena->lock would deadlock against
* arena_vm_close(), which runs with mmap_write_lock held and then
@@ -881,8 +917,14 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
*/
vma = find_vma(mm, vm_start);
if (vma && vma->vm_start == vm_start &&
- vma->vm_file && vma->vm_file->private_data == &arena->map)
- zap_vma_range(vma, uaddr, size);
+ vma->vm_file && vma->vm_file->private_data == &arena->map) {
+ vma_start = (u64)vma->vm_pgoff << PAGE_SHIFT;
+ vma_end = vma_start + vma->vm_end - vma->vm_start;
+ if (start < vma_end && end > vma_start)
+ zap_vma_range(vma, vma->vm_start +
+ (max(start, vma_start) - vma_start),
+ min(end, vma_end) - max(start, vma_start));
+ }
mmap_read_unlock(mm);
mmput(mm);
diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c
index 7837968c0842c..d3dc6f70e32ae 100644
--- a/kernel/bpf/inode.c
+++ b/kernel/bpf/inode.c
@@ -119,8 +119,24 @@ static const struct inode_operations bpf_symlink_iops;
static const struct inode_operations bpf_prog_iops = {
.listxattr = bpf_fs_listxattr,
};
+
+static int bpf_map_setattr(struct mnt_idmap *idmap, struct dentry *dentry,
+ struct iattr *attr)
+{
+ struct inode *inode = d_inode(dentry);
+ struct bpf_map *map = inode->i_private;
+
+ if (map->map_type == BPF_MAP_TYPE_ARENA &&
+ (map->map_flags & BPF_F_ARENA_EXPORT) &&
+ (attr->ia_valid & ATTR_SIZE) && attr->ia_size != i_size_read(inode))
+ return -EINVAL;
+
+ return simple_setattr(idmap, dentry, attr);
+}
+
static const struct inode_operations bpf_map_iops = {
.listxattr = bpf_fs_listxattr,
+ .setattr = bpf_map_setattr,
};
static const struct inode_operations bpf_link_iops = {
.listxattr = bpf_fs_listxattr,
@@ -351,6 +367,68 @@ static const struct file_operations bpffs_map_fops = {
.release = bpffs_map_release,
};
+/*
+ * An arena with BPF_F_ARENA_EXPORT pinned in bpffs can serve as a
+ * shared-memory file. The ordinary map FD is an anonymous inode with no
+ * size and cannot be opened by pathname, so keep the pin's inode and
+ * forward mmap to the map FD implementation. The pinned file uses ordinary
+ * address selection so QEMU can map the pages at another virtual address.
+ */
+static int bpffs_arena_open(struct inode *inode, struct file *file)
+{
+ struct bpf_map *map = inode->i_private;
+ int err;
+
+ err = security_bpf_map(map, file->f_mode);
+ if (err)
+ return err;
+ bpf_map_inc_with_uref(map);
+ file->private_data = map;
+
+ return 0;
+}
+
+static int bpffs_arena_mmap(struct file *file, struct vm_area_struct *vma)
+{
+ /*
+ * Generic mmap clears VM_SHARED for O_RDONLY files, but VM_MAYSHARE
+ * still distinguishes shared mappings from private ones. Restore
+ * VM_SHARED for the map mmap path without granting VM_MAYWRITE.
+ */
+ if (!(file->f_mode & FMODE_WRITE) && (vma->vm_flags & VM_MAYSHARE))
+ vm_flags_set(vma, VM_SHARED);
+
+ return bpf_map_fops.mmap(file, vma);
+}
+
+static int bpffs_arena_release(struct inode *inode, struct file *file)
+{
+ return bpf_map_fops.release(inode, file);
+}
+
+static unsigned long bpffs_arena_get_unmapped_area(struct file *file,
+ unsigned long addr,
+ unsigned long len,
+ unsigned long pgoff,
+ unsigned long flags)
+{
+ return bpf_map_fops.get_unmapped_area(file, addr, len, pgoff, flags);
+}
+
+static loff_t bpffs_arena_llseek(struct file *file, loff_t offset, int whence)
+{
+ return fixed_size_llseek(file, offset, whence, i_size_read(file_inode(file)));
+}
+
+static const struct file_operations bpffs_arena_fops = {
+ .open = bpffs_arena_open,
+ .llseek = bpffs_arena_llseek,
+ .fsync = noop_fsync,
+ .release = bpffs_arena_release,
+ .mmap = bpffs_arena_mmap,
+ .get_unmapped_area = bpffs_arena_get_unmapped_area,
+};
+
static int bpffs_obj_open(struct inode *inode, struct file *file)
{
return -EIO;
@@ -362,7 +440,7 @@ static const struct file_operations bpffs_obj_fops = {
static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
const struct inode_operations *iops,
- const struct file_operations *fops)
+ const struct file_operations *fops, loff_t size)
{
struct inode *dir = dentry->d_parent->d_inode;
struct inode *inode;
@@ -382,6 +460,7 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
inode->i_op = iops;
inode->i_fop = fops;
inode->i_private = raw;
+ i_size_write(inode, size);
bpf_dentry_finalize(dentry, inode, dir);
return 0;
@@ -390,16 +469,20 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
static int bpf_mkprog(struct dentry *dentry, umode_t mode, void *arg)
{
return bpf_mkobj_ops(dentry, mode, arg, &bpf_prog_iops,
- &bpffs_obj_fops);
+ &bpffs_obj_fops, 0);
}
static int bpf_mkmap(struct dentry *dentry, umode_t mode, void *arg)
{
struct bpf_map *map = arg;
+ bool shared_arena = map->map_type == BPF_MAP_TYPE_ARENA &&
+ (map->map_flags & BPF_F_ARENA_EXPORT);
return bpf_mkobj_ops(dentry, mode, arg, &bpf_map_iops,
- bpf_map_support_seq_show(map) ?
- &bpffs_map_fops : &bpffs_obj_fops);
+ shared_arena ? &bpffs_arena_fops :
+ bpf_map_support_seq_show(map) ?
+ &bpffs_map_fops : &bpffs_obj_fops,
+ shared_arena ? (loff_t)map->max_entries * PAGE_SIZE : 0);
}
static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)
@@ -408,7 +491,7 @@ static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)
return bpf_mkobj_ops(dentry, mode, arg, &bpf_link_iops,
bpf_link_is_iter(link) ?
- &bpf_iter_fops : &bpffs_obj_fops);
+ &bpf_iter_fops : &bpffs_obj_fops, 0);
}
static struct dentry *
@@ -483,7 +566,7 @@ static int bpf_iter_link_pin_kernel(struct dentry *parent,
if (IS_ERR(dentry))
return PTR_ERR(dentry);
ret = bpf_mkobj_ops(dentry, mode, link, &bpf_link_iops,
- &bpf_iter_fops);
+ &bpf_iter_fops, 0);
simple_done_creating(dentry);
return ret;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 75101d602bd32..40a5489ac9fea 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1503,6 +1503,9 @@ enum {
/* Keep arena pages allocated until the map is destroyed. */
BPF_F_ARENA_NO_FREE = (1U << 20),
+
+ /* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */
+ BPF_F_ARENA_EXPORT = (1U << 21),
};
/* Flags for BPF_PROG_QUERY. */
--
2.56.0
next prev parent reply other threads:[~2026-10-10 22:49 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-10 22:45 [PATCH bpf-next 0/5] " Andrea Righi
2026-10-10 22:45 ` [PATCH 1/5] bpf: Add an arena flag to retain allocated pages Andrea Righi
2026-10-10 22:45 ` Andrea Righi [this message]
2026-10-10 23:40 ` [PATCH 2/5] bpf: Expose pinned arenas as sized bpffs files bot+bpf-ci
2026-10-10 22:45 ` [PATCH 3/5] docs/bpf: Document pinned arena file mappings Andrea Righi
2026-10-10 23:19 ` bot+bpf-ci
2026-10-10 22:45 ` [PATCH 4/5] selftests/bpf: Test " Andrea Righi
2026-10-10 23:40 ` bot+bpf-ci
2026-10-10 22:46 ` [PATCH 5/5] selftests/bpf: Test two-way arena access with two guests Andrea Righi
2026-10-10 23:40 ` bot+bpf-ci
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20261010224902.1232178-3-arighi@nvidia.com \
--to=arighi@nvidia.com \
--cc=andrii@kernel.org \
--cc=ast@kernel.org \
--cc=balbirs@nvidia.com \
--cc=bpf@vger.kernel.org \
--cc=daniel@iogearbox.net \
--cc=eddyz87@gmail.com \
--cc=emil@etsalapatis.com \
--cc=himadrics@protonmail.com \
--cc=ihor.solodrai@linux.dev \
--cc=john.fastabend@gmail.com \
--cc=jolsa@kernel.org \
--cc=joshdon@google.com \
--cc=linux-kernel@vger.kernel.org \
--cc=martin.lau@linux.dev \
--cc=memxor@gmail.com \
--cc=royfru44@gmail.com \
--cc=skolothumtho@nvidia.com \
--cc=song@kernel.org \
--cc=vineethrp@google.com \
--cc=yonghong.song@linux.dev \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®