* [PATCH 1/5] bpf: Add an arena flag to retain allocated pages
2026-10-10 22:45 [PATCH bpf-next 0/5] bpf: Expose pinned arenas as sized bpffs files Andrea Righi
@ 2026-10-10 22:45 ` Andrea Righi
2026-10-10 22:45 ` [PATCH 2/5] bpf: Expose pinned arenas as sized bpffs files Andrea Righi
` (3 subsequent siblings)
4 siblings, 0 replies; 10+ messages in thread
From: Andrea Righi @ 2026-10-10 22:45 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
Emil Tsalapatis, Ihor Solodrai, John Fastabend, Josh Don,
Vineeth Pillai, Balbir Singh, Shameer Kolothum,
Himadri Chhaya-Shailesh, NchangRoy, bpf, linux-kernel
Add BPF_F_ARENA_NO_FREE to retain arena pages until the map is
destroyed.
Make BPF page-free operations no-ops in both sleepable and non-sleepable
programs when the flag is set. Programs can manage reusable objects
within the retained pages without releasing their backing memory.
This provides a page-retention policy independently of how userspace
accesses the arena. It also supplies stable backing for external
consumers that must not lose pages while they are still using them.
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
include/uapi/linux/bpf.h | 3 +++
kernel/bpf/arena.c | 9 ++++++---
tools/include/uapi/linux/bpf.h | 3 +++
3 files changed, 12 insertions(+), 3 deletions(-)
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index e0ed44b1bbcb7..75101d602bd32 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1500,6 +1500,9 @@ enum {
/* Enable BPF ringbuf overwrite mode */
BPF_F_RB_OVERWRITE = (1U << 19),
+
+ /* Keep arena pages allocated until the map is destroyed. */
+ BPF_F_ARENA_NO_FREE = (1U << 20),
};
/* Flags for BPF_PROG_QUERY. */
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 0ff707707da04..343e84139ec29 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -283,7 +283,8 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
/* BPF_F_MMAPABLE must be set */
!(attr->map_flags & BPF_F_MMAPABLE) ||
/* No unsupported flags present */
- (attr->map_flags & ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE | BPF_F_NO_USER_CONV)))
+ (attr->map_flags & ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE |
+ BPF_F_NO_USER_CONV | BPF_F_ARENA_NO_FREE)))
return ERR_PTR(-EINVAL);
if (attr->map_extra & ~PAGE_MASK)
@@ -1134,7 +1135,8 @@ __bpf_kfunc void bpf_arena_free_pages(void *p__map, void *ptr__ign, u32 page_cnt
struct bpf_map *map = p__map;
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
- if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign)
+ if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign ||
+ (map->map_flags & BPF_F_ARENA_NO_FREE))
return;
arena_free_pages(arena, (long)ptr__ign, page_cnt, true);
}
@@ -1144,7 +1146,8 @@ void bpf_arena_free_pages_non_sleepable(void *p__map, void *ptr__ign, u32 page_c
struct bpf_map *map = p__map;
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
- if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign)
+ if (map->map_type != BPF_MAP_TYPE_ARENA || !page_cnt || !ptr__ign ||
+ (map->map_flags & BPF_F_ARENA_NO_FREE))
return;
arena_free_pages(arena, (long)ptr__ign, page_cnt, false);
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index e0ed44b1bbcb7..75101d602bd32 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1500,6 +1500,9 @@ enum {
/* Enable BPF ringbuf overwrite mode */
BPF_F_RB_OVERWRITE = (1U << 19),
+
+ /* Keep arena pages allocated until the map is destroyed. */
+ BPF_F_ARENA_NO_FREE = (1U << 20),
};
/* Flags for BPF_PROG_QUERY. */
--
2.56.0
^ permalink raw reply [flat|nested] 10+ messages in thread* [PATCH 2/5] bpf: Expose pinned arenas as sized bpffs files
2026-10-10 22:45 [PATCH bpf-next 0/5] bpf: Expose pinned arenas as sized bpffs files Andrea Righi
2026-10-10 22:45 ` [PATCH 1/5] bpf: Add an arena flag to retain allocated pages Andrea Righi
@ 2026-10-10 22:45 ` Andrea Righi
2026-10-10 23:40 ` bot+bpf-ci
2026-10-10 22:45 ` [PATCH 3/5] docs/bpf: Document pinned arena file mappings Andrea Righi
` (2 subsequent siblings)
4 siblings, 1 reply; 10+ messages in thread
From: Andrea Righi @ 2026-10-10 22:45 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
Emil Tsalapatis, Ihor Solodrai, John Fastabend, Josh Don,
Vineeth Pillai, Balbir Singh, Shameer Kolothum,
Himadri Chhaya-Shailesh, NchangRoy, bpf, linux-kernel
Allow userspace consumers to access pinned BPF arena memory through
ordinary file-backed mappings. This enables shared-memory communication
with consumers that require a sized file, such as QEMU guest RAM
backends.
Add BPF_F_ARENA_EXPORT to expose pinned arenas as sized bpffs files that
userspace can open and mmap. Keep the map FD as the canonical BPF
mapping, while exported mappings can use their own virtual addresses and
bounded file offsets and lengths. Resolve faults by file offset and zap
only the overlap with each exported VMA, allowing consumers to map
disjoint slices of one arena.
Require full-capacity canonical mappings for exported arenas so the
initialized address range matches the size reported by the pinned file.
Keep shorter canonical mappings available for arenas without export.
Validate exported mappings against the map capacity and canonical range
before resolving faults. Keep the pin's size fixed while allowing
permission and ownership updates.
Allow read-only shared mappings through O_RDONLY pin FDs by restoring
VM_SHARED when generic mmap retains VM_MAYSHARE, without granting
VM_MAYWRITE. Keep private mappings rejected and preserve existing BPF
map and LSM checks, allowing pin permissions to distinguish readers from
writers.
Require BPF_F_ARENA_NO_FREE with BPF_F_ARENA_EXPORT, so BPF programs
cannot free pages while external consumers may still use them.
Retention-only arenas keep their ordinary pin behavior. File mappings
and map references keep the arena and its pages alive until the last
reference is released.
Initialize the pin's capacity before publishing its inode, and support
bounded SEEK_SET, SEEK_CUR, and SEEK_END operations. Use the canonical
range end to distinguish initialization from a valid mapping at address
zero.
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
include/uapi/linux/bpf.h | 3 ++
kernel/bpf/arena.c | 68 +++++++++++++++++++-----
kernel/bpf/inode.c | 95 +++++++++++++++++++++++++++++++---
tools/include/uapi/linux/bpf.h | 3 ++
4 files changed, 150 insertions(+), 19 deletions(-)
diff --git a/include/uapi/linux/bpf.h b/include/uapi/linux/bpf.h
index 75101d602bd32..40a5489ac9fea 100644
--- a/include/uapi/linux/bpf.h
+++ b/include/uapi/linux/bpf.h
@@ -1503,6 +1503,9 @@ enum {
/* Keep arena pages allocated until the map is destroyed. */
BPF_F_ARENA_NO_FREE = (1U << 20),
+
+ /* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */
+ BPF_F_ARENA_EXPORT = (1U << 21),
};
/* Flags for BPF_PROG_QUERY. */
diff --git a/kernel/bpf/arena.c b/kernel/bpf/arena.c
index 343e84139ec29..98f6dffe0138f 100644
--- a/kernel/bpf/arena.c
+++ b/kernel/bpf/arena.c
@@ -284,7 +284,12 @@ static struct bpf_map *arena_map_alloc(union bpf_attr *attr)
!(attr->map_flags & BPF_F_MMAPABLE) ||
/* No unsupported flags present */
(attr->map_flags & ~(BPF_F_SEGV_ON_FAULT | BPF_F_MMAPABLE |
- BPF_F_NO_USER_CONV | BPF_F_ARENA_NO_FREE)))
+ BPF_F_NO_USER_CONV | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT)))
+ return ERR_PTR(-EINVAL);
+
+ /* Exported memory must remain allocated while consumers use it. */
+ if ((attr->map_flags & BPF_F_ARENA_EXPORT) && !(attr->map_flags & BPF_F_ARENA_NO_FREE))
return ERR_PTR(-EINVAL);
if (attr->map_extra & ~PAGE_MASK)
@@ -496,7 +501,8 @@ static vm_fault_t arena_vm_fault(struct vm_fault *vmf)
int ret;
kbase = bpf_arena_get_kern_vm_start(arena);
- kaddr = kbase + (u32)(vmf->address);
+ /* vmf->pgoff includes the file offset of a bpffs-backed slice. */
+ kaddr = kbase + (u32)(arena->user_vm_start + ((u64)vmf->pgoff << PAGE_SHIFT));
page = vmalloc_to_page((void *)kaddr);
if (!page && !(arena->map.map_flags & BPF_F_SEGV_ON_FAULT)) {
@@ -613,13 +619,23 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
long ret;
+ if (filp->f_op != &bpf_map_fops) {
+ if (!len || pgoff >= map->max_entries ||
+ len > ((u64)map->max_entries - pgoff) << PAGE_SHIFT)
+ return -EINVAL;
+ return mm_get_unmapped_area(filp, addr, len, pgoff, flags);
+ }
+
if (pgoff)
return -EINVAL;
if (len > SZ_4G)
return -E2BIG;
+ if ((map->map_flags & BPF_F_ARENA_EXPORT) &&
+ len != (u64)map->max_entries << PAGE_SHIFT)
+ return -EINVAL;
- /* if user_vm_start was specified at arena creation time */
- if (arena->user_vm_start) {
+ /* Once established, the canonical range cannot change. */
+ if (arena->user_vm_end) {
if (len > arena->user_vm_end - arena->user_vm_start)
return -E2BIG;
if (len != arena->user_vm_end - arena->user_vm_start)
@@ -633,7 +649,7 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
return ret;
if ((ret >> 32) == ((ret + len - 1) >> 32))
return ret;
- if (WARN_ON_ONCE(arena->user_vm_start))
+ if (WARN_ON_ONCE(arena->user_vm_end))
/* checks at map creation time should prevent this */
return -EFAULT;
return round_up(ret, SZ_4G);
@@ -642,9 +658,13 @@ static unsigned long arena_get_unmapped_area(struct file *filp, unsigned long ad
static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
{
struct bpf_arena *arena = container_of(map, struct bpf_arena, map);
+ bool exported = vma->vm_file->f_op != &bpf_map_fops;
guard(mutex)(&arena->lock);
- if (arena->user_vm_start && arena->user_vm_start != vma->vm_start)
+ if (!exported && (map->map_flags & BPF_F_ARENA_EXPORT) &&
+ vma->vm_end - vma->vm_start != (u64)map->max_entries << PAGE_SHIFT)
+ return -EINVAL;
+ if (!exported && arena->user_vm_end && arena->user_vm_start != vma->vm_start)
/*
* If map_extra was not specified at arena creation time then
* 1st user process can do mmap(NULL, ...) to pick user_vm_start
@@ -655,19 +675,31 @@ static int arena_map_mmap(struct bpf_map *map, struct vm_area_struct *vma)
*/
return -EBUSY;
- if (arena->user_vm_end && arena->user_vm_end != vma->vm_end)
+ if (exported && !arena->user_vm_end)
+ return -EINVAL;
+ if (!exported && arena->user_vm_end && arena->user_vm_end != vma->vm_end)
/* all user processes must have the same size of mmap-ed region */
return -EBUSY;
- /* Earlier checks should prevent this */
- if (WARN_ON_ONCE(vma->vm_end - vma->vm_start > SZ_4G || vma->vm_pgoff))
+ if (exported) {
+ u64 page_cnt = (arena->user_vm_end - arena->user_vm_start) >> PAGE_SHIFT;
+
+ if (vma->vm_pgoff >= page_cnt ||
+ (vma->vm_end - vma->vm_start) >> PAGE_SHIFT >
+ page_cnt - vma->vm_pgoff)
+ return -EINVAL;
+ } else if (WARN_ON_ONCE(vma->vm_end - vma->vm_start > SZ_4G || vma->vm_pgoff)) {
+ /* Earlier checks should prevent this for map FD mappings. */
return -EFAULT;
+ }
if (remember_vma(arena, vma))
return -ENOMEM;
- arena->user_vm_start = vma->vm_start;
- arena->user_vm_end = vma->vm_end;
+ if (!exported) {
+ arena->user_vm_start = vma->vm_start;
+ arena->user_vm_end = vma->vm_end;
+ }
/*
* bpf_map_mmap() checks that it's being mmaped as VM_SHARED and
* clears VM_MAYEXEC. Set VM_DONTEXPAND to avoid potential change
@@ -841,8 +873,12 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
struct mm_struct *mm;
struct vma_list *vml;
unsigned long vm_start;
+ u64 start, end, vma_start, vma_end;
u64 my_gen;
+ start = uaddr - arena->user_vm_start;
+ end = start + size;
+
/*
* Taking mmap_read_lock() under arena->lock would deadlock against
* arena_vm_close(), which runs with mmap_write_lock held and then
@@ -881,8 +917,14 @@ static void zap_pages(struct bpf_arena *arena, long uaddr, long page_cnt)
*/
vma = find_vma(mm, vm_start);
if (vma && vma->vm_start == vm_start &&
- vma->vm_file && vma->vm_file->private_data == &arena->map)
- zap_vma_range(vma, uaddr, size);
+ vma->vm_file && vma->vm_file->private_data == &arena->map) {
+ vma_start = (u64)vma->vm_pgoff << PAGE_SHIFT;
+ vma_end = vma_start + vma->vm_end - vma->vm_start;
+ if (start < vma_end && end > vma_start)
+ zap_vma_range(vma, vma->vm_start +
+ (max(start, vma_start) - vma_start),
+ min(end, vma_end) - max(start, vma_start));
+ }
mmap_read_unlock(mm);
mmput(mm);
diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c
index 7837968c0842c..d3dc6f70e32ae 100644
--- a/kernel/bpf/inode.c
+++ b/kernel/bpf/inode.c
@@ -119,8 +119,24 @@ static const struct inode_operations bpf_symlink_iops;
static const struct inode_operations bpf_prog_iops = {
.listxattr = bpf_fs_listxattr,
};
+
+static int bpf_map_setattr(struct mnt_idmap *idmap, struct dentry *dentry,
+ struct iattr *attr)
+{
+ struct inode *inode = d_inode(dentry);
+ struct bpf_map *map = inode->i_private;
+
+ if (map->map_type == BPF_MAP_TYPE_ARENA &&
+ (map->map_flags & BPF_F_ARENA_EXPORT) &&
+ (attr->ia_valid & ATTR_SIZE) && attr->ia_size != i_size_read(inode))
+ return -EINVAL;
+
+ return simple_setattr(idmap, dentry, attr);
+}
+
static const struct inode_operations bpf_map_iops = {
.listxattr = bpf_fs_listxattr,
+ .setattr = bpf_map_setattr,
};
static const struct inode_operations bpf_link_iops = {
.listxattr = bpf_fs_listxattr,
@@ -351,6 +367,68 @@ static const struct file_operations bpffs_map_fops = {
.release = bpffs_map_release,
};
+/*
+ * An arena with BPF_F_ARENA_EXPORT pinned in bpffs can serve as a
+ * shared-memory file. The ordinary map FD is an anonymous inode with no
+ * size and cannot be opened by pathname, so keep the pin's inode and
+ * forward mmap to the map FD implementation. The pinned file uses ordinary
+ * address selection so QEMU can map the pages at another virtual address.
+ */
+static int bpffs_arena_open(struct inode *inode, struct file *file)
+{
+ struct bpf_map *map = inode->i_private;
+ int err;
+
+ err = security_bpf_map(map, file->f_mode);
+ if (err)
+ return err;
+ bpf_map_inc_with_uref(map);
+ file->private_data = map;
+
+ return 0;
+}
+
+static int bpffs_arena_mmap(struct file *file, struct vm_area_struct *vma)
+{
+ /*
+ * Generic mmap clears VM_SHARED for O_RDONLY files, but VM_MAYSHARE
+ * still distinguishes shared mappings from private ones. Restore
+ * VM_SHARED for the map mmap path without granting VM_MAYWRITE.
+ */
+ if (!(file->f_mode & FMODE_WRITE) && (vma->vm_flags & VM_MAYSHARE))
+ vm_flags_set(vma, VM_SHARED);
+
+ return bpf_map_fops.mmap(file, vma);
+}
+
+static int bpffs_arena_release(struct inode *inode, struct file *file)
+{
+ return bpf_map_fops.release(inode, file);
+}
+
+static unsigned long bpffs_arena_get_unmapped_area(struct file *file,
+ unsigned long addr,
+ unsigned long len,
+ unsigned long pgoff,
+ unsigned long flags)
+{
+ return bpf_map_fops.get_unmapped_area(file, addr, len, pgoff, flags);
+}
+
+static loff_t bpffs_arena_llseek(struct file *file, loff_t offset, int whence)
+{
+ return fixed_size_llseek(file, offset, whence, i_size_read(file_inode(file)));
+}
+
+static const struct file_operations bpffs_arena_fops = {
+ .open = bpffs_arena_open,
+ .llseek = bpffs_arena_llseek,
+ .fsync = noop_fsync,
+ .release = bpffs_arena_release,
+ .mmap = bpffs_arena_mmap,
+ .get_unmapped_area = bpffs_arena_get_unmapped_area,
+};
+
static int bpffs_obj_open(struct inode *inode, struct file *file)
{
return -EIO;
@@ -362,7 +440,7 @@ static const struct file_operations bpffs_obj_fops = {
static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
const struct inode_operations *iops,
- const struct file_operations *fops)
+ const struct file_operations *fops, loff_t size)
{
struct inode *dir = dentry->d_parent->d_inode;
struct inode *inode;
@@ -382,6 +460,7 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
inode->i_op = iops;
inode->i_fop = fops;
inode->i_private = raw;
+ i_size_write(inode, size);
bpf_dentry_finalize(dentry, inode, dir);
return 0;
@@ -390,16 +469,20 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
static int bpf_mkprog(struct dentry *dentry, umode_t mode, void *arg)
{
return bpf_mkobj_ops(dentry, mode, arg, &bpf_prog_iops,
- &bpffs_obj_fops);
+ &bpffs_obj_fops, 0);
}
static int bpf_mkmap(struct dentry *dentry, umode_t mode, void *arg)
{
struct bpf_map *map = arg;
+ bool shared_arena = map->map_type == BPF_MAP_TYPE_ARENA &&
+ (map->map_flags & BPF_F_ARENA_EXPORT);
return bpf_mkobj_ops(dentry, mode, arg, &bpf_map_iops,
- bpf_map_support_seq_show(map) ?
- &bpffs_map_fops : &bpffs_obj_fops);
+ shared_arena ? &bpffs_arena_fops :
+ bpf_map_support_seq_show(map) ?
+ &bpffs_map_fops : &bpffs_obj_fops,
+ shared_arena ? (loff_t)map->max_entries * PAGE_SIZE : 0);
}
static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)
@@ -408,7 +491,7 @@ static int bpf_mklink(struct dentry *dentry, umode_t mode, void *arg)
return bpf_mkobj_ops(dentry, mode, arg, &bpf_link_iops,
bpf_link_is_iter(link) ?
- &bpf_iter_fops : &bpffs_obj_fops);
+ &bpf_iter_fops : &bpffs_obj_fops, 0);
}
static struct dentry *
@@ -483,7 +566,7 @@ static int bpf_iter_link_pin_kernel(struct dentry *parent,
if (IS_ERR(dentry))
return PTR_ERR(dentry);
ret = bpf_mkobj_ops(dentry, mode, link, &bpf_link_iops,
- &bpf_iter_fops);
+ &bpf_iter_fops, 0);
simple_done_creating(dentry);
return ret;
}
diff --git a/tools/include/uapi/linux/bpf.h b/tools/include/uapi/linux/bpf.h
index 75101d602bd32..40a5489ac9fea 100644
--- a/tools/include/uapi/linux/bpf.h
+++ b/tools/include/uapi/linux/bpf.h
@@ -1503,6 +1503,9 @@ enum {
/* Keep arena pages allocated until the map is destroyed. */
BPF_F_ARENA_NO_FREE = (1U << 20),
+
+ /* Expose a pinned arena as a sized file; requires BPF_F_ARENA_NO_FREE. */
+ BPF_F_ARENA_EXPORT = (1U << 21),
};
/* Flags for BPF_PROG_QUERY. */
--
2.56.0
^ permalink raw reply [flat|nested] 10+ messages in thread* Re: [PATCH 2/5] bpf: Expose pinned arenas as sized bpffs files
2026-10-10 22:45 ` [PATCH 2/5] bpf: Expose pinned arenas as sized bpffs files Andrea Righi
@ 2026-10-10 23:40 ` bot+bpf-ci
0 siblings, 0 replies; 10+ messages in thread
From: bot+bpf-ci @ 2026-10-10 23:40 UTC (permalink / raw)
To: arighi, ast, daniel, andrii, eddyz87, memxor
Cc: martin.lau, song, yonghong.song, jolsa, emil, ihor.solodrai,
john.fastabend, joshdon, vineethrp, balbirs, skolothumtho,
himadrics, royfru44, bpf, linux-kernel, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 2134 bytes --]
> bpf: Expose pinned arenas as sized bpffs files
TL;DR: Nits only in bpf_mkmap() in kernel/bpf/inode.c: the nested ternary
could be an early return, continuation indentation is off, and the
arena-export predicate duplicated in bpf_map_setattr() could be shared.
> diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c
> index 7837968c0842..d3dc6f70e32a 100644
> --- a/kernel/bpf/inode.c
> +++ b/kernel/bpf/inode.c
[ ... ]
> @@ -390,16 +469,20 @@ static int bpf_mkobj_ops(struct dentry *dentry, umode_t mode, void *raw,
> static int bpf_mkprog(struct dentry *dentry, umode_t mode, void *arg)
> {
> return bpf_mkobj_ops(dentry, mode, arg, &bpf_prog_iops,
> - &bpffs_obj_fops);
> + &bpffs_obj_fops, 0);
> }
>
> static int bpf_mkmap(struct dentry *dentry, umode_t mode, void *arg)
> {
> struct bpf_map *map = arg;
> + bool shared_arena = map->map_type == BPF_MAP_TYPE_ARENA &&
> + (map->map_flags & BPF_F_ARENA_EXPORT);
>
> return bpf_mkobj_ops(dentry, mode, arg, &bpf_map_iops,
> - bpf_map_support_seq_show(map) ?
> - &bpffs_map_fops : &bpffs_obj_fops);
> + shared_arena ? &bpffs_arena_fops :
> + bpf_map_support_seq_show(map) ?
> + &bpffs_map_fops : &bpffs_obj_fops,
> + shared_arena ? (loff_t)map->max_entries * PAGE_SIZE : 0);
> }
This isn't a bug, but would it read more simply to handle the exported arena
case with an early return in bpf_mkmap(), rather than testing shared_arena
twice inside a nested ternary in the argument list?
The continuation lines here are also indented one column short of the open
parenthesis, unlike the matching call in bpf_mklink().
The same 'map->map_type == BPF_MAP_TYPE_ARENA && (map->map_flags &
BPF_F_ARENA_EXPORT)' predicate is repeated in bpf_map_setattr(). Could it be
shared through a small helper, or could bpf_map_setattr() check
inode->i_fop == &bpffs_arena_fops instead?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/38093513563
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 3/5] docs/bpf: Document pinned arena file mappings
2026-10-10 22:45 [PATCH bpf-next 0/5] bpf: Expose pinned arenas as sized bpffs files Andrea Righi
2026-10-10 22:45 ` [PATCH 1/5] bpf: Add an arena flag to retain allocated pages Andrea Righi
2026-10-10 22:45 ` [PATCH 2/5] bpf: Expose pinned arenas as sized bpffs files Andrea Righi
@ 2026-10-10 22:45 ` Andrea Righi
2026-10-10 23:19 ` bot+bpf-ci
2026-10-10 22:45 ` [PATCH 4/5] selftests/bpf: Test " Andrea Righi
2026-10-10 22:46 ` [PATCH 5/5] selftests/bpf: Test two-way arena access with two guests Andrea Righi
4 siblings, 1 reply; 10+ messages in thread
From: Andrea Righi @ 2026-10-10 22:45 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
Emil Tsalapatis, Ihor Solodrai, John Fastabend, Josh Don,
Vineeth Pillai, Balbir Singh, Shameer Kolothum,
Himadri Chhaya-Shailesh, NchangRoy, bpf, linux-kernel
Describe how consumers can access pinned BPF arenas as sized bpffs files
and establish the canonical BPF address range before mapping them.
Explain slice bounds, access permissions, fixed file size, and mapping
lifetime so consumers can use the shared-memory interface correctly.
Document fixed-size seeking and the inherited arena VMA restrictions:
no fork inheritance, splitting, expansion, or relocation.
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
Documentation/bpf/map_arena.rst | 131 ++++++++++++++++++++++++++++++++
1 file changed, 131 insertions(+)
create mode 100644 Documentation/bpf/map_arena.rst
diff --git a/Documentation/bpf/map_arena.rst b/Documentation/bpf/map_arena.rst
new file mode 100644
index 0000000000000..1b9eadd5fb3a1
--- /dev/null
+++ b/Documentation/bpf/map_arena.rst
@@ -0,0 +1,131 @@
+.. SPDX-License-Identifier: GPL-2.0-only
+
+==================
+BPF_MAP_TYPE_ARENA
+==================
+
+A BPF arena provides shared memory that BPF programs and userspace can
+access directly. Create the map with zero key and value sizes and
+``BPF_F_MMAPABLE``. ``max_entries`` specifies its capacity in pages, up to
+4 GiB. The canonical BPF address range must not cross a 4 GiB boundary.
+
+Pinned arena files
+==================
+
+An arena created with ``BPF_F_ARENA_EXPORT`` can be pinned in bpffs and
+opened as a sized file. Its size is ``max_entries * PAGE_SIZE``. Export
+requires ``BPF_F_ARENA_NO_FREE`` so BPF programs cannot release backing
+pages while external consumers may still use them. Creating an arena with
+``BPF_F_ARENA_EXPORT`` alone fails with ``EINVAL``. Arenas without the
+export flag retain their existing bpffs behavior and cannot be opened
+this way, including arenas created with ``BPF_F_ARENA_NO_FREE`` alone.
+
+Establish the canonical BPF address range before mapping the pinned file:
+
+* Set ``map_extra`` to a nonzero, page-aligned address at map creation.
+ The canonical range then covers the full map capacity, even if the map
+ FD has not been mmaped.
+* Alternatively, leave ``map_extra`` zero and mmap the map FD first. This
+ mapping establishes the canonical address and must cover the full map
+ capacity. Shorter canonical mappings fail with ``EINVAL``.
+
+A canonical map FD mapping may start at address zero if the system's
+low-address mapping policy permits it.
+
+Mapping the pinned file before either step fails with ``EINVAL``. Exported
+mappings do not establish or change the canonical BPF address range.
+
+Once the canonical range is established, it covers the entire capacity
+reported by ``fstat()``. Consumers can map the full file or smaller slices.
+Establishing the full canonical mapping reserves virtual address space;
+it does not populate every arena page. Arenas without
+``BPF_F_ARENA_EXPORT`` can still establish shorter canonical mappings.
+
+Open the pin with ``O_RDONLY`` for read-only access or ``O_RDWR`` for
+writable access, and mmap it with ``MAP_SHARED``. An exported
+mapping may use a different virtual address from the canonical mapping.
+Its offset must be page-aligned, and its offset and length must fit within
+both the map capacity and the established canonical range. Consumers can
+map disjoint slices of one arena. Oversized or out-of-range mappings fail
+with ``EINVAL``. ``MAP_PRIVATE`` mappings are unsupported.
+
+Exported mappings share the bytes of the arena without relocating pointers
+stored in those bytes. Arena pointers produced by a BPF program use that
+arena's canonical address representation, based on ``user_vm_start``.
+A consumer mapping a slice at another address must translate such pointers
+before dereferencing them. A guest BPF arena has its own canonical range;
+sharing the backing pages does not make host arena pointers valid in the
+guest arena, or vice versa. Communication structures can use offsets
+relative to the shared slice, with each side validating the offsets and
+translating them to addresses in its own mapping or arena.
+
+A pin opened with ``O_RDONLY`` supports read-only shared mappings, which
+observe updates made by BPF programs and other writable mappings. Requests
+for ``PROT_WRITE`` through that FD fail with ``EACCES``. A mapping made
+without ``PROT_WRITE`` cannot subsequently acquire write permission with
+``mprotect()``, even when the pin was opened with ``O_RDWR``.
+
+Pin permissions can grant readers access independently of writers. For
+example, mode ``0640`` permits the owner to open the pin for writable
+access and the group to open it for read-only access, subject to the
+applicable LSM checks. Changing permissions does not revoke access through
+existing FDs or mappings.
+
+The file has a fixed size. ``truncate()``, ``ftruncate()``, and ``O_TRUNC``
+cannot change it. A truncate request that leaves the size unchanged is
+allowed. Permission and ownership changes remain available.
+Consumers can discover the capacity with ``fstat()`` or
+``lseek(fd, 0, SEEK_END)``. Seeking supports ``SEEK_SET``, ``SEEK_CUR``, and
+``SEEK_END`` within the file bounds; it does not enable ``read()`` or
+``write()`` access to the memory.
+
+``fsync()``, ``fdatasync()``, and ``msync(..., MS_SYNC)`` succeed without
+performing writeback: arena memory is volatile and has no persistent
+backing. These calls do not synchronize access between consumers or BPF
+programs; shared-memory protocols still need appropriate memory ordering.
+
+Exported mappings inherit the arena's mapping lifecycle restrictions.
+They are not inherited across ``fork()``, and ``MADV_DOFORK`` cannot enable
+inheritance. Consumers must create their own mappings in a child process.
+Mappings cannot be split, expanded, or relocated with ``mremap()``.
+Partial unmapping and protection changes that require splitting a mapping
+fail with ``EINVAL``. Unmapping an entire mapping remains supported, as do
+protection changes that do not require a split and satisfy the access
+restrictions above. Consumers needing independently managed regions
+should create separate slice mappings.
+
+Access control and lifetime
+===========================
+
+Opening the pin uses ordinary filesystem permissions and file LSM checks,
+as well as ``security_bpf_map()`` for the requested access mode. Mapping it
+also passes the ordinary mmap LSM checks. Writable mappings remain subject
+to the BPF map's frozen state and program-read-only restrictions.
+
+Consumers use ``open()`` and ``mmap()`` rather than ``bpf(BPF_OBJ_GET)``.
+Consequently, seccomp rules and LSM policies specific to the ``bpf()``
+syscall do not govern this file interface. Grant access through the pin's
+permissions and the applicable file, mmap, and BPF map LSM policies.
+A process prohibited from calling ``bpf()`` can still access the arena
+through this interface if those permissions and checks allow it.
+``unprivileged_bpf_disabled`` restricts unprivileged map and program
+creation; it does not generally prohibit access to existing maps.
+
+``BPF_F_ARENA_NO_FREE`` makes ``bpf_arena_free_pages()`` a no-op in both
+sleepable and non-sleepable BPF programs. The free operation returns no
+indication that pages were retained. Programs using this flag must not
+rely on freeing pages to restore arena allocation capacity. Arena pages
+are retained until the map is destroyed. Open map and pin FDs, mappings,
+and other map references retain the arena, including after the pin is
+unlinked. Removing the pin therefore does not revoke existing mappings or
+free their pages.
+
+The retention flag does not itself expose the pin as a sized file.
+``BPF_F_ARENA_EXPORT`` opts into that interface and must be combined with
+``BPF_F_ARENA_NO_FREE``. Retention-only arenas can still share memory
+through ordinary map FD mappings and manage reusable objects within
+their retained pages.
+
+For example, a consumer such as QEMU can use an exported slice as a shared
+file-backed guest RAM backend. The file interface provides shared memory;
+consumers must supply their own allocation and communication protocol.
--
2.56.0
^ permalink raw reply [flat|nested] 10+ messages in thread
* Re: [PATCH 3/5] docs/bpf: Document pinned arena file mappings
2026-10-10 22:45 ` [PATCH 3/5] docs/bpf: Document pinned arena file mappings Andrea Righi
@ 2026-10-10 23:19 ` bot+bpf-ci
0 siblings, 0 replies; 10+ messages in thread
From: bot+bpf-ci @ 2026-10-10 23:19 UTC (permalink / raw)
To: arighi, ast, daniel, andrii, eddyz87, memxor
Cc: martin.lau, song, yonghong.song, jolsa, emil, ihor.solodrai,
john.fastabend, joshdon, vineethrp, balbirs, skolothumtho,
himadrics, royfru44, bpf, linux-kernel, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 4093 bytes --]
> docs/bpf: Document pinned arena file mappings
TL;DR:
- Doc says a canonical arena mapping may start at address zero, but the
verifier may then reject programs using it
- Doc claim about shorter mappings of non-exported arenas looks wrong when
map_extra is nonzero
- Doc mentions an unreachable read-only restriction
> diff --git a/Documentation/bpf/map_arena.rst b/Documentation/bpf/map_arena.rst
> new file mode 100644
> index 000000000000..1b9eadd5fb3a
> --- /dev/null
> +++ b/Documentation/bpf/map_arena.rst
> @@ -0,0 +1,131 @@
[ ... ]
> +* Alternatively, leave ``map_extra`` zero and mmap the map FD first. This
> + mapping establishes the canonical address and must cover the full map
> + capacity. Shorter canonical mappings fail with ``EINVAL``.
> +
> +A canonical map FD mapping may start at address zero if the system's
> +low-address mapping policy permits it.
Can a BPF program use an arena in that state? The BPF_MAP_TYPE_ARENA
handling in kernel/bpf/verifier.c still treats user_vm_start == 0 as unset:
if (!bpf_arena_get_user_vm_start(env->prog->aux->arena)) {
verbose(env, "arena's user address must be set via map_extra or mmap()\n");
return -EINVAL;
}
Consider an arena created with BPF_F_MMAPABLE, BPF_F_ARENA_NO_FREE and
BPF_F_ARENA_EXPORT, with map_extra == 0, followed by:
mmap(0, capacity, ..., MAP_SHARED | MAP_FIXED_NOREPLACE, map_fd, 0)
With mmap_min_addr == 0 this succeeds, since arena_get_unmapped_area() and
arena_map_mmap() in kernel/bpf/arena.c now test user_vm_end for an
established range. That leaves user_vm_start = 0 and user_vm_end = capacity,
and loading any BPF program that references the arena then fails with
-EINVAL.
Should the verifier check use the same "established" test, or should this
paragraph say that such an arena cannot be used by BPF programs?
[ ... ]
> +Once the canonical range is established, it covers the entire capacity
> +reported by ``fstat()``. Consumers can map the full file or smaller slices.
> +Establishing the full canonical mapping reserves virtual address space;
> +it does not populate every arena page. Arenas without
> +``BPF_F_ARENA_EXPORT`` can still establish shorter canonical mappings.
Is that last sentence true for every arena without BPF_F_ARENA_EXPORT?
For a non-exported arena created with a nonzero map_extra,
arena_map_alloc() sets user_vm_end = map_extra + max_entries * PAGE_SIZE at
creation, so the canonical range is already the full capacity.
arena_get_unmapped_area() then takes the "if (arena->user_vm_end)" branch and
rejects any len != user_vm_end - user_vm_start with -EINVAL (-E2BIG if
larger), and arena_map_mmap() returns -EBUSY when user_vm_end !=
vma->vm_end. A shorter map FD mapping of such an arena fails, which also
conflicts with the earlier bullet saying a nonzero map_extra makes the range
cover the full map capacity.
Should the sentence be qualified with "when map_extra is zero"?
[ ... ]
> +Opening the pin uses ordinary filesystem permissions and file LSM checks,
> +as well as ``security_bpf_map()`` for the requested access mode. Mapping it
> +also passes the ordinary mmap LSM checks. Writable mappings remain subject
> +to the BPF map's frozen state and program-read-only restrictions.
Can the "program-read-only restrictions" apply to an arena?
bpf_map_mmap() does reject writable mappings of maps with BPF_F_RDONLY_PROG
(-EACCES), but arena_map_alloc() only accepts BPF_F_SEGV_ON_FAULT,
BPF_F_MMAPABLE, BPF_F_NO_USER_CONV, BPF_F_ARENA_NO_FREE and
BPF_F_ARENA_EXPORT. Creating an arena with BPF_F_RDONLY_PROG fails with
-EINVAL, and nothing sets the flag afterwards.
Only the frozen state (BPF_MAP_FREEZE, -EPERM on a writable mapping) is a
live restriction for exported arena mappings. Should the phrase be dropped,
or reworded to say that only the frozen state applies?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/38093513563
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 4/5] selftests/bpf: Test pinned arena file mappings
2026-10-10 22:45 [PATCH bpf-next 0/5] bpf: Expose pinned arenas as sized bpffs files Andrea Righi
` (2 preceding siblings ...)
2026-10-10 22:45 ` [PATCH 3/5] docs/bpf: Document pinned arena file mappings Andrea Righi
@ 2026-10-10 22:45 ` Andrea Righi
2026-10-10 23:40 ` bot+bpf-ci
2026-10-10 22:46 ` [PATCH 5/5] selftests/bpf: Test two-way arena access with two guests Andrea Righi
4 siblings, 1 reply; 10+ messages in thread
From: Andrea Righi @ 2026-10-10 22:45 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
Emil Tsalapatis, Ihor Solodrai, John Fastabend, Josh Don,
Vineeth Pillai, Balbir Singh, Shameer Kolothum,
Himadri Chhaya-Shailesh, NchangRoy, bpf, linux-kernel
Add native tests for pinned arena files, covering canonical range setup
via map FD mmap and map_extra, exported slice bounds, full-capacity
canonical mappings, and mapping lifetime after unpinning. Check that
oversized mapping requests fail cleanly and that size changes are
rejected without zapping mappings, while permission updates and
BPF_OBJ_GET still work.
Reject canonical mapping lengths that differ from the exported arena's
capacity and map the entire size reported by fstat(). Preserve short
canonical mappings for ordinary and retention-only arenas. Drop other
mappings, the pin, and open FDs before checking that the exported slice
retains the arena. Change the pin's mode from its creation permissions
to verify that metadata updates take effect.
Test O_RDONLY pin mappings with MAP_SHARED and MAP_SHARED_VALIDATE.
Verify that readers observe writer updates, retain access after closing
the FD, and cannot request writable or private mappings or upgrade a
mapping with mprotect().
Check that ordinary and retention-only arenas keep their existing pin
behavior, while BPF_F_ARENA_EXPORT requires BPF_F_ARENA_NO_FREE.
Check fixed-size seeking, canonical setup at address zero when security
policy permits it, and the exported mapping lifecycle restrictions.
Verify that partial unmap and protection changes, MADV_DOFORK, and
relocation fail, while fork leaves the parent's mapping intact without
inheriting it in the child.
Assisted-by: LLM
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
.../selftests/bpf/prog_tests/arena_pinned.c | 442 ++++++++++++++++++
1 file changed, 442 insertions(+)
create mode 100644 tools/testing/selftests/bpf/prog_tests/arena_pinned.c
diff --git a/tools/testing/selftests/bpf/prog_tests/arena_pinned.c b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c
new file mode 100644
index 0000000000000..8b7ca90211296
--- /dev/null
+++ b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c
@@ -0,0 +1,442 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <test_progs.h>
+#include <fcntl.h>
+#include <sys/mman.h>
+#include <sys/stat.h>
+#include <sys/wait.h>
+
+#define ARENA_PAGES 4
+
+struct arena_file {
+ int map_fd;
+ int file_fd;
+ char pin[128];
+ void *canonical;
+ size_t len;
+ bool pinned;
+};
+
+static void cleanup(struct arena_file *file)
+{
+ if (file->canonical != MAP_FAILED)
+ munmap(file->canonical, file->len);
+ if (file->file_fd >= 0)
+ close(file->file_fd);
+ if (file->pinned)
+ unlink(file->pin);
+ if (file->map_fd >= 0)
+ close(file->map_fd);
+}
+
+static void reject_mapping(int fd, size_t len, int flags, off_t offset,
+ const char *name)
+{
+ void *addr;
+ int err;
+
+ addr = mmap(NULL, len, PROT_READ | PROT_WRITE, flags, fd, offset);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, name))
+ ASSERT_EQ(err, EINVAL, "mmap_errno");
+ else
+ munmap(addr, len);
+}
+
+static bool setup(struct arena_file *file, __u64 map_extra, bool zero)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT,
+ .map_extra = map_extra);
+
+ memset(file, 0, sizeof(*file));
+ file->map_fd = -1;
+ file->file_fd = -1;
+ file->canonical = MAP_FAILED;
+ file->len = ARENA_PAGES * getpagesize();
+ snprintf(file->pin, sizeof(file->pin), "/sys/fs/bpf/arena_pinned_%d", getpid());
+ file->map_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_pinned",
+ 0, 0, ARENA_PAGES, &opts);
+ if (file->map_fd < 0 && errno == EOPNOTSUPP) {
+ test__skip();
+ return false;
+ }
+ if (!ASSERT_GE(file->map_fd, 0, "map_create"))
+ return false;
+ if (!ASSERT_OK(bpf_obj_pin(file->map_fd, file->pin), "obj_pin"))
+ return false;
+ file->pinned = true;
+ file->file_fd = open(file->pin, O_RDWR);
+ if (!ASSERT_GE(file->file_fd, 0, "open_pin"))
+ return false;
+ reject_mapping(file->map_fd, getpagesize(), MAP_SHARED, 0, "short_canonical");
+ reject_mapping(file->map_fd, file->len + getpagesize(), MAP_SHARED,
+ 0, "large_canonical");
+ if (map_extra)
+ return true;
+
+ reject_mapping(file->file_fd, getpagesize(), MAP_SHARED, 0, "uninitialized");
+ file->canonical = mmap(NULL, file->len, PROT_READ | PROT_WRITE,
+ MAP_SHARED | (zero ? MAP_FIXED_NOREPLACE : 0),
+ file->map_fd, 0);
+ if (zero && file->canonical == MAP_FAILED &&
+ (errno == EPERM || errno == EACCES)) {
+ test__skip();
+ return false;
+ }
+ return ASSERT_NEQ(file->canonical, MAP_FAILED, "canonical_mmap");
+}
+
+static void test_mappings(__u64 map_extra)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ struct stat st;
+ char *alias = MAP_FAILED;
+ char *full = MAP_FAILED;
+ int expected = 91;
+
+ if (!setup(&file, map_extra, false))
+ goto out;
+ if (!ASSERT_OK(fstat(file.file_fd, &st), "stat_pin"))
+ goto out;
+ ASSERT_EQ(st.st_size, ARENA_PAGES * ps, "pin_size");
+ reject_mapping(file.file_fd, 8ULL << 30, MAP_SHARED, 0, "oversized");
+ reject_mapping(file.file_fd, ps, MAP_SHARED, ARENA_PAGES * ps, "past_capacity");
+ reject_mapping(file.file_fd, ps * 2, MAP_SHARED,
+ (ARENA_PAGES - 1) * ps, "past_capacity_end");
+ reject_mapping(file.file_fd, ps, MAP_PRIVATE, 0, "private_mapping");
+ full = mmap(NULL, st.st_size, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, 0);
+ if (!ASSERT_NEQ(full, MAP_FAILED, "full_file_mmap"))
+ goto out;
+ full[0] = 17;
+ full[file.len - 1] = 62;
+
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, (ARENA_PAGES - 1) * ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "last_page_slice"))
+ goto out;
+ alias[0] = 91;
+ ASSERT_OK(msync(alias, ps, MS_SYNC), "msync_slice");
+ ASSERT_OK(fsync(file.file_fd), "fsync_pin");
+ ASSERT_OK(fdatasync(file.file_fd), "fdatasync_pin");
+ ASSERT_EQ(alias[ps - 1], 62, "full_file_last_page");
+ if (file.canonical != MAP_FAILED) {
+ char *canonical = file.canonical;
+
+ ASSERT_EQ(canonical[0], 17, "full_file_first_page");
+ ASSERT_EQ(canonical[(ARENA_PAGES - 1) * ps], 91, "alias_write");
+ canonical[(ARENA_PAGES - 1) * ps] = 42;
+ ASSERT_EQ(alias[0], 42, "canonical_write");
+ expected = 42;
+ if (!ASSERT_OK(munmap(file.canonical, file.len), "unmap_canonical"))
+ goto out;
+ file.canonical = MAP_FAILED;
+ }
+ if (!ASSERT_OK(munmap(full, file.len), "unmap_full_file"))
+ goto out;
+ full = MAP_FAILED;
+ if (!ASSERT_OK(unlink(file.pin), "unlink_pin"))
+ goto out;
+ file.pinned = false;
+ close(file.file_fd);
+ file.file_fd = -1;
+ close(file.map_fd);
+ file.map_fd = -1;
+ ASSERT_EQ(alias[0], expected, "unpin_lifetime");
+ alias[0] = 73;
+out:
+ if (full != MAP_FAILED)
+ munmap(full, file.len);
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_unexported_pin(__u32 flags)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | flags);
+ char pin[128];
+ struct stat st;
+ char *canonical = MAP_FAILED;
+ size_t ps = getpagesize();
+ int map_fd, fd, err;
+
+ map_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_unexported",
+ 0, 0, ARENA_PAGES, &opts);
+ if (map_fd < 0 && errno == EOPNOTSUPP) {
+ test__skip();
+ return;
+ }
+ if (!ASSERT_GE(map_fd, 0, "unexported_create"))
+ return;
+ canonical = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, map_fd, 0);
+ if (!ASSERT_NEQ(canonical, MAP_FAILED, "unexported_short_mmap"))
+ goto out;
+ canonical[0] = 42;
+ ASSERT_EQ(canonical[0], 42, "unexported_short_access");
+ snprintf(pin, sizeof(pin), "/sys/fs/bpf/arena_unexported_%d", getpid());
+ if (!ASSERT_OK(bpf_obj_pin(map_fd, pin), "unexported_pin"))
+ goto out;
+ if (ASSERT_OK(stat(pin, &st), "unexported_stat"))
+ ASSERT_EQ(st.st_size, 0, "unexported_size");
+ fd = open(pin, O_RDONLY);
+ err = errno;
+ if (ASSERT_EQ(fd, -1, "unexported_open"))
+ ASSERT_EQ(err, EIO, "unexported_open_errno");
+ else
+ close(fd);
+ fd = bpf_obj_get(pin);
+ if (ASSERT_GE(fd, 0, "unexported_obj_get"))
+ close(fd);
+ ASSERT_OK(unlink(pin), "unexported_unlink");
+out:
+ if (canonical != MAP_FAILED)
+ munmap(canonical, ps);
+ close(map_fd);
+}
+
+static void test_export_requires_no_free(void)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_EXPORT);
+ int fd;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_export",
+ 0, 0, ARENA_PAGES, &opts);
+ if (fd == -EOPNOTSUPP) {
+ test__skip();
+ return;
+ }
+ ASSERT_EQ(fd, -EINVAL, "export_requires_no_free");
+ if (fd >= 0)
+ close(fd);
+}
+
+static void test_read_only(int flags)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *addr;
+ int fd = -1, err, ret;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ if (!ASSERT_OK(fchmod(file.file_fd, 0400), "chmod_read_only"))
+ goto out;
+ fd = open(file.pin, O_RDONLY);
+ if (!ASSERT_GE(fd, 0, "open_read_only"))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ, flags, fd, ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "read_only_mmap"))
+ goto out;
+ ASSERT_EQ(alias[0], 0, "read_only_fault");
+ ((char *)file.canonical)[ps] = 42;
+ ASSERT_EQ(alias[0], 42, "read_only_shared_update");
+
+ addr = mmap(NULL, ps, PROT_READ | PROT_WRITE, flags, fd, ps);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, "read_only_writable_mmap"))
+ ASSERT_EQ(err, EACCES, "read_only_writable_errno");
+ else
+ munmap(addr, ps);
+ ret = mprotect(alias, ps, PROT_READ | PROT_WRITE);
+ err = errno;
+ ASSERT_EQ(ret, -1, "read_only_mprotect");
+ ASSERT_EQ(err, EACCES, "read_only_mprotect_errno");
+ addr = mmap(NULL, ps, PROT_READ, MAP_PRIVATE, fd, ps);
+ err = errno;
+ if (ASSERT_EQ(addr, MAP_FAILED, "read_only_private_mmap"))
+ ASSERT_EQ(err, EINVAL, "read_only_private_errno");
+ else
+ munmap(addr, ps);
+
+ close(fd);
+ fd = -1;
+ ((char *)file.canonical)[ps] = 73;
+ ASSERT_EQ(alias[0], 73, "read_only_after_close");
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ if (fd >= 0)
+ close(fd);
+ cleanup(&file);
+}
+
+static void test_fixed_size(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ struct stat st;
+ unsigned char resident;
+ char *alias = MAP_FAILED;
+ int err, fd, ret;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, file.file_fd, 0);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "alias_mmap"))
+ goto out;
+ alias[0] = 42;
+ ret = ftruncate(file.file_fd, 0);
+ err = errno;
+ ASSERT_EQ(ret, -1, "ftruncate_shrink");
+ ASSERT_EQ(err, EINVAL, "ftruncate_shrink_errno");
+ ret = ftruncate(file.file_fd, file.len + ps);
+ err = errno;
+ ASSERT_EQ(ret, -1, "ftruncate_grow");
+ ASSERT_EQ(err, EINVAL, "ftruncate_grow_errno");
+ ret = truncate(file.pin, 0);
+ err = errno;
+ ASSERT_EQ(ret, -1, "truncate_pin");
+ ASSERT_EQ(err, EINVAL, "truncate_pin_errno");
+ fd = open(file.pin, O_RDWR | O_TRUNC);
+ err = errno;
+ if (ASSERT_EQ(fd, -1, "open_trunc"))
+ ASSERT_EQ(err, EINVAL, "open_trunc_errno");
+ else
+ close(fd);
+ ASSERT_OK(ftruncate(file.file_fd, file.len), "ftruncate_same_size");
+ if (!ASSERT_OK(mincore(alias, ps, &resident), "mincore"))
+ goto out;
+ ASSERT_EQ(resident & 1, 1, "mapping_not_zapped");
+ ASSERT_EQ(alias[0], 42, "mapping_value");
+ ASSERT_OK(fchmod(file.file_fd, 0640), "chmod_pin");
+ if (ASSERT_OK(fstat(file.file_fd, &st), "stat_pin")) {
+ ASSERT_EQ(st.st_size, file.len, "fixed_size");
+ ASSERT_EQ(st.st_mode & 0777, 0640, "pin_mode");
+ }
+ ASSERT_EQ(lseek(file.file_fd, 0, SEEK_END), file.len, "seek_capacity");
+ ASSERT_EQ(lseek(file.file_fd, -(off_t)ps, SEEK_CUR), file.len - ps, "seek_back");
+ ASSERT_EQ(lseek(file.file_fd, ps, SEEK_SET), ps, "seek_offset");
+ errno = 0;
+ ASSERT_EQ(lseek(file.file_fd, file.len + 1, SEEK_SET), -1, "seek_past_end");
+ ASSERT_EQ(errno, EINVAL, "seek_past_end_errno");
+ errno = 0;
+ ASSERT_EQ(lseek(file.file_fd, -1, SEEK_SET), -1, "seek_before_start");
+ ASSERT_EQ(errno, EINVAL, "seek_before_start_errno");
+ fd = bpf_obj_get(file.pin);
+ if (ASSERT_GE(fd, 0, "obj_get_after_chmod"))
+ close(fd);
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_zero_canonical(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *addr;
+
+ if (!setup(&file, 0, true))
+ goto out;
+ if (!ASSERT_EQ((unsigned long)file.canonical, 0, "zero_canonical"))
+ goto out;
+ alias = mmap(NULL, ps, PROT_READ | PROT_WRITE, MAP_SHARED, file.file_fd, ps);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "zero_canonical_export"))
+ goto out;
+ alias[0] = 42;
+ ASSERT_EQ(alias[0], 42, "zero_canonical_fault");
+ addr = mmap((void *)ps, file.len, PROT_READ | PROT_WRITE,
+ MAP_SHARED | MAP_FIXED_NOREPLACE, file.map_fd, 0);
+ if (ASSERT_EQ(addr, MAP_FAILED, "canonical_cannot_move"))
+ ASSERT_EQ(errno, EINVAL, "canonical_cannot_move_errno");
+ else
+ munmap(addr, file.len);
+out:
+ if (alias != MAP_FAILED)
+ munmap(alias, ps);
+ cleanup(&file);
+}
+
+static void test_lifecycle(void)
+{
+ struct arena_file file;
+ size_t ps = getpagesize();
+ char *alias = MAP_FAILED;
+ void *target = MAP_FAILED, *addr;
+ pid_t child;
+ int status, ret, err;
+
+ if (!setup(&file, 0, false))
+ goto out;
+ alias = mmap(NULL, ps * 2, PROT_READ | PROT_WRITE, MAP_SHARED,
+ file.file_fd, 0);
+ if (!ASSERT_NEQ(alias, MAP_FAILED, "lifecycle_mmap"))
+ goto out;
+ alias[0] = 42;
+ ret = munmap(alias, ps);
+ err = errno;
+ ASSERT_EQ(ret, -1, "partial_unmap");
+ ASSERT_EQ(err, EINVAL, "partial_unmap_errno");
+ ret = mprotect(alias, ps, PROT_READ);
+ err = errno;
+ ASSERT_EQ(ret, -1, "partial_mprotect");
+ ASSERT_EQ(err, EINVAL, "partial_mprotect_errno");
+ ret = madvise(alias, ps * 2, MADV_DOFORK);
+ err = errno;
+ ASSERT_EQ(ret, -1, "enable_inheritance");
+ ASSERT_EQ(err, EINVAL, "enable_inheritance_errno");
+ target = mmap(NULL, ps * 2, PROT_NONE, MAP_PRIVATE | MAP_ANONYMOUS, -1, 0);
+ if (!ASSERT_NEQ(target, MAP_FAILED, "reserve_remap_target"))
+ goto out;
+ addr = mremap(alias, ps * 2, ps * 2, MREMAP_MAYMOVE | MREMAP_FIXED, target);
+ err = errno;
+ if (!ASSERT_EQ(addr, MAP_FAILED, "relocate_mapping")) {
+ alias = addr;
+ target = MAP_FAILED;
+ }
+ ASSERT_EQ(err, EINVAL, "relocate_mapping_errno");
+ child = fork();
+ if (!ASSERT_GE(child, 0, "fork"))
+ goto out;
+ if (!child) {
+ unsigned char resident[2];
+
+ ret = mincore(alias, ps * 2, resident);
+ _exit(ret == -1 && errno == ENOMEM ? 0 : 1);
+ }
+ if (ASSERT_EQ(waitpid(child, &status, 0), child, "wait_child") &&
+ ASSERT_TRUE(WIFEXITED(status), "child_exited"))
+ ASSERT_EQ(WEXITSTATUS(status), 0, "mapping_not_inherited");
+ ASSERT_EQ(alias[0], 42, "parent_mapping_retained");
+out:
+ if (target != MAP_FAILED)
+ munmap(target, ps * 2);
+ if (alias != MAP_FAILED)
+ munmap(alias, ps * 2);
+ cleanup(&file);
+}
+
+void serial_test_arena_pinned(void)
+{
+ if (test__start_subtest("ordinary_pin"))
+ test_unexported_pin(0);
+ if (test__start_subtest("retained_pin"))
+ test_unexported_pin(BPF_F_ARENA_NO_FREE);
+ if (test__start_subtest("export_requires_no_free"))
+ test_export_requires_no_free();
+ if (test__start_subtest("mappings"))
+ test_mappings(0);
+ if (test__start_subtest("map_extra"))
+ test_mappings(1ULL << 32);
+ if (test__start_subtest("fixed_size"))
+ test_fixed_size();
+ if (test__start_subtest("read_only"))
+ test_read_only(MAP_SHARED);
+ if (test__start_subtest("read_only_validate"))
+ test_read_only(MAP_SHARED_VALIDATE);
+ if (test__start_subtest("zero_canonical"))
+ test_zero_canonical();
+ if (test__start_subtest("lifecycle"))
+ test_lifecycle();
+}
--
2.56.0
^ permalink raw reply [flat|nested] 10+ messages in thread* Re: [PATCH 4/5] selftests/bpf: Test pinned arena file mappings
2026-10-10 22:45 ` [PATCH 4/5] selftests/bpf: Test " Andrea Righi
@ 2026-10-10 23:40 ` bot+bpf-ci
0 siblings, 0 replies; 10+ messages in thread
From: bot+bpf-ci @ 2026-10-10 23:40 UTC (permalink / raw)
To: arighi, ast, daniel, andrii, eddyz87, memxor
Cc: martin.lau, song, yonghong.song, jolsa, emil, ihor.solodrai,
john.fastabend, joshdon, vineethrp, balbirs, skolothumtho,
himadrics, royfru44, bpf, linux-kernel, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 5106 bytes --]
> selftests/bpf: Test pinned arena file mappings
TL;DR:
- Commit message second paragraph uses imperatives that read like kernel
changes and repeats other paragraphs
- Selftest skips print no reason, so address-0 and unsupported-arena skips
look identical
> selftests/bpf: Test pinned arena file mappings
>
> Add native tests for pinned arena files, covering canonical range setup
> via map FD mmap and map_extra, exported slice bounds, full-capacity
> canonical mappings, and mapping lifetime after unpinning. Check that
> oversized mapping requests fail cleanly and that size changes are
> rejected without zapping mappings, while permission updates and
> BPF_OBJ_GET still work.
>
> Reject canonical mapping lengths that differ from the exported arena's
> capacity and map the entire size reported by fstat(). Preserve short
> canonical mappings for ordinary and retention-only arenas. Drop other
> mappings, the pin, and open FDs before checking that the exported slice
> retains the arena. Change the pin's mode from its creation permissions
> to verify that metadata updates take effect.
>
> Test O_RDONLY pin mappings with MAP_SHARED and MAP_SHARED_VALIDATE.
> Verify that readers observe writer updates, retain access after closing
> the FD, and cannot request writable or private mappings or upgrade a
> mapping with mprotect().
>
> Check that ordinary and retention-only arenas keep their existing pin
> behavior, while BPF_F_ARENA_EXPORT requires BPF_F_ARENA_NO_FREE.
>
> Check fixed-size seeking, canonical setup at address zero when security
> policy permits it, and the exported mapping lifecycle restrictions.
> Verify that partial unmap and protection changes, MADV_DOFORK, and
> relocation fail, while fork leaves the parent's mapping intact without
> inheriting it in the child.
>
> Assisted-by: LLM
> Signed-off-by: Andrea Righi <arighi@nvidia.com>
This isn't a bug, but the second paragraph of the changelog uses
imperatives ("Reject canonical mapping lengths...", "Preserve short
canonical mappings...", "Drop other mappings...", "Change the pin's
mode...") that read like kernel behaviour changes, even though this patch
only adds a selftest.
Could the second paragraph be reworded to say what the tests check, for
example "Check that canonical mappings whose length differs from the
exported capacity are rejected ...", and merged with the first paragraph?
It also repeats items already listed there ("full-capacity canonical
mappings", "permission updates") and in a later paragraph ("ordinary and
retention-only arenas"). Any version-to-version notes could go below the
--- line.
> diff --git a/tools/testing/selftests/bpf/prog_tests/arena_pinned.c b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c
> new file mode 100644
> index 000000000000..8b7ca9021129
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/prog_tests/arena_pinned.c
> @@ -0,0 +1,442 @@
[ ... ]
> +static bool setup(struct arena_file *file, __u64 map_extra, bool zero)
> +{
[ ... ]
> + file->map_fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_pinned",
> + 0, 0, ARENA_PAGES, &opts);
> + if (file->map_fd < 0 && errno == EOPNOTSUPP) {
> + test__skip();
> + return false;
> + }
[ ... ]
> + file->canonical = mmap(NULL, file->len, PROT_READ | PROT_WRITE,
> + MAP_SHARED | (zero ? MAP_FIXED_NOREPLACE : 0),
> + file->map_fd, 0);
> + if (zero && file->canonical == MAP_FAILED &&
> + (errno == EPERM || errno == EACCES)) {
> + test__skip();
> + return false;
> + }
> + return ASSERT_NEQ(file->canonical, MAP_FAILED, "canonical_mmap");
> +}
A subsystem pattern flags this as potentially concerning: a skip that
prints no reason. This isn't a bug, but every skip path in this file is a
bare test__skip(), so test_progs only reports SKIP for the subtest.
For the zero_canonical subtest this skip is taken whenever mmap_min_addr
or an LSM refuses a mapping at address 0, and the log cannot tell that
apart from the arena-unsupported skip taken after bpf_map_create() returns
EOPNOTSUPP. The same bare test__skip() appears after the EOPNOTSUPP
checks in test_unexported_pin() and test_export_requires_no_free().
Other arena tests print the reason first, for example arena_htab.c:
printf("%s:SKIP:compiler doesn't support arena_cast\n", __func__);
test__skip();
and arena_memcg.c:
fprintf(stderr, "%s:SKIP:no memory controller or other env error\n",
__func__);
test__skip();
The bare test__skip() after an arena bpf_map_create() returning
EOPNOTSUPP does match existing usage in arena_direct_value.c, so the main
gap is the environment-dependent zero-canonical skip. Would a short
reason such as "mapping at address 0 not permitted" or "arena not
supported" be worth printing before each test__skip() here?
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/38093513563
^ permalink raw reply [flat|nested] 10+ messages in thread
* [PATCH 5/5] selftests/bpf: Test two-way arena access with two guests
2026-10-10 22:45 [PATCH bpf-next 0/5] bpf: Expose pinned arenas as sized bpffs files Andrea Righi
` (3 preceding siblings ...)
2026-10-10 22:45 ` [PATCH 4/5] selftests/bpf: Test " Andrea Righi
@ 2026-10-10 22:46 ` Andrea Righi
2026-10-10 23:40 ` bot+bpf-ci
4 siblings, 1 reply; 10+ messages in thread
From: Andrea Righi @ 2026-10-10 22:46 UTC (permalink / raw)
To: Alexei Starovoitov, Daniel Borkmann, Andrii Nakryiko,
Eduard Zingerman, Kumar Kartikeya Dwivedi
Cc: Martin KaFai Lau, Song Liu, Yonghong Song, Jiri Olsa,
Emil Tsalapatis, Ihor Solodrai, John Fastabend, Josh Don,
Vineeth Pillai, Balbir Singh, Shameer Kolothum,
Himadri Chhaya-Shailesh, NchangRoy, bpf, linux-kernel
Back a NUMA node in each of two simultaneous virtme-ng guests with a
disjoint slice of one pinned host arena.
Create the host arena with BPF_F_ARENA_EXPORT and BPF_F_ARENA_NO_FREE.
Guest arenas use only BPF_F_ARENA_NO_FREE because they do not export
pins.
Exchange two values in each direction with direct BPF loads and stores.
Check that a write to one slice leaves the other signal untouched. Unpin
the host arena during the exchange, verify that neither BPF side can
free the shared pages and keep one guest active after the other exits.
Let the host launcher own protocol deadlines, so ready guests can wait
while another guest boots or completes its exchange. Bound host helper
calls and register guest ownership before starting output readers.
Handle SIGTERM and SIGINT through normal cleanup. Act as a child
subreaper and kill and reap descendants with pidfds across sessions and
process groups, checking that readers finish after process cleanup.
Inspect compiler support before loading BPF or booting guests, and skip
when arena address-space conversions are unavailable. Supply the new
flag definitions when building against older kernel BTF. Use the
barrier-based acquire/release fallback on PowerPC, whose JIT does not
support these atomic load/store instructions for arena memory.
Honor the compiler's v4 capability check and permissive BPF build mode.
Give shared nodes room for boot allocations and allocator watermarks.
Check the allocated page's NUMA node and skip confirmed fallback, while
keeping physical-address bounds failures fatal on the expected node.
Assisted-by: LLM
Signed-off-by: Andrea Righi <arighi@nvidia.com>
---
tools/testing/selftests/bpf/.gitignore | 2 +
tools/testing/selftests/bpf/Makefile | 27 +-
tools/testing/selftests/bpf/arena_kvm.py | 448 ++++++++++++++++++
.../selftests/bpf/arena_kvm_guest.bpf.c | 89 ++++
tools/testing/selftests/bpf/arena_kvm_guest.c | 224 +++++++++
.../selftests/bpf/arena_kvm_host.bpf.c | 89 ++++
tools/testing/selftests/bpf/arena_kvm_host.c | 147 ++++++
.../testing/selftests/bpf/arena_kvm_shared.h | 60 +++
8 files changed, 1085 insertions(+), 1 deletion(-)
create mode 100755 tools/testing/selftests/bpf/arena_kvm.py
create mode 100644 tools/testing/selftests/bpf/arena_kvm_guest.bpf.c
create mode 100644 tools/testing/selftests/bpf/arena_kvm_guest.c
create mode 100644 tools/testing/selftests/bpf/arena_kvm_host.bpf.c
create mode 100644 tools/testing/selftests/bpf/arena_kvm_host.c
create mode 100644 tools/testing/selftests/bpf/arena_kvm_shared.h
diff --git a/tools/testing/selftests/bpf/.gitignore b/tools/testing/selftests/bpf/.gitignore
index b815bf0d88774..9747ad37d698e 100644
--- a/tools/testing/selftests/bpf/.gitignore
+++ b/tools/testing/selftests/bpf/.gitignore
@@ -47,3 +47,5 @@ verification_cert.h
*.BTF.base
usdt_1
usdt_2
+/arena_kvm_guest-init
+/arena_kvm_host-runner
diff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile
index a22be7efd1fa2..29fffa7857c19 100644
--- a/tools/testing/selftests/bpf/Makefile
+++ b/tools/testing/selftests/bpf/Makefile
@@ -42,7 +42,13 @@ TEST_PROGS := test_kmod.sh \
test_bpftool_build.sh \
test_doc_build.sh \
test_xsk.sh \
- test_xdp_features.sh
+ test_xdp_features.sh \
+ arena_kvm.py
+
+TEST_GEN_FILES += arena_kvm_guest-init arena_kvm_host-runner
+ifneq ($(CLANG_CPUV4),)
+TEST_GEN_FILES += arena_kvm_guest.bpf.o arena_kvm_host.bpf.o
+endif
TEST_PROGS_EXTENDED := \
ima_setup.sh verify_sig_setup.sh
@@ -94,6 +100,25 @@ BPFTOOLDIR := $(TOOLSDIR)/bpf/bpftool
HOST_BPFOBJ := $(HOST_BUILD_DIR)/libbpf/libbpf.a
BPF_TARGET_ENDIAN:=$(if $(IS_LITTLE_ENDIAN),--target=bpfel,--target=bpfeb)
+# These helpers are installed beside arena_kvm.py, including for OUTPUT=.
+$(OUTPUT)/arena_kvm_%.bpf.o: arena_kvm_%.bpf.c arena_kvm_shared.h \
+ libarena/include/bpf_atomic.h \
+ libarena/include/bpf_arena_common.h \
+ $(INCLUDE_DIR)/vmlinux.h $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BPF,,$@)
+ $(Q)$(CLANG) $(BPF_CFLAGS) $(CLANG_CFLAGS) -O2 \
+ $(BPF_TARGET_ENDIAN) -mcpu=v4 -c $< -o $@ $(call skip_on_fail,BPF)
+
+$(OUTPUT)/arena_kvm_guest-init: arena_kvm_guest.c arena_kvm_shared.h \
+ $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BINARY,,$@)
+ $(Q)$(CC) $(CFLAGS) $(LDFLAGS) $< $(BPFOBJ) $(LDLIBS) -lzstd -o $@
+
+$(OUTPUT)/arena_kvm_host-runner: arena_kvm_host.c arena_kvm_shared.h \
+ $(BPFOBJ) | $(OUTPUT)
+ $(call msg,BINARY,,$@)
+ $(Q)$(CC) $(CFLAGS) $(LDFLAGS) $< $(BPFOBJ) $(LDLIBS) -lzstd -o $@
+
NON_CHECK_FEAT_TARGETS := clean docs-clean emit_tests
CHECK_FEAT := $(filter-out $(NON_CHECK_FEAT_TARGETS),$(or $(MAKECMDGOALS), "none"))
ifneq ($(CHECK_FEAT),)
diff --git a/tools/testing/selftests/bpf/arena_kvm.py b/tools/testing/selftests/bpf/arena_kvm.py
new file mode 100755
index 0000000000000..552bfcfb80ac8
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm.py
@@ -0,0 +1,448 @@
+#!/usr/bin/env python3
+# SPDX-License-Identifier: GPL-2.0
+# Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+"""Exchange values between host and guest BPF through NUMA-backed RAM.
+
+Build with the BPF selftests Makefile. Helpers are found in the current
+directory or beside this script, so OUTPUT builds and installed tests work.
+Use --build-dir to select another helper directory and --kernel (or
+ARENA_KVM_KERNEL) to select the kernel booted by virtme-ng. By default, use
+the source tree's kernel when available, otherwise the running release.
+"""
+
+# Dependencies: Python 3.9+, virtme-ng, busybox-static, qemu, util-linux (script).
+
+import argparse
+import array
+import ctypes
+import mmap
+import os
+from pathlib import Path
+import queue
+import re
+import shutil
+import shlex
+import signal
+import socket
+import struct
+import subprocess
+import sys
+import threading
+import time
+
+
+PAGE_SIZE = os.sysconf("SC_PAGE_SIZE")
+KSFT_SKIP = 4
+HELPER_TIMEOUT = 20
+GUEST_MEMORY = 512 * 1024 * 1024
+# RAM starts for virtme-ng's PC/q35, virt (arm64/riscv64), and pseries
+# machines. With 512 MiB of RAM, the last NUMA node follows node 0
+# without crossing a machine's RAM hole. Skip other layouts.
+GUEST_RAM_BASES = {
+ "x86_64": 0,
+ "aarch64": 0x40000000,
+ "arm64": 0x40000000,
+ "riscv64": 0x80000000,
+ "ppc64": 0,
+ "ppc64le": 0,
+}
+# Leave room for guest boot allocations and allocator watermarks.
+# QEMU's pseries machine requires each NUMA node to be 256 MiB aligned.
+SLICE_SIZE = (256 if os.uname().machine in ("ppc64", "ppc64le") else 16) \
+ * 1024 * 1024
+ARENA_SIZE = 2 * SLICE_SIZE
+NUM_PAGES = ARENA_SIZE // PAGE_SIZE
+GUEST_READY = 0x4152454E414B564D
+HOST_FIRST = 0x123456789ABCDEF0
+GUEST_FIRST = 0xFEDCBA9876543210
+HOST_SECOND = 0x1020304050607080
+GUEST_SECOND = 0x8070605040302010
+
+
+def progress(message):
+ print(f"# arena_kvm: {message}", flush=True)
+
+
+class TestSkipped(Exception):
+ pass
+
+
+class TestTerminated(BaseException):
+ def __init__(self, signum):
+ self.signum = signum
+
+
+def terminate(signum, frame):
+ # A second signal must not interrupt cleanup of pins and descendants.
+ signal.signal(signal.SIGTERM, signal.SIG_IGN)
+ signal.signal(signal.SIGINT, signal.SIG_IGN)
+ raise TestTerminated(signum)
+
+
+def enable_subreaper():
+ # Adopt orphaned descendants even when script/vng changes session.
+ libc = ctypes.CDLL(None, use_errno=True)
+ libc.prctl.argtypes = (ctypes.c_int, *([ctypes.c_ulong] * 4))
+ if libc.prctl(36, 1, 0, 0, 0): # PR_SET_CHILD_SUBREAPER
+ err = ctypes.get_errno()
+ raise OSError(err, os.strerror(err))
+
+
+def children(pid):
+ # /proc/PID/task/TID/children requires CONFIG_CHECKPOINT_RESTORE.
+ # PPid is always available and covers children of every parent thread.
+ result = set()
+ for path in Path("/proc").iterdir():
+ if not path.name.isdigit():
+ continue
+ try:
+ status = (path / "status").read_text()
+ except (FileNotFoundError, ProcessLookupError, PermissionError):
+ continue
+ for line in status.splitlines():
+ if line.startswith("PPid:") and int(line.split()[1]) == pid:
+ result.add(int(path.name))
+ break
+ return result
+
+
+def stop_children(processes=()):
+ # Kill direct children, then repeat for descendants adopted by the
+ # subreaper. Sessions and process groups do not affect adoption.
+ # Reaping until ECHILD verifies that the entire owned tree has exited.
+ known = {process.pid: process for process in processes}
+ deadline = time.monotonic() + 10
+ while True:
+ for pid in children(os.getpid()):
+ try:
+ fd = os.pidfd_open(pid)
+ except ProcessLookupError:
+ continue
+ try:
+ signal.pidfd_send_signal(fd, signal.SIGKILL)
+ except ProcessLookupError:
+ pass
+ finally:
+ os.close(fd)
+ while True:
+ try:
+ pid, status = os.waitpid(-1, os.WNOHANG)
+ except ChildProcessError:
+ return
+ if not pid:
+ break
+ if pid in known:
+ known[pid].returncode = os.waitstatus_to_exitcode(status)
+ if time.monotonic() >= deadline:
+ raise TimeoutError("guest descendants did not exit")
+ time.sleep(0.01)
+
+
+def getq(arena, offset):
+ return struct.unpack_from("=Q", arena, offset)[0]
+
+
+def wait_reply(runner, obj, fd, offset, guest, lines):
+ deadline = time.monotonic() + 20
+ while time.monotonic() < deadline:
+ result = run_host_bpf(runner, obj, fd, offset, "check_reply",
+ timeout=max(0.001, deadline - time.monotonic()))
+ if result == 0:
+ return
+ if result != 1:
+ raise RuntimeError(f"guest BPF reply is wrong: result={result}")
+ if guest.poll() is not None:
+ raise RuntimeError("guest exited early:\n" + "".join(lines))
+ time.sleep(0.001)
+ raise TimeoutError(f"timed out waiting for BPF reply at offset {offset}")
+
+
+def read_lines(guest, messages, lines):
+ for line in guest.stdout:
+ lines.append(line)
+ # Serial startup may prefix a message with NULs echoed as ^@.
+ messages.put(re.sub(r"^(?:\x00|\^@)+", "", line.strip()))
+
+
+def wait_message(messages, lines, expected, guest):
+ deadline = time.monotonic() + 45
+ while time.monotonic() < deadline:
+ try:
+ line = messages.get(timeout=0.5)
+ except queue.Empty:
+ if guest.poll() is not None:
+ break
+ continue
+ if expected in line:
+ return line
+ if "GUEST_BPF_FAILED" in line:
+ break
+ if "GUEST_BPF_SKIPPED" in line:
+ raise TestSkipped(line)
+ raise RuntimeError(f"guest did not print {expected} "
+ f"(launcher exit status: {guest.poll()}):\n" +
+ "".join(lines))
+
+
+def run_host_bpf(runner, obj, fd, offset, program="exchange",
+ timeout=HELPER_TIMEOUT):
+ output = subprocess.check_output(
+ [str(runner), str(fd), str(obj), str(offset), program],
+ pass_fds=(fd,), text=True, timeout=timeout)
+ return int(output)
+
+
+def resolve_vng():
+ executable = shutil.which("vng")
+ if not executable or not os.access(executable, os.X_OK):
+ print("SKIP: virtme-ng (vng) is unavailable")
+ return None, None
+ return executable, os.environ.copy()
+
+
+def start_guest(pin, build, kernel, index, vng, env, guests):
+ node0_size = GUEST_MEMORY - SLICE_SIZE
+ base = GUEST_RAM_BASES[os.uname().machine] + node0_size
+ progress(f"Starting guest {index}: {SLICE_SIZE // (1024 * 1024)} MiB "
+ f"slice at file offset {index * SLICE_SIZE:#x}, "
+ f"NUMA node 1 physical base {base:#x}")
+ qemu_opts = (
+ f"-object memory-backend-file,id=signal,size={SLICE_SIZE},share=on,"
+ f"offset={index * SLICE_SIZE},mem-path={pin} "
+ "-numa node,nodeid=1,memdev=signal"
+ )
+ guest_command = [str(build / "arena_kvm_guest-init"),
+ str(build / "arena_kvm_guest.bpf.o")]
+ command = [
+ vng, "--verbose", "--run", str(kernel), "--cpus", "2",
+ "--memory", f"{GUEST_MEMORY // (1024 * 1024)}M",
+ "--numa", str(node0_size),
+ "--exec",
+ shlex.join(guest_command),
+ "--append", f"arena_kvm.base={base:#x} arena_kvm.size={SLICE_SIZE}",
+ f"--qemu-opts={qemu_opts}",
+ ]
+ lines = []
+ messages = queue.Queue()
+ guest = subprocess.Popen(["script", "-e", "-q", "-c", shlex.join(command),
+ "/dev/null"], stdout=subprocess.PIPE,
+ stderr=subprocess.STDOUT, text=True, env=env,
+ start_new_session=True)
+ # Register ownership before thread construction or startup can fail.
+ guests.append((guest, messages, lines, None))
+ reader = threading.Thread(target=read_lines, args=(guest, messages, lines),
+ daemon=True)
+ guests[-1] = guest, messages, lines, reader
+ reader.start()
+
+
+def stop_guests(guests):
+ stop_children(process for process, _, _, _ in guests)
+ for process, _, _, reader in guests:
+ process.wait(timeout=10)
+ if reader is not None and reader.ident is not None:
+ reader.join(timeout=2)
+ if reader.is_alive():
+ raise RuntimeError("guest output reader did not exit")
+ process.stdout.close()
+
+
+def exchange(pin, fd, arena, build, kernel, vng, env):
+ guests = []
+ try:
+ for index in range(2):
+ start_guest(pin, build, kernel, index, vng, env, guests)
+ offsets = []
+ for index, (process, messages, lines, _) in enumerate(guests):
+ progress(f"Waiting for guest {index} to allocate its BPF arena page "
+ "and check that BPF cannot free it")
+ ready = wait_message(messages, lines, "GUEST_BPF_READY", process)
+ match = re.fullmatch(r"GUEST_BPF_READY (\d+) (\d+)", ready)
+ if not match:
+ raise RuntimeError(f"malformed guest page offset: {ready}")
+ relative = int(match.group(1))
+ if int(match.group(2)) != PAGE_SIZE:
+ raise RuntimeError("host and guest page sizes differ")
+ offset = index * SLICE_SIZE + relative
+ if relative >= SLICE_SIZE or relative % PAGE_SIZE or \
+ getq(arena, offset) != GUEST_READY:
+ raise RuntimeError(f"invalid guest {index} page offset {relative}")
+ offsets.append(offset)
+ progress(f"Guest {index} ready: page offset {relative:#x} in slice, "
+ f"{offset:#x} in host arena")
+ # The QEMU VMA and the open map FD retain the arena after unlink.
+ progress("Unpinning the host arena while both guests retain mappings")
+ os.unlink(pin)
+ runner = build / "arena_kvm_host-runner"
+ obj = build / "arena_kvm_host.bpf.o"
+ for offset in offsets:
+ if run_host_bpf(runner, obj, fd, offset, "probe_free") != 0 or \
+ getq(arena, offset) != GUEST_READY:
+ raise RuntimeError("host BPF freed a shared RAM page")
+ progress("Host BPF page-free checks passed for both shared pages")
+
+ for index, offset in enumerate(offsets):
+ progress(f"Round 1: host -> guest {index}, value {HOST_FIRST:#018x}")
+ result = run_host_bpf(runner, obj, fd, offset)
+ if result != 0:
+ raise RuntimeError(f"guest {index} first result={result}")
+ other = offsets[1 - index]
+ if index == 0 and getq(arena, other + 16) != 0:
+ raise RuntimeError("guest 0 write modified guest 1 signal")
+ if index == 0:
+ progress("Guest 1 signal unchanged after host write to guest 0")
+
+ for index, (process, messages, lines, _) in enumerate(guests):
+ offset = offsets[index]
+ wait_reply(runner, obj, fd, offset, process, lines)
+ progress(f"Round 1: guest {index} -> host, "
+ f"verified value {GUEST_FIRST:#018x}")
+
+ # Guest 0 exits while guest 1 keeps its slice mapped and active.
+ for index, (process, messages, lines, _) in enumerate(guests):
+ offset = offsets[index]
+ progress(f"Round 2: host -> guest {index}, value {HOST_SECOND:#018x}")
+ if run_host_bpf(runner, obj, fd, offset) != 0:
+ raise RuntimeError(f"host BPF failed guest {index} second value")
+ wait_reply(runner, obj, fd, offset, process, lines)
+ progress(f"Round 2: guest {index} -> host, "
+ f"verified value {GUEST_SECOND:#018x}")
+ if run_host_bpf(runner, obj, fd, offset) != 2:
+ raise RuntimeError(f"host BPF failed guest {index} reply")
+ wait_message(messages, lines, "GUEST_BPF_EXCHANGED", process)
+ process.wait(timeout=15)
+ if process.returncode:
+ raise RuntimeError(f"guest {index} failed:\n" + "".join(lines))
+ progress(f"Guest {index} completed and exited successfully")
+ if index == 0:
+ progress("Continuing with guest 1 after guest 0 exits")
+ print(f"OK: two guests exchanged BPF values through arena slices {offsets}")
+ finally:
+ progress("Cleaning up guest processes")
+ stop_guests(guests)
+
+
+def skip(reason):
+ print(f"SKIP: {reason}")
+ return KSFT_SKIP
+
+
+def create_arena(pin, build):
+ parent, child = socket.socketpair()
+ with parent, child:
+ result = subprocess.run(
+ [str(build / "arena_kvm_host-runner"), "--create", pin,
+ str(child.fileno()), str(NUM_PAGES)], pass_fds=(child.fileno(),),
+ timeout=HELPER_TIMEOUT)
+ child.close()
+ if result.returncode == KSFT_SKIP:
+ return None
+ result.check_returncode()
+ fds = array.array("i")
+ _, ancillary, flags, _ = parent.recvmsg(
+ 1, socket.CMSG_SPACE(fds.itemsize))
+ for level, kind, data in ancillary:
+ if level == socket.SOL_SOCKET and kind == socket.SCM_RIGHTS:
+ fds.frombytes(data[:len(data) - len(data) % fds.itemsize])
+ if flags & socket.MSG_CTRUNC or len(fds) != 1:
+ for fd in fds:
+ os.close(fd)
+ raise RuntimeError("host helper did not send one arena FD")
+ return fds[0]
+
+
+def main():
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--build-dir", type=Path,
+ help="directory containing the built helpers")
+ parser.add_argument("--kernel", default=os.environ.get("ARENA_KVM_KERNEL"),
+ help="virtme-ng kernel path or release (default: source "
+ "tree kernel, otherwise running kernel)")
+ args = parser.parse_args()
+ if not hasattr(os, "pidfd_open") or not hasattr(signal, "pidfd_send_signal"):
+ return skip("Python 3.9+ with Linux pidfd support is required")
+ machine = os.uname().machine
+ if machine not in GUEST_RAM_BASES:
+ return skip(f"unknown guest RAM layout for architecture {machine}")
+ if struct.calcsize("P") != 8:
+ return skip("BPF arenas require a supported 64-bit architecture")
+ if os.geteuid() != 0:
+ return skip("this test requires root")
+ if SLICE_SIZE % PAGE_SIZE:
+ return skip("arena slices must contain a whole number of pages")
+ if not os.path.ismount("/sys/fs/bpf"):
+ return skip("mount bpffs first")
+ directory = Path(__file__).resolve().parent
+ build = args.build_dir
+ if build is None:
+ build = Path.cwd() if (Path.cwd() / "arena_kvm_host-runner").is_file() \
+ else directory
+ build = build.resolve()
+ kernel = args.kernel
+ if kernel is None:
+ kernel = next((str(p) for p in directory.parents
+ if (p / "vmlinux").is_file()), os.uname().release)
+ vng, env = resolve_vng()
+ if not vng:
+ return KSFT_SKIP
+ qemu_arch = {"arm64": "aarch64", "ppc64le": "ppc64"}.get(machine, machine)
+ for tool in ("script", f"qemu-system-{qemu_arch}"):
+ if not shutil.which(tool, path=env.get("PATH")):
+ return skip(f"{tool} is unavailable")
+ for name in ("arena_kvm_guest-init", "arena_kvm_guest.bpf.o",
+ "arena_kvm_host-runner", "arena_kvm_host.bpf.o"):
+ if not (build / name).exists():
+ return skip(f"missing {build / name}; build the BPF selftests")
+ for name in ("arena_kvm_host.bpf.o", "arena_kvm_guest.bpf.o"):
+ result = subprocess.run([str(build / "arena_kvm_host-runner"),
+ "--check-bpf", str(build / name)],
+ timeout=HELPER_TIMEOUT)
+ if result.returncode == KSFT_SKIP:
+ return skip("BPF compiler lacks arena address-space conversions")
+ result.check_returncode()
+ if subprocess.run([str(build / "arena_kvm_host-runner"),
+ "--check-kvm"], timeout=HELPER_TIMEOUT).returncode:
+ return skip("KVM is unavailable or its API version is unsupported")
+ pin = f"/sys/fs/bpf/arena_kvm_{os.getpid()}"
+ progress(f"Architecture: {machine}; "
+ f"page size: {PAGE_SIZE} bytes")
+ progress(f"Creating and pinning a {ARENA_SIZE // (1024 * 1024)} MiB "
+ f"host arena at {pin}")
+ try:
+ fd = create_arena(pin, build)
+ if fd is None:
+ return skip("pinned BPF arenas are unsupported")
+ try:
+ if os.stat(pin).st_size != ARENA_SIZE:
+ raise RuntimeError("pinned arena has the wrong size")
+ with mmap.mmap(fd, ARENA_SIZE, flags=mmap.MAP_SHARED,
+ prot=mmap.PROT_READ | mmap.PROT_WRITE) as arena:
+ progress(f"Populating {NUM_PAGES} host arena pages before "
+ "starting the guests")
+ # Populate the file pages before QEMU uses them as RAM.
+ for i in range(NUM_PAGES):
+ arena[i * PAGE_SIZE]
+ exchange(pin, fd, arena, build, kernel, vng, env)
+ finally:
+ os.close(fd)
+ finally:
+ if os.path.exists(pin):
+ os.unlink(pin)
+
+
+if __name__ == "__main__":
+ enable_subreaper()
+ signal.signal(signal.SIGTERM, terminate)
+ signal.signal(signal.SIGINT, terminate)
+ try:
+ status = main()
+ except TestSkipped as err:
+ status = skip(str(err))
+ except TestTerminated as err:
+ progress(f"Terminated by signal {err.signum}")
+ status = 128 + err.signum
+ finally:
+ signal.signal(signal.SIGTERM, signal.SIG_IGN)
+ signal.signal(signal.SIGINT, signal.SIG_IGN)
+ stop_children()
+ sys.exit(status)
diff --git a/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c b/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c
new file mode 100644
index 0000000000000..461562149c57f
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_guest.bpf.c
@@ -0,0 +1,89 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+#include "bpf_arena_common.h"
+#ifdef __TARGET_ARCH_powerpc
+/* PowerPC does not support arena load-acquire/store-release instructions. */
+#undef __BPF_FEATURE_LOAD_ACQ_STORE_REL
+#endif
+#include "bpf_atomic.h"
+#endif
+#include "arena_kvm_shared.h"
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARENA);
+ __uint(map_flags, BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE);
+ __uint(max_entries, 1);
+} arena SEC(".maps");
+
+volatile __u32 signal_offset;
+
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+
+SEC("syscall")
+int allocate(void *ctx)
+{
+ struct signal_page __arena *page;
+
+ page = bpf_arena_alloc_pages(&arena, NULL, 1, ARENA_KVM_NODE, 0);
+ if (!page)
+ return 1;
+ signal_offset = (__u32)((__u64)page - (__u64)arena_base(&arena));
+ smp_store_release(&page->ready, GUEST_READY);
+ return 0;
+}
+
+SEC("syscall")
+int probe_free(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ bpf_arena_free_pages(&arena, page, 1);
+ return page->ready == GUEST_READY ? 0 : 1;
+}
+
+SEC("syscall")
+int exchange(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = smp_load_acquire(&page->h2g_seq);
+ __u64 payload;
+
+ if ((seq != 1 && seq != 2) || page->g2h_seq == seq)
+ return 1;
+ payload = page->h2g_payload;
+ if (seq == 1 && payload != HOST_FIRST)
+ return 2;
+ if (seq == 2 && payload != HOST_SECOND)
+ return 3;
+ page->g2h_payload = seq == 1 ? GUEST_FIRST : GUEST_SECOND;
+ smp_store_release(&page->g2h_seq, seq);
+ return 0;
+}
+
+SEC("syscall")
+int complete(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ return smp_load_acquire(&page->h2g_seq) == 3 ? 0 : 1;
+}
+
+#else
+SEC("syscall") int allocate(void *ctx) { return 1; }
+SEC("syscall") int probe_free(void *ctx) { return 1; }
+SEC("syscall") int exchange(void *ctx) { return 1; }
+SEC("syscall") int complete(void *ctx) { return 1; }
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/arena_kvm_guest.c b/tools/testing/selftests/bpf/arena_kvm_guest.c
new file mode 100644
index 0000000000000..da6586ae660d1
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_guest.c
@@ -0,0 +1,224 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ *
+ * Guest helper for a BPF arena allocated from shared NUMA RAM.
+ */
+#define __EXPORTED_HEADERS__
+
+#include <bpf/bpf.h>
+#include <bpf/libbpf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <linux/bpf.h>
+#include <linux/mempolicy.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/mman.h>
+#include <sys/mount.h>
+#include <sys/syscall.h>
+#include <unistd.h>
+
+#include "arena_kvm_shared.h"
+
+static int signal_file_offset(void *arena, uint64_t *offset)
+{
+ char cmdline[8192], *arg, *end;
+ uint64_t entry, pfn, gpa, base = 0, size = 0, *value;
+ ssize_t n;
+ int fd, node, found = 0;
+
+ /* The launcher supplies QEMU's guest physical base for this slice. */
+ fd = open("/proc/cmdline", O_RDONLY);
+ if (fd < 0)
+ return -errno;
+ n = read(fd, cmdline, sizeof(cmdline) - 1);
+ close(fd);
+ if (n <= 0 || n == sizeof(cmdline) - 1)
+ return -EIO;
+ cmdline[n] = '\0';
+ for (arg = strtok(cmdline, "\n "); arg; arg = strtok(NULL, "\n ")) {
+ if (!strncmp(arg, "arena_kvm.base=", 15)) {
+ value = &base;
+ found |= 1;
+ } else if (!strncmp(arg, "arena_kvm.size=", 15)) {
+ value = &size;
+ found |= 2;
+ } else {
+ continue;
+ }
+ errno = 0;
+ *value = strtoull(arg + 15, &end, 0);
+ if (errno || *end || end == arg + 15 || *value % getpagesize())
+ return -EINVAL;
+ }
+ if (found != 3 || size < getpagesize())
+ return -EINVAL;
+
+ /* Fault the BPF-allocated page into the guest user VMA. */
+ if (!*(volatile uint64_t *)arena)
+ return -EINVAL;
+ if (syscall(SYS_get_mempolicy, &node, NULL, 0, arena,
+ MPOL_F_NODE | MPOL_F_ADDR))
+ return -errno;
+ if (node != ARENA_KVM_NODE)
+ return -EXDEV;
+ fd = open("/proc/self/pagemap", O_RDONLY);
+ if (fd < 0)
+ return -errno;
+ n = pread(fd, &entry, sizeof(entry),
+ (uintptr_t)arena / getpagesize() * sizeof(entry));
+ close(fd);
+ if (n != sizeof(entry) || !(entry & (1ULL << 63)))
+ return -EIO;
+ pfn = entry & ((1ULL << 55) - 1);
+ if (!pfn)
+ return -EPERM;
+ gpa = pfn * getpagesize();
+ if (gpa < base || gpa - base >= size || size - (gpa - base) < getpagesize())
+ return -ERANGE;
+ *offset = gpa - base;
+ return 0;
+}
+
+static int create_arena(void)
+{
+ union bpf_attr attr = {
+ .map_type = BPF_MAP_TYPE_ARENA,
+ .max_entries = 1,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE,
+ };
+
+ return syscall(__NR_bpf, BPF_MAP_CREATE, &attr, sizeof(attr));
+}
+
+static int run_prog(int fd, unsigned int *retval)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ int err = bpf_prog_test_run_opts(fd, &opts);
+
+ if (err)
+ return err;
+ *retval = opts.retval;
+ return 0;
+}
+
+static int exchange_once(int fd)
+{
+ unsigned int result = 0;
+ int err;
+
+ /* The host bounds each protocol stage and stops the guest on timeout. */
+ for (;;) {
+ err = run_prog(fd, &result);
+ if (err || result > 1)
+ return err ? err : -EINVAL;
+ if (!result)
+ return 0;
+ usleep(1000);
+ }
+}
+
+int main(int argc, char **argv)
+{
+ struct bpf_object *obj = NULL;
+ struct bpf_program *alloc, *probe_free, *exchange, *complete;
+ struct bpf_map *map;
+ void *arena = MAP_FAILED;
+ uint64_t file_offset;
+ unsigned int result = 0;
+ int fd = -1;
+ int err;
+
+ setbuf(stdout, NULL);
+ if (argc != 2 ||
+ (mount("sysfs", "/sys", "sysfs", 0, NULL) && errno != EBUSY) ||
+ (mount("proc", "/proc", "proc", 0, NULL) && errno != EBUSY))
+ goto fail;
+ if (access("/sys/devices/system/node/node1", F_OK)) {
+ fputs("guest NUMA node 1 is unavailable\n", stderr);
+ goto fail;
+ }
+ fd = create_arena();
+ if (fd < 0) {
+ perror("create guest arena");
+ goto fail;
+ }
+ arena = mmap(NULL, getpagesize(), PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
+ if (arena == MAP_FAILED) {
+ perror("mmap guest arena");
+ goto fail;
+ }
+ obj = bpf_object__open_file(argv[1], NULL);
+ if (!obj)
+ goto fail;
+ err = arena_kvm_check_features(obj);
+ if (err == 4) {
+ puts("GUEST_BPF_SKIPPED: compiler lacks arena address-space casts");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ close(fd);
+ return 4;
+ }
+ if (err)
+ goto fail;
+ map = bpf_object__find_map_by_name(obj, "arena");
+ alloc = bpf_object__find_program_by_name(obj, "allocate");
+ probe_free = bpf_object__find_program_by_name(obj, "probe_free");
+ exchange = bpf_object__find_program_by_name(obj, "exchange");
+ complete = bpf_object__find_program_by_name(obj, "complete");
+ if (!map || !alloc || !probe_free || !exchange || !complete ||
+ bpf_map__reuse_fd(map, fd))
+ goto fail;
+ close(fd);
+ fd = -1;
+ err = bpf_object__load(obj);
+ if (err) {
+ fprintf(stderr, "load guest BPF: %d\n", err);
+ goto fail;
+ }
+ err = run_prog(bpf_program__fd(alloc), &result);
+ if (err || result) {
+ fprintf(stderr, "allocate guest page: err=%d result=%u\n",
+ err, result);
+ goto fail;
+ }
+ err = run_prog(bpf_program__fd(probe_free), &result);
+ if (err || result) {
+ fprintf(stderr, "guest page lifetime check: err=%d result=%u\n",
+ err, result);
+ goto fail;
+ }
+ err = signal_file_offset(arena, &file_offset);
+ if (err == -EXDEV) {
+ puts("GUEST_BPF_SKIPPED: allocation fell back from shared NUMA node");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ return 4;
+ }
+ if (err) {
+ fprintf(stderr, "resolve shared page offset: %d\n", err);
+ goto fail;
+ }
+ printf("GUEST_BPF_READY %llu %d\n",
+ (unsigned long long)file_offset, getpagesize());
+ if (exchange_once(bpf_program__fd(exchange)) ||
+ exchange_once(bpf_program__fd(exchange)) ||
+ exchange_once(bpf_program__fd(complete)))
+ goto fail;
+ puts("GUEST_BPF_EXCHANGED");
+ bpf_object__close(obj);
+ munmap(arena, getpagesize());
+ return 0;
+
+fail:
+ puts("GUEST_BPF_FAILED");
+ bpf_object__close(obj);
+ if (arena != MAP_FAILED)
+ munmap(arena, getpagesize());
+ if (fd >= 0)
+ close(fd);
+ return 1;
+}
diff --git a/tools/testing/selftests/bpf/arena_kvm_host.bpf.c b/tools/testing/selftests/bpf/arena_kvm_host.bpf.c
new file mode 100644
index 0000000000000..9c1c535fe429b
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_host.bpf.c
@@ -0,0 +1,89 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <vmlinux.h>
+#include <bpf/bpf_helpers.h>
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+#include "bpf_arena_common.h"
+#ifdef __TARGET_ARCH_powerpc
+/* PowerPC does not support arena load-acquire/store-release instructions. */
+#undef __BPF_FEATURE_LOAD_ACQ_STORE_REL
+#endif
+#include "bpf_atomic.h"
+#endif
+#include "arena_kvm_shared.h"
+
+struct {
+ __uint(type, BPF_MAP_TYPE_ARENA);
+ __uint(map_flags, BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE | BPF_F_ARENA_EXPORT);
+ /* The host runner reuses the map created with the selected capacity. */
+ __uint(max_entries, 1);
+} arena SEC(".maps");
+
+volatile __u32 signal_offset;
+
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+
+SEC("syscall")
+int exchange(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = page->h2g_seq;
+
+ if (smp_load_acquire(&page->ready) != GUEST_READY)
+ return 4;
+ if (!seq) {
+ page->h2g_payload = HOST_FIRST;
+ smp_store_release(&page->h2g_seq, 1);
+ return 0;
+ }
+ if (seq == 1 && smp_load_acquire(&page->g2h_seq) == 1) {
+ if (page->g2h_payload != GUEST_FIRST)
+ return 3;
+ page->h2g_payload = HOST_SECOND;
+ smp_store_release(&page->h2g_seq, 2);
+ return 0;
+ }
+ if (seq == 2 && smp_load_acquire(&page->g2h_seq) == 2) {
+ if (page->g2h_payload != GUEST_SECOND)
+ return 3;
+ smp_store_release(&page->h2g_seq, 3);
+ return 2;
+ }
+ return 1;
+}
+
+SEC("syscall")
+int check_reply(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+ __u64 seq = page->h2g_seq;
+
+ if ((seq != 1 && seq != 2) || smp_load_acquire(&page->g2h_seq) != seq)
+ return 1;
+ return page->g2h_payload == (seq == 1 ? GUEST_FIRST : GUEST_SECOND) ? 0 : 3;
+}
+
+SEC("syscall")
+int probe_free(void *ctx)
+{
+ struct signal_page __arena *page =
+ (struct signal_page __arena *)((char __arena *)arena_base(&arena) +
+ signal_offset);
+
+ bpf_arena_free_pages(&arena, page, 1);
+ return page->ready == GUEST_READY ? 0 : 1;
+}
+
+#else
+SEC("syscall") int exchange(void *ctx) { return 1; }
+SEC("syscall") int check_reply(void *ctx) { return 1; }
+SEC("syscall") int probe_free(void *ctx) { return 1; }
+#endif
+
+char _license[] SEC("license") = "GPL";
diff --git a/tools/testing/selftests/bpf/arena_kvm_host.c b/tools/testing/selftests/bpf/arena_kvm_host.c
new file mode 100644
index 0000000000000..90a77a54bd6f4
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_host.c
@@ -0,0 +1,147 @@
+// SPDX-License-Identifier: GPL-2.0
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#include <bpf/bpf.h>
+#include <bpf/libbpf.h>
+#include <errno.h>
+#include <fcntl.h>
+#include <linux/kvm.h>
+#include <limits.h>
+#include <stdint.h>
+#include <stdio.h>
+#include <stdlib.h>
+#include <string.h>
+#include <sys/ioctl.h>
+#include <sys/socket.h>
+#include <unistd.h>
+
+#include "arena_kvm_shared.h"
+
+static int create_arena(const char *pin, int socket_fd, unsigned int pages)
+{
+ LIBBPF_OPTS(bpf_map_create_opts, opts,
+ .map_flags = BPF_F_MMAPABLE | BPF_F_ARENA_NO_FREE |
+ BPF_F_ARENA_EXPORT);
+ char control[CMSG_SPACE(sizeof(int))] = {};
+ char byte = 0;
+ struct iovec iov = { .iov_base = &byte, .iov_len = 1 };
+ struct msghdr msg = {
+ .msg_iov = &iov,
+ .msg_iovlen = 1,
+ .msg_control = control,
+ .msg_controllen = sizeof(control),
+ };
+ struct cmsghdr *cmsg = CMSG_FIRSTHDR(&msg);
+ int fd, err = 1;
+
+ fd = bpf_map_create(BPF_MAP_TYPE_ARENA, "arena_kvm", 0, 0,
+ pages, &opts);
+ if (fd < 0) {
+ if (errno == EOPNOTSUPP || errno == EINVAL)
+ return 4; /* KSFT_SKIP: arena type or flags unsupported. */
+ perror("create host arena");
+ return 1;
+ }
+ if (bpf_obj_pin(fd, pin)) {
+ perror("pin host arena");
+ goto out;
+ }
+ cmsg->cmsg_level = SOL_SOCKET;
+ cmsg->cmsg_type = SCM_RIGHTS;
+ cmsg->cmsg_len = CMSG_LEN(sizeof(fd));
+ memcpy(CMSG_DATA(cmsg), &fd, sizeof(fd));
+ if (sendmsg(socket_fd, &msg, 0) == 1)
+ err = 0;
+ else
+ perror("send arena FD");
+out:
+ close(fd);
+ return err;
+}
+
+int main(int argc, char **argv)
+{
+ LIBBPF_OPTS(bpf_test_run_opts, opts);
+ struct bpf_object *obj = NULL;
+ struct bpf_program *prog;
+ struct bpf_map *arena, *bss;
+ uint32_t key = 0, offset;
+ unsigned long parsed;
+ int fd, err;
+ char *end;
+
+ if (argc == 3 && !strcmp(argv[1], "--check-bpf")) {
+ obj = bpf_object__open_file(argv[2], NULL);
+ if (!obj)
+ return 1;
+ err = arena_kvm_check_features(obj);
+ bpf_object__close(obj);
+ return err;
+ }
+
+ if (argc == 2 && !strcmp(argv[1], "--check-kvm")) {
+ fd = open("/dev/kvm", O_RDWR);
+ if (fd < 0)
+ return 4;
+ err = ioctl(fd, KVM_GET_API_VERSION, 0);
+ close(fd);
+ return err == KVM_API_VERSION ? 0 : 4;
+ }
+ if (argc == 5 && !strcmp(argv[1], "--create")) {
+ errno = 0;
+ parsed = strtoul(argv[3], &end, 0);
+ if (errno || *end || end == argv[3] || parsed > INT_MAX)
+ return 1;
+ fd = parsed;
+ errno = 0;
+ parsed = strtoul(argv[4], &end, 0);
+ if (errno || *end || end == argv[4] || !parsed || parsed > UINT_MAX)
+ return 1;
+ return create_arena(argv[2], fd, parsed);
+ }
+ if (argc != 5)
+ return 1;
+ errno = 0;
+ parsed = strtoul(argv[1], &end, 0);
+ if (errno || *end || parsed > INT_MAX)
+ return 1;
+ fd = dup(parsed);
+ if (fd < 0)
+ return 1;
+ errno = 0;
+ parsed = strtoul(argv[3], &end, 0);
+ if (errno || *end || parsed > UINT_MAX || parsed % getpagesize())
+ goto fail;
+ offset = parsed;
+ obj = bpf_object__open_file(argv[2], NULL);
+ if (!obj)
+ goto fail;
+ arena = bpf_object__find_map_by_name(obj, "arena");
+ bss = bpf_object__find_map_by_name(obj, ".bss");
+ prog = bpf_object__find_program_by_name(obj, argv[4]);
+ if (!arena || !bss || !prog || bpf_map__reuse_fd(arena, fd))
+ goto fail;
+ if (offset >= (uint64_t)bpf_map__max_entries(arena) * getpagesize())
+ goto fail;
+ if (bpf_object__load(obj)) {
+ fputs("load host BPF failed\n", stderr);
+ goto fail;
+ }
+ if (bpf_map_update_elem(bpf_map__fd(bss), &key, &offset, BPF_ANY)) {
+ perror("set signal offset");
+ goto fail;
+ }
+ err = bpf_prog_test_run_opts(bpf_program__fd(prog), &opts);
+ if (err)
+ goto fail;
+ printf("%u\n", opts.retval);
+ bpf_object__close(obj);
+ close(fd);
+ return 0;
+
+fail:
+ bpf_object__close(obj);
+ close(fd);
+ return 1;
+}
diff --git a/tools/testing/selftests/bpf/arena_kvm_shared.h b/tools/testing/selftests/bpf/arena_kvm_shared.h
new file mode 100644
index 0000000000000..63aafac25efc3
--- /dev/null
+++ b/tools/testing/selftests/bpf/arena_kvm_shared.h
@@ -0,0 +1,60 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+/*
+ * Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES
+ */
+#ifndef ARENA_KVM_SHARED_H
+#define ARENA_KVM_SHARED_H
+
+/* These flags may be absent from a pre-series kernel's vmlinux.h. */
+#ifndef BPF_F_ARENA_NO_FREE
+#define BPF_F_ARENA_NO_FREE (1U << 20)
+#endif
+#ifndef BPF_F_ARENA_EXPORT
+#define BPF_F_ARENA_EXPORT (1U << 21)
+#endif
+
+struct arena_kvm_features {
+ unsigned int addr_space_cast;
+};
+
+#ifdef __BPF__
+#ifdef __BPF_FEATURE_ADDR_SPACE_CAST
+const volatile struct arena_kvm_features features = { .addr_space_cast = 1 };
+#else
+const volatile struct arena_kvm_features features = {};
+#endif
+#else
+/* Inspect compiler support before loading maps or booting either guest. */
+static inline int arena_kvm_check_features(struct bpf_object *obj)
+{
+ const struct arena_kvm_features *features;
+ struct bpf_map *map;
+ size_t size;
+
+ map = bpf_object__find_map_by_name(obj, ".rodata");
+ if (!map)
+ return 1;
+ features = bpf_map__initial_value(map, &size);
+ if (!features || size != sizeof(*features))
+ return 1;
+ return features->addr_space_cast ? 0 : 4; /* KSFT_SKIP */
+}
+#endif
+
+#define ARENA_KVM_NODE 1
+
+#define GUEST_READY 0x4152454e414b564dULL
+#define HOST_FIRST 0x123456789abcdef0ULL
+#define GUEST_FIRST 0xfedcba9876543210ULL
+#define HOST_SECOND 0x1020304050607080ULL
+#define GUEST_SECOND 0x8070605040302010ULL
+
+struct signal_page {
+ volatile unsigned long long ready;
+ volatile unsigned long long h2g_payload;
+ volatile unsigned long long h2g_seq;
+ volatile unsigned long long g2h_payload;
+ volatile unsigned long long g2h_seq;
+};
+
+#endif
--
2.56.0
^ permalink raw reply [flat|nested] 10+ messages in thread* Re: [PATCH 5/5] selftests/bpf: Test two-way arena access with two guests
2026-10-10 22:46 ` [PATCH 5/5] selftests/bpf: Test two-way arena access with two guests Andrea Righi
@ 2026-10-10 23:40 ` bot+bpf-ci
0 siblings, 0 replies; 10+ messages in thread
From: bot+bpf-ci @ 2026-10-10 23:40 UTC (permalink / raw)
To: arighi, ast, daniel, andrii, eddyz87, memxor
Cc: martin.lau, song, yonghong.song, jolsa, emil, ihor.solodrai,
john.fastabend, joshdon, vineethrp, balbirs, skolothumtho,
himadrics, royfru44, bpf, linux-kernel, ast, andrii, daniel,
martin.lau, eddyz87, yonghong.song, mason, ihor.solodrai
[-- Attachment #1: Type: text/plain, Size: 4284 bytes --]
> selftests/bpf: Test two-way arena access with two guests
TL;DR:
- Unconditional -lzstd in the new Makefile link rules may break the whole
selftests/bpf build where libzstd is missing
- arena_kvm.py doesn't use ksft.py reporting (minor; results still counted)
> diff --git a/tools/testing/selftests/bpf/Makefile b/tools/testing/selftests/bpf/Makefile
> index a22be7efd1fa..29fffa7857c1 100644
> --- a/tools/testing/selftests/bpf/Makefile
> +++ b/tools/testing/selftests/bpf/Makefile
[ ... ]
> @@ -94,6 +100,25 @@ BPFTOOLDIR := $(TOOLSDIR)/bpf/bpftool
> HOST_BPFOBJ := $(HOST_BUILD_DIR)/libbpf/libbpf.a
> BPF_TARGET_ENDIAN:=$(if $(IS_LITTLE_ENDIAN),--target=bpfel,--target=bpfeb)
>
> +# These helpers are installed beside arena_kvm.py, including for OUTPUT=.
[ ... ]
> +$(OUTPUT)/arena_kvm_guest-init: arena_kvm_guest.c arena_kvm_shared.h \
> + $(BPFOBJ) | $(OUTPUT)
> + $(call msg,BINARY,,$@)
> + $(Q)$(CC) $(CFLAGS) $(LDFLAGS) $< $(BPFOBJ) $(LDLIBS) -lzstd -o $@
> +
> +$(OUTPUT)/arena_kvm_host-runner: arena_kvm_host.c arena_kvm_shared.h \
> + $(BPFOBJ) | $(OUTPUT)
> + $(call msg,BINARY,,$@)
> + $(Q)$(CC) $(CFLAGS) $(LDFLAGS) $< $(BPFOBJ) $(LDLIBS) -lzstd -o $@
> +
> NON_CHECK_FEAT_TARGETS := clean docs-clean emit_tests
> CHECK_FEAT := $(filter-out $(NON_CHECK_FEAT_TARGETS),$(or $(MAKECMDGOALS), "none"))
> ifneq ($(CHECK_FEAT),)
Is it intended to pass -lzstd unconditionally in these two link rules?
Neither arena_kvm_guest.c nor arena_kvm_host.c includes a zstd header or
calls into zstd, and no other userspace link line in selftests/bpf uses
-lzstd. test_progs, for example, links with plain $(LDLIBS), which is
$(LIBELF_LIBS) -lz -lrt -lpthread.
Elsewhere in tools/, libzstd is optional. bpftool only adds it when the
feature-libelf-zstd probe passes, and libelf builds without zstd are a
supported configuration.
Both binaries are in TEST_GEN_FILES without a CLANG_CPUV4 or feature guard,
and these link rules do not use $(call skip_on_fail,...). On a host or
sysroot that has libelf and zlib development files but no linkable libzstd,
wouldn't 'make -C tools/testing/selftests/bpf' now fail with "cannot find
-lzstd", even with BPF_STRICT_BUILD=0? That would block all of
selftests/bpf, not just this test.
Would dropping -lzstd be enough, since a dynamic libelf pulls in its own
dependencies? If it is needed for a static libelf, it could be made
conditional the same way bpftool does it.
> diff --git a/tools/testing/selftests/bpf/arena_kvm.py b/tools/testing/selftests/bpf/arena_kvm.py
> new file mode 100755
> index 000000000000..552bfcfb80ac
> --- /dev/null
> +++ b/tools/testing/selftests/bpf/arena_kvm.py
> @@ -0,0 +1,448 @@
[ ... ]
> +KSFT_SKIP = 4
[ ... ]
> + print(f"OK: two guests exchanged BPF values through arena slices {offsets}")
[ ... ]
> +def skip(reason):
> + print(f"SKIP: {reason}")
> + return KSFT_SKIP
[ ... ]
> +if __name__ == "__main__":
[ ... ]
> + stop_children()
> + sys.exit(status)
A subsystem guideline flags this as potentially concerning: a test that
prints its own ad hoc result text and exits directly, instead of going
through the kselftest reporting helpers.
arena_kvm.py defines its own KSFT_SKIP, prints hand-written "SKIP: ..." and
"OK: ..." lines, and then calls sys.exit(status) directly.
tools/testing/selftests/kselftest/ksft.py provides KSFT_SKIP,
test_result_pass(), test_result_skip(), test_result_fail(), exit_pass() and
exit_fail(), and none of them are used here, so the output is not KTAP.
Should this script report through ksft.py instead?
For context, when the kselftest runner executes a TEST_PROGS entry, it
prefixes the script's output with '# ' and turns the exit code into the TAP
result line (exit 4 becomes '# SKIP'), so the one result this script
produces is still counted correctly. The other TEST_PROGS scripts in this
directory (test_kmod.sh, test_xsk.sh, test_xdp_features.sh) also print their
own results. The guideline text is written about C test binaries and does
not name the Python helper.
---
AI reviewed your patch. Please fix the bug or email reply why it's not a bug.
See: https://github.com/kernel-patches/vmtest/blob/master/ci/claude/README.md
CI run summary: https://github.com/kernel-patches/bpf/actions/runs/38093513563
^ permalink raw reply [flat|nested] 10+ messages in thread