From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-oa1-f48.google.com (mail-oa1-f48.google.com [209.85.160.48]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1E564428490 for ; Fri, 6 Feb 2026 18:58:10 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.160.48 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770404291; cv=none; b=UIvVISq6rRiK/NetfGLS86p+uyoZa6cbGpD4z1+cDI7i9UvbfoLTn8g1RSjN0z6QVQq6hcYK6+pdeTsa3DyEmoNGk4bI00iogRyg2dlCl8sWNkSJdJoqLWsrV9vrCoTHbYKZmfrIk0qwCZyfDw08GVyayqJEeVfEt3prK0Op9v8= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1770404291; c=relaxed/simple; bh=Ov5+ShbkfbVH1cj72wn0K96kZfFVsGXrUgq7a4ISWlI=; h=Message-ID:Date:MIME-Version:From:Subject:To:Cc:Content-Type; b=MC2/9MrAI2nZfVDX70lmB49CVWPMvUcT6qJ7VZROv2MBXYW0ldw03p9wYPA1+119wNaTsmxC7g40zDsVMdIgN/1Ifrkd4n0za/UF1H/dnpJogb7rSH9EqoKRfwDNz3699zlZCv/1kJGWp67ZIGbEfszWdnSu9euK4+T4TwJ573Y= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk; spf=pass smtp.mailfrom=kernel.dk; dkim=pass (2048-bit key) header.d=kernel-dk.20230601.gappssmtp.com header.i=@kernel-dk.20230601.gappssmtp.com header.b=HTuvgBSb; arc=none smtp.client-ip=209.85.160.48 Authentication-Results: smtp.subspace.kernel.org; dmarc=none (p=none dis=none) header.from=kernel.dk Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=kernel.dk Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel-dk.20230601.gappssmtp.com header.i=@kernel-dk.20230601.gappssmtp.com header.b="HTuvgBSb" Received: by mail-oa1-f48.google.com with SMTP id 586e51a60fabf-4044d3ff57bso409455fac.0 for ; Fri, 06 Feb 2026 10:58:10 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=kernel-dk.20230601.gappssmtp.com; s=20230601; t=1770404290; x=1771009090; darn=vger.kernel.org; h=content-transfer-encoding:content-language:cc:to:subject:from :user-agent:mime-version:date:message-id:from:to:cc:subject:date :message-id:reply-to; bh=NFayfXFCr8uTtagrI+irxnDsr7Q6K8DeRpjr+PMPA0A=; b=HTuvgBSbr82UG3wGIbG7czDW2LIii11nrWrrvshU3YPaGv+dxrMrpQ5FYO25XHc/e6 nBcQYYIstZVaAq8FhG3gMOsyUWWZDxl7b2mwp+XmaFDQ970JYXMhBKOcJy1Pw7oIBzoL Evquc/FT4FPbh6IbSN2uwqgfs7/hFxFHazbz5sfQ/DtxcNQj2pUfPAAUkazojt9glafn iQ2s5q4C54+6AlCE+unI/O5YZoUXVkf3K635cAbWQR6nl2XyE4qIw1TfgWbiDfwuwSM+ Aew5+t58v1+IUh0ToQbWzArbRuv0Ng1FZD8DwnqWBk84ErU8TwMCgYY3tZWyuI9AADcG lQmw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1770404290; x=1771009090; h=content-transfer-encoding:content-language:cc:to:subject:from :user-agent:mime-version:date:message-id:x-gm-gg:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=NFayfXFCr8uTtagrI+irxnDsr7Q6K8DeRpjr+PMPA0A=; b=UonoEG0xBNidl4F/iVXoBJUl4YBlKVF+K+ysX3FkvMb3P4HH5o1JSlNiuTa8A2HCMb LszQ536Ba4ZBO8nIXIyW6vr+UXB4q6a6kMnxWFae/cCXw6earHcQOXZP4cpv0boT/2KJ i8LqkrLI+t5JjvPo9/YOMbkY+3SPjsKQbHQr0fx0TW/096I57tRdT8vyHesEK6JI3Tlr u/FmIHOpOnQPeV0xspPbIXJwbWtpY+c43Uzwk2TcsNq0p7J7cUgG/R9y7QRrT+L0XPPw vJe0iob91G7FOBvoV26Ke+koq+uTMw3xgO3zinuf2jS3rdG4afejnzlW1+CyPH2hhAw+ cyJw== X-Forwarded-Encrypted: i=1; AJvYcCXUtqpP8JEtfsHFdHef8TZF5PkhL0qeeWVrwH/PuWfPyRfM3vHKPidh5tIjtU0OdUMrTUPI4H0TkETDryE=@vger.kernel.org X-Gm-Message-State: AOJu0Yy/QuEwTk+WnNyHNAs4ofOnd6GZ8+WWfMeJIEBvHH9Da1Gt5NU3 I79weISAMRrY8rClvmxb7D80yLOg8J8qjYn8Oy0MsadWi8g6HoWYD1a9k7y+sc29sXQ= X-Gm-Gg: AZuq6aK18e0nzZtXa47b9pNwi5OX2d6OSKR2IZu9p3XkJW1H1rJMUFfmoJ8CbwykX3P HI1D28fmJZ1zOci09C9nA4uZbIdVfa3Zc2XpTUZlwm0/WTzdyTbF6YrpGnHhQXTyq9ZxfCWG+lS Pf01JUTmLJLsD4Hbu44xZrtZ6TVhHDtsi9tFRA2IRTH0rOP7ELUiVuzQ24Bj2dgiGDjhPSL3YLW Dc+PrvbJHm/jYGCT9prkNeeAPOllSVOzLNeGsmOQPsEzfASi7DFPVr27/f5UoKY1hukZ6VXA7hF L0tFVxzPDcyXH4UAsPpAWXP5Ns1LauYM8SwQYlCexPHZ+kmVIgE0CDNRPX51sy9CxdSGZEdLjwd DLCeueDe/wzChQdMVGjq48wWcHR4sSo9T6BQKqgU0k09BQKMjEwk6TUJA6NJHBd+JfWCZZkb3AR awAOuP6s/YIkplI8wM8DeM+i1OD9wj239NgNIk96vkSnoSxgPsbpV/jraiCrrYdUHkJlKW X-Received: by 2002:a05:6870:8dcd:b0:3ec:4f18:9c79 with SMTP id 586e51a60fabf-40a96ca7334mr2072534fac.13.1770404289918; Fri, 06 Feb 2026 10:58:09 -0800 (PST) Received: from [192.168.1.102] ([96.43.243.2]) by smtp.gmail.com with ESMTPSA id 586e51a60fabf-40a99787786sm2432762fac.19.2026.02.06.10.58.08 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Fri, 06 Feb 2026 10:58:09 -0800 (PST) Message-ID: Date: Fri, 6 Feb 2026 11:58:08 -0700 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird From: Jens Axboe Subject: [GIT PULL] io_uring cBPF filter support To: Linus Torvalds Cc: io-uring , LKML , Christian Brauner Content-Language: en-US Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit Hi Linus, On top of the core io_uring changes, this adds support for both cBPF filters for io_uring, as well as task inherited restrictions and filters. seccomp and io_uring don't play along nicely, as most of the interesting data to filter on resides somewhat out-of-band, in the submission queue ring. As a result, things like containers and systemd that apply seccomp filters, can't filter io_uring operations. That leaves them with just one choice if filtering is critical - filter the actual io_uring_setup(2) system call to simply disallow io_uring. That's rather unfortunate, and has limited us because of it. io_uring already has some filtering support. It requires the ring to be setup in a disabled state, and then a filter set can be applied. This filter set is completely bi-modal - an opcode is either enabled or it's not. Once a filter set is registered, the ring can be enabled. This is very restrictive, and it's not useful at all to systemd or containers which really want both broader and more specific control. This patchset first adds support for cBPF filters for opcodes, which enables tighter control over what exactly a specific opcode may do. As examples, specific support is added for IORING_OP_OPENAT/OPENAT2, allowing filtering on resolve flags. And another example is added for IORING_OP_SOCKET, allowing filtering on domain/type/protocol. These are both common use cases. cBPF was chosen rather than eBPF, because the latter is often restricted in containers as well. These filters are run post the init phase of the request, which allows filters to even dip into data that is being passed in struct in user memory, as the init side of requests make that data stable by bringing it into the kernel. This allows filtering without needing to copy this data twice, or have filters etc know about the exact layout of the user data. The filters get the already copied and sanitized data passed. On top of that support is added for per-task filters, meaning that any ring created with a task that has a per-task filter will get those filters applied when it's created. These filters are inherited across fork as well. Once a filter has been registered, any further added filters may only further restrict what operations are permitted. Filters cannot change the return value of an operation, they can only permit or deny it based on the contents. Please pull! The following changes since commit 0105b0562a5ed6374f06e5cd4246a3f1311a65a0: io_uring: split out CQ waiting code into wait.c (2026-01-22 09:21:16 -0700) are available in the Git repository at: https://git.kernel.org/pub/scm/linux/kernel/git/axboe/linux.git tags/io_uring-bpf-restrictions.4-20260206 for you to fetch changes up to ed82f35b926b2e505c14b7006473614b8f58b4f4: io_uring: allow registration of per-task restrictions (2026-02-06 07:29:19 -0700) ---------------------------------------------------------------- io_uring-bpf-restrictions.4-20260206 ---------------------------------------------------------------- Jens Axboe (7): io_uring: add support for BPF filtering for opcode restrictions io_uring/net: allow filtering on IORING_OP_SOCKET data io_uring/bpf_filter: allow filtering on contents of struct open_how io_uring/bpf_filter: cache lookup table in ctx->bpf_filters io_uring/bpf_filter: add ref counts to struct io_bpf_filter io_uring: add task fork hook io_uring: allow registration of per-task restrictions include/linux/io_uring.h | 14 +- include/linux/io_uring_types.h | 13 + include/linux/sched.h | 1 + include/uapi/linux/io_uring.h | 10 + include/uapi/linux/io_uring/bpf_filter.h | 62 +++++ io_uring/Kconfig | 5 + io_uring/Makefile | 1 + io_uring/bpf_filter.c | 430 +++++++++++++++++++++++++++++++ io_uring/bpf_filter.h | 48 ++++ io_uring/io_uring.c | 48 ++++ io_uring/io_uring.h | 1 + io_uring/net.c | 9 + io_uring/net.h | 6 + io_uring/openclose.c | 9 + io_uring/openclose.h | 3 + io_uring/register.c | 91 +++++++ io_uring/tctx.c | 42 ++- kernel/fork.c | 6 + 18 files changed, 789 insertions(+), 10 deletions(-) create mode 100644 include/uapi/linux/io_uring/bpf_filter.h create mode 100644 io_uring/bpf_filter.c create mode 100644 io_uring/bpf_filter.h -- Jens Axboe