From: "Serge E. Hallyn" <serge@hallyn.com>
To: Josef Bacik <josef@toxicpanda.com>
Cc: Paul Moore <paul@paul-moore.com>,
Christian Brauner <brauner@kernel.org>,
James Morris <jmorris@namei.org>,
David Howells <dhowells@redhat.com>,
Jarkko Sakkinen <jarkko@kernel.org>,
"Andrew G. Morgan" <morgan@kernel.org>,
Serge Hallyn <sergeh@kernel.org>,
linux-security-module@vger.kernel.org,
linux-kernel@vger.kernel.org, keyrings@vger.kernel.org,
linux-fsdevel@vger.kernel.org
Subject: Re: [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
Date: Thu, 8 Oct 2026 08:57:56 -0500 [thread overview]
Message-ID: <asehZEWAwsJcbqvM@hallyn.com> (raw)
In-Reply-To: <20261006-b4-setfcap-userns-v1-0-f47e7ed66072@toxicpanda.com>
On Tue, Oct 06, 2026 at 03:44:16PM +0000, Josef Bacik wrote:
> Hello,
>
> Commit db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
> stops a root task that has given up CAP_SETFCAP from creating a user
> namespace that maps uid 0 and then writing file capabilities in it that
> the initial namespace honours. The check only looks at the task that
> creates the namespace, so if somebody who did have CAP_SETFCAP created
> one, there are still several ways around it:
>
> - setns() into their namespace and write uid_map from inside (patch 2)
> - write "0 0 1" through a uid_map fd that they opened (patch 3)
> - setns() into a namespace of theirs that already maps uid 0 and set
> security.capability there, no map write needed (patch 4)
> - ptrace one of their tasks and have it do any of the above (patch 5)
>
> On an unmodified kernel we took a uid 0 task with CAP_SETFCAP dropped
> from its permitted, effective and bounding sets and, through each of
> these, ended up with a file that a uid 1000 user execs with
> CAP_SYS_ADMIN in its effective set.
>
> Patch 1 adds cred->setfcap_level, which records how far up the
> namespace tree a task's CAP_SETFCAP reached when it entered its
> namespace, and patches 2-5 check it. Patch 4 is the check that closes
> the class, the map patches make the uid 0 map rule mean what
> db2e718a4798 meant it to, and patch 5 keeps ptrace from borrowing what
> the target is entitled to.
>
> This does change behaviour. Everything new is -EPERM:
>
> - a task that entered a namespace without CAP_SETFCAP outside can't map
> uid 0 of the outside or write fscaps for that root user anymore
> - a uid_map fd opened by a task with CAP_SETFCAP can't be used by a
> task without it to map uid 0
> - if a privileged task maps "0 0 1" from the parent for a namespace
> created by a task without CAP_SETFCAP, that namespace can no longer
> write fscaps honoured outside
> - PTRACE_ATTACH and PTRACE_TRACEME fail when the tracer gave up
> CAP_SETFCAP and CAP_SYS_PTRACE, the target didn't, and they share a
> root user
>
> Rootless containers, privileged runtimes writing the map from the
> parent, nested unprivileged namespaces and containers that don't map
> host uid 0 aren't affected. The ptrace check is one compare for
> targets in the initial namespace and in namespaces entered without
> the capability.
>
> Testing: a set of flows run on the base and patched kernels, every
> bypass route above gets -EPERM with the series and the 16 legitimate
> flows behave the same. The capabilities, namespaces, ptrace, pidfd and
> proc selftests give the same results before and after. Thanks,
>
> Josef
>
> ---
> Josef Bacik (5):
> cred: record how far up CAP_SETFCAP reaches
> userns: don't let setns() lend the right to map uid 0
> userns: check the writer too before mapping uid 0
> capabilities: limit fscaps to where CAP_SETFCAP reaches
> capabilities: don't let ptrace borrow CAP_SETFCAP
Thanks, Josef.
Would you mind describing what other solutions you considered? I've been
looking over this set since Tuesday, and finding it hard to reason about.
(Part of that is certainly the nature of the problem, and it's possible
that this is the best/simplest solution.)
If we replaced the userns->parent_could_setfcap bool with a ref to the
creator's cred, then at both setns and write we could check the actor's
credentials, right? There are probably issues with that specific idea,
but that's why it would be good to see what else you've considered.
Of course UID 0 will always continue to carry privileges even with an
empty cap_eff. Here we're stopping it from writing filecaps to uid 0
owned files, but if it can open a 0 owned file on the host, like
/bin/sh or a systemd init file, or ptrace a process (in a child ns that
maps parent uid 0) doing so, it can still cause damage. My point being,
we do need to keep in mind the tradeoff of keeping the code simple
versus the realistic threat of the problem being addressed.
> include/linux/capability.h | 4 ++
> include/linux/cred.h | 1 +
> kernel/user_namespace.c | 40 +++++++++++----
> security/commoncap.c | 119 +++++++++++++++++++++++++++++++++++++++++--
> security/keys/process_keys.c | 1 +
> 5 files changed, 150 insertions(+), 15 deletions(-)
> ---
> base-commit: 7909a3e30a05e40bbc8bfb7f5629ed642abeaab8
> change-id: 20261006-b4-setfcap-userns-d63185a31ce3
next prev parent reply other threads:[~2026-10-08 13:57 UTC|newest]
Thread overview: 10+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-06 15:44 Josef Bacik
2026-10-06 15:44 ` [PATCH 1/5] cred: record how far up CAP_SETFCAP reaches Josef Bacik
2026-10-06 15:44 ` [PATCH 2/5] userns: don't let setns() lend the right to map uid 0 Josef Bacik
2026-10-06 15:44 ` [PATCH 3/5] userns: check the writer too before mapping " Josef Bacik
2026-10-06 15:44 ` [PATCH 4/5] capabilities: limit fscaps to where CAP_SETFCAP reaches Josef Bacik
2026-10-06 15:44 ` [PATCH 5/5] capabilities: don't let ptrace borrow CAP_SETFCAP Josef Bacik
2026-10-08 13:57 ` Serge E. Hallyn [this message]
2026-10-08 15:39 ` [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Josef Bacik
2026-10-08 16:13 ` Serge E. Hallyn
2026-10-09 7:41 ` Christian Brauner
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=asehZEWAwsJcbqvM@hallyn.com \
--to=serge@hallyn.com \
--cc=brauner@kernel.org \
--cc=dhowells@redhat.com \
--cc=jarkko@kernel.org \
--cc=jmorris@namei.org \
--cc=josef@toxicpanda.com \
--cc=keyrings@vger.kernel.org \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-security-module@vger.kernel.org \
--cc=morgan@kernel.org \
--cc=paul@paul-moore.com \
--cc=sergeh@kernel.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®