* [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
@ 2026-10-06 15:44 Josef Bacik
2026-10-06 15:44 ` [PATCH 1/5] cred: record how far up CAP_SETFCAP reaches Josef Bacik
` (5 more replies)
0 siblings, 6 replies; 9+ messages in thread
From: Josef Bacik @ 2026-10-06 15:44 UTC (permalink / raw)
To: Serge Hallyn, Paul Moore, Christian Brauner
Cc: James Morris, David Howells, Jarkko Sakkinen, Andrew G. Morgan,
Serge Hallyn, linux-security-module, linux-kernel, keyrings,
linux-fsdevel, Josef Bacik
Hello,
Commit db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
stops a root task that has given up CAP_SETFCAP from creating a user
namespace that maps uid 0 and then writing file capabilities in it that
the initial namespace honours. The check only looks at the task that
creates the namespace, so if somebody who did have CAP_SETFCAP created
one, there are still several ways around it:
- setns() into their namespace and write uid_map from inside (patch 2)
- write "0 0 1" through a uid_map fd that they opened (patch 3)
- setns() into a namespace of theirs that already maps uid 0 and set
security.capability there, no map write needed (patch 4)
- ptrace one of their tasks and have it do any of the above (patch 5)
On an unmodified kernel we took a uid 0 task with CAP_SETFCAP dropped
from its permitted, effective and bounding sets and, through each of
these, ended up with a file that a uid 1000 user execs with
CAP_SYS_ADMIN in its effective set.
Patch 1 adds cred->setfcap_level, which records how far up the
namespace tree a task's CAP_SETFCAP reached when it entered its
namespace, and patches 2-5 check it. Patch 4 is the check that closes
the class, the map patches make the uid 0 map rule mean what
db2e718a4798 meant it to, and patch 5 keeps ptrace from borrowing what
the target is entitled to.
This does change behaviour. Everything new is -EPERM:
- a task that entered a namespace without CAP_SETFCAP outside can't map
uid 0 of the outside or write fscaps for that root user anymore
- a uid_map fd opened by a task with CAP_SETFCAP can't be used by a
task without it to map uid 0
- if a privileged task maps "0 0 1" from the parent for a namespace
created by a task without CAP_SETFCAP, that namespace can no longer
write fscaps honoured outside
- PTRACE_ATTACH and PTRACE_TRACEME fail when the tracer gave up
CAP_SETFCAP and CAP_SYS_PTRACE, the target didn't, and they share a
root user
Rootless containers, privileged runtimes writing the map from the
parent, nested unprivileged namespaces and containers that don't map
host uid 0 aren't affected. The ptrace check is one compare for
targets in the initial namespace and in namespaces entered without
the capability.
Testing: a set of flows run on the base and patched kernels, every
bypass route above gets -EPERM with the series and the 16 legitimate
flows behave the same. The capabilities, namespaces, ptrace, pidfd and
proc selftests give the same results before and after. Thanks,
Josef
---
Josef Bacik (5):
cred: record how far up CAP_SETFCAP reaches
userns: don't let setns() lend the right to map uid 0
userns: check the writer too before mapping uid 0
capabilities: limit fscaps to where CAP_SETFCAP reaches
capabilities: don't let ptrace borrow CAP_SETFCAP
include/linux/capability.h | 4 ++
include/linux/cred.h | 1 +
kernel/user_namespace.c | 40 +++++++++++----
security/commoncap.c | 119 +++++++++++++++++++++++++++++++++++++++++--
security/keys/process_keys.c | 1 +
5 files changed, 150 insertions(+), 15 deletions(-)
---
base-commit: 7909a3e30a05e40bbc8bfb7f5629ed642abeaab8
change-id: 20261006-b4-setfcap-userns-d63185a31ce3
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH 1/5] cred: record how far up CAP_SETFCAP reaches
2026-10-06 15:44 [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Josef Bacik
@ 2026-10-06 15:44 ` Josef Bacik
2026-10-06 15:44 ` [PATCH 2/5] userns: don't let setns() lend the right to map uid 0 Josef Bacik
` (4 subsequent siblings)
5 siblings, 0 replies; 9+ messages in thread
From: Josef Bacik @ 2026-10-06 15:44 UTC (permalink / raw)
To: Serge Hallyn, Paul Moore, Christian Brauner
Cc: James Morris, David Howells, Jarkko Sakkinen, Andrew G. Morgan,
Serge Hallyn, linux-security-module, linux-kernel, keyrings,
linux-fsdevel, Josef Bacik
File capabilities are tied to the kuid of the root user of a user
namespace, and a namespace that maps uid 0 of its parent shares that kuid
with the parent. So CAP_SETFCAP in such a namespace is worth as much as
CAP_SETFCAP in the parent. But every task gets a full capability set when
it enters a user namespace, and from then on its credentials say nothing
about what it was allowed to do outside.
The only trace that is kept is ns->parent_could_setfcap, which describes
the task that created the namespace and not the ones that use it.
Add cred->setfcap_level: CAP_SETFCAP of these credentials counts in
cred->user_ns and in its ancestors down to that ->level. It is 0 in the
initial namespace, where the capability sets speak for themselves. When
credentials enter a namespace, set_cred_user_ns() keeps the value if they
hold CAP_SETFCAP, and otherwise raises it to the first namespace on the
way down over which they have it: the child they own, or else the
namespace that is being entered. cap_setfcap_level() computes that. The
value can only grow, and it is inherited over fork() and execve() like the
rest of the cred. key_change_session_keyring() copies field by field and
has to carry it over.
Nothing looks at the new field yet.
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
---
include/linux/capability.h | 3 +++
include/linux/cred.h | 1 +
kernel/user_namespace.c | 3 +++
security/commoncap.c | 26 ++++++++++++++++++++++++++
security/keys/process_keys.c | 1 +
5 files changed, 34 insertions(+)
diff --git a/include/linux/capability.h b/include/linux/capability.h
index 622137f66f09..026974a5111b 100644
--- a/include/linux/capability.h
+++ b/include/linux/capability.h
@@ -34,6 +34,7 @@ struct cpu_vfs_cap_data {
#define _USER_CAP_HEADER_SIZE (sizeof(struct __user_cap_header_struct))
#define _KERNEL_CAP_T_SIZE (sizeof(kernel_cap_t))
+struct cred;
struct file;
struct inode;
struct dentry;
@@ -222,4 +223,6 @@ int get_vfs_caps_from_disk(const struct mnt_idmap *idmap,
int cap_convert_nscap(const struct mnt_idmap *idmap, struct dentry *dentry,
const void **ivalue, size_t size);
+int cap_setfcap_level(const struct cred *cred, struct user_namespace *ns);
+
#endif /* !_LINUX_CAPABILITY_H */
diff --git a/include/linux/cred.h b/include/linux/cred.h
index 6ef1750c93e2..b95119099993 100644
--- a/include/linux/cred.h
+++ b/include/linux/cred.h
@@ -123,6 +123,7 @@ struct cred {
kuid_t fsuid; /* UID for VFS ops */
kgid_t fsgid; /* GID for VFS ops */
unsigned securebits; /* SUID-less security management */
+ int setfcap_level; /* how far up CAP_SETFCAP counts */
kernel_cap_t cap_inheritable; /* caps our children can inherit */
kernel_cap_t cap_permitted; /* caps we're permitted */
kernel_cap_t cap_effective; /* caps we can actually use */
diff --git a/kernel/user_namespace.c b/kernel/user_namespace.c
index 1b23d819d398..6b45df3a8d82 100644
--- a/kernel/user_namespace.c
+++ b/kernel/user_namespace.c
@@ -44,6 +44,9 @@ static void dec_user_namespaces(struct ucounts *ucounts)
static void set_cred_user_ns(struct cred *cred, struct user_namespace *user_ns)
{
+ /* The last chance to see what we can do outside of the new namespace. */
+ cred->setfcap_level = cap_setfcap_level(cred, user_ns->parent);
+
/* Start with the same capabilities as init but useless for doing
* anything as the capabilities are bound to the new user namespace.
*/
diff --git a/security/commoncap.c b/security/commoncap.c
index d47ab3022343..513473f859be 100644
--- a/security/commoncap.c
+++ b/security/commoncap.c
@@ -131,6 +131,32 @@ int cap_capable(const struct cred *cred, struct user_namespace *target_ns,
return ret;
}
+/**
+ * cap_setfcap_level - Determine how far up CAP_SETFCAP of a cred reaches
+ * @cred: The credentials to use
+ * @ns: The user namespace of @cred or one of its descendants
+ *
+ * File capabilities belong to the kuid of a namespace's root user, and the
+ * same kuid can be the root user of ancestors of that namespace. Every task
+ * gets CAP_SETFCAP when it enters a user namespace, so having it there says
+ * nothing about those ancestors. cred->setfcap_level does: it is handed down
+ * from namespace to namespace for as long as the capability is held.
+ *
+ * Return: the ->level of the topmost namespace, from @ns upwards, for which
+ * CAP_SETFCAP of @cred counts; @ns->level + 1 if it doesn't even over @ns.
+ */
+int cap_setfcap_level(const struct cred *cred, struct user_namespace *ns)
+{
+ if (cap_capable(cred, ns, CAP_SETFCAP, CAP_OPT_NOAUDIT))
+ return ns->level + 1;
+
+ if (cap_raised(cred->cap_effective, CAP_SETFCAP))
+ return cred->setfcap_level;
+
+ /* All we have is that we own a child of our namespace. */
+ return cred->user_ns->level + 1;
+}
+
/**
* cap_settime - Determine whether the current process may set the system clock
* @ts: The time to set
diff --git a/security/keys/process_keys.c b/security/keys/process_keys.c
index a63c46bb2d14..f9cf3ce2426f 100644
--- a/security/keys/process_keys.c
+++ b/security/keys/process_keys.c
@@ -939,6 +939,7 @@ void key_change_session_keyring(struct callback_head *twork)
new->group_info = get_group_info(old->group_info);
new->securebits = old->securebits;
+ new->setfcap_level = old->setfcap_level;
new->cap_inheritable = old->cap_inheritable;
new->cap_permitted = old->cap_permitted;
new->cap_effective = old->cap_effective;
--
2.55.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH 2/5] userns: don't let setns() lend the right to map uid 0
2026-10-06 15:44 [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Josef Bacik
2026-10-06 15:44 ` [PATCH 1/5] cred: record how far up CAP_SETFCAP reaches Josef Bacik
@ 2026-10-06 15:44 ` Josef Bacik
2026-10-06 15:44 ` [PATCH 3/5] userns: check the writer too before mapping " Josef Bacik
` (3 subsequent siblings)
5 siblings, 0 replies; 9+ messages in thread
From: Josef Bacik @ 2026-10-06 15:44 UTC (permalink / raw)
To: Serge Hallyn, Paul Moore, Christian Brauner
Cc: James Morris, David Howells, Jarkko Sakkinen, Andrew G. Morgan,
Serge Hallyn, linux-security-module, linux-kernel, keyrings,
linux-fsdevel, Josef Bacik
Mapping uid 0 of the parent into a user namespace requires CAP_SETFCAP,
because a task in such a namespace can write file capabilities that are
honoured in the parent. For a writer that already sits in the new
namespace nothing can be read from its capability sets any more, so
create_user_ns() records in ns->parent_could_setfcap whether the creator
had the capability and verify_root_map() trusts that.
The flag describes the creator, but it is applied to whoever opens
/proc/self/uid_map from inside the namespace, and setns() lets other
tasks in:
task A: uid 0, full caps task B: uid 0, no CAP_SETFCAP
unshare(CLONE_NEWUSER)
ns->parent_could_setfcap = 1
setns(A's user ns)
allowed, euid == ns->owner
write "0 0 1" to /proc/self/uid_map
map_ns == file_ns and
parent_could_setfcap -> allowed
B could not have written that map from the parent namespace, nor into a
namespace it unshared itself.
The check for an opener in the parent namespace has the same weakness one
level up. If B enters a namespace that already maps uid 0, it has
CAP_SETFCAP there, may unshare again and map uid 0 once more, and the
kuid behind that is still the root user of the initial namespace.
Use cred->setfcap_level, which tells what the opener itself was allowed to
do before it entered its namespace: CAP_SETFCAP has to reach up to the
topmost namespace that has the same root user as the parent of the
namespace being mapped. The old conditions stay, so nothing becomes
allowed that was refused before.
The change in behaviour is that a task which entered a namespace while
lacking CAP_SETFCAP outside gets -EPERM when it tries to pass on uid 0 of
the outside. The creator of a namespace, its children, tasks that join
with the capability, helpers in the parent namespace and unprivileged users
nesting namespaces below one of their own are not affected.
Fixes: db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
---
include/linux/capability.h | 1 +
kernel/user_namespace.c | 31 ++++++++++++++++++-------------
security/commoncap.c | 23 +++++++++++++++++++++++
3 files changed, 42 insertions(+), 13 deletions(-)
diff --git a/include/linux/capability.h b/include/linux/capability.h
index 026974a5111b..90f9f976d216 100644
--- a/include/linux/capability.h
+++ b/include/linux/capability.h
@@ -224,5 +224,6 @@ int cap_convert_nscap(const struct mnt_idmap *idmap, struct dentry *dentry,
const void **ivalue, size_t size);
int cap_setfcap_level(const struct cred *cred, struct user_namespace *ns);
+int cap_root_level(kuid_t kuid, struct user_namespace *ns);
#endif /* !_LINUX_CAPABILITY_H */
diff --git a/kernel/user_namespace.c b/kernel/user_namespace.c
index 6b45df3a8d82..bd5f9cea7430 100644
--- a/kernel/user_namespace.c
+++ b/kernel/user_namespace.c
@@ -900,7 +900,7 @@ static bool verify_root_map(const struct file *file,
struct user_namespace *map_ns,
struct uid_gid_map *new_map)
{
- int idx;
+ int idx, level;
const struct user_namespace *file_ns = file->f_cred->user_ns;
struct uid_gid_extent *extent0 = NULL;
@@ -918,24 +918,29 @@ static bool verify_root_map(const struct file *file,
if (!extent0)
return true;
+ /* The parent may in turn share its root user with its ancestors. */
+ level = cap_root_level(make_kuid(map_ns->parent, 0), map_ns->parent);
+
if (map_ns == file_ns) {
- /* The process unshared its ns and is writing to its own
+ /* The process is in the new ns and is writing to its own
* /proc/self/uid_map. User already has full capabilites in
- * the new namespace. Verify that the parent had CAP_SETFCAP
- * when it unshared.
- * */
+ * the new namespace. Verify that the creator had CAP_SETFCAP
+ * when it unshared, and that the opener, which may have come
+ * in later with setns(), had it as well when it entered.
+ */
if (!file_ns->parent_could_setfcap)
return false;
- } else {
- /* Process p1 is writing to uid_map of p2, who is in a child
- * user namespace to p1's. Verify that the opener of the map
- * file has CAP_SETFCAP against the parent of the new map
- * namespace */
- if (!file_ns_capable(file, map_ns->parent, CAP_SETFCAP))
- return false;
+ return file->f_cred->setfcap_level <= level;
}
- return true;
+ /* Process p1 is writing to uid_map of p2, who is in a child
+ * user namespace to p1's. Verify that the opener of the map
+ * file has CAP_SETFCAP against the parent of the new map
+ * namespace, and not just because it entered that.
+ */
+ if (!file_ns_capable(file, map_ns->parent, CAP_SETFCAP))
+ return false;
+ return cap_setfcap_level(file->f_cred, map_ns->parent) <= level;
}
static ssize_t map_write(struct file *file, const char __user *buf,
diff --git a/security/commoncap.c b/security/commoncap.c
index 513473f859be..b406ede2fadc 100644
--- a/security/commoncap.c
+++ b/security/commoncap.c
@@ -157,6 +157,29 @@ int cap_setfcap_level(const struct cred *cred, struct user_namespace *ns)
return cred->user_ns->level + 1;
}
+/**
+ * cap_root_level - Find the topmost namespace in which a kuid is the root user
+ * @kuid: The kuid to look for
+ * @ns: The user namespace to start from
+ *
+ * Return: the lowest ->level among @ns and its ancestors in which @kuid is
+ * uid 0, which is how far up file capabilities with that root user are
+ * honoured; INT_MAX if there is no such namespace.
+ */
+int cap_root_level(kuid_t kuid, struct user_namespace *ns)
+{
+ int level = INT_MAX;
+
+ for (;; ns = ns->parent) {
+ if (from_kuid(ns, kuid) == 0)
+ level = ns->level;
+ if (ns == &init_user_ns)
+ break;
+ }
+
+ return level;
+}
+
/**
* cap_settime - Determine whether the current process may set the system clock
* @ts: The time to set
--
2.55.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH 3/5] userns: check the writer too before mapping uid 0
2026-10-06 15:44 [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Josef Bacik
2026-10-06 15:44 ` [PATCH 1/5] cred: record how far up CAP_SETFCAP reaches Josef Bacik
2026-10-06 15:44 ` [PATCH 2/5] userns: don't let setns() lend the right to map uid 0 Josef Bacik
@ 2026-10-06 15:44 ` Josef Bacik
2026-10-06 15:44 ` [PATCH 4/5] capabilities: limit fscaps to where CAP_SETFCAP reaches Josef Bacik
` (2 subsequent siblings)
5 siblings, 0 replies; 9+ messages in thread
From: Josef Bacik @ 2026-10-06 15:44 UTC (permalink / raw)
To: Serge Hallyn, Paul Moore, Christian Brauner
Cc: James Morris, David Howells, Jarkko Sakkinen, Andrew G. Morgan,
Serge Hallyn, linux-security-module, linux-kernel, keyrings,
linux-fsdevel, Josef Bacik
verify_root_map() only looks at file->f_cred. For every other privileged
id mapping new_idmap_permitted() wants the capability from the opener of
the map file and from the task that calls write(), but a map that consists
of the single line "0 0 1" takes the unprivileged branch, where only the
opener's euid counts. So an open uid_map file is a token for mapping
uid 0:
task A: uid 0, full caps task B: uid 0, no CAP_SETFCAP
fd = open("/proc/<pid>/uid_map")
B gets fd by inheritance
or SCM_RIGHTS
write(fd, "0 0 1")
opener had CAP_SETFCAP -> allowed
B set up a mapping that it is not allowed to set up.
Apply to the writer what is applied to the opener: in the namespace that
is being mapped, its credentials must have come in with CAP_SETFCAP over
the parent; anywhere else it needs CAP_SETFCAP over the parent now.
This only affects maps that contain uid 0 of the parent namespace, and
only when the file was opened by a task with CAP_SETFCAP and is written to
by one without. Those writes now fail with -EPERM.
Fixes: db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
---
kernel/user_namespace.c | 32 +++++++++++++++++++++-----------
1 file changed, 21 insertions(+), 11 deletions(-)
diff --git a/kernel/user_namespace.c b/kernel/user_namespace.c
index bd5f9cea7430..421769e2d24f 100644
--- a/kernel/user_namespace.c
+++ b/kernel/user_namespace.c
@@ -891,8 +891,9 @@ EXPORT_SYMBOL_IF_KUNIT(uid_gid_map_sort);
* @new_map: requested idmap
*
* If a process requests mapping parent uid 0 into the new ns, verify that the
- * process writing the map had the CAP_SETFCAP capability as the target process
- * will be able to write fscaps that are valid in ancestor user namespaces.
+ * process that opened the map file and the process writing the map had the
+ * CAP_SETFCAP capability as the target process will be able to write fscaps
+ * that are valid in ancestor user namespaces.
*
* Return: true if the mapping is allowed, false if not.
*/
@@ -928,19 +929,28 @@ static bool verify_root_map(const struct file *file,
* when it unshared, and that the opener, which may have come
* in later with setns(), had it as well when it entered.
*/
- if (!file_ns->parent_could_setfcap)
+ if (!file_ns->parent_could_setfcap ||
+ file->f_cred->setfcap_level > level)
+ return false;
+ } else {
+ /* Process p1 is writing to uid_map of p2, who is in a child
+ * user namespace to p1's. Verify that the opener of the map
+ * file has CAP_SETFCAP against the parent of the new map
+ * namespace, and not just because it entered that.
+ */
+ if (!file_ns_capable(file, map_ns->parent, CAP_SETFCAP) ||
+ cap_setfcap_level(file->f_cred, map_ns->parent) > level)
return false;
- return file->f_cred->setfcap_level <= level;
}
- /* Process p1 is writing to uid_map of p2, who is in a child
- * user namespace to p1's. Verify that the opener of the map
- * file has CAP_SETFCAP against the parent of the new map
- * namespace, and not just because it entered that.
+ /* The file may have been handed to someone else since it was opened,
+ * so the same goes for the process that is doing the write.
*/
- if (!file_ns_capable(file, map_ns->parent, CAP_SETFCAP))
- return false;
- return cap_setfcap_level(file->f_cred, map_ns->parent) <= level;
+ if (map_ns == current_user_ns())
+ return current_cred()->setfcap_level <= level;
+
+ return ns_capable(map_ns->parent, CAP_SETFCAP) &&
+ cap_setfcap_level(current_cred(), map_ns->parent) <= level;
}
static ssize_t map_write(struct file *file, const char __user *buf,
--
2.55.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH 4/5] capabilities: limit fscaps to where CAP_SETFCAP reaches
2026-10-06 15:44 [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Josef Bacik
` (2 preceding siblings ...)
2026-10-06 15:44 ` [PATCH 3/5] userns: check the writer too before mapping " Josef Bacik
@ 2026-10-06 15:44 ` Josef Bacik
2026-10-06 15:44 ` [PATCH 5/5] capabilities: don't let ptrace borrow CAP_SETFCAP Josef Bacik
2026-10-08 13:57 ` [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Serge E. Hallyn
5 siblings, 0 replies; 9+ messages in thread
From: Josef Bacik @ 2026-10-06 15:44 UTC (permalink / raw)
To: Serge Hallyn, Paul Moore, Christian Brauner
Cc: James Morris, David Howells, Jarkko Sakkinen, Andrew G. Morgan,
Serge Hallyn, linux-security-module, linux-kernel, keyrings,
linux-fsdevel, Josef Bacik
Commit db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
keeps a root process without CAP_SETFCAP from creating a user namespace
that maps uid 0, in which it could attach capabilities to a file that are
then honoured outside. It does not have to create one, though. Any
namespace will do that maps uid 0 and that it may enter, and it may enter
all those that were created by uid 0:
task A: uid 0, full caps task B: uid 0, no CAP_SETFCAP
unshare(CLONE_NEWUSER)
write "0 0 1" to uid_map, gid_map
setns(A's user ns)
full capability set
setxattr(file, "security.capability")
rootid = kuid 0
(back in the initial namespace)
execve(file)
file capabilities apply
The map does not even have to show uid 0 of the parent as uid 0, since a
v3 xattr can name any mapped uid as the root user.
Check at the place where it matters, in cap_convert_nscap(). On an
idmapped mount the root user named in the xattr has two identities: the
kuid seen through this mount, and the kuid that is stored and seen
through every other mount of the filesystem. If either of them is the
root user of an ancestor of the caller's namespace too, CAP_SETFCAP has
to reach up to there, which cred->setfcap_level tells.
Tasks in the initial namespace, tasks that held CAP_SETFCAP when they
entered their namespace, and file capabilities for a root user that only
exists in namespaces the caller came into with CAP_SETFCAP or below are
not affected; that includes everything an unprivileged user does in his
own namespaces. The rest gets -EPERM.
Fixes: 8db6c34f1dbc ("Introduce v3 namespaced file capabilities")
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
---
security/commoncap.c | 14 ++++++++++++++
1 file changed, 14 insertions(+)
diff --git a/security/commoncap.c b/security/commoncap.c
index b406ede2fadc..26798d62e6b1 100644
--- a/security/commoncap.c
+++ b/security/commoncap.c
@@ -644,10 +644,24 @@ int cap_convert_nscap(const struct mnt_idmap *idmap, struct dentry *dentry,
if (!vfsuid_valid(vfsrootid))
return -EINVAL;
+ /*
+ * The root user may be the root user of ancestors of our namespace as
+ * well. CAP_SETFCAP that we got for entering it doesn't cover those.
+ * On an idmapped mount the root user is vfsrootid as seen through
+ * this mount and rootid as seen through every other mount of the
+ * filesystem, so both have to stay within reach.
+ */
+ if (cap_root_level(vfsuid_into_kuid(vfsrootid), task_ns) <
+ current_cred()->setfcap_level)
+ return -EPERM;
+
rootid = from_vfsuid(idmap, fs_ns, vfsrootid);
if (!uid_valid(rootid))
return -EINVAL;
+ if (cap_root_level(rootid, task_ns) < current_cred()->setfcap_level)
+ return -EPERM;
+
nsrootid = from_kuid(fs_ns, rootid);
if (nsrootid == -1)
return -EINVAL;
--
2.55.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* [PATCH 5/5] capabilities: don't let ptrace borrow CAP_SETFCAP
2026-10-06 15:44 [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Josef Bacik
` (3 preceding siblings ...)
2026-10-06 15:44 ` [PATCH 4/5] capabilities: limit fscaps to where CAP_SETFCAP reaches Josef Bacik
@ 2026-10-06 15:44 ` Josef Bacik
2026-10-08 13:57 ` [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Serge E. Hallyn
5 siblings, 0 replies; 9+ messages in thread
From: Josef Bacik @ 2026-10-06 15:44 UTC (permalink / raw)
To: Serge Hallyn, Paul Moore, Christian Brauner
Cc: James Morris, David Howells, Jarkko Sakkinen, Andrew G. Morgan,
Serge Hallyn, linux-security-module, linux-kernel, keyrings,
linux-fsdevel, Josef Bacik
A task that entered a user namespace while it held CAP_SETFCAP outside may
map uid 0 of the parent and write file capabilities that are honoured
there. That is the one privilege it keeps outside of its namespace, and
the ptrace checks don't know about it: for a target in another user
namespace cap_ptrace_access_check() is satisfied with CAP_SYS_PTRACE over
the target's namespace, which the owner rule hands to any task with the
right euid, and inside the namespace everybody has a full set.
task A: uid 0, full caps task B: uid 0, no CAP_SETFCAP
unshare(CLONE_NEWUSER)
ptrace(PTRACE_ATTACH, A)
uids match, euid == ns->owner
make A write "0 0 1" to its uid_map,
or set file capabilities
A is entitled -> allowed
The same goes for /proc/<pid>/mem, process_vm_writev() and pidfd_getfd().
Require for PTRACE_MODE_ATTACH and PTRACE_TRACEME that CAP_SETFCAP of the
tracer reaches as far up as the target can make use of: to the topmost
ancestor, within the target's setfcap_level, whose root user is mapped
into the target's namespace, or can still be mapped because the namespace
has no uid map yet. CAP_SYS_PTRACE over that ancestor is accepted too,
since it gives control of tasks there that have CAP_SETFCAP anyway.
Targets in the initial namespace, in namespaces that were entered without
CAP_SETFCAP (what unprivileged users create) and in namespaces that don't
map the root user of an ancestor (the usual container) are not affected.
What remains is a tracer that gave up CAP_SETFCAP and CAP_SYS_PTRACE and a
target that didn't, in a namespace that shares its root user with the
tracer's. PTRACE_MODE_READ is unchanged.
Fixes: db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
Assisted-by: LLM
Signed-off-by: Josef Bacik <josef@toxicpanda.com>
---
security/commoncap.c | 56 ++++++++++++++++++++++++++++++++++++++++++++++++----
1 file changed, 52 insertions(+), 4 deletions(-)
diff --git a/security/commoncap.c b/security/commoncap.c
index 26798d62e6b1..7263c78b6397 100644
--- a/security/commoncap.c
+++ b/security/commoncap.c
@@ -195,14 +195,52 @@ int cap_settime(const struct timespec64 *ts, const struct timezone *tz)
return 0;
}
+/*
+ * CAP_SETFCAP of a task can count in ancestors of its user namespace, see
+ * cap_setfcap_level(). That is a privilege outside of the namespace, so
+ * capabilities over the namespace are not enough to take control of the task.
+ */
+static bool cap_covers_setfcap(const struct cred *cred,
+ const struct cred *child_cred)
+{
+ struct user_namespace *ns = child_cred->user_ns, *seen, *p, *top = NULL;
+
+ if (child_cred->setfcap_level >= ns->level)
+ return true;
+
+ /*
+ * It is of use only for root users that are mapped into the namespace,
+ * or that can still be because there is no map yet.
+ */
+ seen = READ_ONCE(ns->uid_map.nr_extents) ? ns : ns->parent;
+ for (p = ns->parent; p; p = p->parent) {
+ if (p->level < child_cred->setfcap_level)
+ break;
+ if (kuid_has_mapping(seen, make_kuid(p, 0)))
+ top = p;
+ }
+ if (!top)
+ return true;
+
+ if (cred->user_ns == ns)
+ return cred->setfcap_level <= top->level;
+
+ /* CAP_SYS_PTRACE up there gives control of tasks with CAP_SETFCAP. */
+ return cap_setfcap_level(cred, ns->parent) <= top->level ||
+ !cap_capable(cred, top, CAP_SYS_PTRACE, CAP_OPT_NOAUDIT);
+}
+
/**
* cap_ptrace_access_check - Determine whether the current process may access
* another
* @child: The process to be accessed
* @mode: The mode of attachment.
*
- * If we are in the same or an ancestor user_ns and have all the target
- * task's capabilities, then ptrace access is allowed.
+ * For PTRACE_MODE_ATTACH, if the target task may make use of CAP_SETFCAP in
+ * an ancestor of its user_ns that our CAP_SETFCAP or CAP_SYS_PTRACE doesn't
+ * reach, then ptrace access is denied.
+ * Otherwise, if we are in the same or an ancestor user_ns and have all the
+ * target task's capabilities, then ptrace access is allowed.
* If we have the ptrace capability to the target user_ns, then ptrace
* access is allowed.
* Else denied.
@@ -223,11 +261,15 @@ int cap_ptrace_access_check(struct task_struct *child, unsigned int mode)
caller_caps = &cred->cap_effective;
else
caller_caps = &cred->cap_permitted;
+ if ((mode & PTRACE_MODE_ATTACH) &&
+ !cap_covers_setfcap(cred, child_cred))
+ goto deny;
if (cred->user_ns == child_cred->user_ns &&
cap_issubset(child_cred->cap_permitted, *caller_caps))
goto out;
if (ns_capable(child_cred->user_ns, CAP_SYS_PTRACE))
goto out;
+deny:
ret = -EPERM;
out:
rcu_read_unlock();
@@ -238,8 +280,11 @@ int cap_ptrace_access_check(struct task_struct *child, unsigned int mode)
* cap_ptrace_traceme - Determine whether another process may trace the current
* @parent: The task proposed to be the tracer
*
- * If parent is in the same or an ancestor user_ns and has all current's
- * capabilities, then ptrace access is allowed.
+ * If current may make use of CAP_SETFCAP in an ancestor of its user_ns that
+ * parent's CAP_SETFCAP or CAP_SYS_PTRACE doesn't reach, then ptrace access
+ * is denied.
+ * Otherwise, if parent is in the same or an ancestor user_ns and has all
+ * current's capabilities, then ptrace access is allowed.
* If parent has the ptrace capability to current's user_ns, then ptrace
* access is allowed.
* Else denied.
@@ -255,11 +300,14 @@ int cap_ptrace_traceme(struct task_struct *parent)
rcu_read_lock();
cred = __task_cred(parent);
child_cred = current_cred();
+ if (!cap_covers_setfcap(cred, child_cred))
+ goto deny;
if (cred->user_ns == child_cred->user_ns &&
cap_issubset(child_cred->cap_permitted, cred->cap_permitted))
goto out;
if (has_ns_capability(parent, child_cred->user_ns, CAP_SYS_PTRACE))
goto out;
+deny:
ret = -EPERM;
out:
rcu_read_unlock();
--
2.55.0
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
2026-10-06 15:44 [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Josef Bacik
` (4 preceding siblings ...)
2026-10-06 15:44 ` [PATCH 5/5] capabilities: don't let ptrace borrow CAP_SETFCAP Josef Bacik
@ 2026-10-08 13:57 ` Serge E. Hallyn
2026-10-08 15:39 ` Josef Bacik
5 siblings, 1 reply; 9+ messages in thread
From: Serge E. Hallyn @ 2026-10-08 13:57 UTC (permalink / raw)
To: Josef Bacik
Cc: Paul Moore, Christian Brauner, James Morris, David Howells,
Jarkko Sakkinen, Andrew G. Morgan, Serge Hallyn,
linux-security-module, linux-kernel, keyrings, linux-fsdevel
On Tue, Oct 06, 2026 at 03:44:16PM +0000, Josef Bacik wrote:
> Hello,
>
> Commit db2e718a4798 ("capabilities: require CAP_SETFCAP to map uid 0")
> stops a root task that has given up CAP_SETFCAP from creating a user
> namespace that maps uid 0 and then writing file capabilities in it that
> the initial namespace honours. The check only looks at the task that
> creates the namespace, so if somebody who did have CAP_SETFCAP created
> one, there are still several ways around it:
>
> - setns() into their namespace and write uid_map from inside (patch 2)
> - write "0 0 1" through a uid_map fd that they opened (patch 3)
> - setns() into a namespace of theirs that already maps uid 0 and set
> security.capability there, no map write needed (patch 4)
> - ptrace one of their tasks and have it do any of the above (patch 5)
>
> On an unmodified kernel we took a uid 0 task with CAP_SETFCAP dropped
> from its permitted, effective and bounding sets and, through each of
> these, ended up with a file that a uid 1000 user execs with
> CAP_SYS_ADMIN in its effective set.
>
> Patch 1 adds cred->setfcap_level, which records how far up the
> namespace tree a task's CAP_SETFCAP reached when it entered its
> namespace, and patches 2-5 check it. Patch 4 is the check that closes
> the class, the map patches make the uid 0 map rule mean what
> db2e718a4798 meant it to, and patch 5 keeps ptrace from borrowing what
> the target is entitled to.
>
> This does change behaviour. Everything new is -EPERM:
>
> - a task that entered a namespace without CAP_SETFCAP outside can't map
> uid 0 of the outside or write fscaps for that root user anymore
> - a uid_map fd opened by a task with CAP_SETFCAP can't be used by a
> task without it to map uid 0
> - if a privileged task maps "0 0 1" from the parent for a namespace
> created by a task without CAP_SETFCAP, that namespace can no longer
> write fscaps honoured outside
> - PTRACE_ATTACH and PTRACE_TRACEME fail when the tracer gave up
> CAP_SETFCAP and CAP_SYS_PTRACE, the target didn't, and they share a
> root user
>
> Rootless containers, privileged runtimes writing the map from the
> parent, nested unprivileged namespaces and containers that don't map
> host uid 0 aren't affected. The ptrace check is one compare for
> targets in the initial namespace and in namespaces entered without
> the capability.
>
> Testing: a set of flows run on the base and patched kernels, every
> bypass route above gets -EPERM with the series and the 16 legitimate
> flows behave the same. The capabilities, namespaces, ptrace, pidfd and
> proc selftests give the same results before and after. Thanks,
>
> Josef
>
> ---
> Josef Bacik (5):
> cred: record how far up CAP_SETFCAP reaches
> userns: don't let setns() lend the right to map uid 0
> userns: check the writer too before mapping uid 0
> capabilities: limit fscaps to where CAP_SETFCAP reaches
> capabilities: don't let ptrace borrow CAP_SETFCAP
Thanks, Josef.
Would you mind describing what other solutions you considered? I've been
looking over this set since Tuesday, and finding it hard to reason about.
(Part of that is certainly the nature of the problem, and it's possible
that this is the best/simplest solution.)
If we replaced the userns->parent_could_setfcap bool with a ref to the
creator's cred, then at both setns and write we could check the actor's
credentials, right? There are probably issues with that specific idea,
but that's why it would be good to see what else you've considered.
Of course UID 0 will always continue to carry privileges even with an
empty cap_eff. Here we're stopping it from writing filecaps to uid 0
owned files, but if it can open a 0 owned file on the host, like
/bin/sh or a systemd init file, or ptrace a process (in a child ns that
maps parent uid 0) doing so, it can still cause damage. My point being,
we do need to keep in mind the tradeoff of keeping the code simple
versus the realistic threat of the problem being addressed.
> include/linux/capability.h | 4 ++
> include/linux/cred.h | 1 +
> kernel/user_namespace.c | 40 +++++++++++----
> security/commoncap.c | 119 +++++++++++++++++++++++++++++++++++++++++--
> security/keys/process_keys.c | 1 +
> 5 files changed, 150 insertions(+), 15 deletions(-)
> ---
> base-commit: 7909a3e30a05e40bbc8bfb7f5629ed642abeaab8
> change-id: 20261006-b4-setfcap-userns-d63185a31ce3
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
2026-10-08 13:57 ` [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Serge E. Hallyn
@ 2026-10-08 15:39 ` Josef Bacik
2026-10-08 16:13 ` Serge E. Hallyn
0 siblings, 1 reply; 9+ messages in thread
From: Josef Bacik @ 2026-10-08 15:39 UTC (permalink / raw)
To: Serge E. Hallyn
Cc: Paul Moore, Christian Brauner, James Morris, David Howells,
Jarkko Sakkinen, Andrew G. Morgan, Serge Hallyn,
linux-security-module, linux-kernel, keyrings, linux-fsdevel
On Thu, Oct 08, 2026 at 08:57:56AM -0500, Serge E. Hallyn wrote:
> Would you mind describing what other solutions you considered? I've been
> looking over this set since Tuesday, and finding it hard to reason about.
> (Part of that is certainly the nature of the problem, and it's possible
> that this is the best/simplest solution.)
Everything we looked at kept the state in the cred. We model checked
the variants before writing the code, and these fell over:
- a bool per cred for "had CAP_SETFCAP over the parent when it entered".
It breaks on two hops: setns() into a namespace that maps 0, unshare
again, and the bool says yes for the second namespace. Hence the
level.
- checking only at uid_map write time. That misses setxattr of
security.capability in a namespace that already maps 0, hence patch 4.
We didn't look at keeping the state on the namespace.
> If we replaced the userns->parent_could_setfcap bool with a ref to the
> creator's cred, then at both setns and write we could check the actor's
> credentials, right? There are probably issues with that specific idea,
> but that's why it would be good to see what else you've considered.
Checking at write time alone doesn't work: once a task is inside the
namespace its cred says nothing about what it could do outside, so a
joiner and the creator look the same.
It does work if setns() refuses to join a namespace that maps, or can
still map, the parent's uid 0 unless the joiner has CAP_SETFCAP over the
parent. Then everybody in a namespace has the same reach, and it can be
a level stored on the namespace at create time instead of a cred ref,
which would pin keyrings and the rest for the life of the namespace.
The checks would be setns(), the map write (opener and writer are in the
parent, so a plain capable check), setxattr of security.capability, and
ptrace.
The difference in behaviour is that the -EPERM moves to setns(): a root
task without CAP_SETFCAP couldn't enter a root-owned container that maps
host uid 0 at all, where with this series it can enter and is refused
only for the map, fscaps and ptrace. I can prototype it if you prefer
that.
> Of course UID 0 will always continue to carry privileges even with an
> empty cap_eff. Here we're stopping it from writing filecaps to uid 0
> owned files, but if it can open a 0 owned file on the host, like
> /bin/sh or a systemd init file, or ptrace a process (in a child ns that
> maps parent uid 0) doing so, it can still cause damage. My point being,
> we do need to keep in mind the tradeoff of keeping the code simple
> versus the realistic threat of the problem being addressed.
On the same kernels, the restricted root task can copy a binary and
chmod 4755 it (it owns it, no capability needed), and a uid 1000 user
runs it with a full CapEff. With SECBIT_NOROOT that setuid copy gives
uid 1000 nothing, while the fscap file still gives it what's in the
xattr on an unpatched kernel. So SECBIT_NOROOT is the case the series
adds anything for, the same case db2e718a4798 covers.
A smaller version is patches 1, 4 and 5, with cap_root_level() moved
from 2 into 4. I built that and ran the same flows: every route that
ends in a file capability still gets -EPERM and uid 1000 gets nothing.
The uid 0 map writes refused by 2 and 3 go through again, but the fscap
write after them is refused.
Let me know which way you'd like to go and I'll rework it.
Thanks,
Josef
^ permalink raw reply [flat|nested] 9+ messages in thread
* Re: [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0
2026-10-08 15:39 ` Josef Bacik
@ 2026-10-08 16:13 ` Serge E. Hallyn
0 siblings, 0 replies; 9+ messages in thread
From: Serge E. Hallyn @ 2026-10-08 16:13 UTC (permalink / raw)
To: Josef Bacik
Cc: Paul Moore, Christian Brauner, James Morris, David Howells,
Jarkko Sakkinen, Andrew G. Morgan, Serge Hallyn,
linux-security-module, linux-kernel, keyrings, linux-fsdevel
On Thu, Oct 08, 2026 at 03:39:30PM +0000, Josef Bacik wrote:
> On Thu, Oct 08, 2026 at 08:57:56AM -0500, Serge E. Hallyn wrote:
> > Would you mind describing what other solutions you considered? I've been
> > looking over this set since Tuesday, and finding it hard to reason about.
> > (Part of that is certainly the nature of the problem, and it's possible
> > that this is the best/simplest solution.)
>
> Everything we looked at kept the state in the cred. We model checked
> the variants before writing the code, and these fell over:
>
> - a bool per cred for "had CAP_SETFCAP over the parent when it entered".
> It breaks on two hops: setns() into a namespace that maps 0, unshare
> again, and the bool says yes for the second namespace. Hence the
> level.
> - checking only at uid_map write time. That misses setxattr of
> security.capability in a namespace that already maps 0, hence patch 4.
>
> We didn't look at keeping the state on the namespace.
>
> > If we replaced the userns->parent_could_setfcap bool with a ref to the
> > creator's cred, then at both setns and write we could check the actor's
> > credentials, right? There are probably issues with that specific idea,
> > but that's why it would be good to see what else you've considered.
>
> Checking at write time alone doesn't work: once a task is inside the
> namespace its cred says nothing about what it could do outside, so a
> joiner and the creator look the same.
>
> It does work if setns() refuses to join a namespace that maps, or can
> still map, the parent's uid 0 unless the joiner has CAP_SETFCAP over the
> parent. Then everybody in a namespace has the same reach, and it can be
Yeah, that's what I was thinking. Or even stricter: ensure that to join
any user namespace, a process must have a superset of the namespace
creator's capabilities.
> a level stored on the namespace at create time instead of a cred ref,
> which would pin keyrings and the rest for the life of the namespace.
> The checks would be setns(), the map write (opener and writer are in the
> parent, so a plain capable check), setxattr of security.capability, and
> ptrace.
>
> The difference in behaviour is that the -EPERM moves to setns(): a root
> task without CAP_SETFCAP couldn't enter a root-owned container that maps
> host uid 0 at all, where with this series it can enter and is refused
> only for the map, fscaps and ptrace. I can prototype it if you prefer
> that.
Sorry let me think about it (or let us talk about it) a bit more.
> > Of course UID 0 will always continue to carry privileges even with an
> > empty cap_eff. Here we're stopping it from writing filecaps to uid 0
> > owned files, but if it can open a 0 owned file on the host, like
> > /bin/sh or a systemd init file, or ptrace a process (in a child ns that
> > maps parent uid 0) doing so, it can still cause damage. My point being,
> > we do need to keep in mind the tradeoff of keeping the code simple
> > versus the realistic threat of the problem being addressed.
>
> On the same kernels, the restricted root task can copy a binary and
> chmod 4755 it (it owns it, no capability needed), and a uid 1000 user
> runs it with a full CapEff. With SECBIT_NOROOT that setuid copy gives
> uid 1000 nothing, while the fscap file still gives it what's in the
> xattr on an unpatched kernel. So SECBIT_NOROOT is the case the series
> adds anything for, the same case db2e718a4798 covers.
>
> A smaller version is patches 1, 4 and 5, with cap_root_level() moved
> from 2 into 4. I built that and ran the same flows: every route that
> ends in a file capability still gets -EPERM and uid 1000 gets nothing.
> The uid 0 map writes refused by 2 and 3 go through again, but the fscap
> write after them is refused.
>
> Let me know which way you'd like to go and I'll rework it.
>
> Thanks,
> Josef
^ permalink raw reply [flat|nested] 9+ messages in thread
end of thread, other threads:[~2026-10-08 16:14 UTC | newest]
Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-06 15:44 [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Josef Bacik
2026-10-06 15:44 ` [PATCH 1/5] cred: record how far up CAP_SETFCAP reaches Josef Bacik
2026-10-06 15:44 ` [PATCH 2/5] userns: don't let setns() lend the right to map uid 0 Josef Bacik
2026-10-06 15:44 ` [PATCH 3/5] userns: check the writer too before mapping " Josef Bacik
2026-10-06 15:44 ` [PATCH 4/5] capabilities: limit fscaps to where CAP_SETFCAP reaches Josef Bacik
2026-10-06 15:44 ` [PATCH 5/5] capabilities: don't let ptrace borrow CAP_SETFCAP Josef Bacik
2026-10-08 13:57 ` [PATCH 0/5] capabilities: close the ways around the CAP_SETFCAP rule for uid 0 Serge E. Hallyn
2026-10-08 15:39 ` Josef Bacik
2026-10-08 16:13 ` Serge E. Hallyn
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®