mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Frederic Weisbecker <frederic@kernel.org>
To: Thomas Gleixner <tglx@kernel.org>
Cc: LKML <linux-kernel@vger.kernel.org>,
	Hyunwoo Kim <imv4bel@gmail.com>, Oleg Nesterov <oleg@redhat.com>,
	Christian Brauner <brauner@kernel.org>,
	Peter Zijlstra <peterz@infradead.org>,
	John Stultz <jstultz@google.com>, Ingo Molnar <mingo@kernel.org>,
	Alexander Viro <viro@zeniv.linux.org.uk>,
	"Eric W. Biederman" <ebiederm@xmission.com>,
	Alan Stern <stern@rowland.harvard.edu>,
	stable@vger.kernel.org
Subject: Re: [patch V3 1/8] signal: Prevent exec() race
Date: Fri, 11 Sep 2026 15:08:45 +0200	[thread overview]
Message-ID: <aqP9XYckv9mhLuQf@localhost.localdomain> (raw)
In-Reply-To: <20260911090541.572536604@kernel.org>

Le Fri, Sep 11, 2026 at 11:09:12AM +0200, Thomas Gleixner a écrit :
> From: Thomas Gleixner <tglx@kernel.org>
> 
> Hyunwoo debugged the following KASAN UAF splat:
> 
>   BUG: KASAN: slab-use-after-free in __send_signal_locked+0xb27/0xba0
>   Write of size 8 at addr ffff888007ed80c8 by task poc/79
>   ...
>   Call Trace:
>    __send_signal_locked+0xb27/0xba0
>    do_send_sig_info+0xa7/0x160
>    do_send_specific+0x76/0xa0
>    __x64_sys_tgkill+0x193/0x270
>   ...
>   Allocated by task 80:
>    do_timer_create+0x1a4/0x1030
>    __x64_sys_timer_create+0x145/0x190
>   ...
>   Freed by task 12:
>    kmem_cache_free_bulk+0x1f8/0x4a0
>    kvfree_rcu_bulk+0x14f/0x1c0
>    kfree_rcu_work+0x128/0x1a0
>   ...
>   Last potentially related work creation:
>    kvfree_call_rcu+0x39/0x390
>    __flush_itimer_signals+0x211/0x320
>    flush_itimer_signals+0x47/0x90
>    begin_new_exec+0xa6b/0x28c0
> 
> It turned out that this happens with a non-leader exec() as Hyunwoo
> explained:
> 
> de_thread() calls exchange_tids() before release_task(leader), so the
> struct pid held by a SIGEV_THREAD_ID timer created against the leader's tid
> now points to the thread which called execve(). pid_task() returns that
> thread and lock_task_sighand() on it succeeds.
> 
> If the timer signal is blocked, its sigqueue stays queued on the leader's
> task::pending. The next expiry of that timer can then run while
> release_task() flushes the queue.
> 
> posixtimer_send_sigqueue() checks whether the sigqueue is already queued
> with a plain list_empty(), which only reads list_head::next.
> list_del_init() is not atomic and INIT_LIST_HEAD() stores list_head::next
> before list_head::prev, so the check can pass in between. list_add_tail()
> queues the entry on the task::pending of the live thread, and the
> list_head::prev store from the flush then overwrites the list_head::prev
> link that list_add_tail() has just set.
> 
> __flush_itimer_signals() does not undo that either. With list_head::prev
> pointing at the entry itself, its list_del_init() only stores the same
> values again, so the entry is not removed from the list. It is still there
> after the last reference is dropped and the timer is freed by RCU, and the
> list_add_tail() of a later tgkill() follows that list_head::prev into the
> freed timer.
> 
> This problem surfaced with the recent commit which moved the sigqueue flush
> out of the sighand lock held region.
> 
> Hyonwoo proposed to fix this by using list_del_init_careful(), but that
> just papers over the problem. After some disucssions and various attempts
> to solve it, Eric pointed out that there is no reason to flush
> task::pending late in release_task() and it should be done in
> exit_signals() already.
> 
> As nothing can collect and deliver signals which are queued in a dying
> task's pending queue, there is no reason to delay it further.
> 
> But it has to be ensured that no signals can be queued into it after that
> point. exit_signals() sets PF_EXITING in task::flags, which can be used as
> an indicator for this.
> 
> Cure it by:
> 
>   - Preventing signal queueing for task private signals (PIDTYPE_PID) when
>     the task has PF_EXITING set in __send_signal_locked() and in
>     posixtimer_send_sigqueue().
> 
>   - Protecting the unlocked setting of PF_EXITING in exit_signals() for the
>     task group empty and the group exit case with sighand lock
> 
>   - Flushing task::pending signals right there.
> 
>     Optimize that by moving the whole pending list to an on-stack list head
>     under sighand lock and free the signals without the lock held.
> 
> There has been quite some discussion about the lockless flush and the
> non-leader exec case on weakly ordered systems. The problem is that a third
> party which tries to send a posix timer signal relies on the PID lookup to
> find the target task and that lookup might result in the new leader when
> the signal was originaly directed to the old leader. In case that the
> signal was queued on the old leader then the lockless flush raised a
> concern over the following situation:
> 
>    old_leader		new_leader              third party
> 
> A: flush_list()	// list_del_init() stores to sigqueue
> 
>    LOCK (tasklist)
>    old_leader->exit_state = EXIT_ZOMBIE;
> B: UNLOCK (tasklist)
> 
> C:			LOCK (tasklist)
> 			if (old_leader->exit_state)
> 			   transfer_tids()
> D:			     store PID
> 						posix_timer_send_sigqueue()
> 						// Observes #D so t = new_leader
> E:						t = get_target()
> 
> F:						LOCK (sighand)
> 
> G:						   if (list_empty(sigqueue))
> 							list_add(sigqueue)
> 
> The concern was that the third party might observe #D but not observe #A
> and therefore would proceed to #G while the list_del() stores (#A) in
> flush_list() are not visible yet, which could result in list corruption.
> 
> That would be possible if looking at it solely from a RELEASE+ACQUIRE
> ordering point of view, but B-C is a UNLOCK+LOCK hand-over, which is not
> the same as RELEASE+ACQUIRE:
> 
>   RELEASE+ACQUIRE: RCpc,  only the CPUs involved agree on the ordering
>   UNLOCK+LOCK:     RCtso, the hand-over is store-ordering
> 
> As B-C is UNLOCK+LOCK, which is RCtso and that does impose store order,
> A stores must happen before the D store.
> 
> Combine with E-F, which has a data dependency from the LOAD to the LOCK and
> thereby constraints later LOADs, those sigqueue loads in G that come after
> F must in fact observe the A stores.
> 
> Fixes: fb3bbcfe344e ("exit: change the release_task() paths to call flush_sigqueue() lockless")
> Reported-by: Hyunwoo Kim <imv4bel@gmail.com>
> Debugged-by: Hyunwoo Kim <imv4bel@gmail.com>
> Suggested-by: "Eric W. Biederman" <ebiederm@xmission.com>
> Signed-off-by: Thomas Gleixner <tglx@kernel.org>
> Reviewed-by: Oleg Nesterov <oleg@redhat.com>
> Cc: stable@vger.kernel.org
> Closes: https://patch.msgid.link/aok1rdkBgZsynHZB@v4bel

Reviewed-by: Frederic Weisbecker <frederic@kernel.org>

-- 
Frederic Weisbecker
SUSE Labs

  reply	other threads:[~2026-09-11 13:08 UTC|newest]

Thread overview: 19+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-11  9:09 [patch V3 0/8] exec/exit: POSIX timer related bugfixes and related cleanups Thomas Gleixner
2026-09-11  9:09 ` [patch V3 1/8] signal: Prevent exec() race Thomas Gleixner
2026-09-11 13:08   ` Frederic Weisbecker [this message]
2026-09-11  9:09 ` [patch V3 2/8] exec: Cleanup POSIX timers right after de_thread() Thomas Gleixner
2026-09-11  9:09 ` [patch V3 3/8] posix-timers: Move posixtimer_exec_cleanup() out of exec.c Thomas Gleixner
2026-09-11 13:58   ` Oleg Nesterov
2026-09-11  9:09 ` [patch V3 4/8] posix-timers: Move POSIX timer group exit related code out of do_exit() Thomas Gleixner
2026-09-11 13:59   ` Oleg Nesterov
2026-09-11  9:09 ` [patch V3 5/8] posix-cpu-timers: Move inlines out of public header Thomas Gleixner
2026-09-11 13:59   ` Oleg Nesterov
2026-09-11  9:09 ` [patch V3 6/8] posix-cpu-timers: Use PF_EXITING to indicate exit Thomas Gleixner
2026-09-11 13:15   ` Frederic Weisbecker
2026-09-11 14:00   ` Oleg Nesterov
2026-09-11  9:09 ` [patch V3 7/8] posix-cpu-timers: Prevent enqueueing when PF_EXITING is set Thomas Gleixner
2026-09-12 21:09   ` Frederic Weisbecker
2026-09-11  9:09 ` [patch V3 8/8] posix-timers: Handle exit in do_exit() completely Thomas Gleixner
2026-09-11 14:02   ` Oleg Nesterov
2026-09-12 22:17   ` Frederic Weisbecker
2026-09-12 14:20 ` [patch V3 0/8] exec/exit: POSIX timer related bugfixes and related cleanups Kijo Park

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=aqP9XYckv9mhLuQf@localhost.localdomain \
    --to=frederic@kernel.org \
    --cc=brauner@kernel.org \
    --cc=ebiederm@xmission.com \
    --cc=imv4bel@gmail.com \
    --cc=jstultz@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mingo@kernel.org \
    --cc=oleg@redhat.com \
    --cc=peterz@infradead.org \
    --cc=stable@vger.kernel.org \
    --cc=stern@rowland.harvard.edu \
    --cc=tglx@kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®