mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Jiayuan Chen <jiayuan.chen@linux.dev>
To: Kuniyuki Iwashima <kuniyu@google.com>, bot+bpf-ci@kernel.org
Cc: bpf@vger.kernel.org, daniel@iogearbox.net,
	john.fastabend@gmail.com, sdf@fomichev.me, martin.lau@linux.dev,
	ast@kernel.org, andrii@kernel.org, eddyz87@gmail.com,
	memxor@gmail.com, song@kernel.org, yonghong.song@linux.dev,
	jolsa@kernel.org, emil@etsalapatis.com, ihor.solodrai@linux.dev,
	davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, horms@kernel.org, ncardwell@google.com,
	shuah@kernel.org, aditi.ghag@isovalent.com,
	netdev@vger.kernel.org, linux-kernel@vger.kernel.org,
	linux-kselftest@vger.kernel.org, martin.lau@kernel.org,
	mason@kernel.org
Subject: Re: [PATCH bpf v2 2/3] tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context
Date: Tue, 8 Sep 2026 20:22:02 +0800	[thread overview]
Message-ID: <c6c4fd69-3074-4d6d-804b-cf8ff62359da@linux.dev> (raw)
In-Reply-To: <a44c1e4d-e9a8-4df3-aa25-20875cd4d30e@linux.dev>


On 9/8/26 4:07 PM, Jiayuan Chen wrote:
>
> On 9/8/26 7:34 AM, Kuniyuki Iwashima wrote:
>> On Sun, Sep 6, 2026 at 1:23 AM <bot+bpf-ci@kernel.org> wrote:
>>>> tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context
>>>>
>>>> bpf_sock_destroy() runs from the tcp iterator, under 
>>>> rcu_read_lock(). If
>>>> the sock is a listener that still has children in its accept queue,
>>>> tcp_abort() ends up in inet_csk_listen_stop() and the cond_resched()
>>>> there trips the debug check:
>>>>
>>>> BUG: sleeping function called from invalid context at 
>>>> net/ipv4/inet_connection_sock.c:1523
>>>> in_atomic(): 0, irqs_disabled(): 0, non_block: 0, pid: 628, name: 
>>>> test_progs
>>>> preempt_count: 0, expected: 0
>>>> RCU nest depth: 1, expected: 0
>>>> locks held by test_progs/628: 3, last CPU#3:
>>>>   #0: ffff8881158cee18 (&p->lock){+.+.}-{4:4}, at: 
>>>> bpf_seq_read+0x56/0x1210
>>>>   #1: ffff8881106bb858 (sk_lock-AF_INET6){+.+.}-{0:0}, at: 
>>>> bpf_iter_tcp_seq_show+0x32b/0x4b0
>>>>   #2: ffffffffb435af20 (rcu_read_lock){....}-{1:3}, at: 
>>>> bpf_iter_run_prog+0x46b/0xde0
>>>> CPU: 3 UID: 0 PID: 628 Comm: test_progs Tainted: G W           
>>>> 7.2.0+ #65 PREEMPT
>>>> Tainted: [W]=WARN
>>>> Call Trace:
>>>>   <TASK>
>>>>   dump_stack_lvl+0xc1/0xf0
>>>>   dump_stack+0x10/0x20
>>>>   __might_resched+0x3d2/0x610
>>>>   inet_csk_listen_stop+0x7b/0xbf0
>>>>   tcp_abort+0x23b/0x3b0
>>>>   bpf_sock_destroy+0xfc/0x140
>>>>   bpf_prog_448133d24601754f_iter_tcp6_server+0x81/0x8a
>>>>   bpf_iter_run_prog+0x538/0xde0
>>>>   bpf_iter_tcp_seq_show+0x26b/0x4b0
>>>>   bpf_seq_read+0x424/0x1210
>>>>   vfs_read+0x197/0xe40
>>>>   ksys_read+0x119/0x240
>>>>   __x64_sys_read+0x72/0xc0
>>>>   x64_sys_call+0x647/0x27e0
>>>>   do_syscall_64+0xe5/0x610
>>>>   entry_SYSCALL_64_after_hwframe+0x76/0x7e
>>>> RIP: 0033:0x7fad39b28aca
>>>> RSP: 002b:00007ffc381c61c0 EFLAGS: 00000246 ORIG_RAX: 0000000000000000
>>>> RAX: ffffffffffffffda RBX: 00007ffc381c6a88 RCX: 00007fad39b28aca
>>>> RDX: 0000000000000032 RSI: 00007ffc381c6250 RDI: 0000000000000014
>>>> RBP: 00007ffc381c61e0 R08: 0000000000000000 R09: 0000000000000000
>>>> R10: 0000000000000000 R11: 0000000000000246 R12: 0000000000000003
>>>> R13: 0000000000000000 R14: 000055f077c1bbb0 R15: 00007fad3a0f3000
>>>>   </TASK>
>>>>
>>>> The commit that added the kfunc already guards lock_sock() in 
>>>> tcp_abort()
>>>> and udp_abort() with has_current_bpf_ctx(), but missed the listener 
>>>> path.
>>>> Do the same for the cond_resched(), it can't reschedule there anyway.
>>>>
>>>> Fixes: 4ddbcb886268 ("bpf: Add bpf_sock_destroy kfunc")
>>>> Signed-off-by: Jiayuan Chen <jiayuan.chen@linux.dev>
>>> Is the justification "it can't reschedule there anyway" accurate?
>>>
>>> With CONFIG_PREEMPT_DYNAMIC=y booted with preempt=none or 
>>> preempt=voluntary,
>>> cond_resched() expands to __cond_resched() which can actually 
>>> reschedule.
>>> The splat in the commit message confirms preempt_count is 0 while 
>>> RCU nest
>>> depth is 1. With preempt_count==0, should_resched(0) can be true and
>>> __cond_resched() will call preempt_schedule_common() for a real 
>>> reschedule.
>>>
>>> Additionally, in configurations with CONFIG_PREEMPT_RCU=n where
>>> rcu_read_lock() is preempt_disable(), __cond_resched() falls through to
>>> rcu_all_qs() which calls rcu_qs() to report a quiescent state from 
>>> inside
>>> an RCU read-side critical section. That's a correctness problem 
>>> beyond just
>>> the debug check.
>>>
>>> So the call can either reschedule (PREEMPT_DYNAMIC none/voluntary) 
>>> or report
>>> a bogus quiescent state (non-preemptible RCU). Could the 
>>> justification be
>>> reworded to explain that the loop runs inside the iterator's RCU 
>>> read-side
>>> critical section and must not reschedule or report a quiescent state 
>>> there?
>>> The code change itself is correct and matches the existing pattern in
>>> tcp_abort() and udp_abort().
>> or maybe simply remove cond_resched(), hoping 7dadeaa6e851 would
>> resolve the scheduling issue.
>>
>
> Good suggestion. cond_resched has become old practice under 
> CONFIG_PREEMPT_LAZY


After reconsideration, I think it's not a good idea to remove it if we 
treat it as a fix and the fix will be backported to LTS.

Or we just drop both cond_resched and Fixes tag.



  reply	other threads:[~2026-09-08 12:22 UTC|newest]

Thread overview: 9+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-06  7:41 [PATCH bpf v2 0/3] bpf,tcp: Fix bpf_sock_destroy() on TIME_WAIT and listener socks Jiayuan Chen
2026-09-06  7:41 ` [PATCH bpf v2 1/3] bpf: Fix out-of-bounds read of sk_protocol in bpf_sock_destroy() Jiayuan Chen
2026-09-07 23:23   ` Kuniyuki Iwashima
2026-09-06  7:41 ` [PATCH bpf v2 2/3] tcp: Skip cond_resched() in inet_csk_listen_stop() under BPF context Jiayuan Chen
2026-09-06  8:23   ` bot+bpf-ci
2026-09-07 23:34     ` Kuniyuki Iwashima
2026-09-08  8:07       ` Jiayuan Chen
2026-09-08 12:22         ` Jiayuan Chen [this message]
2026-09-06  7:41 ` [PATCH bpf v2 3/3] selftests/bpf: Test bpf_sock_destroy() on TIME_WAIT and listener socks Jiayuan Chen

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=c6c4fd69-3074-4d6d-804b-cf8ff62359da@linux.dev \
    --to=jiayuan.chen@linux.dev \
    --cc=aditi.ghag@isovalent.com \
    --cc=andrii@kernel.org \
    --cc=ast@kernel.org \
    --cc=bot+bpf-ci@kernel.org \
    --cc=bpf@vger.kernel.org \
    --cc=daniel@iogearbox.net \
    --cc=davem@davemloft.net \
    --cc=eddyz87@gmail.com \
    --cc=edumazet@google.com \
    --cc=emil@etsalapatis.com \
    --cc=horms@kernel.org \
    --cc=ihor.solodrai@linux.dev \
    --cc=john.fastabend@gmail.com \
    --cc=jolsa@kernel.org \
    --cc=kuba@kernel.org \
    --cc=kuniyu@google.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=martin.lau@kernel.org \
    --cc=martin.lau@linux.dev \
    --cc=mason@kernel.org \
    --cc=memxor@gmail.com \
    --cc=ncardwell@google.com \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=sdf@fomichev.me \
    --cc=shuah@kernel.org \
    --cc=song@kernel.org \
    --cc=yonghong.song@linux.dev \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®