mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: "wangwensheng (C)" <wangwensheng4@huawei.com>
To: Saravana Kannan <saravanak@google.com>,
	Greg KH <gregkh@linuxfoundation.org>
Cc: <rafael@kernel.org>, <dakr@kernel.org>, <robh@kernel.org>,
	<broonie@kernel.org>, <linux-kernel@vger.kernel.org>,
	<chenjun102@huawei.com>
Subject: Re: [PATCH v2] driver core: Fix concurrent problem of deferred_probe_extend_timeout()
Date: Fri, 15 Aug 2025 09:56:29 +0800	[thread overview]
Message-ID: <40fa16cf-950b-4ca7-9935-dbce75e46eb9@huawei.com> (raw)
In-Reply-To: <CAGETcx_-otRyDknVs4SFVWzf5-Zi07TiKUEpetDJJ3r0BTVqmw@mail.gmail.com>



在 2025/8/15 2:16, Saravana Kannan 写道:
> On Thu, Aug 14, 2025 at 7:20 AM Greg KH <gregkh@linuxfoundation.org> wrote:
>>
>> On Thu, Aug 14, 2025 at 08:29:49PM +0800, Wang Wensheng wrote:
>>> The deferred_probe_timeout_work may be canceled forever unexpected when
>>> deferred_probe_extend_timeout() executes concurrently. Start with
>>> deferred_probe_timeout_work pending, and the problem would
>>> occur after the following sequence.
>>>
>>>           CPU0                                 CPU1
>>> deferred_probe_extend_timeout
>>>    -> cancel_delayed_work => true
>>>                                       deferred_probe_extend_timeout
>>>                                         -> cancel_delayed_wrok
>>>                                           -> __cancel_work
>>>                                             -> try_grab_pending
>>>    -> schedule_delayed_work
>>>     -> queue_delayed_work_on
>>> since pending bit is grabbed,
>>> just return without doing anything
>>>                                          -> set_work_pool_and_clear_pending
>>>                                       this __cancel_work return false and
>>>                                       the work would never be queued again
>>>
>>> The root cause is that the PENDING_BIT of the work_struct would be set
>>> temporaily in __cancel_work and this bit could prevent the work_struct
>>> to be queued in another CPU.
> 
> This feels more like a workqueue API issue (this isn't too obvious
> from the documentation) or me misusing the workqueue API.
> 
> Is this issue still there if you use cancel_delayed_work_sync()
> instead of cancel_delayed_work()? If so, just switch to that and add
> proper comment on why it needs to by "sync".
> 
> -Saravana
> 
cancel_delayed_work_sync() cannot solve the issue. Becasue this issue is 
to do with the interaction between cancel and queue operations for a 
work. The synchronization of the single cancel operation doesn't matter.

>>>
>>> Use deferred_probe_mutex to protect the cancel and queue operations for
>>> the deferred_probe_timeout_work to fix this problem.
>>>
>>> Fixes: 2b28a1a84a0e ("driver core: Extend deferred probe timeout on driver registration")
>>> Cc: <stable@vger.kernel.org>
>>> Signed-off-by: Wang Wensheng <wangwensheng4@huawei.com>
>>> ---
>>>   drivers/base/dd.c | 1 +
>>>   1 file changed, 1 insertion(+)
>>>
>>> diff --git a/drivers/base/dd.c b/drivers/base/dd.c
>>> index 13ab98e033ea..00419d2ee910 100644
>>> --- a/drivers/base/dd.c
>>> +++ b/drivers/base/dd.c
>>> @@ -323,6 +323,7 @@ static DECLARE_DELAYED_WORK(deferred_probe_timeout_work, deferred_probe_timeout_
>>>
>>>   void deferred_probe_extend_timeout(void)
>>>   {
>>> +     guard(mutex)(&deferred_probe_mutex);
>>
>> But if you grab the lock here, in the probe timeout function, the lock
>> will be grabbed again, causing a deadlock, right?  If not, why not?
>>
>> Have you run this patch with lockdep enabled?
>>
>> This feels broken to me, what am I missing?
>>
>> thanks,
>>
>> greg k-h
> 


  reply	other threads:[~2025-08-15  1:56 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2025-08-14 12:29 Wang Wensheng
2025-08-14 14:19 ` Greg KH
2025-08-14 18:16   ` Saravana Kannan
2025-08-15  1:56     ` wangwensheng (C) [this message]
2025-08-15  1:43   ` wangwensheng (C)

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=40fa16cf-950b-4ca7-9935-dbce75e46eb9@huawei.com \
    --to=wangwensheng4@huawei.com \
    --cc=broonie@kernel.org \
    --cc=chenjun102@huawei.com \
    --cc=dakr@kernel.org \
    --cc=gregkh@linuxfoundation.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=rafael@kernel.org \
    --cc=robh@kernel.org \
    --cc=saravanak@google.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®