mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Cen Zhang <zzzccc427@gmail.com>
To: mark@fasheh.com, jlbec@evilplan.org, joseph.qi@linux.alibaba.com,
	akpm@linux-foundation.org, moonafterrain@outlook.com,
	brauner@kernel.org, rppt@kernel.org, kees@kernel.org,
	julia.lawall@inria.fr, kurt.hackel@oracle.com
Cc: ocfs2-devel@lists.linux.dev, linux-kernel@vger.kernel.org,
	baijiaju1990@gmail.com, jjzuming@gmail.com, zzzccc427@gmail.com
Subject: [PATCH] ocfs2: drain domain handlers before destroying the DLM worker
Date: Fri,  9 Oct 2026 16:11:23 +0800	[thread overview]
Message-ID: <pm-ocfs2-objects-candidate-0056-v2-57239c657ac017b50fa1@gmail.com> (raw)

The domain workqueue must remain usable until every network handler that
can submit work has finished. dlm_unregister_domain_handlers() removes
handlers from the o2net lookup tree, but an already referenced handler can
still run. Its dlm_grab() reference pins the context, not dlm_worker.

On the final local domain disconnect, a migration receive callback can
have passed dlm_joined() before shutdown changes the domain state, yet
still be preparing its work item when teardown destroys the queue:

  o2net receive worker             Final domain teardown
  --------------------             ---------------------
  Get the handler reference
  dlm_grab(); pass dlm_joined()
                                   Unregister domain handlers
                                   Stop the DLM threads
                                   destroy_workqueue(dlm_worker)
                                   dlm_worker = NULL
  Publish the migration work item
  queue_work(dlm_worker, ...)

The handler lookup reference lets the callback continue after unregister.
The work-list lock does not protect the queue lifetime, and destroying the
queue drains submitted work without waiting for this producer. The late
queue_work() therefore passes NULL to __queue_work() and crashes.

Use o2net_unregister_and_flush_handler_list() at the existing domain
handler unregister point. Removing the handlers prevents new lookups,
and flushing o2net receive work waits for already referenced callbacks
and their post handlers. They can finish submitting while dlm_worker is
still live; the subsequent destroy_workqueue() drains those submissions.
The same unregister helper covers failed registration and join cleanup,
without changing the DLM thread or workqueue teardown order.

KASAN report as below:

    Oops: general protection fault, probably for non-canonical address 0xdffffc0000000038: 0000 [#1] SMP KASAN NOPTI
    KASAN: null-ptr-deref in range [0x00000000000001c0-0x00000000000001c7]
    CPU: 0 UID: 0 PID: 14 Comm: kworker/u8:1 Not tainted 7.3.0-rc4-next-20260921-pmb-bt-functional-v1+ #1 PREEMPT(lazy)

[Hardware details omitted.]

    Workqueue: o2net o2net_rx_until_empty
    RIP: 0010:__queue_work+0x9b/0x1600

[Instruction and register dump omitted.]

    Call Trace:
     <TASK>
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? __pfx___queue_work+0x10/0x10
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? clear_pending_if_disabled+0x83/0x1c0
     ? __pfx_clear_pending_if_disabled+0x10/0x10
     ? __pfx_pmbd_probe_hit_cookie+0x10/0x10
     ? dlm_mig_lockres_handler+0x984/0x1500
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? lock_release+0xc8/0x290
     queue_work_on+0xda/0xf0
     dlm_mig_lockres_handler+0x9d5/0x1500
     ? percpu_rwsem_wake_function+0x10/0x480
     ? __pfx_dlm_mig_lockres_handler+0x10/0x10
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? trace_hardirqs_on+0x18/0x160
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? kvm_clock_get_cycles+0x31/0x60
     ? srso_alias_return_thunk+0x5/0xfbef5
     o2net_rx_until_empty+0x1a55/0x32f0
     ? reacquire_held_locks+0xdd/0x200
     ? __pfx_o2net_rx_until_empty+0x10/0x10
     ? lock_acquire+0x190/0x300
     ? process_one_work+0x935/0x1b40
     ? process_one_work+0x834/0x1b40
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? lock_release+0xc8/0x290
     ? srso_alias_return_thunk+0x5/0xfbef5
     process_one_work+0x9a8/0x1b40
     ? __pfx_process_one_work+0x10/0x10
     ? lock_acquire+0x190/0x300
     ? lock_is_held_type+0x8f/0x100
     ? srso_alias_return_thunk+0x5/0xfbef5
     worker_thread+0x65c/0xe40
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? __kthread_parkme+0x177/0x220
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? __pfx_worker_thread+0x10/0x10
     kthread+0x351/0x460
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? __pfx_kthread+0x10/0x10
     ret_from_fork+0x659/0x940
     ? __pfx_ret_from_fork+0x10/0x10
     ? srso_alias_return_thunk+0x5/0xfbef5
     ? __switch_to+0x74f/0xf70
     ? __pfx_kthread+0x10/0x10
     ret_from_fork_asm+0x1a/0x30
     </TASK>

[Empty module list omitted.]

    ---[ end trace 0000000000000000 ]---

Fixes: 3156d2670166 ("ocfs2: move dlm work to a private work queue")
Assisted-by: LLM
Signed-off-by: Cen Zhang <zzzccc427@gmail.com>
---

diff --git a/fs/ocfs2/dlm/dlmdomain.c b/fs/ocfs2/dlm/dlmdomain.c
index 97bb9400e24bbf6badea1339dd8de1623fe5bac5..b69baf99a4a6e4d5a5c51664625bafa905dc3bf5 100644
--- a/fs/ocfs2/dlm/dlmdomain.c
+++ b/fs/ocfs2/dlm/dlmdomain.c
@@ -1708,7 +1708,7 @@ static void dlm_unregister_domain_handlers(struct dlm_ctxt *dlm)
 {
 	o2hb_unregister_callback(dlm->name, &dlm->dlm_hb_up);
 	o2hb_unregister_callback(dlm->name, &dlm->dlm_hb_down);
-	o2net_unregister_handler_list(&dlm->dlm_domain_handlers);
+	o2net_unregister_and_flush_handler_list(&dlm->dlm_domain_handlers);
 }
 
 static int dlm_register_domain_handlers(struct dlm_ctxt *dlm)

                 reply	other threads:[~2026-10-09  8:11 UTC|newest]

Thread overview: [no followups] expand[flat|nested]  mbox.gz  Atom feed

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=pm-ocfs2-objects-candidate-0056-v2-57239c657ac017b50fa1@gmail.com \
    --to=zzzccc427@gmail.com \
    --cc=akpm@linux-foundation.org \
    --cc=baijiaju1990@gmail.com \
    --cc=brauner@kernel.org \
    --cc=jjzuming@gmail.com \
    --cc=jlbec@evilplan.org \
    --cc=joseph.qi@linux.alibaba.com \
    --cc=julia.lawall@inria.fr \
    --cc=kees@kernel.org \
    --cc=kurt.hackel@oracle.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mark@fasheh.com \
    --cc=moonafterrain@outlook.com \
    --cc=ocfs2-devel@lists.linux.dev \
    --cc=rppt@kernel.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®