mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Mahanta Jambigi <mjambigi@linux.ibm.com>
To: Chengfeng Ye <nicoyip.dev@gmail.com>,
	"D . Wythe" <alibuda@linux.alibaba.com>,
	Dust Li <dust.li@linux.alibaba.com>,
	Sidraya Jayagond <sidraya@linux.ibm.com>
Cc: Tony Lu <tonylu@linux.alibaba.com>,
	Wen Gu <guwen@linux.alibaba.com>,
	"David S . Miller" <davem@davemloft.net>,
	Eric Dumazet <edumazet@kernel.org>,
	Jakub Kicinski <kuba@kernel.org>, Paolo Abeni <pabeni@redhat.com>,
	Simon Horman <horms@kernel.org>,
	linux-rdma@vger.kernel.org, linux-s390@vger.kernel.org,
	netdev@vger.kernel.org, linux-kernel@vger.kernel.org,
	stable@vger.kernel.org
Subject: Re: [PATCH net v2] net/smc: serialize clcsock access with its release
Date: Thu, 8 Oct 2026 13:01:49 +0530	[thread overview]
Message-ID: <4bb7f15e-2cd8-4598-9bab-1ea5f5d2bb59@linux.ibm.com> (raw)
In-Reply-To: <20261003183325.2289707-1-nicoyip.dev@gmail.com>



On 04/10/26 12:03 am, Chengfeng Ye wrote:
> Link-group termination can release the CLC socket through
> smc_close_active_abort() while the SMC socket is still open, for example
> after shutdown(SHUT_WR). A file reference keeps the SMC socket alive but
> does not prevent this asynchronous release of its CLC socket.
> 
> smc_getname() can race the pointer removal and sock_release(). The same
> missing lifetime synchronization affects smc_set_keepalive(), diagnostic
> address copying and an in-flight SMC-R CDC receiver. smc_shutdown() can
> also reach its final CLC shutdown after its close helper drops the socket
> lock and a concurrent abort closes the socket. Holding the SMC socket
> lock alone is insufficient because CLC release runs outside that lock.
> 
> KASAN reported the getname failure:
> 
>   BUG: KASAN: slab-use-after-free in smc_getname+0x19e/0x1b0
>   Read of size 8 at addr ffff888109abb4e0 by task poc/103
>   Call Trace:
>    smc_getname+0x19e/0x1b0
>    do_getsockname+0xe5/0x170
>    __sys_getsockname+0x8c/0x100
> 
>   Allocated by task 95:
>    sock_alloc_inode+0x1e/0x280
>    sock_alloc+0x3d/0x240
>    __sock_create+0x7e/0x430
>    smc_create+0x121/0x240
> 
>   Freed by task 0:
>    kmem_cache_free+0xcc/0x340
>    rcu_core+0x50a/0x1850
> 
>   Last potentially related work creation:
>    evict+0x446/0x6c0
>    smc_clcsock_release+0xa8/0xd0
>    smc_close_active_abort+0x26a/0x3a0
>    __smc_lgr_terminate.part.0+0x137/0x2e0

Hi Chengfeng,

Thanks for the KASAN report and the fix. The UAF in smc_getname is real
and needs to go to net and stable. However, I think the v2 approach of
adding a new spinlock and extending clcsock_release_lock to cover more
readers is treating symptoms rather than the root cause. Let me explain
what I think should happen instead. What to keep from your patch.

Please send a v3 with only the smc_getname fix:

int smc_getname(struct socket *sock, struct sockaddr *addr,
		int peer)
{
	struct smc_sock *smc;
+	int rc = -EBADF;

	if (peer && (sock->sk->sk_state != SMC_ACTIVE) &&
	    (sock->sk->sk_state != SMC_APPCLOSEWAIT1))
		return -ENOTCONN;

	smc = smc_sk(sock->sk);

-	return smc->clcsock->ops->getname(smc->clcsock, addr, peer);
+	mutex_lock(&smc->clcsock_release_lock);
+	if (smc->clcsock)
+		rc = smc->clcsock->ops->getname(smc->clcsock, addr, peer);
+	mutex_unlock(&smc->clcsock_release_lock);
+	return rc;
}

That is 4 lines against the confirmed KASAN-reported UAF, uses
infrastructure that already exists (clcsock_release_lock is already held
by smc_clcsock_release() when it frees clcsock, and already initialized
in smc_sk_init()), and is a clean candidate for stable. Nothing else
from v2 is needed for this specific bug.

Drop the clcsock_lock spinlock, the CDC change, the diag change, the
shutdown change, the connect-abort change, and the smc_accept_dequeue
change. Those races are real but I will address them with a proper
structural fix as Me & Dust Li have already discussed on LKML in August[1].

Why the other races exist and what the right fix is

Every race in your v2 — keepalive, CDC, diag, shutdown, the
connect-abort path — has the same root cause: clcsock can be freed while
the SMC socket is still alive. Several close paths call
sock_release(clcsock) before the SMC socket's own refcount reaches zero:

1) smc_close_active_abort() — for PEERCLOSEWAIT*, PROCESSABORT,
APPFINCLOSEWAIT states
2) smc_close_passive_work() — when the passive close work transitions to
SMC_CLOSED
3) __smc_release() — when sk_state == SMC_CLOSED

Every access site that can race with those releases is then forced to
take clcsock_release_lock and check if (!smc->clcsock). Your v2 adds a
second lock on top of this for the BH/atomic readers that cannot take a
mutex. This complexity is unnecessary because none of those early paths
actually need to destroy the socket — they only need to stop it.
tcp_abort() and kernel_sock_shutdown() are sufficient for that, and both
are safe to call more than once. sock_release() is the exception: it
frees memory and must happen exactly once.

The fix is to move that single sock_release() call to smc_destruct() —
the sk->sk_destruct callback that fires from __sk_free() when the last
sock reference drops. At that point no concurrent user can exist:

1) fd users are gone: smc_release() calls sock_orphan() before dropping
its reference, so no file descriptor can reach the socket after that point
2) workqueue contexts (close_work, smc_listen_work) hold a sock_hold()
and therefore keep smc_destruct() from running while they are active
3) accept-queue entries hold a sock_hold() via smc_accept_enqueue() for
the same reason

This gives us a simple invariant: clcsock is non-NULL for the entire
lifetime of the SMC socket. With that invariant every reader becomes
trivially safe — no lock needed, no NULL check needed, the race
condition simply cannot occur.

[1] https://lore.kernel.org/netdev/ao5bB9OCbJ5PQbEp@linux.alibaba.com/

  parent reply	other threads:[~2026-10-08  7:32 UTC|newest]

Thread overview: 4+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-03 18:33 Chengfeng Ye
2026-10-08  2:17 ` Jakub Kicinski
2026-10-08  7:31 ` Mahanta Jambigi [this message]
2026-10-08  8:47   ` Chengfeng Ye

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=4bb7f15e-2cd8-4598-9bab-1ea5f5d2bb59@linux.ibm.com \
    --to=mjambigi@linux.ibm.com \
    --cc=alibuda@linux.alibaba.com \
    --cc=davem@davemloft.net \
    --cc=dust.li@linux.alibaba.com \
    --cc=edumazet@kernel.org \
    --cc=guwen@linux.alibaba.com \
    --cc=horms@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-rdma@vger.kernel.org \
    --cc=linux-s390@vger.kernel.org \
    --cc=netdev@vger.kernel.org \
    --cc=nicoyip.dev@gmail.com \
    --cc=pabeni@redhat.com \
    --cc=sidraya@linux.ibm.com \
    --cc=stable@vger.kernel.org \
    --cc=tonylu@linux.alibaba.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®