From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from lgeamrelo11.lge.com (lgeamrelo11.lge.com [156.147.23.51]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id F210D3D1CB2 for ; Mon, 20 Jul 2026 09:12:01 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=156.147.23.51 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784538726; cv=none; b=EJk6Vy8q0iuEm7YGzV1Urq9fx4GFrkkUT73FV3CYK5bE7jfV3dEilF4/FFr/Djy2ictoNtXar4fb4K37RMOb749HsDAW21NmR2JSmZzqhBU1LO4gsp6482mwQysqZe5N8Rd1nykjs3+4la8tYhZcWDf6DDQxPO4YtpeSc2VCHXk= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1784538726; c=relaxed/simple; bh=9LuwYAX+veg+96PECbolo8yRo3lxHB7dR1ltRSvUxIQ=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=rhofHJuP8tZc1GtfkDfwUr72LyNHrkKjOJfA/+SUv/WFKk2h/jk9yK7SGL41vExCWOYJX23HSV0Tbtjc5KJdj8ePM5Fgbm5aNVlvA5hDzCHeLgR6H56jXOTISAWykXfoLtbTRRiVJ66k6UlnMG9SeS6j6p77OauhUeWpee60Vjw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=lge.com; spf=pass smtp.mailfrom=lge.com; arc=none smtp.client-ip=156.147.23.51 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=lge.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=lge.com Received: from unknown (HELO lgeamrelo01.lge.com) (156.147.1.125) by 156.147.23.51 with ESMTP; 20 Jul 2026 18:11:54 +0900 X-Original-SENDERIP: 156.147.1.125 X-Original-MAILFROM: youngjun.park@lge.com Received: from unknown (HELO yjaykim-PowerEdge-T330) (10.177.112.156) by 156.147.1.125 with ESMTP; 20 Jul 2026 18:11:54 +0900 X-Original-SENDERIP: 10.177.112.156 X-Original-MAILFROM: youngjun.park@lge.com Date: Mon, 20 Jul 2026 18:11:54 +0900 From: Youngjun Park To: Kemeng Shi Cc: chrisl@kernel.org, kasong@tencent.com, nphamcs@gmail.com, baoquan.he@linux.dev, baohua@kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: [PATCH 1/4] mm, swap: Fix potential NULL dereference when trying a sleep table allocation Message-ID: References: <20260720071342.50742-1-shikemeng@huaweicloud.com> <20260720071342.50742-2-shikemeng@huaweicloud.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <20260720071342.50742-2-shikemeng@huaweicloud.com> On Mon, Jul 20, 2026 at 03:13:39PM +0800, Kemeng Shi wrote: Hello Kemeng Shi, Good catch! it indeed looks like a possible race condition (though I haven't verified it at runtime either). > The root cause of this issue is because multi-tables are updated in non > atomic context. To be more specific, the issue could be triggerred as > following: Here the cluster is isolated (CLUSTER_FLAG_NONE). > swap_alloc_fast swap_cluster_populate() > /* Try a sleep allocation */ > spin_unlock(&ci->lock); And from here, the table becomes visible, > swap_cluster_alloc_table() > rcu_assign_pointer(ci->table, table); All entry on this cluster is freed(e.g process dead) , so this might be the free cluster. > ci = swap_cluster_lock(si, offset) > cluster_is_usable(ci, order) Since it's a normal cluster (CLUSTER_FLAG_NONE) and the table exists, it passes here... > if (!cluster_table_is_alloced(ci)) // ok > alloc_swap_scan_cluster() > cluster_scan_range() > __swap_table_get() > > /* free table when more table allocation fails */ > ci->memcg_table = kzalloc_obj(*ci->memcg_table, > gfp); > if (!ci->memcg_table) > swap_cluster_free_table() Nullified > rcu_assign_pointer(ci->table, NULL); Now it happens. > table = rcu_dereference_check(ci->table, lockdep_is_held(&ci->lock)); > atomic_long_read(&table[off]); // NULL dereference > > Fix the issue by updating allocated tables in atomic context. > > Fixes: 2fe7a6f5024b8 ("mm/memcg, swap: store cgroup id in cluster table directly") > Signed-off-by: Kemeng Shi > --- > mm/swapfile.c | 21 ++++++++++++++++++++- > 1 file changed, 20 insertions(+), 1 deletion(-) > > diff --git a/mm/swapfile.c b/mm/swapfile.c > index 615d90867111..d29062d9c3cd 100644 > --- a/mm/swapfile.c > +++ b/mm/swapfile.c > @@ -490,6 +490,20 @@ static int swap_cluster_alloc_table(struct swap_cluster_info *ci, gfp_t gfp) > return 0; > } > > +static void swap_cluster_copy_table(struct swap_cluster_info *d_ci, > + struct swap_cluster_info *s_ci) > +{ > + rcu_assign_pointer(d_ci->table, rcu_access_pointer(s_ci->table)); > + > +#ifdef CONFIG_MEMCG > + d_ci->memcg_table = s_ci->memcg_table; > +#endif > + > +#if !SWAP_TABLE_HAS_ZEROFLAG > + d_ci->zero_bitmap = s_ci->zero_bitmap; > +#endif > +} > + > /* > * Sanity check to ensure nothing leaked, and the specified range is empty. > * One special case is that bad slots can't be freed, so check the number of > @@ -527,6 +541,7 @@ static struct swap_cluster_info * > swap_cluster_populate(struct swap_info_struct *si, > struct swap_cluster_info *ci) > { > + struct swap_cluster_info tmp_ci; > int ret; IMHO, How about resolving everything inside swap_cluster_alloc_table() instead? We could consider the following two things. - Assign the table at the very end inside swap_cluster_alloc_table(). - Handle the table freeing properly if subsequent allocations fail. Rather than allocating a temp_ci (which adds special handling for this case), wouldn't it be better to maintain the original intention of the swap_cluster_alloc_table() function? Although this is not a hot path and table allocation will usually succeed with the ATOMIC allocator, this approach would also save the memcpy overhead and stack memory usage. (Another option might be checking memcg_table and zero_bitmap inside cluster_table_is_alloced(). However, I personally dislike this idea as it might introduce side effects regarding its coverage.) Your current approach is also a good direction, but please review this idea as well :) Thanks, Youngjun