From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-alma10-1.taild15c8.ts.net [100.103.45.18]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id C6F81389100 for ; Fri, 21 Aug 2026 17:40:45 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=100.103.45.18 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787334052; cv=none; b=R6y9PBd+ud+4IUCMs5P45wpuC7p/NqguOq7CcYffuibiHmFKNJTqDE8lPx0KHfwyZR1N97vqd5Oiv9ASRRULdDPiq61Ua8Bx8pqqSY0Mll4sBOHGAsYCRHs3SH0xWcX0c6wJR+gwB57+8d2fx8COIZt4Z8RuRjOKP1PyNV3b25Q= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787334052; c=relaxed/simple; bh=Y738IklUAyhahdIqmzReU7xUaDTIZdc8MQbmE4S9qFs=; h=Date:From:To:Cc:Subject:Message-Id:In-Reply-To:References: Mime-Version:Content-Type; b=fZeVdTJ+hWVi9/YY+L4eZBRp60KU8Ii26/RVGoMbQXBOTB7KrhRGiXbpirTDd7faMV3/FoK0S5Nfz+Y2FoHY9201tLN6H2poOK0dZe9AYCJC1cbX+0Km7cDz5ZlGFN5yWVngb1ej9S7R6g5fXxXFFVU9Rf0WXtB8O+TvVrw4nhE= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b=LsjNT4Gx; arc=none smtp.client-ip=100.103.45.18 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=linux-foundation.org header.i=@linux-foundation.org header.b="LsjNT4Gx" Received: by smtp.kernel.org (Postfix) with ESMTPSA id EA46A1F000E9; Fri, 21 Aug 2026 17:40:43 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=linux-foundation.org; s=korg; t=1787334044; bh=hpocrNT/gbphpP2uFGBQ6vPM0EP71cQkKN+OiupwOJs=; h=Date:From:To:Cc:Subject:In-Reply-To:References; b=LsjNT4GxUrx/fs/tLKpvdqHBFGBdL/p4ML74ZQcFzP3t86xFFPWcoHvzctu/VHV7L amJ2q7KxElIo65xyGPV86jcqVSV97pGVDwGJJYgEl2MWaZy01R7fvQHcBB7PU4BEcJ NtT0tGAfvYR5lm7Yl9/lGj6CDy8EdkaXsgSgzeNs= Date: Fri, 21 Aug 2026 10:40:43 -0700 From: Andrew Morton To: Eric Dumazet Cc: linux-kernel , syzbot+0dbf6d295b3350944f0b@syzkaller.appspotmail.com, David Hildenbrand , Zi Yan , Matthew Brost , Joshua Hahn , Rakie Kim , Byungchul Park , Gregory Price , Ying Huang , Alistair Popple , linux-mm@kvack.org Subject: Re: [PATCH] mm/mempolicy: Fix sleeping allocation in alloc_pages_bulk_weighted_interleave() Message-Id: <20260821104043.f692421fec915c0c5bc1fbe6@linux-foundation.org> In-Reply-To: <20260821170407.3721004-1-edumazet@google.com> References: <20260821170407.3721004-1-edumazet@google.com> X-Mailer: Sylpheed 3.8.0beta1 (GTK+ 2.24.33; x86_64-pc-linux-gnu) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Fri, 21 Aug 2026 17:04:07 +0000 Eric Dumazet wrote: > syzbot reported a sleeping function called from invalid context splat > in bucket_table_alloc(). That was quick (7 minutes!). I was just looking at this. > When rhashtable_insert_slow() rehashes the table under rcu_read_lock(), > it calls bucket_table_alloc(..., GFP_ATOMIC | __GFP_NOWARN). > If the bucket table allocation uses vmalloc, __vmalloc_node_range_noprof() > invokes vm_area_alloc_pages() -> alloc_pages_bulk_mempolicy_noprof() with > the passed GFP_ATOMIC flags. > > If the current task has an MPOL_WEIGHTED_INTERLEAVE mempolicy, > alloc_pages_bulk_weighted_interleave() is called and currently hardcodes > GFP_KERNEL when allocating the temporary weights array, triggering > a might_alloc() splat in atomic/RCU contexts. 2 years ago. Why are we discovering this now? > Pass the gfp flags (masked with GFP_RECLAIM_MASK to strip page-allocator > zone modifiers like __GFP_HIGHMEM) received by > alloc_pages_bulk_weighted_interleave() to kmalloc() instead of > hardcoding GFP_KERNEL. Since the weights buffer is immediately > initialized in full, kmalloc() is sufficient. > > Fixes: fa3bea4e1f82 ("mm/mempolicy: introduce MPOL_WEIGHTED_INTERLEAVE for weighted interleaving") I'll add cc:stable > --- a/mm/mempolicy.c > +++ b/mm/mempolicy.c > @@ -2688,7 +2688,7 @@ static unsigned long alloc_pages_bulk_weighted_interleave(gfp_t gfp, > prev_node = node; > > /* create a local copy of node weights to operate on outside rcu */ > - weights = kzalloc(nr_node_ids, GFP_KERNEL); > + weights = kmalloc(nr_node_ids, gfp & GFP_RECLAIM_MASK); lgtm, thanks. I wonder if we *really* need the local copy of state->iw_table. Perhaps with appropriate care we can directly use state->iw_table in here. How much would it hurt to expand the rcu_read_lock() coverage? A local array of MAX_NUMNODES bytes isn't attractive - 1k of stack. A spinlock-protected static array would work, if super-rare slowpath. > if (!weights) > return total_allocated;