From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-ej1-f41.google.com (mail-ej1-f41.google.com [209.85.218.41]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 42979306B0A for ; Tue, 3 Mar 2026 08:28:31 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.218.41 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772526512; cv=none; b=H55Hm5Rz+B59cNctN8L32g3yR9mPOzwUD1gww/neSRwOHhyCHtntLcAxB78N7A6EjQmjOh7UZ+1gd83q3QwnO4SS56Bd6O9sJme7wuleIQsaRYpj0owLB5g/V7hPeopNfXrQbcg+I9aqhWNxSCXGoJDzrwoNjDYFojZ74I9827g= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1772526512; c=relaxed/simple; bh=Aj5EdWjVF2jxAv9z+0TobEuciQMJWsagsW0+QMnh1Sw=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=XU83CwAXusB5/UHrEXTTY1HXzPvq0/mrloQ/Q5mDBP+oUZdaXQYU41k7i7BL+cYE/tCna4oJs6EVO1NUc7pUrQUjuydRC+xs90bxHOgXf3ooMyDv5iIvrCLL5qs1cCrdUQeBft9EvARj90+EfUe/A19zhabXFTrk85QsPNKxozI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=auuNx0TI; arc=none smtp.client-ip=209.85.218.41 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="auuNx0TI" Received: by mail-ej1-f41.google.com with SMTP id a640c23a62f3a-b8fa449e618so772549866b.0 for ; Tue, 03 Mar 2026 00:28:31 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1772526510; x=1773131310; darn=vger.kernel.org; h=user-agent:in-reply-to:content-disposition:mime-version:references :reply-to:message-id:subject:cc:to:from:date:from:to:cc:subject:date :message-id:reply-to; bh=tvQ9ogjqY7GpWZR0jhfLALG/vdl5wh/KcgbqbjmLgoM=; b=auuNx0TITUBFiDejhfcWAgaKfZGi66w9epcPluFifmQQVMHeAbWmAbxxpkd00+voNs VaLtWNSYtcrnH1IfNzIocbnswtPjEnfpYwfmd1gqQ9s8eMB7uuISwzOjt9fotju89BYB /h0GTQR6udvguK8EEgTsIZtJPn35KqpS+Vfm0Hhh7gsKpTcggxfebbDMl/EnL9M+m+xA 4G0eAdK05eb0D/9BZdnub7EdoCnrSVgfuVVe7xGi9/z2mxcjz+2stb0un8I4hcUoVGxw bLf/Br2IUi7heuQHX3Pfk3TBLK3EsG5LzoS4pki/7ia0j0qT52QWIdTYGgaJ1bqX/+0+ HETg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1772526510; x=1773131310; h=user-agent:in-reply-to:content-disposition:mime-version:references :reply-to:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=tvQ9ogjqY7GpWZR0jhfLALG/vdl5wh/KcgbqbjmLgoM=; b=lIb0kb4OB20DuTNzVlLhkrqTaUqCrRHd9qMtylCWrq/dAn12Ap6C8p3FIZBKC5Ta0C zKjxfeRs//4F5gVvH52fU7ZNI9UHKH8UzcMuWOpLjA5Qhk7TmBz22U2IW2N5tNWwELzN 1BaijS1lX7XFwycm35cZeFo3JSZ8XDE/kFDxXkryuV0h8Wm4h2JSgxQPpXmYuJ8bBSxS LOK0DLhUXh+xhABrGIpGlf41Pqv0+n2q3e8zK4XI8Iz9qsc79UmeY3UlhVYuACOWZicO EDvZMGI5CKZ0UN9YD3YIzTwqHJhPi/K6x6R9+WiW4VJB53QHHc93I/yO63EhqlinZPhG ezXQ== X-Forwarded-Encrypted: i=1; AJvYcCUDCHK25g/h7fk3Eo0/jldm3c73bIkHMAvBnQN3xH/5smDsj36Nbbdz2Mxg9nSRmgMpbMk85UND3OYIIzM=@vger.kernel.org X-Gm-Message-State: AOJu0YwxT6vJLXO7rpn0qCeat3zKakqIWPTtY/RNvT6NpmnYATW552YU FGAv7EFNhUTfOjQ+NI/sBzgi/dittggxsCW8M/d5u+w4WRZqELh6Q5+8 X-Gm-Gg: ATEYQzwylSN4u4CY53S1ioZUZQLynhTpAuNIXThT5PvAdwjSvqbSSueX/tmnpBG7IrN 4UQTaBqdXBlBXQh2B0GiC0lqBLAZJUGspoiS8ldCYXtP5oC56LXEA80uYeXA/1emULb5oDcQV3Q Zr19BngomrNcrYK27wUoVAA6dYO/ikkNoYBYkkmS3YPprjTgSF3/nD6aZyH04pwWFlF1uwkl08f pX/0B3XquZWkoDTggYhKxKxVCHwg77PdFNOCHkwXSLNEOH3oFqQBcK7jlR6jI4m3qqBvUGt7PoH 0v2MYFtgTP0145C+DSreOEK3uvxoVEQAl0XSgTvE56oNXkjeh5dN9x2shbWzQaBQ5QqANH6SI/V Y2XCoC9P3C1L7wpJ3j55587YhJ5CndHOmsPJk25tDAMXdjSISqAmyKuD0UkxWo85KgCT7i2P/S5 Xcaj/DfLmBfIc6nV0Bfx3MYQ== X-Received: by 2002:a17:907:9492:b0:b93:94b9:26fe with SMTP id a640c23a62f3a-b9394b98ceamr771288166b.52.1772526509327; Tue, 03 Mar 2026 00:28:29 -0800 (PST) Received: from localhost ([185.92.221.13]) by smtp.gmail.com with ESMTPSA id a640c23a62f3a-b935ac70b01sm559271466b.23.2026.03.03.00.28.28 (version=TLS1_2 cipher=ECDHE-ECDSA-CHACHA20-POLY1305 bits=256/256); Tue, 03 Mar 2026 00:28:28 -0800 (PST) Date: Tue, 3 Mar 2026 08:28:28 +0000 From: Wei Yang To: Zi Yan Cc: Andrew Morton , David Hildenbrand , Lorenzo Stoakes , Hugh Dickins , Baolin Wang , "Liam R. Howlett" , Nico Pache , Ryan Roberts , Dev Jain , Barry Song , Lance Yang , Matthew Wilcox , Bas van Dijk , Eero Kelly , Andrew Battat , Adam Bratschi-Kaye , linux-mm@kvack.org, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, stable@vger.kernel.org Subject: Re: [PATCH v2] mm/huge_memory: fix a folio_split() race condition with folio_try_get() Message-ID: <20260303082828.x2gypytceqn6pb6x@master> Reply-To: Wei Yang References: <20260302203159.3208341-1-ziy@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260302203159.3208341-1-ziy@nvidia.com> User-Agent: NeoMutt/20170113 (1.7.2) On Mon, Mar 02, 2026 at 03:31:59PM -0500, Zi Yan wrote: >During a pagecache folio split, the values in the related xarray should not >be changed from the original folio at xarray split time until all >after-split folios are well formed and stored in the xarray. Current use >of xas_try_split() in __split_unmapped_folio() lets some after-split folios >show up at wrong indices in the xarray. When these misplaced after-split >folios are unfrozen, before correct folios are stored via __xa_store(), and >grabbed by folio_try_get(), they are returned to userspace at wrong file >indices, causing data corruption. More detailed explanation is at the >bottom. > >The reproducer is at: https://github.com/dfinity/thp-madv-remove-test >It >1. creates a memfd, >2. forks, >3. in the child process, maps the file with large folios (via shmem code > path) and reads the mapped file continuously with 16 threads, >4. in the parent process, uses madvise(MADV_REMOVE) to punch poles in the > large folio. > >Data corruption can be observed without the fix. Basically, data from a >wrong page->index is returned. > >Fix it by using the original folio in xas_try_split() calls, so that >folio_try_get() can get the right after-split folios after the original >folio is unfrozen. > >Uniform split, split_huge_page*(), is not affected, since it uses >xas_split_alloc() and xas_split() only once and stores the original folio >in the xarray. Change xas_split() used in uniform split branch to use >the original folio to avoid confusion. > >Fixes below points to the commit introduces the code, but folio_split() is >used in a later commit 7460b470a131f ("mm/truncate: use folio_split() in >truncate operation"). > >More details: > >For example, a folio f is split non-uniformly into f, f2, f3, f4 like >below: >+----------------+---------+----+----+ >| f | f2 | f3 | f4 | >+----------------+---------+----+----+ >but the xarray would look like below after __split_unmapped_folio() is >done: >+----------------+---------+----+----+ >| f | f2 | f3 | f3 | >+----------------+---------+----+----+ > Thanks for the detailed explanation, I finally realized it behaves like this. >After __split_unmapped_folio(), the code changes the xarray and unfreezes >after-split folios: > >1. unfreezes f2, __xa_store(f2) >2. unfreezes f3, __xa_store(f3) >3. unfreezes f4, __xa_store(f4), which overwrites the second f3 to f4. >4. unfreezes f. > >Meanwhile, a parallel filemap_get_entry() can read the second f3 from the >xarray and use folio_try_get() on it at step 2 when f3 is unfrozen. Then, >f3 is wrongly returned to user. > >After the fix, the xarray looks like below after __split_unmapped_folio(): >+----------------+---------+----+----+ >| f | f | f | f | >+----------------+---------+----+----+ >so that the race window no longer exists. Since we unfreeze f at last. > >Fixes: 00527733d0dc8 ("mm/huge_memory: add two new (not yet used) functions for folio_split()") >Signed-off-by: Zi Yan >Reported-by: Bas van Dijk >Closes: https://lore.kernel.org/all/CAKNNEtw5_kZomhkugedKMPOG-sxs5Q5OLumWJdiWXv+C9Yct0w@mail.gmail.com/ >Tested-by: Lance Yang >Cc: So thanks for the fix. Reviewed-by: Wei Yang -- Wei Yang Help you, Help me