From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-yw1-f173.google.com (mail-yw1-f173.google.com [209.85.128.173]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E635233B6C4 for ; Wed, 19 Aug 2026 20:33:57 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.173 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787171639; cv=none; b=gV4MoMOjkXhwnV4zf+WR9pnRC6ZUksNRBEfxES9HwyblDb0WLmBr16kbVIQga/ccUj4jDt7YaYWGT/ooO8k4LlmIam+lHJAuinZF+QP5EmfFJhKia96v1HwfyeoS4LK/5Va8L3d92JVHkE2aVq/KipWSOcHuwksnQTl5xYtYZ3k= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1787171639; c=relaxed/simple; bh=bvijR6UOVgGrT1+tBf081NUxEFiDBwRj6/iaXp2oNgM=; h=Date:From:To:cc:Subject:In-Reply-To:Message-ID:References: MIME-Version:Content-Type; b=nCCPgnrOTJN6SqPysC/Y4icmeGgGmQVl8oxUcSy1jPOab9AgBrj+QUlR1F9JSWvgc1mYqGnUmSQLrbdK8VcRPm7x6ddk44OU9sVi6twj8HQxo50sb7OlqLRB8msIAu/iP1ghj+eBb00JRsGCUGuKIkvBnqK0Chpvea3Vu+dRPos= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=ozPJuFDo; arc=none smtp.client-ip=209.85.128.173 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="ozPJuFDo" Received: by mail-yw1-f173.google.com with SMTP id 00721157ae682-836c590b61eso25231697b3.0 for ; Wed, 19 Aug 2026 13:33:57 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1787171637; x=1787776437; darn=vger.kernel.org; h=content-type:mime-version:references:message-id:in-reply-to:subject :cc:to:from:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=JJDZ3zG6BST+HSWSWwE8GsdZoVMnM7ou/ROzox6VqsE=; b=ozPJuFDov2/LqWqL1UOQtOwU3dLIdQN7ZyILPofc15tchfbFofFAeH2HvFRL0317ZH CHZjcp2eBLOTbJAUcUcaSTWgWsWEV4mMpMMKUXWPpzz4FE5ZFbf3wd28+NjmO+3huy/4 9pxGvy2k/y708aQpSNK6JsBmakyNb3kLFuX+LHaOtCdJS2H0EDynzq0PsaL4NUR1B835 GGmgR/hhEiloYBgBUyAE8lRuW9KpWBaP882u7ua/k9r1+4/B50+U0B9CqXfFWAfYRIGI QWQ36Dzs3UK6no73kbvo8P444j8G3Z8zKMOfvgvxQF2NBNc/IxHR8TeCSY0jBz9B3rdu CTvw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1787171637; x=1787776437; h=content-type:mime-version:references:message-id:in-reply-to:subject :cc:to:from:date:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=JJDZ3zG6BST+HSWSWwE8GsdZoVMnM7ou/ROzox6VqsE=; b=Da6KDROVtjf0ZZZvPSCvJ+efdVnsyMooPwKputES3Ze/DBuNkc4Eq0W44B3NyWffzc SL8Q0ZVz5L000qg6gG2GXUHdutNJ6yGJfwV6E2GUB+7um3oGtbsBHecXsKUDUvaHd2qy ncUNo3sS1dewar8gQNeSUR8pTdrDW8exVtQgRlvXEuYwwDLfQvpEgUJ/piQuhkcgdJ8Z RUAzJnTXnL0n5+Xuo01GKvsFywZFZW+2Ai4NAKXbak5pAszMkh7q0JZrZrxZrbAzi7AF WiXs0AiQ0mXIPJNhSAXdisSN5N44QZMzkg4Amwgf+GjwPdakaI7aqcg3aFGV/AX63yBF IZ6g== X-Forwarded-Encrypted: i=1; AHgh+RoUagYAYZvjUfa/YI+WhxkDDBIsWqUKVPFNivguaRZ8BpQvdCRZTOoxukKD73x/xLw7gV8XjFCXXImbPFs=@vger.kernel.org X-Gm-Message-State: AFuF++le4Q3+OJkOKGhixSifwOXtpRmbdfkK4Mo2B0eI8KF2Z/tgN/rn lsqBk2Sz9UN8p8TBonYkVTcFCcfEC+7HW+39eE+HWazCe8ByOQQQ/5RSKzkjSuoNkg== X-Gm-Gg: AR+sD10/slee8tDsMry/tOLJa1JkJvgA+CJtvo8R4QZRlai7SgYkIcUUCM91WPv8eOm ghWmhP1QtjmYPtU0LHZczLlcmFXo3Gvvfa521vV5PcVTSje8WyEemEwNsp3qzms63UEjRDXKXCu 6ewgCx58T6zX6enlJNC/G2JK0wN/8AqJFetq7tReYPcTzurKqCx4xMddGI2jY3I76wuEKWluKZZ u+0SIxNbYSB2qcqtsJBgRB4zRIlbJAopXHAIsBgxsrD5HgX3I5CVuq2Y08VZluVW87FlCavsks0 HqraZYmPinj/gQ3h/wdSOmcgLCHfC1zWlycKND8hoZ6zfDUBCVOaxQORyf4dnLbSnCHv5Pfygqg 2sUHv3nMWOI+RLZZpaxD5mihnhmcZVi/+vPXZnXOcGz9jGNa8el5cuhYgJOPIAcUsByoipdtsWk AoaVzxe6OXHJQUi8x33nHldhrMtq+JW2/WFkINSK5M6NSiepmuV9FUv29jkhW3epG2GogUpHalQ L8oj42cFtaAIMk8IyG3GUJGHScJH72zhh/e87r4NH2FR2Z6 X-Received: by 2002:a05:690e:1a46:b0:663:9a0a:7b80 with SMTP id 956f58d0204a3-66ccb61cec0mr2093887d50.19.1787171636227; Wed, 19 Aug 2026 13:33:56 -0700 (PDT) Received: from darker.attlocal.net (172-10-233-147.lightspeed.sntcca.sbcglobal.net. [172.10.233.147]) by smtp.gmail.com with ESMTPSA id 00721157ae682-84512e347e8sm15186657b3.20.2026.08.19.13.33.53 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 19 Aug 2026 13:33:55 -0700 (PDT) Date: Wed, 19 Aug 2026 13:33:42 -0700 (PDT) From: Hugh Dickins To: Kiryl Shutsemau cc: Usama Arif , Hugh Dickins , Andrew Morton , baohua@kernel.org, baolin.wang@linux.alibaba.com, david@kernel.org, dev.jain@arm.com, lance.yang@linux.dev, liam@infradead.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, ljs@kernel.org, nico.pache@linux.dev, ryan.roberts@arm.com, ziy@nvidia.com, nphamcs@gmail.com, hannes@cmpxchg.org, riel@surriel.com, shakeel.butt@linux.dev, kernel-team@meta.com, stable@vger.kernel.org Subject: Re: [PATCH] mm/huge_memory: transfer the pmd dirty bit to the folio on zap In-Reply-To: Message-ID: <5bd5c5fe-0d55-9fec-aadf-477891428887@google.com> References: <20260819101222.3732660-1-usama.arif@linux.dev> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII On Wed, 19 Aug 2026, Kiryl Shutsemau wrote: > On Wed, Aug 19, 2026 at 03:12:22AM -0700, Usama Arif wrote: > > zap_huge_pmd_folio() propagates the pmd young bit to the folio for the > > file case, but not the dirty bit. The pte path does propagate it, in > > zap_present_folio_ptes() and so does the pmd split path, in > > __split_huge_pmd_locked(). > > > > For most file mappings the omission is harmless, because writing to a > > shared file mapping goes through page_mkwrite(), which dirties the > > folio. tmpfs is different: it has no page_mkwrite(), and > > vma_wants_writenotify() is false for it, so a *read* fault on a > > MAP_SHARED tmpfs mapping installs a writable pmd via do_read_fault(). > > do_read_fault() does not call fault_dirty_shared_page(), so subsequent > > stores through that mapping set only the hardware dirty bit in the pmd > > and never call folio_mark_dirty(). > > > > A shmem folio allocated by a fault > > is marked uptodate but not dirty (see the clear: block in > > shmem_get_folio_gfp()), so PG_dirty is never set at all. > > > > Unmapping such a folio - munmap(), or exit_mmap() when the process dies > > - then loses the only record that it was written, because zap_huge_pmd() > > drops the pmd without transferring the dirty bit. Reclaim afterwards > > sees a clean shmem folio: the whole swap-out block in > > shrink_folio_list() is inside "if (folio_test_dirty(folio))", so > > pageout() is skipped and the folio falls into __remove_mapping(). > > There, folio_is_file_lru() is false for a swapbacked folio, so no shadow > > entry is created and __filemap_remove_folio(folio, NULL) simply empties > > the i_pages slot. The data is freed without ever being written to swap, > > and the next fault on that index returns a freshly zeroed folio. > > > > This is silent data loss for any process that keeps state in a > > MAP_SHARED tmpfs segment across an unmap - for example a cache handed > > from one process generation to the next through /dev/shm. It requires > > the folio to be PMD-mapped, so it only shows up once shmem THP is > > enabled (which is what we did in Meta fleet and started noticing crashes); > > with THP off the pte path transfers the dirty bit correctly. > > It also only becomes visible when swap is enabled, because with no swap > > device shmem folios (which are on the anon LRU) are not scanned by > > reclaim at all, so the clean folio is never dropped. > > > > Reproduced on x86_64 with a tmpfs mounted huge=within_size: read-fault a > > 2MB-backed region, write a known pattern through the resulting mapping, > > munmap, force reclaim of the cgroup, then re-map and read back. Without > > this patch the region reads back as zeros and vmstat shows zswpout 0 - > > the data was discarded rather than swapped. With this patch the region > > reads back correctly and the pages are swapped out as expected. With > > huge=never, or when the first touch is a write, the test passes either > > way. > > +Hugh. > > Oopsie. How ghastly! Thanks for finding and fixing, Usama. > > I'm confused why it took a decade to discover the bug... > Maybe read ahead of write for shmem is too rare, I donno. It isn't entirely clear from the report, but this is all about modifying a 2MiB+ *hole* in a shmem file through an mmap thereof, with first fault a read fault not a write fault. I suppose only a few proceed in that way (though truncating an empty file to some size and then mmap'ing that size is very normal). Or everybody who tried to report this bug, wrote their report into a 2MiB hole in a shmem file through an mmap thereof. Other than holes, all shmem folios are dirty throughout (and re-marked dirty as soon as brought back from swap): so for most, it doesn't matter what the pmd says. This raised a dim memory, took a while to locate what I was remembering: e1f1b1572e8d ("mm/huge_memory.c: fix data loss when splitting a file pmd") from 2018. Not quite the same; but what a pity that one didn't prompt any of us to look further and find what Usama now has. > > > > > Fixes: 800d8c63b2e9 ("shmem: add huge pages support") > > This would be more precise: b5072380eb61 ("thp: support file pages in zap_huge_pmd()") > > Reviewed-by: Kiryl Shutsemau > > > Cc: > > Signed-off-by: Usama Arif Acked-by: Hugh Dickins