mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
@ 2026-10-04  3:03 Lance Yang
  2026-10-05 10:37 ` David Hildenbrand (Arm)
  0 siblings, 1 reply; 4+ messages in thread
From: Lance Yang @ 2026-10-04  3:03 UTC (permalink / raw)
  To: pjw, palmer, aou
  Cc: alex, akpm, zhangchunyan, rppt, kas, david, andrew+kernel,
	rmclure, debug, baolin.wang, usama.arif, wangruikang, namcao,
	linux-riscv, linux-kernel, alexghiti, viro, ajones, arnd,
	axelrasmussen, brauner, conor.dooley, conor, jack, liam, ljs,
	mhocko, paul.walmsley, peterx, robh, surenb, vbabka, yuanchu,
	stable, pasha.tatashin, linux-mm, me, Lance Yang

RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
a problem for PMD migration entries, since pmd_present() also checks
_PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
cleared during splitting.

When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
migration PMD as a present THP! The fault handler skips
pmd_migration_entry_wait(), and a write fault can end up in
do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
as a mapped PFN.

Move the swap soft-dirty bit to bit 12 and start the swap offset at
bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
in migration PMDs and lets us keep the existing pmd_present() check
for invalidated THPs. Leave the offset at bit 12 when soft-dirty
tracking is disabled.

Fixes: 2a3ebad4db63 ("riscv: mm: add soft-dirty page tracking support")
Cc: stable@vger.kernel.org
Signed-off-by: Lance Yang <lance.yang@linux.dev>
---
I found this while reviewing Usama's PMD-level swap entries series [1]
and comparing the swap flag bits with the PMD presence checks across
architectures.

[1] https://lore.kernel.org/all/20261002095503.3585565-1-usama.arif@linux.dev/

 arch/riscv/include/asm/pgtable-bits.h |  6 +++---
 arch/riscv/include/asm/pgtable.h      | 10 ++++++----
 2 files changed, 9 insertions(+), 7 deletions(-)

diff --git a/arch/riscv/include/asm/pgtable-bits.h b/arch/riscv/include/asm/pgtable-bits.h
index d5a86b4df3ce6..d84b459e010c1 100644
--- a/arch/riscv/include/asm/pgtable-bits.h
+++ b/arch/riscv/include/asm/pgtable-bits.h
@@ -27,12 +27,12 @@
 	((riscv_has_extension_unlikely(RISCV_ISA_EXT_SVRSW60T59B)) ?	\
 	 (1UL << 59) : 0)
 /*
- * Bit 3 is always zero for swap entry computation, so we
- * can borrow it for swap page soft-dirty tracking.
+ * Bit 12 is reserved for swap soft-dirty tracking. The swap offset
+ * starts at bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled.
  */
 #define _PAGE_SWP_SOFT_DIRTY						\
 	((riscv_has_extension_unlikely(RISCV_ISA_EXT_SVRSW60T59B)) ?	\
-	 _PAGE_EXEC : 0)
+	 (1UL << 12) : 0)
 #else
 #define _PAGE_SOFT_DIRTY	0
 #define _PAGE_SWP_SOFT_DIRTY	0
diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
index b644db16bda94..acaf6d3b54d04 100644
--- a/arch/riscv/include/asm/pgtable.h
+++ b/arch/riscv/include/asm/pgtable.h
@@ -1178,18 +1178,20 @@ static inline pud_t pud_modify(pud_t pud, pgprot_t newprot)
  *
  * Format of swap PTE:
  *	bit            0:	_PAGE_PRESENT (zero)
- *	bit       1 to 2:	(zero)
- *	bit            3:	_PAGE_SWP_SOFT_DIRTY
+ *	bit       1 to 3:	_PAGE_LEAF (zero)
  *	bit            4:	_PAGE_SWP_UFFD
  *	bit            5:	_PAGE_PROT_NONE (zero)
  *	bit            6:	exclusive marker
  *	bits      7 to 11:	swap type
- *	bits 12 to XLEN-1:	swap offset
+ *	bit           12:	_PAGE_SWP_SOFT_DIRTY (CONFIG_MEM_SOFT_DIRTY)
+ *	bits 12/13 to XLEN-1:	swap offset (without/with CONFIG_MEM_SOFT_DIRTY)
  */
 #define __SWP_TYPE_SHIFT	7
 #define __SWP_TYPE_BITS		5
 #define __SWP_TYPE_MASK		((1UL << __SWP_TYPE_BITS) - 1)
-#define __SWP_OFFSET_SHIFT	(__SWP_TYPE_BITS + __SWP_TYPE_SHIFT)
+#define __SWP_OFFSET_SHIFT \
+	(__SWP_TYPE_BITS + __SWP_TYPE_SHIFT + \
+	 IS_ENABLED(CONFIG_MEM_SOFT_DIRTY))
 
 #define MAX_SWAPFILES_CHECK()	\
 	BUILD_BUG_ON(MAX_SWAPFILES_SHIFT > __SWP_TYPE_BITS)
-- 
2.49.0


^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
  2026-10-04  3:03 [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present Lance Yang
@ 2026-10-05 10:37 ` David Hildenbrand (Arm)
  2026-10-05 13:46   ` Lance Yang
  0 siblings, 1 reply; 4+ messages in thread
From: David Hildenbrand (Arm) @ 2026-10-05 10:37 UTC (permalink / raw)
  To: Lance Yang, pjw, palmer, aou
  Cc: alex, akpm, zhangchunyan, rppt, kas, andrew+kernel, rmclure,
	debug, baolin.wang, usama.arif, wangruikang, namcao, linux-riscv,
	linux-kernel, alexghiti, viro, ajones, arnd, axelrasmussen,
	brauner, conor.dooley, conor, jack, liam, ljs, mhocko,
	paul.walmsley, peterx, robh, surenb, vbabka, yuanchu, stable,
	pasha.tatashin, linux-mm, me

On 10/4/26 05:03, Lance Yang wrote:
> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
> a problem for PMD migration entries, since pmd_present() also checks
> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
> cleared during splitting.

I'm curious: why do we have to set leaf indications for non-present things? The
HW sure will ignore it, right?

Is this a sw problem? Who needs that?

> 
> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
> migration PMD as a present THP! The fault handler skips
> pmd_migration_entry_wait(), and a write fault can end up in
> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
> as a mapped PFN.

That sounds bad.

> 
> Move the swap soft-dirty bit to bit 12 and start the swap offset at
> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
> in migration PMDs and lets us keep the existing pmd_present() check
> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
> tracking is disabled.

That reduces the effective swap size (and PFN we can store). Could that be a
problem?

> 
> Fixes: 2a3ebad4db63 ("riscv: mm: add soft-dirty page tracking support")
> Cc: stable@vger.kernel.org
> Signed-off-by: Lance Yang <lance.yang@linux.dev>
> ---
> I found this while reviewing Usama's PMD-level swap entries series [1]
> and comparing the swap flag bits with the PMD presence checks across
> architectures.
> 
> [1] https://lore.kernel.org/all/20261002095503.3585565-1-usama.arif@linux.dev/
> 
>  arch/riscv/include/asm/pgtable-bits.h |  6 +++---
>  arch/riscv/include/asm/pgtable.h      | 10 ++++++----
>  2 files changed, 9 insertions(+), 7 deletions(-)
> 
> diff --git a/arch/riscv/include/asm/pgtable-bits.h b/arch/riscv/include/asm/pgtable-bits.h
> index d5a86b4df3ce6..d84b459e010c1 100644
> --- a/arch/riscv/include/asm/pgtable-bits.h
> +++ b/arch/riscv/include/asm/pgtable-bits.h
> @@ -27,12 +27,12 @@
>  	((riscv_has_extension_unlikely(RISCV_ISA_EXT_SVRSW60T59B)) ?	\
>  	 (1UL << 59) : 0)
>  /*
> - * Bit 3 is always zero for swap entry computation, so we
> - * can borrow it for swap page soft-dirty tracking.
> + * Bit 12 is reserved for swap soft-dirty tracking. The swap offset
> + * starts at bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled.
>   */
>  #define _PAGE_SWP_SOFT_DIRTY						\
>  	((riscv_has_extension_unlikely(RISCV_ISA_EXT_SVRSW60T59B)) ?	\
> -	 _PAGE_EXEC : 0)
> +	 (1UL << 12) : 0)
>  #else
>  #define _PAGE_SOFT_DIRTY	0
>  #define _PAGE_SWP_SOFT_DIRTY	0
> diff --git a/arch/riscv/include/asm/pgtable.h b/arch/riscv/include/asm/pgtable.h
> index b644db16bda94..acaf6d3b54d04 100644
> --- a/arch/riscv/include/asm/pgtable.h
> +++ b/arch/riscv/include/asm/pgtable.h
> @@ -1178,18 +1178,20 @@ static inline pud_t pud_modify(pud_t pud, pgprot_t newprot)
>   *
>   * Format of swap PTE:
>   *	bit            0:	_PAGE_PRESENT (zero)
> - *	bit       1 to 2:	(zero)
> - *	bit            3:	_PAGE_SWP_SOFT_DIRTY
> + *	bit       1 to 3:	_PAGE_LEAF (zero)
>   *	bit            4:	_PAGE_SWP_UFFD
>   *	bit            5:	_PAGE_PROT_NONE (zero)
>   *	bit            6:	exclusive marker
>   *	bits      7 to 11:	swap type
> - *	bits 12 to XLEN-1:	swap offset
> + *	bit           12:	_PAGE_SWP_SOFT_DIRTY (CONFIG_MEM_SOFT_DIRTY)
> + *	bits 12/13 to XLEN-1:	swap offset (without/with CONFIG_MEM_SOFT_DIRTY)
>   */
>  #define __SWP_TYPE_SHIFT	7
>  #define __SWP_TYPE_BITS		5
>  #define __SWP_TYPE_MASK		((1UL << __SWP_TYPE_BITS) - 1)
> -#define __SWP_OFFSET_SHIFT	(__SWP_TYPE_BITS + __SWP_TYPE_SHIFT)
> +#define __SWP_OFFSET_SHIFT \
> +	(__SWP_TYPE_BITS + __SWP_TYPE_SHIFT + \
> +	 IS_ENABLED(CONFIG_MEM_SOFT_DIRTY))
>  
>  #define MAX_SWAPFILES_CHECK()	\
>  	BUILD_BUG_ON(MAX_SWAPFILES_SHIFT > __SWP_TYPE_BITS)


-- 
Cheers,

David

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
  2026-10-05 10:37 ` David Hildenbrand (Arm)
@ 2026-10-05 13:46   ` Lance Yang
  2026-10-05 14:23     ` Lance Yang
  0 siblings, 1 reply; 4+ messages in thread
From: Lance Yang @ 2026-10-05 13:46 UTC (permalink / raw)
  To: david
  Cc: lance.yang, pjw, palmer, aou, alex, akpm, zhangchunyan, rppt,
	kas, andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
	wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
	ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
	liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
	yuanchu, stable, pasha.tatashin, linux-mm, me


On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>On 10/4/26 05:03, Lance Yang wrote:
>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>> a problem for PMD migration entries, since pmd_present() also checks
>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>> cleared during splitting.
>
>I'm curious: why do we have to set leaf indications for non-present things? The
>HW sure will ignore it, right?

Yeah, that surprised me too :) Still wrapping my head around the details
...

>Is this a sw problem? Who needs that?

IIUC, it's for software during a PMD split.

__split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
before installing the PTE table. Software still needs pmd_present() and
pmd_trans_huge() to recognize the THP in between.

RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
to hardware but still identifiable as a THP by software.

Hopefully I didn't miss something.

>> 
>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>> migration PMD as a present THP! The fault handler skips
>> pmd_migration_entry_wait(), and a write fault can end up in
>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>> as a mapped PFN.
>
>That sounds bad.

YES, looks a bit off ...

>> 
>> Move the swap soft-dirty bit to bit 12 and start the swap offset at
>> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
>> in migration PMDs and lets us keep the existing pmd_present() check
>> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
>> tracking is disabled.
>
>That reduces the effective swap size (and PFN we can store). Could that be a
>problem?

We don't need all 52 bits of the swap offset.

RV64 PFNs only need 44 bits for migration entries, and actual swap is
already limited to about 16 TiB per area with 4 KiB pages by
last_page (__u32) and swap_info_struct.max (unsigned int).

So there's room to reserve a bit without reducing the supported swap
size or PFN range.

CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
offset.

[...]

Cheers, Lance

^ permalink raw reply	[flat|nested] 4+ messages in thread

* Re: [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present
  2026-10-05 13:46   ` Lance Yang
@ 2026-10-05 14:23     ` Lance Yang
  0 siblings, 0 replies; 4+ messages in thread
From: Lance Yang @ 2026-10-05 14:23 UTC (permalink / raw)
  To: david
  Cc: pjw, palmer, aou, alex, akpm, zhangchunyan, rppt, kas,
	andrew+kernel, rmclure, debug, baolin.wang, usama.arif,
	wangruikang, namcao, linux-riscv, linux-kernel, alexghiti, viro,
	ajones, arnd, axelrasmussen, brauner, conor.dooley, conor, jack,
	liam, ljs, mhocko, paul.walmsley, peterx, robh, surenb, vbabka,
	yuanchu, stable, pasha.tatashin, linux-mm, me



On 2026/10/5 21:46, Lance Yang wrote:
> 
> On Mon, Oct 05, 2026 at 12:37:20PM +0200, David Hildenbrand (Arm) wrote:
>> On 10/4/26 05:03, Lance Yang wrote:
>>> RISC-V uses _PAGE_EXEC for swap soft-dirty tracking when
>>> CONFIG_MEM_SOFT_DIRTY is enabled and Svrsw60t59b is available. That's
>>> a problem for PMD migration entries, since pmd_present() also checks
>>> _PAGE_LEAF (R/W/X) to recognize THPs with _PAGE_PRESENT temporarily
>>> cleared during splitting.
>>
>> I'm curious: why do we have to set leaf indications for non-present things? The
>> HW sure will ignore it, right?
> 
> Yeah, that surprised me too :) Still wrapping my head around the details
> ...
> 
>> Is this a sw problem? Who needs that?
> 
> IIUC, it's for software during a PMD split.
> 
> __split_huge_pmd_locked() invalidates the huge PMD and flushes the TLB
> before installing the PTE table. Software still needs pmd_present() and
> pmd_trans_huge() to recognize the THP in between.

Also, x86 keeps _PAGE_PSE to identify the huge PMD. RISC-V doesn't
have a separate leaf bit, so it keeps the R/W/X bits to identify
the PMD as a leaf entry :)

> 
> RISC-V clears V but keeps the R/W/X bits for that, so the entry is invalid
> to hardware but still identifiable as a THP by software.
> 
> Hopefully I didn't miss something.
> 
>>>
>>> When a soft-dirty THP is migrated, set_pmd_migration_entry() preserves
>>> soft-dirty with pmd_swp_mksoft_dirty(), setting the X bit in the
>>> migration PMD. Even with _PAGE_PRESENT clear, we end up treating a
>>> migration PMD as a present THP! The fault handler skips
>>> pmd_migration_entry_wait(), and a write fault can end up in
>>> do_huge_pmd_wp_page(), where pmd_page() decodes the migration entry
>>> as a mapped PFN.
>>
>> That sounds bad.
> 
> YES, looks a bit off ...
> 
>>>
>>> Move the swap soft-dirty bit to bit 12 and start the swap offset at
>>> bit 13 when CONFIG_MEM_SOFT_DIRTY is enabled. This keeps R/W/X clear
>>> in migration PMDs and lets us keep the existing pmd_present() check
>>> for invalidated THPs. Leave the offset at bit 12 when soft-dirty
>>> tracking is disabled.
>>
>> That reduces the effective swap size (and PFN we can store). Could that be a
>> problem?
> 
> We don't need all 52 bits of the swap offset.
> 
> RV64 PFNs only need 44 bits for migration entries, and actual swap is
> already limited to about 16 TiB per area with 4 KiB pages by
> last_page (__u32) and swap_info_struct.max (unsigned int).
> 
> So there's room to reserve a bit without reducing the supported swap
> size or PFN range.
> 
> CONFIG_MEM_SOFT_DIRTY is only available on RV64, so RV32 keeps its 20-bit
> offset.
> 
> [...]
> 
> Cheers, Lance


^ permalink raw reply	[flat|nested] 4+ messages in thread

end of thread, other threads:[~2026-10-05 14:24 UTC | newest]

Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-10-04  3:03 [PATCH 1/1] riscv/mm: fix soft-dirty migration PMDs being treated as present Lance Yang
2026-10-05 10:37 ` David Hildenbrand (Arm)
2026-10-05 13:46   ` Lance Yang
2026-10-05 14:23     ` Lance Yang

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®