* [PATCH v4 0/3] Optimize S2 hugepage splitting, introduce skip-level flags
@ 2026-09-30 17:22 Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 1/3] KVM: arm64: Avoid re-testing walk_continue Leonardo Bras
` (2 more replies)
0 siblings, 3 replies; 4+ messages in thread
From: Leonardo Bras @ 2026-09-30 17:22 UTC (permalink / raw)
To: Marc Zyngier, Oliver Upton, Fuad Tabba, Joey Gouly,
Steffen Eiden, Suzuki K Poulose, Zenghui Yu, Catalin Marinas,
Will Deacon, Mark Rutland, Leonardo Bras, Raghavendra Rao Ananta
Cc: linux-arm-kernel, kvmarm, linux-kernel
While playing with dirty-bit tracking, I decided to take a look on how page
splitting works. Found out all entries are walked, even though we don't need
to walk the level-3 entries, as they don't need to be split.
This patches' idea is to introduce new walking flags to skip pagetable
levels 0-3.
Optimization measured on two scenarios involving eager-splitting on a
VM with 32 memslot of 2GB (total 64GB), and vcpu per slot:
- Scenario 1: No manual protect, whole memslot split at dirty-track enable
(KVM_SET_USER_MEMORY_REGION2 ioctl with KVM_MEM_LOG_DIRTY_PAGES)
- Split happens only once, whole region
- Evalutes improved batch performance of splitting
- Scenario 2: Manual protect, split happens during every dirty-bit clean
(KVM_CLEAR_DIRTY_LOG ioctl), average for 2 iterations.
- Split called multiple times, for smaller 64-page sections.
- Evaluate improved performance for multiple calls
Scenario 1, improvement on dirty-track enable ioctl for the memslot:
- Memory was already split (4k pages): -47.82% runtime
- THP backed memory: -27.50% runtime
- 64x1GB hugetlb memory: -28.47% runtime
Scenario 2, improvement on dirty-log clean ioctl for the memslot:
- Memory was already split (4k pages): -44.72% runtime
- THP backed memory: -29.80% runtime
- 64x1GB hugetlb memory: -30.01% runtime
For collecting above numbers, the following script was ran in both vanilla
and patched kernels, with kernel parameter 'default_hugepagesz=1G', on an
TX2 with 128GB RAM.
--- dirty_test.sh
#!/bin/bash
filename=$(uname -r |cut -d'-' -f 4-)
run_test(){
base_test="./dirty_log_perf_test -b 2G -v 32 -m 6 -m 8"
# Manual cleaning disable
${base_test} -g
${base_test} -g -s anonymous_thp
echo 64 > /proc/sys/vm/nr_hugepages
${base_test} -g -s shared_hugetlb
echo 0 > /proc/sys/vm/nr_hugepages
# Manual cleaning enable
${base_test}
${base_test} -s anonymous_thp
echo 64 > /proc/sys/vm/nr_hugepages
${base_test} -s shared_hugetlb
echo 0 > /proc/sys/vm/nr_hugepages
}
run_test 2>&1 | tee ${filename}
---
Above dirty_log_perf_test command is the standard kvm selftest found in the
kernel tree. It tested the following guest modes:
Testing guest mode: PA-bits:40, VA-bits:48, 4K pages
Testing guest mode: PA-bits:40, VA-bits:48, 64K pages
(Modes with PA-bits:36 were discarted in this version, given the amount
of RAM being used for testing, and the similarity of previous results)
Performance numbers from above modes were used to calculate average showed
in the optimization improvements.
Changes since v3:
- Check if root level should be skipped,
- Improve commit messages (Marc)
- Improve skip_level documentation (Marc & Wei Lin)
- Improved skip_level code (Marc)
- Dropped skip_children explanation in the cover letter (Dev)
- Rebased on top of v7.3-rc5
v3 Link: https://lore.kernel.org/all/20260708134101.2514759-1-leo.bras@arm.com/
Changes since v2:
- Rebased on top of v7.2-rc1
- Improved testing, added more memory, re-tested
- Now: 32 vcpus @ total of 64G
- Before: 1cpu @ 16G
v2 Link: https://lore.kernel.org/all/20260618131447.764085-1-leo.bras@arm.com/
Changes since v1:
- Fixed inverted flag verification priority (Sashiko)
- Fixed incorrectly skipping POST call if level was skipped (Sashiko), and to that
- New pre-patch that changes goto-out -> return to avoid re-testing walk_continue
v1 Link: https://lore.kernel.org/lkml/20260610202112.2695205-2-leo.bras@arm.com/
Changes since RFC:
- Changed approach from return value to walk flags (Will Deacon)
- Discarted skip_child approach (Oliver Upton)
- Measured in real hardware, and from userspace perspective (Marc Zyngier)
- Better explanation of what and how numbers were collected
RFC Link: https://lore.kernel.org/all/20260515195904.2466381-1-leo.bras@arm.com/
Thanks!
Leo
Leonardo Bras (3):
KVM: arm64: Avoid re-testing walk_continue
KVM: arm64: Introduce KVM_PGTABLE_WALK_SKIP_LEVEL* walk flags
KVM: arm64: Make stage2_split_walker() skip unnecessary walks
arch/arm64/include/asm/kvm_pgtable.h | 14 ++++++++++++++
arch/arm64/kvm/hyp/pgtable.c | 25 ++++++++++++++++++-------
2 files changed, 32 insertions(+), 7 deletions(-)
base-commit: 72d3fcf802c45d00b300f25b848a93c3a2bd7c7e
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH v4 1/3] KVM: arm64: Avoid re-testing walk_continue
2026-09-30 17:22 [PATCH v4 0/3] Optimize S2 hugepage splitting, introduce skip-level flags Leonardo Bras
@ 2026-09-30 17:22 ` Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 2/3] KVM: arm64: Introduce KVM_PGTABLE_WALK_SKIP_LEVEL* walk flags Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 3/3] KVM: arm64: Make stage2_split_walker() skip unnecessary walks Leonardo Bras
2 siblings, 0 replies; 4+ messages in thread
From: Leonardo Bras @ 2026-09-30 17:22 UTC (permalink / raw)
To: Marc Zyngier, Oliver Upton, Fuad Tabba, Joey Gouly,
Steffen Eiden, Suzuki K Poulose, Zenghui Yu, Catalin Marinas,
Will Deacon, Mark Rutland, Leonardo Bras, Raghavendra Rao Ananta
Cc: linux-arm-kernel, kvmarm, linux-kernel
__kvm_pgtable_visit() performs a bunch of calls to
kvm_pgtable_walk_continue() to find out whether the walk can continue
further, and if not, 'goto out', which retests the possibility of
continuing the walk before exiting.
Given that it's testing the same ret variable again, there is no reason
the result of kvm_pgtable_walk_continue() would have changed since the
previous check. So turn this goto into an early return, simplifying the
code and paving the way for further rework."
Signed-off-by: Leonardo Bras <leo.bras@arm.com>
---
arch/arm64/kvm/hyp/pgtable.c | 4 ++--
1 file changed, 2 insertions(+), 2 deletions(-)
diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
index b74dd5ce1efd..7a51f78b6446 100644
--- a/arch/arm64/kvm/hyp/pgtable.c
+++ b/arch/arm64/kvm/hyp/pgtable.c
@@ -183,32 +183,32 @@ static inline int __kvm_pgtable_visit(struct kvm_pgtable_walk_data *data,
* Reload the page table after invoking the walker callback for leaf
* entries or after pre-order traversal, to allow the walker to descend
* into a newly installed or replaced table.
*/
if (reload) {
ctx.old = READ_ONCE(*ptep);
table = kvm_pte_table(ctx.old, level);
}
if (!kvm_pgtable_walk_continue(data->walker, ret))
- goto out;
+ return ret;
if (!table) {
data->addr = ALIGN_DOWN(data->addr, kvm_granule_size(level));
data->addr += kvm_granule_size(level);
goto out;
}
childp = (kvm_pteref_t)kvm_pte_follow(ctx.old, mm_ops);
ret = __kvm_pgtable_walk(data, mm_ops, childp, level + 1);
if (!kvm_pgtable_walk_continue(data->walker, ret))
- goto out;
+ return ret;
if (ctx.flags & KVM_PGTABLE_WALK_TABLE_POST)
ret = kvm_pgtable_visitor_cb(data, &ctx, KVM_PGTABLE_WALK_TABLE_POST);
out:
if (kvm_pgtable_walk_continue(data->walker, ret))
return 0;
return ret;
}
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH v4 2/3] KVM: arm64: Introduce KVM_PGTABLE_WALK_SKIP_LEVEL* walk flags
2026-09-30 17:22 [PATCH v4 0/3] Optimize S2 hugepage splitting, introduce skip-level flags Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 1/3] KVM: arm64: Avoid re-testing walk_continue Leonardo Bras
@ 2026-09-30 17:22 ` Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 3/3] KVM: arm64: Make stage2_split_walker() skip unnecessary walks Leonardo Bras
2 siblings, 0 replies; 4+ messages in thread
From: Leonardo Bras @ 2026-09-30 17:22 UTC (permalink / raw)
To: Marc Zyngier, Oliver Upton, Fuad Tabba, Joey Gouly,
Steffen Eiden, Suzuki K Poulose, Zenghui Yu, Catalin Marinas,
Will Deacon, Mark Rutland, Leonardo Bras, Raghavendra Rao Ananta
Cc: linux-arm-kernel, kvmarm, linux-kernel
Add new walking flags KVM_PGTABLE_WALK_SKIP_LEVEL{0-3}.
When used, they make the kvm_pgtable_walk() skip the given level, as well
as levels higher than that.
Being able to skip levels substancially decreases the runtime of a
pagetable walk when the callback function does nothing on that level.
Signed-off-by: Leonardo Bras <leo.bras@arm.com>
---
arch/arm64/include/asm/kvm_pgtable.h | 14 ++++++++++++++
arch/arm64/kvm/hyp/pgtable.c | 16 +++++++++++++---
2 files changed, 27 insertions(+), 3 deletions(-)
diff --git a/arch/arm64/include/asm/kvm_pgtable.h b/arch/arm64/include/asm/kvm_pgtable.h
index 41a8687938eb..5f37ceef8d63 100644
--- a/arch/arm64/include/asm/kvm_pgtable.h
+++ b/arch/arm64/include/asm/kvm_pgtable.h
@@ -311,31 +311,45 @@ typedef bool (*kvm_pgtable_force_pte_cb_t)(u64 addr, u64 end,
* @KVM_PGTABLE_WALK_SHARED: Indicates the page-tables may be shared
* with other software walkers.
* @KVM_PGTABLE_WALK_IGNORE_EAGAIN: Don't terminate the walk early if
* the walker returns -EAGAIN.
* @KVM_PGTABLE_WALK_SKIP_BBM_TLBI: Visit and update table entries
* without Break-before-make's
* TLB invalidation.
* @KVM_PGTABLE_WALK_SKIP_CMO: Visit and update table entries
* without Cache maintenance
* operations required.
+ * @KVM_PGTABLE_WALK_SKIP_LEVEL0: Skip visiting level 0+ entries
+ * @KVM_PGTABLE_WALK_SKIP_LEVEL1: Skip visiting level 1+ entries
+ * @KVM_PGTABLE_WALK_SKIP_LEVEL2: Skip visiting level 2+ entries
+ * @KVM_PGTABLE_WALK_SKIP_LEVEL3: Skip visiting level 3 entries
*/
enum kvm_pgtable_walk_flags {
KVM_PGTABLE_WALK_LEAF = BIT(0),
KVM_PGTABLE_WALK_TABLE_PRE = BIT(1),
KVM_PGTABLE_WALK_TABLE_POST = BIT(2),
KVM_PGTABLE_WALK_SHARED = BIT(3),
KVM_PGTABLE_WALK_IGNORE_EAGAIN = BIT(4),
KVM_PGTABLE_WALK_SKIP_BBM_TLBI = BIT(5),
KVM_PGTABLE_WALK_SKIP_CMO = BIT(6),
+ /* SKIP_LEVEL* flags should not be split nor reordered */
+ KVM_PGTABLE_WALK_SKIP_LEVEL0 = BIT(7),
+ KVM_PGTABLE_WALK_SKIP_LEVEL1 = BIT(8),
+ KVM_PGTABLE_WALK_SKIP_LEVEL2 = BIT(9),
+ KVM_PGTABLE_WALK_SKIP_LEVEL3 = BIT(10),
};
+#define KVM_PGTABLE_WALK_SKIP_LEVELS (KVM_PGTABLE_WALK_SKIP_LEVEL0 | \
+ KVM_PGTABLE_WALK_SKIP_LEVEL1 | \
+ KVM_PGTABLE_WALK_SKIP_LEVEL2 | \
+ KVM_PGTABLE_WALK_SKIP_LEVEL3 )
+
struct kvm_pgtable_visit_ctx {
kvm_pte_t *ptep;
kvm_pte_t old;
void *arg;
struct kvm_pgtable_mm_ops *mm_ops;
u64 start;
u64 addr;
u64 end;
s8 level;
enum kvm_pgtable_walk_flags flags;
diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
index 7a51f78b6446..80744478c810 100644
--- a/arch/arm64/kvm/hyp/pgtable.c
+++ b/arch/arm64/kvm/hyp/pgtable.c
@@ -137,20 +137,27 @@ static bool kvm_pgtable_walk_continue(const struct kvm_pgtable_walker *walker,
* Ignore the return code altogether for walkers outside a fault handler
* (e.g. write protecting a range of memory) and chug along with the
* page table walk.
*/
if (r == -EAGAIN)
return walker->flags & KVM_PGTABLE_WALK_IGNORE_EAGAIN;
return !r;
}
+static bool kvm_pgtable_skip_level(s8 level, enum kvm_pgtable_walk_flags flags)
+{
+ u32 skip = FIELD_GET(KVM_PGTABLE_WALK_SKIP_LEVELS, flags);
+
+ return unlikely(skip && level >= ffs(skip) - 1);
+}
+
static int __kvm_pgtable_walk(struct kvm_pgtable_walk_data *data,
struct kvm_pgtable_mm_ops *mm_ops, kvm_pteref_t pgtable, s8 level);
static inline int __kvm_pgtable_visit(struct kvm_pgtable_walk_data *data,
struct kvm_pgtable_mm_ops *mm_ops,
kvm_pteref_t pteref, s8 level)
{
enum kvm_pgtable_walk_flags flags = data->walker->flags;
kvm_pte_t *ptep = kvm_dereference_pteref(data->walker, pteref);
struct kvm_pgtable_visit_ctx ctx = {
@@ -185,35 +192,35 @@ static inline int __kvm_pgtable_visit(struct kvm_pgtable_walk_data *data,
* into a newly installed or replaced table.
*/
if (reload) {
ctx.old = READ_ONCE(*ptep);
table = kvm_pte_table(ctx.old, level);
}
if (!kvm_pgtable_walk_continue(data->walker, ret))
return ret;
- if (!table) {
+ if (!table || kvm_pgtable_skip_level(level + 1, ctx.flags)) {
data->addr = ALIGN_DOWN(data->addr, kvm_granule_size(level));
data->addr += kvm_granule_size(level);
goto out;
}
childp = (kvm_pteref_t)kvm_pte_follow(ctx.old, mm_ops);
ret = __kvm_pgtable_walk(data, mm_ops, childp, level + 1);
if (!kvm_pgtable_walk_continue(data->walker, ret))
return ret;
- if (ctx.flags & KVM_PGTABLE_WALK_TABLE_POST)
+out:
+ if (table && ctx.flags & KVM_PGTABLE_WALK_TABLE_POST)
ret = kvm_pgtable_visitor_cb(data, &ctx, KVM_PGTABLE_WALK_TABLE_POST);
-out:
if (kvm_pgtable_walk_continue(data->walker, ret))
return 0;
return ret;
}
static int __kvm_pgtable_walk(struct kvm_pgtable_walk_data *data,
struct kvm_pgtable_mm_ops *mm_ops, kvm_pteref_t pgtable, s8 level)
{
u32 idx;
@@ -242,20 +249,23 @@ static int _kvm_pgtable_walk(struct kvm_pgtable *pgt, struct kvm_pgtable_walk_da
u32 idx;
int ret = 0;
u64 limit = BIT(pgt->ia_bits);
if (data->addr > limit || data->end > limit)
return -ERANGE;
if (!pgt->pgd)
return -EINVAL;
+ if (kvm_pgtable_skip_level(pgt->start_level, data->walker->flags))
+ return 0;
+
for (idx = kvm_pgd_page_idx(pgt, data->addr); data->addr < data->end; ++idx) {
kvm_pteref_t pteref = &pgt->pgd[idx * PTRS_PER_PTE];
ret = __kvm_pgtable_walk(data, pgt->mm_ops, pteref, pgt->start_level);
if (ret)
break;
}
return ret;
}
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread
* [PATCH v4 3/3] KVM: arm64: Make stage2_split_walker() skip unnecessary walks
2026-09-30 17:22 [PATCH v4 0/3] Optimize S2 hugepage splitting, introduce skip-level flags Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 1/3] KVM: arm64: Avoid re-testing walk_continue Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 2/3] KVM: arm64: Introduce KVM_PGTABLE_WALK_SKIP_LEVEL* walk flags Leonardo Bras
@ 2026-09-30 17:22 ` Leonardo Bras
2 siblings, 0 replies; 4+ messages in thread
From: Leonardo Bras @ 2026-09-30 17:22 UTC (permalink / raw)
To: Marc Zyngier, Oliver Upton, Fuad Tabba, Joey Gouly,
Steffen Eiden, Suzuki K Poulose, Zenghui Yu, Catalin Marinas,
Will Deacon, Mark Rutland, Leonardo Bras, Raghavendra Rao Ananta
Cc: linux-arm-kernel, kvmarm, linux-kernel
Currently, when splitting, all the child nodes will be walked, with the
walker just returning earlier if there is nothing to do. This means all
pagetable entries in the splitting range get a callback from the walker
function, even if it was a level-3 entry.
Optimize splitting by skipping all level-3 entries, as they are already the
smallest block size and can't be split any further.
(i.e. set flag KVM_PGTABLE_WALK_SKIP_LEVEL3)
Signed-off-by: Leonardo Bras <leo.bras@arm.com>
---
arch/arm64/kvm/hyp/pgtable.c | 3 ++-
1 file changed, 2 insertions(+), 1 deletion(-)
diff --git a/arch/arm64/kvm/hyp/pgtable.c b/arch/arm64/kvm/hyp/pgtable.c
index 80744478c810..f75cb666a103 100644
--- a/arch/arm64/kvm/hyp/pgtable.c
+++ b/arch/arm64/kvm/hyp/pgtable.c
@@ -1565,21 +1565,22 @@ static int stage2_split_walker(const struct kvm_pgtable_visit_ctx *ctx,
new = kvm_init_table_pte(childp, mm_ops);
stage2_make_pte(ctx, new);
return 0;
}
int kvm_pgtable_stage2_split(struct kvm_pgtable *pgt, u64 addr, u64 size,
struct kvm_mmu_memory_cache *mc)
{
struct kvm_pgtable_walker walker = {
.cb = stage2_split_walker,
- .flags = KVM_PGTABLE_WALK_LEAF,
+ .flags = KVM_PGTABLE_WALK_LEAF |
+ KVM_PGTABLE_WALK_SKIP_LEVEL3,
.arg = mc,
};
int ret;
ret = kvm_pgtable_walk(pgt, addr, size, &walker);
dsb(ishst);
return ret;
}
int __kvm_pgtable_stage2_init(struct kvm_pgtable *pgt, struct kvm_s2_mmu *mmu,
--
2.55.0
^ permalink raw reply [flat|nested] 4+ messages in thread
end of thread, other threads:[~2026-09-30 17:23 UTC | newest]
Thread overview: 4+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-30 17:22 [PATCH v4 0/3] Optimize S2 hugepage splitting, introduce skip-level flags Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 1/3] KVM: arm64: Avoid re-testing walk_continue Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 2/3] KVM: arm64: Introduce KVM_PGTABLE_WALK_SKIP_LEVEL* walk flags Leonardo Bras
2026-09-30 17:22 ` [PATCH v4 3/3] KVM: arm64: Make stage2_split_walker() skip unnecessary walks Leonardo Bras
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®