From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pg1-f197.google.com (mail-pg1-f197.google.com [209.85.215.197]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DB1D52BD033 for ; Sat, 3 Oct 2026 00:21:50 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.215.197 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790986923; cv=none; b=lApdi80Vc8eOcHo3+of5aMKYBuy8E5U7b4mJoX5uM+RIPwJvphBQlJPspqivl8zirG83VC9mubTAhMqLqwM4CUKHiIyFJCah2vJdFAd1eIH5Jf2kR2UI5bY+NR+GNkI+me06XaC6j78P+O2fL8gHmzfgqVB9dD3Sy7AVvHwoO4Q= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790986923; c=relaxed/simple; bh=t29MAHfS9B6ywTxJUcQV+1kfVv1uqq5lmgl52OdcW9c=; h=Date:In-Reply-To:Mime-Version:References:Message-ID:Subject:From: To:Cc:Content-Type; b=sl3EBT/WIgxkx8F9wcuxwUDuMPD6KDM6AKDhp+YVyc/Gcf8jNm9Td/rUPRqN4auj0U4NhgjFZVYMoZuYnIZVt8zl9beQGw5/Ve36AQOkYYYy3zAp7zYCUpACcYdvjbHy6I94/WzrfOVZ80xNz/KYE/IlNG4vxOZP3jteCGnAkcI= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=flex--jthoughton.bounces.google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=nbjjWQ8T; arc=none smtp.client-ip=209.85.215.197 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=flex--jthoughton.bounces.google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="nbjjWQ8T" Received: by mail-pg1-f197.google.com with SMTP id 41be03b00d2f7-cc7e6720b31so73397a12.3 for ; Fri, 02 Oct 2026 17:21:50 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790986908; x=1791591708; darn=vger.kernel.org; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:from:to:cc:subject:date:message-id:reply-to :content-type; bh=el7RdsuPQmPVCa/coAHdRCW5aUIql/HQngm8YsWnGFg=; b=nbjjWQ8TXgIfRIjK+k+Ju6R1PEMT+L3gjuAWokGeD9BjeNDRwYe5rN1MWRLv10/Nb2 EObKNQlpafTyASZTTnbsCWgj5nDsFH9Ezg168Od2A/p3LWKpgIWfm/dEZPWieGWPkNN5 euLKxT/g2ry5aG72mA51S4pMQJce0M67pRcJzA9E6zJShhD7SIUeUaIPms8kdoQyXD6i vCkP4cSmMaFOzGsXJ6RVyBA0U3rBCfZUJ6aeg3ofSjSb9XkpXcq8KOTPFDKP4uTrf0n4 P90blixGWPIa3HRHX8DpaminvjAqdn8WHmH6hTraT7k6k+X0r/4WhIv4ZsLE1vmvhhWH pUGQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790986908; x=1791591708; h=content-type:cc:to:from:subject:message-id:references:mime-version :in-reply-to:date:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=el7RdsuPQmPVCa/coAHdRCW5aUIql/HQngm8YsWnGFg=; b=h7Swc4vHlsHH40KrSK7GD840EsFB92kq2KgDCprapbTJa1iWdql5iVwnPDBicpt4Nw 5ZKpIbIPsVYSjIYVp1k4+0G4RxSZGmE3nHVqFMU6Yy7B49/tXsAZOuo8u//5F6iXu3Ig 2vBx3lNhqppZV3sn0Mke/VoR+kkCDfFnuPcc4jNTqtFP2fpQYVfguPZ+zqi6GXzaCCJd FXkJXeBmr6srSPDJuye5NsmfGsyujjVm+rctchHSMTLSL0l0dB6m9buLV0KPKzd9YGpp aR4/YgIla2S0G+aGdmskusrrFBTtLIAePWdHB3GI3v7Ss9jlhIMl57rDhZZZGUqEi436 FXbQ== X-Forwarded-Encrypted: i=1; AKwUvBxWjEH0gWU26Y42vqjPDDdKBlN2qCNIahDq54OvaguZEXa/v6tzHtfuRPWRIf+uRE1WIUSo1zvj1BHCBgU=@vger.kernel.org X-Gm-Message-State: AFq9FYJ9fpnvOo2lOsxcypovo4iRQYZJihGq9SDuEMU/19KlhXiDykXW Mk26Q+0RcG0DZgFYFXzu9cLEkNMWwgH55uZaXLKaExxfY80FVFOE5PdJvZRuklfwwjhwHDyvAAq mFicCZuIxR/dzQ6e0czWl4w== X-Received: from plhb14.prod.google.com ([2002:a17:903:228e:b0:2dd:bff7:1e38]) (user=jthoughton job=prod-delivery.src-stubby-dispatcher) by 2002:a17:903:191:b0:2dd:c053:d741 with SMTP id d9443c01a7336-2e49b684cffmr41681905ad.40.1790986907988; Fri, 02 Oct 2026 17:21:47 -0700 (PDT) Date: Sat, 3 Oct 2026 00:21:18 +0000 In-Reply-To: <20261003002123.505555-1-jthoughton@google.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: Mime-Version: 1.0 References: <20261003002123.505555-1-jthoughton@google.com> X-Mailer: git-send-email 2.56.0.rc1.315.gc6ed9934b7-goog Message-ID: <20261003002123.505555-16-jthoughton@google.com> Subject: [PATCH v2 15/20] selftests/mm: Add HugeTLB vmemmap optimization stress test From: James Houghton To: Will Deacon , Catalin Marinas , Muchun Song , Oscar Salvador , Andrew Morton Cc: Nikos Nikoleris , Linu Cherian , Mark Rutland , David Hildenbrand , Ryan Roberts , Nanyong Sun , Yu Zhao , Frank van der Linden , David Rientjes , James Houghton , linux-kernel@vger.kernel.org, linux-arm-kernel@lists.infradead.org, linux-mm@kvack.org Content-Type: text/plain; charset="UTF-8" Add a script that stresses HugeTLB vmemmap optimization (HVO), in particular on architectures that update the vmemmap in place while it may be concurrently accessed: - Check that optimizing N folios frees exactly N * (vmemmap pages per folio - 1) vmemmap pages (per nr_memmap_pages and nr_memmap_boot_pages), and that restoring them gives them back. - Repeatedly optimize and restore folios by resizing the hugepage pool, while concurrently reading struct pages through /proc/kpageflags and /proc/kpagecount, compacting memory, and optionally reading page_owner and offlining/onlining memory blocks. - Do the same with fail_hugetlb_vmemmap_pte fault injection enabled, if available, to exercise the rollback and partially-optimized folio paths. After each phase, the pool must shrink back to its original size and the memmap accounting must return to its baseline. Finally, the kernel log must not contain warnings or oopses. Assisted-by: LLM Signed-off-by: James Houghton --- tools/testing/selftests/mm/Makefile | 2 + .../selftests/mm/hugetlb_vmemmap_stress.sh | 347 ++++++++++++++++++ .../selftests/mm/ksft_hugetlb_vmemmap.sh | 4 + tools/testing/selftests/mm/run_vmtests.sh | 4 + 4 files changed, 357 insertions(+) create mode 100755 tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh create mode 100755 tools/testing/selftests/mm/ksft_hugetlb_vmemmap.sh diff --git a/tools/testing/selftests/mm/Makefile b/tools/testing/selftests/mm/Makefile index beacc0f87304..51d8fba80c03 100644 --- a/tools/testing/selftests/mm/Makefile +++ b/tools/testing/selftests/mm/Makefile @@ -149,6 +149,7 @@ TEST_PROGS += ksft_cow.sh TEST_PROGS += ksft_gup_test.sh TEST_PROGS += ksft_hmm.sh TEST_PROGS += ksft_hugetlb.sh +TEST_PROGS += ksft_hugetlb_vmemmap.sh TEST_PROGS += ksft_hugevm.sh TEST_PROGS += ksft_kmemleak_confirm.sh TEST_PROGS += ksft_kmemleak_dedup.sh @@ -180,6 +181,7 @@ TEST_FILES += test_hmm.sh TEST_FILES += va_high_addr_switch.sh TEST_FILES += charge_reserved_hugetlb.sh TEST_FILES += hugetlb_reparenting_test.sh +TEST_FILES += hugetlb_vmemmap_stress.sh TEST_FILES += test_page_frag.sh TEST_FILES += run_vmtests.sh diff --git a/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh b/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh new file mode 100755 index 000000000000..94357a4a45c2 --- /dev/null +++ b/tools/testing/selftests/mm/hugetlb_vmemmap_stress.sh @@ -0,0 +1,347 @@ +#!/bin/bash +# SPDX-License-Identifier: GPL-2.0 +# +# Stress test for HugeTLB vmemmap optimization (HVO). +# +# Phases: +# 1. accounting: allocate N hugepages, check that nr_memmap_pages + +# nr_memmap_boot_pages drops by exactly N * (vmemmap pages per folio - 1), +# then free them and check it returns to the baseline. +# 2. stress: churn the pool (optimize/restore) while concurrently reading +# struct pages (/proc/kpageflags, /proc/kpagecount), compacting memory, +# and optionally reading page_owner and offlining/onlining memory. +# 3. pte-inject: like 2, with fail_hugetlb_vmemmap_pte enabled; this hits +# both the optimize rollback and the restore (partial HVO) paths. +# +# After each phase the pool is drained back to its original size, memmap +# accounting must be back at the phase's baseline, and the kernel log must +# not contain warnings/oopses. +# +# The fault-injection phases require CONFIG_FAIL_HUGETLB_VMEMMAP and +# CONFIG_FAULT_INJECTION_DEBUG_FS, and are skipped otherwise. + +set -u + +KSFT_PASS=0 +KSFT_FAIL=1 +KSFT_SKIP=4 + +size_kb= +nr=16 +duration=60 +prob=20 +readers=4 +struct_page_size=64 +do_page_owner=0 +do_hotplug=0 + +usage() { + cat </dev/null + +orig_nr=$(cat "$hp_dir/nr_hugepages") +target_nr=$(( orig_nr + nr )) +marker="hvo-stress-$$-$(date +%s)" +tmpdir=$(mktemp -d) +pids=() + +log "hugepage size ${size_kb}kB, page size $page_size" +log "struct page size $struct_page_size" +log "vmemmap pages/folio: $vmemmap_pages ($freed_per_folio freed by HVO)" +log "pool: $orig_nr initially, churning between $orig_nr and $target_nr" +[ "$orig_nr" -eq 0 ] || + log "WARNING: pool not initially empty, accounting may be inexact" + +memmap_total() { + awk '/^nr_memmap_(boot_)?pages / {s += $2} END {print s + 0}' \ + /proc/vmstat +} + +set_nr() { + echo "$1" > "$hp_dir/nr_hugepages" 2>/dev/null + cat "$hp_dir/nr_hugepages" +} + +# Shrink the pool back to orig_nr. Restore may transiently fail (e.g. with +# fault injection active), so retry for a while. +drain() { + local i cur surplus + + for i in $(seq 20); do + # Pages whose vmemmap could not be restored are kept as free + # surplus pages, which shrinking nr_hugepages does not free. + # Writing the current size converts them back to persistent + # pages first. + set_nr "$(cat "$hp_dir/nr_hugepages")" > /dev/null + cur=$(set_nr "$orig_nr") + [ "$cur" -eq "$orig_nr" ] && return 0 + sleep 1 + done + surplus=$(cat "$hp_dir/surplus_hugepages") + fail "pool stuck at $cur hugepages ($surplus surplus)," \ + "expected $orig_nr" + return 1 +} + +fa_dir() { echo "$dbgfs/$1"; } + +fa_enable() { + local d + d=$(fa_dir "$1") + echo 0 > "$d/verbose" + echo N > "$d/task-filter" + echo 1 > "$d/interval" + echo 1000000 > "$d/times" + echo "$prob" > "$d/probability" +} + +# Disable and print the number of injected failures. +fa_disable() { + local d left + d=$(fa_dir "$1") + echo 0 > "$d/probability" + left=$(cat "$d/times") + echo 0 > "$d/times" + echo $(( 1000000 - left )) +} + +check_dmesg() { + local bad pat + + pat='WARNING:|BUG[: ]|Oops|Unable to handle kernel' + pat+='|Internal error|KASAN:|UBSAN:' + pat+='|list_(add|del) corruption|page dumped because' + bad=$(dmesg | sed -n "/$marker/,\$p" | grep -E "$pat") + if [ -n "$bad" ]; then + fail "kernel log reports problems:" + echo "$bad" | head -20 | sed 's/^/# /' + fi +} + +## Workers + +churn() { + while :; do + set_nr "$target_nr" > /dev/null + set_nr "$orig_nr" > /dev/null + done +} + +kpage_reader() { + while :; do + dd if=/proc/kpageflags of=/dev/null bs=4M status=none + dd if=/proc/kpagecount of=/dev/null bs=4M status=none + done +} + +compactor() { + while :; do + echo 1 > /proc/sys/vm/compact_memory + sleep 1 + done +} + +page_owner_reader() { + while :; do + cat $dbgfs/page_owner > /dev/null + done +} + +hotplugger() { + local blk state removable + + while :; do + for blk in /sys/devices/system/memory/memory*; do + state=$(cat "$blk/state" 2>/dev/null) + removable=$(cat "$blk/removable" 2>/dev/null || echo 1) + [ "$state" = online ] || continue + [ "$removable" = 1 ] || continue + echo "$blk" >> "$tmpdir/hotplug" + # A signal aborts a pending offline_pages(). + timeout 10 sh -c "echo offline > $blk/state" 2>/dev/null + echo online > "$blk/state" 2>/dev/null + sleep 1 + done + done +} + +start_workers() { + local i + + churn & pids+=($!) + for i in $(seq "$readers"); do + kpage_reader & pids+=($!) + done + compactor & pids+=($!) + if [ "$do_page_owner" -eq 1 ]; then + if [ -r $dbgfs/page_owner ]; then + page_owner_reader & pids+=($!) + else + log "page_owner not available, not reading it" + fi + fi + [ "$do_hotplug" -eq 1 ] && { hotplugger & pids+=($!); } +} + +stop_workers() { + [ "${#pids[@]}" -gt 0 ] || return 0 + kill "${pids[@]}" 2>/dev/null + wait "${pids[@]}" 2>/dev/null + pids=() +} + +cleanup() { + local blk t + + stop_workers + for t in fail_hugetlb_vmemmap_pte; do + [ -d "$(fa_dir $t)" ] && fa_disable $t > /dev/null + done + if [ -f "$tmpdir/hotplug" ]; then + sort -u "$tmpdir/hotplug" | while read -r blk; do + [ "$(cat "$blk/state")" = online ] || + echo online > "$blk/state" 2>/dev/null + done + fi + set_nr "$orig_nr" > /dev/null + rm -rf "$tmpdir" +} +trap cleanup EXIT +trap 'exit $KSFT_FAIL' INT TERM + +# Run a stress phase. $1: name, $2: fault attr to enable ("" for none). +stress_phase() { + local name=$1 fa=$2 base after injected + + if [ -n "$fa" ] && [ ! -d "$(fa_dir "$fa")" ]; then + echo "SKIP: $name (no $(fa_dir "$fa"))" + return + fi + + log "phase $name: ${duration}s" + base=$(memmap_total) + [ -n "$fa" ] && fa_enable "$fa" + start_workers + sleep "$duration" + stop_workers + if [ -n "$fa" ]; then + injected=$(fa_disable "$fa") + log "$name: injected $injected failures" + [ "$injected" -gt 0 ] || + log "WARNING: $name: no failures injected" + fi + + drain || return + after=$(memmap_total) + if [ "$after" -ne "$base" ]; then + fail "$name: memmap pages $after after drain, expected $base" + else + pass "$name" + fi +} + +accounting_phase() { + local base got added after expect + + log "phase accounting" + base=$(memmap_total) + got=$(set_nr "$target_nr") + added=$(( got - orig_nr )) + if [ "$added" -le 0 ]; then + fail "accounting: could not allocate any ${size_kb}kB hugepages" + return + fi + [ "$added" -eq "$nr" ] || log "accounting: only allocated $added of $nr" + + after=$(memmap_total) + expect=$(( base - added * freed_per_folio )) + if [ "$after" -ne "$expect" ]; then + fail "accounting: memmap pages $after after optimizing" \ + "$added folios, expected $expect (baseline $base)" + else + pass "accounting: optimize freed $(( base - after ))" \ + "vmemmap pages" + fi + + drain || return + after=$(memmap_total) + if [ "$after" -ne "$base" ]; then + fail "accounting: memmap pages $after after restore," \ + "expected $base" + else + pass "accounting: restore" + fi +} + +echo "$marker" > /dev/kmsg + +accounting_phase +stress_phase stress "" +stress_phase pte-inject fail_hugetlb_vmemmap_pte + +check_dmesg + +if [ "$failures" -ne 0 ]; then + echo "FAILED: $failures check(s)" + exit $KSFT_FAIL +fi +echo "OK" +exit $KSFT_PASS diff --git a/tools/testing/selftests/mm/ksft_hugetlb_vmemmap.sh b/tools/testing/selftests/mm/ksft_hugetlb_vmemmap.sh new file mode 100755 index 000000000000..905b75b2cfb4 --- /dev/null +++ b/tools/testing/selftests/mm/ksft_hugetlb_vmemmap.sh @@ -0,0 +1,4 @@ +#!/bin/sh -e +# SPDX-License-Identifier: GPL-2.0 + +./run_vmtests.sh -t hugetlb_vmemmap diff --git a/tools/testing/selftests/mm/run_vmtests.sh b/tools/testing/selftests/mm/run_vmtests.sh index a1b45a3dedae..8dcdee7be501 100755 --- a/tools/testing/selftests/mm/run_vmtests.sh +++ b/tools/testing/selftests/mm/run_vmtests.sh @@ -77,6 +77,8 @@ separated by spaces: test transparent huge pages - hugetlb test hugetlbfs huge pages +- hugetlb_vmemmap + test the hugetlb vmemmap optimization - migration invoke move_pages(2) to exercise the migration entry code paths in the kernel @@ -312,6 +314,8 @@ echo "$enable_soft_offline" > /proc/sys/vm/enable_soft_offline CATEGORY="hugetlb" run_test ./hugetlb-read-hwpoison fi +CATEGORY="hugetlb_vmemmap" run_test ./hugetlb_vmemmap_stress.sh + if [ $VADDR64 -ne 0 ]; then # va high address boundary switch test CATEGORY="hugevm" run_test bash ./va_high_addr_switch.sh -- 2.56.0.rc1.315.gc6ed9934b7-goog