* [PATCH v2 1/3] mm: page_counter: add page_counter_protection struct and init API
2026-09-09 9:44 [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of struct page_counter linuszeng via B4 Relay
@ 2026-09-09 9:44 ` linuszeng via B4 Relay
2026-09-09 9:44 ` [PATCH v2 2/3] mm: page_counter: track protection state in page_counter_protection linuszeng via B4 Relay
` (2 subsequent siblings)
3 siblings, 0 replies; 6+ messages in thread
From: linuszeng via B4 Relay @ 2026-09-09 9:44 UTC (permalink / raw)
To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt,
Muchun Song, Andrew Morton, David Hildenbrand, Lorenzo Stoakes,
Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Maarten Lankhorst, Maxime Ripard,
Natalie Vock, Tejun Heo, Michal Koutný,
Oscar Salvador, Jingxiang Zeng
Cc: Michal Hocko, cgroups, linux-mm, linux-kernel, dri-devel, linuszeng
From: linuszeng <linuszeng@tencent.com>
This commit extracts the hierarchical protection state (memory.min and
memory.low) from struct page_counter into a new page_counter_protection
structure. It introduces page_counter_init_protection() to attach this
context, saving space for counters that don't support protection.
The dmem pool allocator now points its counter at the embedded
protection context, and the pool fix-up path in get_cg_pool_locked()
links the new prot->parent the same way it links cnt.parent, so pools
created bottom-up do not lose hierarchical protection.
No functional change.
---
include/linux/memcontrol.h | 7 ++++++
include/linux/page_counter.h | 59 +++++++++++++++++++++++++++++++++++++++-----
kernel/cgroup/dmem.c | 9 ++++---
mm/hugetlb_cgroup.c | 4 +--
mm/memcontrol.c | 21 ++++++++++------
mm/page_counter.c | 2 +-
6 files changed, 82 insertions(+), 20 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index 058ebd73ff16..ed863f4ed233 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -195,6 +195,13 @@ struct mem_cgroup {
/* Accounted resources */
struct page_counter memory; /* Both v1 & v2 */
+ /*
+ * Hierarchical memory.min/memory.low protection tracking for the
+ * memory page counter. swap/memsw, kmem and tcpmem counters do not
+ * support protection and have no such context.
+ */
+ struct page_counter_protection memory_prot;
+
union {
struct page_counter swap; /* v2 only */
struct page_counter memsw; /* v1 only */
diff --git a/include/linux/page_counter.h b/include/linux/page_counter.h
index 07b7cb12249c..b81f16702764 100644
--- a/include/linux/page_counter.h
+++ b/include/linux/page_counter.h
@@ -7,6 +7,32 @@
#include <linux/limits.h>
#include <asm/page.h>
+/*
+ * Hierarchical protection (memory.min / memory.low) tracking.
+ *
+ * Only the memory page counter (and dmem pools) participate in protection.
+ * swap/memsw, kmem and tcpmem page counters never do, so the protection
+ * fields are kept out of struct page_counter in this separate structure to
+ * save space in the common case. struct page_counter links to it via ->prot,
+ * which is NULL for counters without protection support.
+ */
+struct page_counter_protection {
+ struct page_counter_protection *parent;
+
+ /* effective memory.min and memory.min usage tracking */
+ unsigned long emin;
+ atomic_long_t min_usage;
+ atomic_long_t children_min_usage;
+
+ /* effective memory.low and memory.low usage tracking */
+ unsigned long elow;
+ atomic_long_t low_usage;
+ atomic_long_t children_low_usage;
+
+ unsigned long min;
+ unsigned long low;
+};
+
struct page_counter {
/*
* Make sure 'usage' does not share cacheline with any other field in
@@ -41,6 +67,12 @@ struct page_counter {
unsigned long high;
unsigned long max;
struct page_counter *parent;
+
+ /*
+ * Hierarchical protection context, NULL for counters that do not
+ * support memory.min/memory.low (swap, memsw, kmem, tcpmem, ...).
+ */
+ struct page_counter_protection *prot;
} ____cacheline_internodealigned_in_smp;
#if BITS_PER_LONG == 32
@@ -49,18 +81,33 @@ struct page_counter {
#define PAGE_COUNTER_MAX (LONG_MAX / PAGE_SIZE)
#endif
-/*
- * Protection is supported only for the first counter (with id 0).
- */
static inline void page_counter_init(struct page_counter *counter,
- struct page_counter *parent,
- bool protection_support)
+ struct page_counter *parent)
{
counter->usage = (atomic_long_t)ATOMIC_LONG_INIT(0);
counter->max = PAGE_COUNTER_MAX;
counter->parent = parent;
- counter->protection_support = protection_support;
counter->track_failcnt = false;
+ counter->prot = NULL;
+}
+
+/*
+ * Enable hierarchical protection (memory.min/memory.low) on @counter.
+ * @prot and @parent are the protection contexts of @counter and its
+ * parent page counter respectively. Only the memory page counter (and
+ * dmem pools) call this.
+ *
+ * The remaining members of @prot (emin, elow and the usage counters) are
+ * expected to be zero already, so @prot must come from zeroed memory.
+ */
+static inline void page_counter_init_protection(struct page_counter *counter,
+ struct page_counter_protection *prot,
+ struct page_counter_protection *parent)
+{
+ counter->prot = prot;
+ prot->parent = parent;
+ prot->min = 0;
+ prot->low = 0;
}
static inline unsigned long page_counter_read(struct page_counter *counter)
diff --git a/kernel/cgroup/dmem.c b/kernel/cgroup/dmem.c
index 4683f3d68022..a4bac0d5ac3b 100644
--- a/kernel/cgroup/dmem.c
+++ b/kernel/cgroup/dmem.c
@@ -88,6 +88,7 @@ struct dmem_cgroup_pool_state {
struct rcu_head rcu;
struct page_counter cnt;
+ struct page_counter_protection prot;
struct dmem_cgroup_pool_state *parent;
refcount_t ref;
@@ -426,8 +427,9 @@ alloc_pool_single(struct dmemcg_state *dmemcs, struct dmem_cgroup_region *region
if (parent)
ppool = find_cg_pool_locked(parent, region);
- page_counter_init(&pool->cnt,
- ppool ? &ppool->cnt : NULL, true);
+ page_counter_init(&pool->cnt, ppool ? &ppool->cnt : NULL);
+ page_counter_init_protection(&pool->cnt, &pool->prot,
+ ppool ? &ppool->prot : NULL);
reset_all_resource_limits(pool);
refcount_set(&pool->ref, 1);
kref_get(®ion->ref);
@@ -480,8 +482,9 @@ get_cg_pool_locked(struct dmemcg_state *dmemcs, struct dmem_cgroup_region *regio
/* ppool was created if it didn't exist by above loop. */
ppool = find_cg_pool_locked(pp, region);
- /* Fix up parent links, mark as inited. */
+ /* Fix up parent links (counter and protection), mark as inited. */
pool->cnt.parent = &ppool->cnt;
+ pool->prot.parent = &ppool->prot;
if (ppool && !pool->parent) {
pool->parent = ppool;
dmemcg_pool_get(ppool);
diff --git a/mm/hugetlb_cgroup.c b/mm/hugetlb_cgroup.c
index ecb6e0b7819a..7fdae504cfc6 100644
--- a/mm/hugetlb_cgroup.c
+++ b/mm/hugetlb_cgroup.c
@@ -108,8 +108,8 @@ static void hugetlb_cgroup_init(struct hugetlb_cgroup *h_cgroup,
fault = hugetlb_cgroup_counter_from_cgroup(h_cgroup, idx);
rsvd = hugetlb_cgroup_counter_from_cgroup_rsvd(h_cgroup, idx);
- page_counter_init(fault, fault_parent, false);
- page_counter_init(rsvd, rsvd_parent, false);
+ page_counter_init(fault, fault_parent);
+ page_counter_init(rsvd, rsvd_parent);
if (!cgroup_subsys_on_dfl(hugetlb_cgrp_subsys)) {
fault->track_failcnt = true;
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index 86ff580c7018..ffa1ced3baae 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -4267,25 +4267,30 @@ mem_cgroup_css_alloc(struct cgroup_subsys_state *parent_css)
#endif
page_counter_set_high(&memcg->swap, PAGE_COUNTER_MAX);
if (parent) {
- page_counter_init(&memcg->memory, &parent->memory, memcg_on_dfl);
- page_counter_init(&memcg->swap, &parent->swap, false);
+ page_counter_init(&memcg->memory, &parent->memory);
+ if (memcg_on_dfl)
+ page_counter_init_protection(&memcg->memory, &memcg->memory_prot,
+ &parent->memory_prot);
+ page_counter_init(&memcg->swap, &parent->swap);
#ifdef CONFIG_MEMCG_V1
WRITE_ONCE(memcg->swappiness, mem_cgroup_swappiness(parent));
memcg->memory.track_failcnt = !memcg_on_dfl;
memcg->memsw.track_failcnt = !memcg_on_dfl;
WRITE_ONCE(memcg->oom_kill_disable, READ_ONCE(parent->oom_kill_disable));
- page_counter_init(&memcg->kmem, &parent->kmem, false);
- page_counter_init(&memcg->tcpmem, &parent->tcpmem, false);
+ page_counter_init(&memcg->kmem, &parent->kmem);
+ page_counter_init(&memcg->tcpmem, &parent->tcpmem);
memcg->tcpmem.track_failcnt = !memcg_on_dfl;
#endif
} else {
init_memcg_stats();
init_memcg_events();
- page_counter_init(&memcg->memory, NULL, true);
- page_counter_init(&memcg->swap, NULL, false);
+ page_counter_init(&memcg->memory, NULL);
+ page_counter_init_protection(&memcg->memory, &memcg->memory_prot,
+ NULL);
+ page_counter_init(&memcg->swap, NULL);
#ifdef CONFIG_MEMCG_V1
- page_counter_init(&memcg->kmem, NULL, false);
- page_counter_init(&memcg->tcpmem, NULL, false);
+ page_counter_init(&memcg->kmem, NULL);
+ page_counter_init(&memcg->tcpmem, NULL);
#endif
root_mem_cgroup = memcg;
return &memcg->css;
diff --git a/mm/page_counter.c b/mm/page_counter.c
index 450543f4b318..38cb99f5f50e 100644
--- a/mm/page_counter.c
+++ b/mm/page_counter.c
@@ -15,7 +15,7 @@
static bool track_protection(struct page_counter *c)
{
- return c->protection_support;
+ return c->prot != NULL;
}
static void propagate_protected_usage(struct page_counter *c,
--
2.43.7
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH v2 2/3] mm: page_counter: track protection state in page_counter_protection
2026-09-09 9:44 [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of struct page_counter linuszeng via B4 Relay
2026-09-09 9:44 ` [PATCH v2 1/3] mm: page_counter: add page_counter_protection struct and init API linuszeng via B4 Relay
@ 2026-09-09 9:44 ` linuszeng via B4 Relay
2026-09-09 9:44 ` [PATCH v2 3/3] mm: page_counter: drop protection fields from struct page_counter linuszeng via B4 Relay
2026-09-11 15:17 ` [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of " Michal Koutný
3 siblings, 0 replies; 6+ messages in thread
From: linuszeng via B4 Relay @ 2026-09-09 9:44 UTC (permalink / raw)
To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt,
Muchun Song, Andrew Morton, David Hildenbrand, Lorenzo Stoakes,
Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Maarten Lankhorst, Maxime Ripard,
Natalie Vock, Tejun Heo, Michal Koutný,
Oscar Salvador, Jingxiang Zeng
Cc: Michal Hocko, cgroups, linux-mm, linux-kernel, dri-devel, linuszeng
From: linuszeng <linuszeng@tencent.com>
Move the read/write side of hierarchical protection from struct
page_counter to struct page_counter_protection: propagate_protected_usage()
updates the protection context of the parent, page_counter_set_min()/low()
and page_counter_calculate_protection() operate on it, and memcg and dmem
accessors (including dmem_cgroup_below_min()/below_low()) read
min/low/emin/elow and children_*_usage from it.
struct page_counter keeps its now-unused protection fields for now; they
are removed in a follow-up commit.
No functional change.
---
include/linux/memcontrol.h | 8 +++----
kernel/cgroup/dmem.c | 12 +++++-----
mm/memcontrol.c | 8 +++----
mm/page_counter.c | 59 +++++++++++++++++++++++++++++-----------------
4 files changed, 52 insertions(+), 35 deletions(-)
diff --git a/include/linux/memcontrol.h b/include/linux/memcontrol.h
index ed863f4ed233..44065001a66a 100644
--- a/include/linux/memcontrol.h
+++ b/include/linux/memcontrol.h
@@ -591,8 +591,8 @@ static inline void mem_cgroup_protection(struct mem_cgroup *root,
if (root == memcg)
return;
- *min = READ_ONCE(memcg->memory.emin);
- *low = READ_ONCE(memcg->memory.elow);
+ *min = READ_ONCE(memcg->memory_prot.emin);
+ *low = READ_ONCE(memcg->memory_prot.elow);
}
void mem_cgroup_calculate_protection(struct mem_cgroup *root,
@@ -616,7 +616,7 @@ static inline bool mem_cgroup_below_low(struct mem_cgroup *target,
if (mem_cgroup_unprotected(target, memcg))
return false;
- return READ_ONCE(memcg->memory.elow) >=
+ return READ_ONCE(memcg->memory_prot.elow) >=
page_counter_read(&memcg->memory);
}
@@ -626,7 +626,7 @@ static inline bool mem_cgroup_below_min(struct mem_cgroup *target,
if (mem_cgroup_unprotected(target, memcg))
return false;
- return READ_ONCE(memcg->memory.emin) >=
+ return READ_ONCE(memcg->memory_prot.emin) >=
page_counter_read(&memcg->memory);
}
diff --git a/kernel/cgroup/dmem.c b/kernel/cgroup/dmem.c
index a4bac0d5ac3b..4027d3d309c8 100644
--- a/kernel/cgroup/dmem.c
+++ b/kernel/cgroup/dmem.c
@@ -212,12 +212,12 @@ set_resource_max(struct dmem_cgroup_pool_state *pool, u64 val, bool nonblock)
static u64 get_resource_low(struct dmem_cgroup_pool_state *pool)
{
- return pool ? READ_ONCE(pool->cnt.low) : 0;
+ return pool ? READ_ONCE(pool->cnt.prot->low) : 0;
}
static u64 get_resource_min(struct dmem_cgroup_pool_state *pool)
{
- return pool ? READ_ONCE(pool->cnt.min) : 0;
+ return pool ? READ_ONCE(pool->cnt.prot->min) : 0;
}
static u64 get_resource_max(struct dmem_cgroup_pool_state *pool)
@@ -388,13 +388,13 @@ bool dmem_cgroup_state_evict_valuable(struct dmem_cgroup_pool_state *limit_pool,
dmem_cgroup_calculate_protection(limit_pool, test_pool);
used = page_counter_read(ctest);
- min = READ_ONCE(ctest->emin);
+ min = READ_ONCE(ctest->prot->emin);
if (used <= min)
return false;
if (!ignore_low) {
- low = READ_ONCE(ctest->elow);
+ low = READ_ONCE(ctest->prot->elow);
if (used > low)
return true;
@@ -787,7 +787,7 @@ bool dmem_cgroup_below_min(struct dmem_cgroup_pool_state *root,
* here.
*/
dmem_cgroup_calculate_protection(root, test);
- return page_counter_read(&test->cnt) <= READ_ONCE(test->cnt.emin);
+ return page_counter_read(&test->cnt) <= READ_ONCE(test->cnt.prot->emin);
}
EXPORT_SYMBOL_GPL(dmem_cgroup_below_min);
@@ -818,7 +818,7 @@ bool dmem_cgroup_below_low(struct dmem_cgroup_pool_state *root,
* here.
*/
dmem_cgroup_calculate_protection(root, test);
- return page_counter_read(&test->cnt) <= READ_ONCE(test->cnt.elow);
+ return page_counter_read(&test->cnt) <= READ_ONCE(test->cnt.prot->elow);
}
EXPORT_SYMBOL_GPL(dmem_cgroup_below_low);
diff --git a/mm/memcontrol.c b/mm/memcontrol.c
index ffa1ced3baae..b4c01a0dfd4f 100644
--- a/mm/memcontrol.c
+++ b/mm/memcontrol.c
@@ -4823,7 +4823,7 @@ static ssize_t memory_peak_write(struct kernfs_open_file *of, char *buf,
static int memory_min_show(struct seq_file *m, void *v)
{
return seq_puts_memcg_tunable(m,
- READ_ONCE(mem_cgroup_from_seq(m)->memory.min));
+ READ_ONCE(mem_cgroup_from_seq(m)->memory_prot.min));
}
static ssize_t memory_min_write(struct kernfs_open_file *of,
@@ -4846,7 +4846,7 @@ static ssize_t memory_min_write(struct kernfs_open_file *of,
static int memory_low_show(struct seq_file *m, void *v)
{
return seq_puts_memcg_tunable(m,
- READ_ONCE(mem_cgroup_from_seq(m)->memory.low));
+ READ_ONCE(mem_cgroup_from_seq(m)->memory_prot.low));
}
static ssize_t memory_low_write(struct kernfs_open_file *of,
@@ -6271,6 +6271,6 @@ void mem_cgroup_show_protected_memory(struct mem_cgroup *memcg)
memcg = root_mem_cgroup;
pr_warn("Memory cgroup min protection %lukB -- low protection %lukB",
- K(atomic_long_read(&memcg->memory.children_min_usage)),
- K(atomic_long_read(&memcg->memory.children_low_usage)));
+ K(atomic_long_read(&memcg->memory_prot.children_min_usage)),
+ K(atomic_long_read(&memcg->memory_prot.children_low_usage)));
}
diff --git a/mm/page_counter.c b/mm/page_counter.c
index 38cb99f5f50e..401201c8e390 100644
--- a/mm/page_counter.c
+++ b/mm/page_counter.c
@@ -21,28 +21,29 @@ static bool track_protection(struct page_counter *c)
static void propagate_protected_usage(struct page_counter *c,
unsigned long usage)
{
+ struct page_counter_protection *prot = c->prot;
unsigned long protected, old_protected;
long delta;
- if (!c->parent)
+ if (!prot || !prot->parent)
return;
- protected = min(usage, READ_ONCE(c->min));
- old_protected = atomic_long_read(&c->min_usage);
+ protected = min(usage, READ_ONCE(prot->min));
+ old_protected = atomic_long_read(&prot->min_usage);
if (protected != old_protected) {
- old_protected = atomic_long_xchg(&c->min_usage, protected);
+ old_protected = atomic_long_xchg(&prot->min_usage, protected);
delta = protected - old_protected;
if (delta)
- atomic_long_add(delta, &c->parent->children_min_usage);
+ atomic_long_add(delta, &prot->parent->children_min_usage);
}
- protected = min(usage, READ_ONCE(c->low));
- old_protected = atomic_long_read(&c->low_usage);
+ protected = min(usage, READ_ONCE(prot->low));
+ old_protected = atomic_long_read(&prot->low_usage);
if (protected != old_protected) {
- old_protected = atomic_long_xchg(&c->low_usage, protected);
+ old_protected = atomic_long_xchg(&prot->low_usage, protected);
delta = protected - old_protected;
if (delta)
- atomic_long_add(delta, &c->parent->children_low_usage);
+ atomic_long_add(delta, &prot->parent->children_low_usage);
}
}
@@ -257,7 +258,10 @@ void page_counter_set_min(struct page_counter *counter, unsigned long nr_pages)
{
struct page_counter *c;
- WRITE_ONCE(counter->min, nr_pages);
+ if (!counter->prot)
+ return;
+
+ WRITE_ONCE(counter->prot->min, nr_pages);
for (c = counter; c; c = c->parent)
propagate_protected_usage(c, atomic_long_read(&c->usage));
@@ -274,7 +278,10 @@ void page_counter_set_low(struct page_counter *counter, unsigned long nr_pages)
{
struct page_counter *c;
- WRITE_ONCE(counter->low, nr_pages);
+ if (!counter->prot)
+ return;
+
+ WRITE_ONCE(counter->prot->low, nr_pages);
for (c = counter; c; c = c->parent)
propagate_protected_usage(c, atomic_long_read(&c->usage));
@@ -445,9 +452,18 @@ void page_counter_calculate_protection(struct page_counter *root,
struct page_counter *counter,
bool recursive_protection)
{
+ struct page_counter_protection *prot = counter->prot;
+ struct page_counter_protection *parent_prot;
unsigned long usage, parent_usage;
struct page_counter *parent = counter->parent;
+ /*
+ * Only counters with protection support (memory, dmem pools) are
+ * ever passed here, but guard anyway.
+ */
+ if (!prot)
+ return;
+
/*
* Effective values of the reclaim targets are ignored so they
* can be stale. Have a look at mem_cgroup_protection for more
@@ -463,23 +479,24 @@ void page_counter_calculate_protection(struct page_counter *root,
return;
if (parent == root) {
- counter->emin = READ_ONCE(counter->min);
- counter->elow = READ_ONCE(counter->low);
+ prot->emin = READ_ONCE(prot->min);
+ prot->elow = READ_ONCE(prot->low);
return;
}
+ parent_prot = parent->prot;
parent_usage = page_counter_read(parent);
- WRITE_ONCE(counter->emin, effective_protection(usage, parent_usage,
- READ_ONCE(counter->min),
- READ_ONCE(parent->emin),
- atomic_long_read(&parent->children_min_usage),
+ WRITE_ONCE(prot->emin, effective_protection(usage, parent_usage,
+ READ_ONCE(prot->min),
+ READ_ONCE(parent_prot->emin),
+ atomic_long_read(&parent_prot->children_min_usage),
recursive_protection));
- WRITE_ONCE(counter->elow, effective_protection(usage, parent_usage,
- READ_ONCE(counter->low),
- READ_ONCE(parent->elow),
- atomic_long_read(&parent->children_low_usage),
+ WRITE_ONCE(prot->elow, effective_protection(usage, parent_usage,
+ READ_ONCE(prot->low),
+ READ_ONCE(parent_prot->elow),
+ atomic_long_read(&parent_prot->children_low_usage),
recursive_protection));
}
#endif /* CONFIG_MEMCG || CONFIG_CGROUP_DMEM */
--
2.43.7
^ permalink raw reply [flat|nested] 6+ messages in thread* [PATCH v2 3/3] mm: page_counter: drop protection fields from struct page_counter
2026-09-09 9:44 [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of struct page_counter linuszeng via B4 Relay
2026-09-09 9:44 ` [PATCH v2 1/3] mm: page_counter: add page_counter_protection struct and init API linuszeng via B4 Relay
2026-09-09 9:44 ` [PATCH v2 2/3] mm: page_counter: track protection state in page_counter_protection linuszeng via B4 Relay
@ 2026-09-09 9:44 ` linuszeng via B4 Relay
2026-09-11 15:17 ` [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of " Michal Koutný
3 siblings, 0 replies; 6+ messages in thread
From: linuszeng via B4 Relay @ 2026-09-09 9:44 UTC (permalink / raw)
To: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt,
Muchun Song, Andrew Morton, David Hildenbrand, Lorenzo Stoakes,
Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Maarten Lankhorst, Maxime Ripard,
Natalie Vock, Tejun Heo, Michal Koutný,
Oscar Salvador, Jingxiang Zeng
Cc: Michal Hocko, cgroups, linux-mm, linux-kernel, dri-devel, linuszeng
From: linuszeng <linuszeng@tencent.com>
The protection state now lives in struct page_counter_protection, so
remove the emin/min_usage/children_min_usage, elow/low_usage/
children_low_usage, min, low and protection_support fields from struct
page_counter.
Also drop the now-orphaned _pad2_ padding and its comment: with the
protection fields gone it no longer separates the read-mostly fields
from anything, and the structure's cacheline alignment already pads the
tail out.
swap/memsw, kmem, tcpmem and hugetlb counters no longer carry this
unused state: on 64-bit the structure shrinks from three cache lines to
two, saving one cache line.
---
include/linux/page_counter.h | 16 ----------------
1 file changed, 16 deletions(-)
diff --git a/include/linux/page_counter.h b/include/linux/page_counter.h
index b81f16702764..0007960ba5a4 100644
--- a/include/linux/page_counter.h
+++ b/include/linux/page_counter.h
@@ -43,27 +43,11 @@ struct page_counter {
CACHELINE_PADDING(_pad1_);
- /* effective memory.min and memory.min usage tracking */
- unsigned long emin;
- atomic_long_t min_usage;
- atomic_long_t children_min_usage;
-
- /* effective memory.low and memory.low usage tracking */
- unsigned long elow;
- atomic_long_t low_usage;
- atomic_long_t children_low_usage;
-
unsigned long watermark;
/* Latest cg2 reset watermark */
unsigned long local_watermark;
- /* Keep all the read most fields in a separete cacheline. */
- CACHELINE_PADDING(_pad2_);
-
- bool protection_support;
bool track_failcnt;
- unsigned long min;
- unsigned long low;
unsigned long high;
unsigned long max;
struct page_counter *parent;
--
2.43.7
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of struct page_counter
2026-09-09 9:44 [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of struct page_counter linuszeng via B4 Relay
` (2 preceding siblings ...)
2026-09-09 9:44 ` [PATCH v2 3/3] mm: page_counter: drop protection fields from struct page_counter linuszeng via B4 Relay
@ 2026-09-11 15:17 ` Michal Koutný
2026-09-17 6:29 ` jingxiang zeng
3 siblings, 1 reply; 6+ messages in thread
From: Michal Koutný @ 2026-09-11 15:17 UTC (permalink / raw)
To: linuszeng
Cc: Johannes Weiner, Michal Hocko, Roman Gushchin, Shakeel Butt,
Muchun Song, Andrew Morton, David Hildenbrand, Lorenzo Stoakes,
Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Maarten Lankhorst, Maxime Ripard,
Natalie Vock, Tejun Heo, Oscar Salvador, Jingxiang Zeng,
Michal Hocko, cgroups, linux-mm, linux-kernel, dri-devel
[-- Attachment #1: Type: text/plain, Size: 565 bytes --]
Hi.
On Wed, Sep 09, 2026 at 05:44:18PM +0800, linuszeng via B4 Relay <devnull+linuszeng.tencent.com@kernel.org> wrote:
> No functional change is intended: protection semantics and the cgroup
> v1/v2 behaviour are preserved.
It is not clear from the description what is the intention then :-)
Do you have any measurements that the reduced cache footprint
changes performance for setups without protection?
And what is the positive impact on protected scenarios where the
counters are in (possibly) different cacheline and one indirection
further?
Thanks,
Michal
[-- Attachment #2: signature.asc --]
[-- Type: application/pgp-signature, Size: 265 bytes --]
^ permalink raw reply [flat|nested] 6+ messages in thread* Re: [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of struct page_counter
2026-09-11 15:17 ` [PATCH v2 0/3] mm: page_counter: move hierarchical protection out of " Michal Koutný
@ 2026-09-17 6:29 ` jingxiang zeng
0 siblings, 0 replies; 6+ messages in thread
From: jingxiang zeng @ 2026-09-17 6:29 UTC (permalink / raw)
To: Michal Koutný
Cc: linuszeng, Johannes Weiner, Michal Hocko, Roman Gushchin,
Shakeel Butt, Muchun Song, Andrew Morton, David Hildenbrand,
Lorenzo Stoakes, Liam R. Howlett, Vlastimil Babka, Mike Rapoport,
Suren Baghdasaryan, Maarten Lankhorst, Maxime Ripard,
Natalie Vock, Tejun Heo, Oscar Salvador, Michal Hocko, cgroups,
linux-mm, linux-kernel, dri-devel
On Fri, 11 Sept 2026 at 23:17, Michal Koutný <mkoutny@suse.com> wrote:
>
> Hi.
>
> On Wed, Sep 09, 2026 at 05:44:18PM +0800, linuszeng via B4 Relay <devnull+linuszeng.tencent.com@kernel.org> wrote:
> > No functional change is intended: protection semantics and the cgroup
> > v1/v2 behaviour are preserved.
>
> It is not clear from the description what is the intention then :-)
You are right, and that is my fault for posting these three patches
without the context they came from.
They implement what Johannes asked for when I last tried to make a
combined memory+swap limit available on the default hierarchy [1]:
My suggestion is to factor out from struct page_counter all the stuff
that is not necessary for all users, and then have separate counters
for swap and memsw.
The protection stuff is long overdue for this. It makes up nearly half
of the struct's members, but is only used by the memory counter. Even
before your patches this is unnecessary bloat in the swap/memsw, kmem
and tcpmem counters.
Fix that and having separate counters is a non-issue.
and, in the same thread, about the cost of a second counter [2]:
It seems like a good opportunity to refactor struct page_counter.
So the intention is not the cache footprint on its own.
The combined memory+swap counter is the one v1 exposes as
memsw.limit_in_bytes. On the default hierarchy it does not exist:
struct mem_cgroup keeps swap and memsw in a union, because v1 only ever
charges memsw and v2 only ever charges swap, so the two never needed to
be live at the same time. Making the combined limit available on v2
means charging both on both hierarchies, which means giving them separate
page counters.
That is where struct page_counter comes in. Adding a counter costs a
cgroup one more of them, and at 192 bytes each that is 192 bytes per
cgroup for a feature most of them will not use. Trimming the counter to
128 bytes first frees 128 bytes per cgroup, which is very nearly what the
new counter then costs, so the combined limit becomes close to free in
struct mem_cgroup rather than something every cgroup pays for. This is
the "good opportunity to refactor struct page_counter" from [2], and it
is why the preparation is a prerequisite rather than a cleanup I happened
to do on the side.
The follow-up is written and tested; I should have posted it together
with these patches instead of sending the preparation on its own, and I
will do that now (details at the end).
>
> Do you have any measurements that the reduced cache footprint
> changes performance for setups without protection?
No. So far I have only used pahole to look at the cache line footprint
of struct mem_cgroup and struct page_counter (pahole, x86_64,
64-byte cache lines, CONFIG_MEMCG_V1=y)::
struct page_counter 192 -> 128 bytes (3 -> 2 cache lines)
struct page_counter_protection - -> 72 bytes
struct mem_cgroup 2176 -> 2048 bytes
The four embedded counters lose 64 bytes each and the one protection
context takes 72 back, which nets out to 128 bytes per cgroup. The
counters that never participate in protection - swap/memsw, kmem, tcpmem
and hugetlb - are also down from three cache lines to two.
The reason I need these patches is the prerequisite
above, and those 128 bytes are exactly what the follow-up needs to afford
splitting the swap and memsw counters.
> And what is the positive impact on protected scenarios where the
> counters are in (possibly) different cacheline and one indirection
> further?
There is none, and in the form I posted it was worse than before. Your
reading of the layout was correct.
propagate_protected_usage() runs once per level on every charge and
uncharge, and touches min, low, min_usage, low_usage and the parent's
children_{min,low}_usage. Counting the cache lines each level touches:
before this series page_counter 3 + parent 1 = 4, no indirection
v2 as posted page_counter 2 + prot 2 + parent prot 1 = 5
with the fix below page_counter 2 + prot 1 + parent prot 1 = 4
In v2 the new structure kept the field order of the old one, which put
min at offset 56 and low at 64. The two values the propagation path
reads together ended up on either side of a cache line boundary, while
emin and elow - which are only recomputed by
page_counter_calculate_protection() during reclaim - occupied the first
line. That is how the count got to 5.
The seven fields the charge path touches are 56 bytes and do fit in one
line, so I have reordered the structure to parent, min, low, min_usage,
low_usage, children_min_usage, children_low_usage, then emin and elow
last. Both embedders already place the context on a cache line boundary
- offset 384 in struct mem_cgroup, 192 in the dmem pool state - so no
alignment attribute is needed and the structure stays 72 bytes. That
brings the per-level line count back to what it was before the series.
The dependent load of ->prot stays; it cannot be removed while the state
lives outside the counter. The pointer shares a line with ->parent and
->local_watermark, which the same loop reads anyway, so it costs an
address dependency rather than an extra miss.
To summarise honestly: this series is size-neutral for a cgroup and,
after the reorder, layout-neutral for protected charging. It earns its
place only as groundwork for the combined limit.
So rather than reposting these three patches on their own, I am going to
send them as the first half of
[PATCH 0/6] mm/memcontrol: implement the memsw limit on cgroup v2
which is not posted yet; it will follow shortly after this reply, with
the reordered protection context folded into patch 1. The second half
builds directly on them: it splits the swap and memsw page counters out
of their union - which is what needs struct page_counter to have stopped
carrying the protection state - maintains the combined counter on both
hierarchies, and adds memory.memsw.current and memory.memsw.max to the
default hierarchy.
[1] https://lore.kernel.org/all/20250320144722.GH1876369@cmpxchg.org/
[2] https://lore.kernel.org/all/20250320142846.GG1876369@cmpxchg.org/
Thanks for looking at this.
>
> Thanks,
> Michal
^ permalink raw reply [flat|nested] 6+ messages in thread