* [RFC PATCH 0/6] f2fs: retain clean node blocks in a compressed cache
@ 2026-09-29 7:29 Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 1/6] f2fs: generalize metadata cache shrinking to explicit lists Wenjie Qi
` (5 more replies)
0 siblings, 6 replies; 7+ messages in thread
From: Wenjie Qi @ 2026-09-29 7:29 UTC (permalink / raw)
To: Jaegeuk Kim, Chao Yu
Cc: Barry Song, linux-f2fs-devel, linux-kernel, Wenjie Qi
Hi,
This RFC adds a reclaimable compressed representation for clean node-cache
entries on top of the F2FS metadata cache.
The node cache avoids repeated NAT and node-block reads, but retaining a raw
node still consumes one filesystem block. A large raw cache therefore has a
substantial memory footprint, while reclaiming entries early causes later
accesses to issue node reads again.
This series provides a middle ground:
- active or recently accessed nodes remain in the raw representation;
- a background worker compresses colder clean nodes into small private-slab
objects;
- a compressed entry keeps its NID identity;
- access restores the raw block before normal node validation;
- compressed data remains a discardable optimization, and restore or
validation failure falls back to the ordinary disk-read path; and
- both raw and compressed entries remain reclaimable by the F2FS shrinker.
Compressed representation
=========================
When the feature is enabled, NODE_CACHE uses an extended entry to retain the
compressed payload length, private-slab allocation size, CRC of the original
node block, and compression-context ownership.
Compressed payloads use per-superblock private slabs with three object sizes:
256 bytes
512 bytes
1024 bytes
A result that exceeds either the configured threshold or the 1024-byte limit
is not retained. Raw entries stay on the existing node LRU. Compressed
entries are placed on separate 256-, 512-, and 1024-byte queues according to
their allocation size.
Ownership and restore
=====================
The conversion path may briefly hold both the raw block and a newly allocated
compressed object. After publication, however, the entry owns only one
representation. Compressed bytes are never submitted directly to BIO.
Restore runs while the entry lock and lookup reference still exclude normal
concurrent byte users. It validates:
allocation size
compressed payload length
LZ4 output length
raw-block CRC
node footer
inode checksum
requested node type
A raw-buffer allocation failure leaves the compressed representation intact
so that a caller may retry. A malformed payload or later semantic validation
failure discards the optimization and uses the normal disk-read path. The
compressed cache therefore does not replace the on-disk source of truth.
Background worker
=================
The series reuses the existing per-superblock cache thread instead of adding
a new kernel thread. Writeback and compression keep separate due times, but
execute serially in that shared thread.
Each worker pass is bounded by:
maximum raw scan: 1024 entries
maximum candidates: 256 entries
reschedule batch: 32 entries
The series adds two sysfs controls:
node_compress_threshold
node_compress_interval
The threshold is the maximum retained compressed payload as a percentage of
the original block; zero disables compression. The interval range is
100--30000 ms, with a default of 1000 ms.
Reclaim policy
==============
The series lets the node shrinker reclaim raw and compressed queues
independently, so every compressed entry created by the worker remains
reclaimable.
The shrinker uses fixed reclaim weights. A raw entry has a larger weight
because it occupies a full block. Compressed entries are retained
preferentially, but remain reclaimable under sustained pressure.
The same weights are used for:
- the effective shrinker count; and
- per-queue scan quotas.
Quota rounding credit is carried across shrinker invocations. If a queue
reaches its population cap, unused quota is redistributed to the other
queues.
REFERENCED second chance
========================
The worker and shrinker use the same REFERENCED rule:
test REFERENCED
clear REFERENCED
move the entry to the tail of its current queue
skip it in the current pass
The worker remembers the original raw-list tail so entries moved during a
pass are not revisited in that pass.
If an entry is accessed after compression completes but before publication,
the temporary compressed object is discarded and the entry receives the same
second chance.
NVMe-backed QEMU test
====================
The test environment was:
x86_64 KVM
4 vCPUs
3 GiB guest memory
2 GiB raw F2FS image
host backing filesystem: /dev/nvme0n1p2
guest block device: /dev/nvme0n1
QEMU device: nvme
cache=none
aio=native
Each test suite created its own empty, formatted 2 GiB seed image. For every
guest boot, the host copied that seed into an independent sparse image, and
the guest rebuilt the same file set:
/mnt/bench/
|-- files/
| |-- d-0/f-0 ... f-63
| |-- d-1/f-0 ... f-63
| `-- d-255/f-0 ... f-63
|-- shaped/
| |-- b-1-0 ... b-1-3
| |-- b-32-0 ... b-32-3
| |-- b-80-0 ... b-80-3
| |-- b-128-0 ... b-128-3
| |-- b-192-0 ... b-192-3
| |-- b-256-0 ... b-256-3
| |-- b-384-0 ... b-384-3
| `-- b-512-0 ... b-512-3
`-- sentinel
The files/ tree contains 256 directories with 64 files each, for 16,384
replay files. These files build a fixed inode, dentry, and node-page
population and are accessed later in a fixed manifest order.
The shaped/ tree contains 32 zero-filled files ranging from 1 to 512 blocks.
Their different extent and block-address densities diversify node-page
contents and compression outcomes. They contribute to the initial cache
population and subsequent compression and reclaim, but are excluded from the
timed replay, which remains limited to the 16,384 uniform small files. The
sentinel is used only for clean unmount, remount, and content verification
after the measured workload.
The workload compares memory use and replay latency at equal node counts, and
retained nodes, node reads, and replay latency at equal node-cache memory.
For the results below, compression-off means compression admission is disabled
at runtime, while compression-on means background conversion is enabled.
Attributed node-cache memory is the sum of cache-entry bytes, raw block
buffers or compressed private-slab slots, and known fixed context bytes. It
is an accounting metric, not measured physical RSS, and excludes allocator
metadata.
At equal node counts, both modes retained 16,676 nodes:
compression-off attributed memory: 69,792,344 bytes
compression-on attributed memory: 6,008,152 bytes
memory reduction: 91.391%
The paired median foreground replay regression was 0.845%.
At equal node counts, the compressed representation substantially reduced
attributed memory while typical replay latency remained close.
At an approximately 3 MiB attributed node-cache budget:
compression-off: 3,199,800 bytes, 760 nodes
compression-on: 3,207,992 bytes, 9,208 nodes
retained-node ratio: 12.116x
Node reads were:
compression-off: 15,668 reads / 64,176,128 bytes
compression-on: 7,209 reads / 29,528,064 bytes
reduction: 53.989%
The compression-on configuration also completed 9,175 compressed restores.
For the one-pass foreground replay, those restores plus 7,209 node reads equal
the 16,384 per-file node retrievals. These counters cover F2FS node
retrievals, not all guest filesystem I/O.
The paired median foreground replay delta was -24.984%.
At the same node-cache memory budget, the compressed representation retained
more nodes, reduced node reads, and improved replay latency in this workload.
Android phone test
==================
The mechanism was also exercised on a real Android phone with application,
memory-pressure, and cache-lifecycle workloads.
Each successful trial performed:
reboot
launch 18 applications for a seed pass
pre-pressure normalization
allocate and hold 1 GiB of memory pressure for 30 seconds
release pressure
launch 18 applications for measured pass 1
launch the same 18 applications for measured pass 2
restore the original policy
clean up and run an offline audit
The seed pass was excluded from the reported measured launch-time results. It
populated the post-reboot F2FS node cache with application-related node
entries so that background compression, pressure reclaim, and later restore
operated on a non-empty cache.
The application set included messaging, short-video, shopping, social-media,
browser, and video applications. Each application record retained the
package, Activity, PID, process start time, launch timing, and F2FS node-cache
counters.
In the snapshot statistics below, attached node entries are the current raw
and compressed cache entries. Attached raw bytes count 4 KiB raw buffers,
while compressed slot bytes count allocated private-slab bucket sizes.
Node-cache logical bytes combine entry, raw-buffer, and slot accounting; they
are not measured physical RSS.
After restoring the original sysfs policy, the final snapshots showed these
trial-level means. Restoring the policy did not flush existing cache entries:
compression-off:
attached node entries: 14,509
compressed entries: 0
attached raw bytes: 59,430,229
node-cache logical bytes: 60,823,125
compression-on:
attached node entries: 25,331
compressed entries: 16,402
attached raw bytes: 36,573,184
compressed slot bytes: 8,769,877
node-cache logical bytes: 47,774,805
At the same final snapshots, the mean compressed-entry distribution per trial,
rounded to the nearest entry, was:
256-byte: 373 entries ( 2.27%)
512-byte: 15,116 entries (92.16%)
1024-byte: 913 entries ( 5.57%)
Under the F2FS node-cache accounting model, compression-on changed the above
values by:
retained node entries: +74.59%
raw bytes: -38.46%
logical cache bytes: -21.45%
The node-read count is the filesystem-wide F2FS FS_NODE_READ_IO count for the
/data mount. Each event represents a submitted 4 KiB node-block read. It
includes activity from the measured applications and concurrent Android
services, and is neither a per-application attribution nor a count of all
storage I/O. The counts were:
compression-off mean: 309,360
compression-on mean: 287,417
difference: -7.09%
The compression-on trials also exercised raw-to-compressed conversion,
compressed restore, compressed-queue reclaim, and REFERENCED clear-and-move.
Each compression-on trial completed about 70,441--76,786 conversions and
retained a mean of 16,402 compressed entries in the final snapshots after
policy restoration.
In this Android workload, the compressed representation reduced node-cache
accounting memory, retained more node entries, and showed a lower filesystem-
wide node-read counter. Raw and compressed entries continued to be reclaimed
under pressure, followed by restore and recompression during application
replay. The observed node-read reduction describes the direction in this
workload rather than a strict per-application causal measurement.
Patch layout
============
Patch 1 generalizes metadata-cache shrinking to accept explicit lists while
preserving the behavior of existing caches.
Patch 2 adds the extended node entry, private slabs, compressed queues,
ownership, and memory accounting.
Patch 3 connects queue-aware reclaim, fixed reclaim weights, and carried
rounding credit.
Patch 4 adds restore, length and CRC validation, complete node validation,
and disk fallback.
Patch 5 adds the bounded background worker to the cache thread and exposes
the threshold and interval controls.
Patch 6 gives the worker and shrinker the same REFERENCED clear-and-move
behavior.
This series is based on F2FS dev-test commit:
1b629035be4a ("f2fs: rename nr_pages_to_skip with nr_caches_to_skip")
I would especially appreciate feedback on:
1. whether per-superblock 256-, 512-, and 1024-byte private slabs are an
appropriate allocator boundary;
2. whether fixed reclaim weights for raw blocks and compressed slots are a
reasonable initial reclaim policy; and
3. whether the threshold and interval provide a sufficient initial control
interface.
Thanks,
Wenjie
Wenjie Qi (6):
f2fs: generalize metadata cache shrinking to explicit lists
f2fs: add compressed clean node cache representation
f2fs: bias node cache reclaim toward raw entries
f2fs: restore compressed node cache entries before access
f2fs: compress clean node cache entries in background
f2fs: give referenced node cache entries a shared second chance
Documentation/ABI/testing/sysfs-fs-f2fs | 22 +
fs/f2fs/Kconfig | 12 +
fs/f2fs/Makefile | 2 +
fs/f2fs/cache.c | 145 +++-
fs/f2fs/cache.h | 52 +-
fs/f2fs/data.c | 5 +-
fs/f2fs/debug.c | 47 +-
fs/f2fs/f2fs.h | 17 +
fs/f2fs/inode.c | 20 +-
fs/f2fs/node.c | 55 +-
fs/f2fs/node_cache_compress.c | 915 ++++++++++++++++++++++++
fs/f2fs/node_cache_compress.h | 144 ++++
fs/f2fs/node_cache_policy.c | 211 ++++++
fs/f2fs/node_cache_policy.h | 54 ++
fs/f2fs/shrinker.c | 3 +-
fs/f2fs/super.c | 4 +
fs/f2fs/sysfs.c | 59 ++
17 files changed, 1713 insertions(+), 54 deletions(-)
create mode 100644 fs/f2fs/node_cache_compress.c
create mode 100644 fs/f2fs/node_cache_compress.h
create mode 100644 fs/f2fs/node_cache_policy.c
create mode 100644 fs/f2fs/node_cache_policy.h
base-commit: 1b629035be4a085f59d47edc62f4e695c763a94d
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH 1/6] f2fs: generalize metadata cache shrinking to explicit lists
2026-09-29 7:29 [RFC PATCH 0/6] f2fs: retain clean node blocks in a compressed cache Wenjie Qi
@ 2026-09-29 7:29 ` Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 2/6] f2fs: add compressed clean node cache representation Wenjie Qi
` (4 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Wenjie Qi @ 2026-09-29 7:29 UTC (permalink / raw)
To: Jaegeuk Kim, Chao Yu
Cc: Barry Song, linux-f2fs-devel, linux-kernel, Wenjie Qi
The metadata-cache shrinker currently operates on cache->lru_list.
Node-cache compression needs separate raw and compressed queues, while
keeping the existing isolation, truncation and refcount rules.
Move the common shrink loop to f2fs_shrink_cache_list(), and pass it an
explicit list, population and scan budget. Existing caches continue to
use their current LRU lists.
Also return the number of examined entries for queue-specific reclaim
accounting.
Signed-off-by: Wenjie Qi <qiwenjie@xiaomi.com>
---
fs/f2fs/cache.c | 30 +++++++++++++++++++++---------
fs/f2fs/cache.h | 6 +++++-
2 files changed, 26 insertions(+), 10 deletions(-)
diff --git a/fs/f2fs/cache.c b/fs/f2fs/cache.c
index 38fc5eb17f92..8f81aed132ed 100644
--- a/fs/f2fs/cache.c
+++ b/fs/f2fs/cache.c
@@ -547,8 +547,10 @@ void f2fs_destroy_cache(struct f2fs_cached_block_list *cache)
goto next;
}
-static unsigned long f2fs_do_shrink_cache(struct f2fs_cached_block_list *cache,
- unsigned long nr_to_scan)
+unsigned long
+f2fs_shrink_cache_list(struct f2fs_cached_block_list *cache,
+ struct list_head *head, unsigned long nr_entries,
+ unsigned long nr_to_scan, unsigned long *nr_scanned)
{
struct f2fs_cached_block *entry, *next;
LIST_HEAD(dispose_list);
@@ -558,15 +560,15 @@ static unsigned long f2fs_do_shrink_cache(struct f2fs_cached_block_list *cache,
/* Phase 1: Isolate candidate entries from LRU list into dispose_list */
spin_lock(&cache->list_lock);
- list_for_each_entry_safe(entry, next, &cache->lru_list, list) {
- if (scanned >= cache->num_entries)
- break;
- if (scanned++ >= nr_to_scan)
+ nr_entries = min(nr_entries, cache->num_entries);
+ list_for_each_entry_safe(entry, next, head, list) {
+ if (scanned >= nr_entries || scanned >= nr_to_scan)
break;
+ scanned++;
/* If accessed, give it a second chance to rotate to tail */
if (f2fs_cache_test_and_clear_referenced(entry)) {
- list_move_tail(&entry->list, &cache->lru_list);
+ list_move_tail(&entry->list, head);
continue;
}
@@ -620,16 +622,26 @@ static unsigned long f2fs_do_shrink_cache(struct f2fs_cached_block_list *cache,
freed++;
}
- /* Phase 3: Splice un-reclaimed entries back onto cache->lru_list */
+ /* Phase 3: Splice un-reclaimed entries back onto the scanned list */
if (!list_empty(&keep_list)) {
spin_lock(&cache->list_lock);
- list_splice_tail(&keep_list, &cache->lru_list);
+ list_splice_tail(&keep_list, head);
spin_unlock(&cache->list_lock);
}
+ if (nr_scanned)
+ *nr_scanned = scanned;
return freed;
}
+static unsigned long
+f2fs_do_shrink_cache(struct f2fs_cached_block_list *cache,
+ unsigned long nr_to_scan)
+{
+ return f2fs_shrink_cache_list(cache, &cache->lru_list, ULONG_MAX,
+ nr_to_scan, NULL);
+}
+
unsigned long f2fs_shrink_cache(struct f2fs_sb_info *sbi,
unsigned long nr_to_scan)
{
diff --git a/fs/f2fs/cache.h b/fs/f2fs/cache.h
index 6c4db910d767..c4c3d09a0008 100644
--- a/fs/f2fs/cache.h
+++ b/fs/f2fs/cache.h
@@ -222,7 +222,11 @@ void f2fs_drop_cache_range(struct f2fs_cached_block_list *cache,
f2fs_drop_cache_range(NODE_CACHE(sbi), start, len, true)
unsigned long f2fs_shrink_cache(struct f2fs_sb_info *sbi,
- unsigned long nr_to_scan);
+ unsigned long nr_to_scan);
+unsigned long
+f2fs_shrink_cache_list(struct f2fs_cached_block_list *cache,
+ struct list_head *head, unsigned long nr_entries,
+ unsigned long nr_to_scan, unsigned long *nr_scanned);
#define DEF_DIRTY_CACHE_TIMEOUT 5000
#define MIN_DIRTY_CACHE_TIMEOUT 100
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH 2/6] f2fs: add compressed clean node cache representation
2026-09-29 7:29 [RFC PATCH 0/6] f2fs: retain clean node blocks in a compressed cache Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 1/6] f2fs: generalize metadata cache shrinking to explicit lists Wenjie Qi
@ 2026-09-29 7:29 ` Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 3/6] f2fs: bias node cache reclaim toward raw entries Wenjie Qi
` (3 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Wenjie Qi @ 2026-09-29 7:29 UTC (permalink / raw)
To: Jaegeuk Kim, Chao Yu
Cc: Barry Song, linux-f2fs-devel, linux-kernel, Wenjie Qi
A clean node-cache entry currently retains a full block until it is
reclaimed. Add an optional compressed representation that keeps the entry
and its NID in memory with a smaller payload.
Use an extended NODE_CACHE entry and per-mount 256, 512 and 1024-byte
private slabs. Raw entries remain on the existing LRU, while compressed
entries use one queue per allocation size.
Keep raw and compressed ownership exclusive, and make detach and teardown
aware of the compressed queues. If the compression resources cannot be
initialized, leave node compression disabled instead of failing the mount.
The generic entry layout remains unchanged when the feature is disabled.
Signed-off-by: Wenjie Qi <qiwenjie@xiaomi.com>
---
fs/f2fs/Kconfig | 12 ++
fs/f2fs/Makefile | 1 +
fs/f2fs/cache.c | 27 ++-
fs/f2fs/cache.h | 8 +
fs/f2fs/debug.c | 10 +-
fs/f2fs/f2fs.h | 7 +
fs/f2fs/node_cache_compress.c | 303 ++++++++++++++++++++++++++++++++++
fs/f2fs/node_cache_compress.h | 78 +++++++++
fs/f2fs/node_cache_policy.h | 23 +++
fs/f2fs/super.c | 4 +
10 files changed, 462 insertions(+), 11 deletions(-)
create mode 100644 fs/f2fs/node_cache_compress.c
create mode 100644 fs/f2fs/node_cache_compress.h
create mode 100644 fs/f2fs/node_cache_policy.h
diff --git a/fs/f2fs/Kconfig b/fs/f2fs/Kconfig
index 5916a02fb46d..5aa54916e594 100644
--- a/fs/f2fs/Kconfig
+++ b/fs/f2fs/Kconfig
@@ -99,6 +99,18 @@ config F2FS_FS_COMPRESSION
Enable filesystem-level compression on f2fs regular files,
multiple back-end compression algorithms are supported.
+config F2FS_FS_NODE_CACHE_COMPRESSION
+ bool "F2FS clean node cache compression"
+ depends on F2FS_FS
+ select LZ4_COMPRESS
+ select LZ4_DECOMPRESS
+ default n
+ help
+ Retain clean node-cache blocks in a private compressed store.
+ This is independent of file-data compression and COMPRESS_CACHE.
+ The store uses 256, 512 and 1024 byte slab buckets.
+ Compression is disabled at runtime until a nonzero threshold is set.
+
config F2FS_FS_LZO
bool "LZO compression support"
depends on F2FS_FS_COMPRESSION
diff --git a/fs/f2fs/Makefile b/fs/f2fs/Makefile
index fbf49c30b066..e1e4b98707e9 100644
--- a/fs/f2fs/Makefile
+++ b/fs/f2fs/Makefile
@@ -9,4 +9,5 @@ f2fs-$(CONFIG_F2FS_FS_XATTR) += xattr.o
f2fs-$(CONFIG_F2FS_FS_POSIX_ACL) += acl.o
f2fs-$(CONFIG_FS_VERITY) += verity.o
f2fs-$(CONFIG_F2FS_FS_COMPRESSION) += compress.o
+f2fs-$(CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION) += node_cache_compress.o
f2fs-$(CONFIG_F2FS_IOSTAT) += iostat.o
diff --git a/fs/f2fs/cache.c b/fs/f2fs/cache.c
index 8f81aed132ed..3d1d530cdd82 100644
--- a/fs/f2fs/cache.c
+++ b/fs/f2fs/cache.c
@@ -17,6 +17,7 @@
#include "node.h"
#include <trace/events/f2fs.h>
#include "segment.h"
+#include "node_cache_compress.h"
void f2fs_cache_wait_writeback_cond(struct f2fs_cached_block *entry,
enum page_type type)
@@ -127,7 +128,7 @@ static int f2fs_cache_refcount(struct f2fs_cached_block *entry)
static void f2fs_do_free_cache(struct f2fs_cached_block *entry)
{
- kfree(entry->data);
+ f2fs_nc_free_data(entry);
kfree(entry);
}
@@ -159,6 +160,7 @@ static struct f2fs_cached_block *f2fs_create_cache(
{
struct f2fs_sb_info *sbi = cache->sbi;
struct f2fs_cached_block *entry;
+ size_t entry_size = f2fs_nc_entry_alloc_size(cache);
unsigned int flags = GFP_NOFS;
if (index == ULONG_MAX)
@@ -166,10 +168,10 @@ static struct f2fs_cached_block *f2fs_create_cache(
if (nofail) {
flags |= __GFP_NOFAIL;
- entry = kzalloc_obj(*entry, flags);
+ entry = kzalloc(entry_size, flags);
entry->data = kmalloc(sbi->blocksize, flags);
} else {
- entry = f2fs_kzalloc(sbi, sizeof(*entry), flags);
+ entry = f2fs_kzalloc(sbi, entry_size, flags);
if (!entry)
return ERR_PTR(-ENOMEM);
@@ -220,6 +222,7 @@ static struct f2fs_cached_block *f2fs_insert_cache(
f2fs_bug_on(cache->sbi, !list_empty(&e->list));
list_add_tail(&e->list, &cache->lru_list);
cache->num_entries++;
+ f2fs_nc_entry_attached(e);
}
f2fs_cache_get(e);
spin_unlock_irqrestore(&cache->tree_lock, flags);
@@ -429,8 +432,9 @@ static void f2fs_do_truncate_cache(struct f2fs_cached_block *entry,
if (!radix_tree_delete(&cache->root, entry->index))
f2fs_bug_on(cache->sbi, !entry->cache);
- entry->cache = NULL;
cache->num_entries--;
+ f2fs_nc_entry_detached(entry);
+ entry->cache = NULL;
atomic_dec(&entry->refcount);
f2fs_bug_on(cache->sbi, !f2fs_cache_refcount(entry));
@@ -514,14 +518,24 @@ int f2fs_init_cache(struct f2fs_sb_info *sbi,
void f2fs_destroy_cache(struct f2fs_cached_block_list *cache)
{
- struct list_head *head = &cache->lru_list;
+ struct list_head *head;
struct f2fs_cached_block *entry;
unsigned long flags;
+ unsigned int queue;
f2fs_cache_wait_on_all_writeback(cache);
next:
spin_lock(&cache->list_lock);
- if (list_empty(head)) {
+ head = NULL;
+ for (queue = 0; queue < F2FS_NC_NR_QUEUES; queue++) {
+ struct list_head *candidate = f2fs_nc_queue_head(cache, queue);
+
+ if (candidate && !list_empty(candidate)) {
+ head = candidate;
+ break;
+ }
+ }
+ if (!head) {
spin_unlock(&cache->list_lock);
return;
}
@@ -530,6 +544,7 @@ void f2fs_destroy_cache(struct f2fs_cached_block_list *cache)
spin_lock_irqsave(&cache->tree_lock, flags);
radix_tree_delete(&cache->root, entry->index);
cache->num_entries--;
+ f2fs_nc_entry_detached(entry);
list_del_init(&entry->list);
spin_unlock_irqrestore(&cache->tree_lock, flags);
diff --git a/fs/f2fs/cache.h b/fs/f2fs/cache.h
index c4c3d09a0008..8a1eab71b1ae 100644
--- a/fs/f2fs/cache.h
+++ b/fs/f2fs/cache.h
@@ -64,6 +64,9 @@ enum f2fs_cached_state {
F2FS_BLOCK_WRITEBACK, /* cache data is writeback state */
F2FS_BLOCK_REFERENCED, /* cache was accessed recently, shrinker will skip it for once */
F2FS_BLOCK_INLINE_DATA, /* indicate inline data */
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+ F2FS_BLOCK_COMPRESSED,
+#endif
};
enum {
@@ -144,6 +147,11 @@ F2FS_CACHE_FLAG_CLEAR_FUNC(inline, INLINE_DATA);
F2FS_CACHE_FLAG_TEST_FUNC(referenced, REFERENCED);
F2FS_CACHE_FLAG_SET_FUNC(referenced, REFERENCED);
F2FS_CACHE_FLAG_TEST_AND_CLEAR_FUNC(referenced, REFERENCED);
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+F2FS_CACHE_FLAG_TEST_FUNC(compressed, COMPRESSED);
+F2FS_CACHE_FLAG_SET_FUNC(compressed, COMPRESSED);
+F2FS_CACHE_FLAG_CLEAR_FUNC(compressed, COMPRESSED);
+#endif
static inline void *cache_address(const struct f2fs_cached_block *entry)
{
diff --git a/fs/f2fs/debug.c b/fs/f2fs/debug.c
index 6fe606e2c70d..a6096537b495 100644
--- a/fs/f2fs/debug.c
+++ b/fs/f2fs/debug.c
@@ -19,6 +19,7 @@
#include "node.h"
#include "segment.h"
#include "gc.h"
+#include "node_cache_compress.h"
static LIST_HEAD(f2fs_stat_list);
static DEFINE_SPINLOCK(f2fs_stat_lock);
@@ -297,6 +298,7 @@ static void update_general_status(struct f2fs_sb_info *sbi)
static void update_mem_info(struct f2fs_sb_info *sbi)
{
struct f2fs_stat_info *si = F2FS_STAT(sbi);
+ struct f2fs_nc_memory node_memory;
int i;
if (si->base_mem)
@@ -386,11 +388,9 @@ static void update_mem_info(struct f2fs_sb_info *sbi)
si->cache_data_mem[F2FS_META_CACHE] =
(unsigned long long)META_CACHE(sbi)->num_entries * sbi->blocksize;
- si->cache_entry_mem[F2FS_NODE_CACHE] =
- (unsigned long long)NODE_CACHE(sbi)->num_entries *
- sizeof(struct f2fs_cached_block);
- si->cache_data_mem[F2FS_NODE_CACHE] =
- (unsigned long long)NODE_CACHE(sbi)->num_entries * sbi->blocksize;
+ f2fs_nc_memory_usage(sbi, &node_memory);
+ si->cache_entry_mem[F2FS_NODE_CACHE] = node_memory.entry_bytes;
+ si->cache_data_mem[F2FS_NODE_CACHE] = node_memory.data_bytes;
si->cache_mem += si->cache_entry_mem[F2FS_META_CACHE] +
si->cache_entry_mem[F2FS_NODE_CACHE];
diff --git a/fs/f2fs/f2fs.h b/fs/f2fs/f2fs.h
index 089a62c054ea..16cc050124dd 100644
--- a/fs/f2fs/f2fs.h
+++ b/fs/f2fs/f2fs.h
@@ -26,6 +26,10 @@
#include <linux/part_stat.h>
#include <linux/rw_hint.h>
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+struct f2fs_nc_ctx;
+#endif
+
#include <linux/fscrypt.h>
#include <linux/fsverity.h>
@@ -2106,6 +2110,9 @@ struct f2fs_sb_info {
struct f2fs_cached_block_list meta_blocks;
struct f2fs_cached_block_list node_blocks;
struct f2fs_cached_block_list compress_blocks;
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+ struct f2fs_nc_ctx *node_compress;
+#endif
/* internal cache flush thread */
struct f2fs_cache_kthread cache_thread;
diff --git a/fs/f2fs/node_cache_compress.c b/fs/f2fs/node_cache_compress.c
new file mode 100644
index 000000000000..486da4ed6082
--- /dev/null
+++ b/fs/f2fs/node_cache_compress.c
@@ -0,0 +1,303 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <linux/atomic.h>
+#include <linux/f2fs_fs.h>
+#include <linux/refcount.h>
+#include <linux/slab.h>
+
+#include "f2fs.h"
+#include "node_cache_compress.h"
+
+static const u32 f2fs_nc_bucket_sizes[] = {
+ F2FS_NC_BUCKET_256_SIZE,
+ F2FS_NC_BUCKET_512_SIZE,
+ F2FS_NC_BUCKET_1024_SIZE,
+};
+
+/* One private slab and active-object count for each compressed bucket. */
+struct f2fs_nc_store {
+ struct kmem_cache *caches[ARRAY_SIZE(f2fs_nc_bucket_sizes)];
+ atomic_long_t objects[ARRAY_SIZE(f2fs_nc_bucket_sizes)];
+};
+
+static int f2fs_nc_store_bucket_from_size(u32 alloc_size)
+{
+ int i;
+
+ for (i = 0; i < ARRAY_SIZE(f2fs_nc_bucket_sizes); i++)
+ if (alloc_size == f2fs_nc_bucket_sizes[i])
+ return i;
+ return -EINVAL;
+}
+
+static struct f2fs_nc_store *f2fs_nc_store_create(struct f2fs_sb_info *sbi)
+{
+ struct f2fs_nc_store *store;
+ char name[64];
+ int i;
+
+ store = kzalloc_obj(*store, GFP_NOFS);
+ if (!store)
+ return NULL;
+ for (i = 0; i < ARRAY_SIZE(f2fs_nc_bucket_sizes); i++) {
+ if (snprintf(name, sizeof(name), "f2fs-nc-%u-%u-%u",
+ MAJOR(sbi->sb->s_dev), MINOR(sbi->sb->s_dev),
+ f2fs_nc_bucket_sizes[i]) >= sizeof(name))
+ goto fail;
+ store->caches[i] = kmem_cache_create(name,
+ f2fs_nc_bucket_sizes[i], 0,
+ SLAB_RECLAIM_ACCOUNT | SLAB_NO_MERGE, NULL);
+ if (!store->caches[i])
+ goto fail;
+ }
+ return store;
+
+fail:
+ while (i--)
+ kmem_cache_destroy(store->caches[i]);
+ kfree(store);
+ return NULL;
+}
+
+static void f2fs_nc_store_destroy(struct f2fs_nc_store *store)
+{
+ int i;
+
+ if (!store)
+ return;
+ for (i = ARRAY_SIZE(f2fs_nc_bucket_sizes); i-- > 0;) {
+ WARN_ON_ONCE(atomic_long_read(&store->objects[i]));
+ kmem_cache_destroy(store->caches[i]);
+ }
+ kfree(store);
+}
+
+static void f2fs_nc_store_free(struct f2fs_nc_store *store, void *object,
+ u32 alloc_size)
+{
+ int bucket = f2fs_nc_store_bucket_from_size(alloc_size);
+
+ if (WARN_ON_ONCE(!store || !object || bucket < 0))
+ return;
+ kmem_cache_free(store->caches[bucket], object);
+ atomic_long_dec(&store->objects[bucket]);
+}
+
+/*
+ * Per-superblock compression state. Compressed objects hold references, so
+ * the context can outlive its mount until every detached object is freed.
+ */
+struct f2fs_nc_ctx {
+ struct f2fs_sb_info *sbi;
+ struct f2fs_nc_store *store;
+ /* Compressed queues only; raw entries use NODE_CACHE()->lru_list. */
+ struct list_head queues[F2FS_NC_NR_QUEUES - 1];
+ /* Current population and compressed bytes by queue. */
+ atomic_long_t attached[F2FS_NC_NR_QUEUES];
+ atomic64_t attached_payload[F2FS_NC_NR_QUEUES - 1];
+ atomic64_t attached_slot_bytes[F2FS_NC_NR_QUEUES - 1];
+ /* Cumulative detach statistics since mount. */
+ atomic64_t detached[F2FS_NC_NR_QUEUES];
+ /* Mount reference plus references held by compressed objects. */
+ refcount_t refs;
+};
+
+static struct f2fs_node_cached_block *f2fs_nc_node_entry(struct f2fs_cached_block *entry)
+{
+ return container_of(entry, struct f2fs_node_cached_block, base);
+}
+
+static void f2fs_nc_ctx_release(struct f2fs_nc_ctx *ctx)
+{
+ int i;
+
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++)
+ WARN_ON_ONCE(atomic_long_read(&ctx->attached[i]));
+ f2fs_nc_store_destroy(ctx->store);
+ kfree(ctx);
+}
+
+static void f2fs_nc_ctx_put(struct f2fs_nc_ctx *ctx)
+{
+ if (refcount_dec_and_test(&ctx->refs))
+ f2fs_nc_ctx_release(ctx);
+}
+
+static void f2fs_nc_account_add(struct f2fs_nc_ctx *ctx, unsigned int queue,
+ u32 len, u32 alloc_size)
+{
+ atomic_long_inc(&ctx->attached[queue]);
+ if (queue == F2FS_NC_RAW)
+ return;
+ atomic64_add(len, &ctx->attached_payload[queue - 1]);
+ atomic64_add(alloc_size, &ctx->attached_slot_bytes[queue - 1]);
+}
+
+static void f2fs_nc_account_del(struct f2fs_nc_ctx *ctx, unsigned int queue,
+ u32 len, u32 alloc_size)
+{
+ WARN_ON_ONCE(atomic_long_read(&ctx->attached[queue]) <= 0);
+ atomic_long_dec(&ctx->attached[queue]);
+ if (queue == F2FS_NC_RAW)
+ return;
+ WARN_ON_ONCE(atomic64_read(&ctx->attached_payload[queue - 1]) < len);
+ WARN_ON_ONCE(atomic64_read(&ctx->attached_slot_bytes[queue - 1]) <
+ alloc_size);
+ atomic64_sub(len, &ctx->attached_payload[queue - 1]);
+ atomic64_sub(alloc_size, &ctx->attached_slot_bytes[queue - 1]);
+}
+
+size_t f2fs_nc_entry_alloc_size(struct f2fs_cached_block_list *cache)
+{
+ if (IS_NODE_CACHE(cache) && cache->sbi->node_compress)
+ return sizeof(struct f2fs_node_cached_block);
+ return sizeof(struct f2fs_cached_block);
+}
+
+void f2fs_nc_init(struct f2fs_sb_info *sbi)
+{
+ struct f2fs_nc_ctx *ctx;
+ int i;
+
+ ctx = kzalloc_obj(*ctx, GFP_NOFS);
+ if (!ctx)
+ goto fail_open;
+ ctx->sbi = sbi;
+ ctx->store = f2fs_nc_store_create(sbi);
+ if (!ctx->store)
+ goto fail_open;
+ for (i = 0; i < ARRAY_SIZE(ctx->queues); i++)
+ INIT_LIST_HEAD(&ctx->queues[i]);
+ refcount_set(&ctx->refs, 1);
+ sbi->node_compress = ctx;
+ return;
+
+fail_open:
+ if (ctx)
+ f2fs_nc_ctx_release(ctx);
+ f2fs_warn_ratelimited(sbi,
+ "node cache compression resources are unavailable");
+}
+
+void f2fs_nc_destroy(struct f2fs_sb_info *sbi)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+
+ if (!ctx)
+ return;
+ sbi->node_compress = NULL;
+ f2fs_nc_ctx_put(ctx);
+}
+
+struct list_head *f2fs_nc_queue_head(struct f2fs_cached_block_list *cache,
+ unsigned int queue)
+{
+ struct f2fs_nc_ctx *ctx;
+
+ if (queue == F2FS_NC_RAW)
+ return &cache->lru_list;
+ if (!IS_NODE_CACHE(cache) || queue >= F2FS_NC_NR_QUEUES)
+ return NULL;
+ ctx = cache->sbi->node_compress;
+ return ctx ? &ctx->queues[queue - 1] : NULL;
+}
+
+unsigned int f2fs_nc_entry_queue(const struct f2fs_cached_block *entry)
+{
+ const struct f2fs_node_cached_block *node;
+ int bucket;
+
+ if (!f2fs_cache_test_compressed(entry))
+ return F2FS_NC_RAW;
+ node = container_of(entry, struct f2fs_node_cached_block, base);
+ bucket = f2fs_nc_store_bucket_from_size(node->compressed_alloc_size);
+ if (WARN_ON_ONCE(bucket < 0))
+ return F2FS_NC_RAW;
+ return bucket + 1;
+}
+
+void f2fs_nc_entry_attached(struct f2fs_cached_block *entry)
+{
+ struct f2fs_nc_ctx *ctx;
+
+ if (!entry->cache || !IS_NODE_CACHE(entry->cache))
+ return;
+ ctx = entry->cache->sbi->node_compress;
+ if (ctx)
+ f2fs_nc_account_add(ctx, F2FS_NC_RAW, 0, 0);
+}
+
+void f2fs_nc_entry_detached(struct f2fs_cached_block *entry)
+{
+ struct f2fs_node_cached_block *node;
+ struct f2fs_nc_ctx *ctx;
+ unsigned int queue;
+ u32 len = 0, alloc_size = 0;
+
+ if (!entry->cache || !IS_NODE_CACHE(entry->cache))
+ return;
+ ctx = entry->cache->sbi->node_compress;
+ if (!ctx)
+ return;
+ queue = f2fs_nc_entry_queue(entry);
+ if (queue != F2FS_NC_RAW) {
+ node = f2fs_nc_node_entry(entry);
+ len = node->compressed_len;
+ alloc_size = node->compressed_alloc_size;
+ }
+ atomic64_inc(&ctx->detached[queue]);
+ f2fs_nc_account_del(ctx, queue, len, alloc_size);
+}
+
+void f2fs_nc_free_data(struct f2fs_cached_block *entry)
+{
+ struct f2fs_node_cached_block *node;
+ struct f2fs_nc_ctx *ctx;
+ void *object;
+ u32 alloc_size;
+
+ if (!f2fs_cache_test_compressed(entry)) {
+ kfree(entry->data);
+ return;
+ }
+ node = f2fs_nc_node_entry(entry);
+ ctx = node->owner;
+ object = entry->data;
+ alloc_size = node->compressed_alloc_size;
+ if (WARN_ON_ONCE(!ctx || !object ||
+ f2fs_nc_store_bucket_from_size(alloc_size) < 0))
+ return;
+ f2fs_cache_clear_compressed(entry);
+ entry->data = NULL;
+ node->owner = NULL;
+ node->compressed_len = 0;
+ node->compressed_alloc_size = 0;
+ node->compressed_crc = 0;
+ f2fs_nc_store_free(ctx->store, object, alloc_size);
+ f2fs_nc_ctx_put(ctx);
+}
+
+void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
+ struct f2fs_nc_memory *memory)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ u64 entries = 0, data = 0;
+ unsigned int i;
+
+ if (!ctx) {
+ memory->entry_bytes = (u64)NODE_CACHE(sbi)->num_entries *
+ sizeof(struct f2fs_cached_block);
+ memory->data_bytes = (u64)NODE_CACHE(sbi)->num_entries *
+ sbi->blocksize;
+ return;
+ }
+ spin_lock(&NODE_CACHE(sbi)->list_lock);
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++)
+ entries += atomic_long_read(&ctx->attached[i]);
+ data = (u64)atomic_long_read(&ctx->attached[F2FS_NC_RAW]) *
+ sbi->blocksize;
+ for (i = 0; i < F2FS_NC_NR_QUEUES - 1; i++)
+ data += atomic64_read(&ctx->attached_slot_bytes[i]);
+ spin_unlock(&NODE_CACHE(sbi)->list_lock);
+ memory->entry_bytes = entries * sizeof(struct f2fs_node_cached_block);
+ memory->data_bytes = data;
+}
diff --git a/fs/f2fs/node_cache_compress.h b/fs/f2fs/node_cache_compress.h
new file mode 100644
index 000000000000..e1a9b45d67e3
--- /dev/null
+++ b/fs/f2fs/node_cache_compress.h
@@ -0,0 +1,78 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef __F2FS_NODE_CACHE_COMPRESS_H__
+#define __F2FS_NODE_CACHE_COMPRESS_H__
+
+#include "cache.h"
+#include "node_cache_policy.h"
+
+struct f2fs_nc_ctx;
+
+/* Attributed cache memory, excluding worker workspace and slab metadata. */
+struct f2fs_nc_memory {
+ u64 entry_bytes; /* Cache-entry storage. */
+ u64 data_bytes; /* Raw buffers plus compressed slab slots. */
+};
+
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+/* Extended NODE_CACHE entry; base must remain the first member. */
+struct f2fs_node_cached_block {
+ struct f2fs_cached_block base;
+ struct f2fs_nc_ctx *owner; /* Context owning compressed data. */
+ u32 compressed_len; /* Valid compressed payload bytes. */
+ u32 compressed_alloc_size; /* Private-slab slot size. */
+ u32 compressed_crc; /* CRC of the original node block. */
+};
+
+static_assert(offsetof(struct f2fs_node_cached_block, base) == 0);
+
+size_t f2fs_nc_entry_alloc_size(struct f2fs_cached_block_list *cache);
+void f2fs_nc_init(struct f2fs_sb_info *sbi);
+void f2fs_nc_destroy(struct f2fs_sb_info *sbi);
+void f2fs_nc_free_data(struct f2fs_cached_block *entry);
+struct list_head *f2fs_nc_queue_head(struct f2fs_cached_block_list *cache,
+ unsigned int queue);
+unsigned int f2fs_nc_entry_queue(const struct f2fs_cached_block *entry);
+void f2fs_nc_entry_attached(struct f2fs_cached_block *entry);
+void f2fs_nc_entry_detached(struct f2fs_cached_block *entry);
+void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
+ struct f2fs_nc_memory *memory);
+#else
+static inline size_t f2fs_nc_entry_alloc_size(struct f2fs_cached_block_list *cache)
+{
+ return sizeof(struct f2fs_cached_block);
+}
+
+static inline void f2fs_nc_init(struct f2fs_sb_info *sbi) { }
+
+static inline void f2fs_nc_destroy(struct f2fs_sb_info *sbi) { }
+
+static inline void f2fs_nc_free_data(struct f2fs_cached_block *entry)
+{
+ kfree(entry->data);
+}
+
+static inline struct list_head *f2fs_nc_queue_head(struct f2fs_cached_block_list *cache,
+ unsigned int queue)
+{
+ return queue == F2FS_NC_RAW ? &cache->lru_list : NULL;
+}
+
+static inline unsigned int
+f2fs_nc_entry_queue(const struct f2fs_cached_block *entry)
+{
+ return F2FS_NC_RAW;
+}
+
+static inline void f2fs_nc_entry_attached(struct f2fs_cached_block *entry) { }
+static inline void f2fs_nc_entry_detached(struct f2fs_cached_block *entry) { }
+
+static inline void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
+ struct f2fs_nc_memory *memory)
+{
+ memory->entry_bytes = (u64)NODE_CACHE(sbi)->num_entries *
+ sizeof(struct f2fs_cached_block);
+ memory->data_bytes = (u64)NODE_CACHE(sbi)->num_entries * sbi->blocksize;
+}
+#endif
+
+#endif /* __F2FS_NODE_CACHE_COMPRESS_H__ */
diff --git a/fs/f2fs/node_cache_policy.h b/fs/f2fs/node_cache_policy.h
new file mode 100644
index 000000000000..efe906bfcc58
--- /dev/null
+++ b/fs/f2fs/node_cache_policy.h
@@ -0,0 +1,23 @@
+/* SPDX-License-Identifier: GPL-2.0 */
+#ifndef __F2FS_NODE_CACHE_POLICY_H__
+#define __F2FS_NODE_CACHE_POLICY_H__
+
+#include <linux/types.h>
+
+#define F2FS_NC_BUCKET_256_SIZE 256U
+#define F2FS_NC_BUCKET_512_SIZE 512U
+#define F2FS_NC_BUCKET_1024_SIZE 1024U
+#define F2FS_NC_MAX_OBJECT_SIZE F2FS_NC_BUCKET_1024_SIZE
+
+/*
+ * Keep compressed queues in the same order as f2fs_nc_bucket_sizes[].
+ * Raw entries remain on NODE_CACHE(sbi)->lru_list.
+ */
+enum f2fs_nc_queue {
+ F2FS_NC_RAW,
+ F2FS_NC_256,
+ F2FS_NC_512,
+ F2FS_NC_1024,
+ F2FS_NC_NR_QUEUES,
+};
+#endif /* __F2FS_NODE_CACHE_POLICY_H__ */
diff --git a/fs/f2fs/super.c b/fs/f2fs/super.c
index 294f6f2c28a5..461209293dba 100644
--- a/fs/f2fs/super.c
+++ b/fs/f2fs/super.c
@@ -32,6 +32,7 @@
#include <linux/fserror.h>
#include "f2fs.h"
+#include "node_cache_compress.h"
#include "node.h"
#include "segment.h"
#include "xattr.h"
@@ -2071,6 +2072,7 @@ static void f2fs_put_super(struct super_block *sb)
f2fs_destroy_cache(COMPRESS_CACHE(sbi));
f2fs_destroy_cache(NODE_CACHE(sbi));
+ f2fs_nc_destroy(sbi);
f2fs_destroy_cache(META_CACHE(sbi));
/* Should check the page counts after dropping all node/meta pages */
@@ -5263,6 +5265,7 @@ static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc)
f2fs_init_cache(sbi, META_CACHE(sbi), F2FS_META_CACHE);
f2fs_init_cache(sbi, NODE_CACHE(sbi), F2FS_NODE_CACHE);
f2fs_init_cache(sbi, COMPRESS_CACHE(sbi), F2FS_COMPRESS_CACHE);
+ f2fs_nc_init(sbi);
err = f2fs_get_valid_checkpoint(sbi);
if (err) {
@@ -5581,6 +5584,7 @@ static int f2fs_fill_super(struct super_block *sb, struct fs_context *fc)
free_compress_cache:
f2fs_destroy_cache(COMPRESS_CACHE(sbi));
f2fs_destroy_cache(NODE_CACHE(sbi));
+ f2fs_nc_destroy(sbi);
f2fs_destroy_cache(META_CACHE(sbi));
f2fs_destroy_page_array_cache(sbi);
free_percpu:
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH 3/6] f2fs: bias node cache reclaim toward raw entries
2026-09-29 7:29 [RFC PATCH 0/6] f2fs: retain clean node blocks in a compressed cache Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 1/6] f2fs: generalize metadata cache shrinking to explicit lists Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 2/6] f2fs: add compressed clean node cache representation Wenjie Qi
@ 2026-09-29 7:29 ` Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 4/6] f2fs: restore compressed node cache entries before access Wenjie Qi
` (2 subsequent siblings)
5 siblings, 0 replies; 7+ messages in thread
From: Wenjie Qi @ 2026-09-29 7:29 UTC (permalink / raw)
To: Jaegeuk Kim, Chao Yu
Cc: Barry Song, linux-f2fs-devel, linux-kernel, Wenjie Qi
Raw and compressed node-cache entries have very different memory costs.
Counting every entry equally would overstate pressure from compressed
entries and reclaim them too aggressively.
Count and scan the raw and compressed queues separately. Give raw entries
their full block weight and compressed entries 25 percent of their
allocation size. Use the same weights for the shrinker count and the
per-queue scan quotas.
Carry rounding credit across shrinker calls, and redistribute unused quota
when a queue reaches its population limit. This favors reclaiming raw
entries while keeping every compressed queue reclaimable.
Signed-off-by: Wenjie Qi <qiwenjie@xiaomi.com>
---
fs/f2fs/Makefile | 1 +
fs/f2fs/cache.c | 2 +-
fs/f2fs/node_cache_compress.c | 64 ++++++++++++-
fs/f2fs/node_cache_compress.h | 15 ++++
fs/f2fs/node_cache_policy.c | 163 ++++++++++++++++++++++++++++++++++
fs/f2fs/node_cache_policy.h | 9 ++
fs/f2fs/shrinker.c | 3 +-
7 files changed, 254 insertions(+), 3 deletions(-)
create mode 100644 fs/f2fs/node_cache_policy.c
diff --git a/fs/f2fs/Makefile b/fs/f2fs/Makefile
index e1e4b98707e9..79252f510daf 100644
--- a/fs/f2fs/Makefile
+++ b/fs/f2fs/Makefile
@@ -10,4 +10,5 @@ f2fs-$(CONFIG_F2FS_FS_POSIX_ACL) += acl.o
f2fs-$(CONFIG_FS_VERITY) += verity.o
f2fs-$(CONFIG_F2FS_FS_COMPRESSION) += compress.o
f2fs-$(CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION) += node_cache_compress.o
+f2fs-$(CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION) += node_cache_policy.o
f2fs-$(CONFIG_F2FS_IOSTAT) += iostat.o
diff --git a/fs/f2fs/cache.c b/fs/f2fs/cache.c
index 3d1d530cdd82..480412f6c729 100644
--- a/fs/f2fs/cache.c
+++ b/fs/f2fs/cache.c
@@ -666,7 +666,7 @@ unsigned long f2fs_shrink_cache(struct f2fs_sb_info *sbi,
if (freed >= nr_to_scan)
return freed;
- freed += f2fs_do_shrink_cache(NODE_CACHE(sbi), nr_to_scan - freed);
+ freed += f2fs_nc_shrink_nodes(sbi, nr_to_scan - freed);
if (freed >= nr_to_scan)
return freed;
diff --git a/fs/f2fs/node_cache_compress.c b/fs/f2fs/node_cache_compress.c
index 486da4ed6082..50e9c820eb81 100644
--- a/fs/f2fs/node_cache_compress.c
+++ b/fs/f2fs/node_cache_compress.c
@@ -95,8 +95,12 @@ struct f2fs_nc_ctx {
atomic_long_t attached[F2FS_NC_NR_QUEUES];
atomic64_t attached_payload[F2FS_NC_NR_QUEUES - 1];
atomic64_t attached_slot_bytes[F2FS_NC_NR_QUEUES - 1];
- /* Cumulative detach statistics since mount. */
+ /* Cumulative reclaim and detach statistics since mount. */
+ atomic64_t shrink_scanned[F2FS_NC_NR_QUEUES];
+ atomic64_t shrink_freed[F2FS_NC_NR_QUEUES];
atomic64_t detached[F2FS_NC_NR_QUEUES];
+ /* Signed quota history carried between shrinker calls. */
+ s64 shrink_credit[F2FS_NC_NR_QUEUES];
/* Mount reference plus references held by compressed objects. */
refcount_t refs;
};
@@ -276,6 +280,64 @@ void f2fs_nc_free_data(struct f2fs_cached_block *entry)
f2fs_nc_ctx_put(ctx);
}
+static void f2fs_nc_population_snapshot(struct f2fs_nc_ctx *ctx,
+ unsigned long nr[F2FS_NC_NR_QUEUES])
+{
+ struct f2fs_cached_block_list *cache = NODE_CACHE(ctx->sbi);
+ unsigned int i;
+
+ spin_lock(&cache->list_lock);
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++)
+ nr[i] = atomic_long_read(&ctx->attached[i]);
+ spin_unlock(&cache->list_lock);
+}
+
+unsigned long f2fs_nc_count_nodes(struct f2fs_sb_info *sbi)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ unsigned long nr[F2FS_NC_NR_QUEUES];
+
+ if (!ctx)
+ return NODE_CACHE(sbi)->num_entries;
+ f2fs_nc_population_snapshot(ctx, nr);
+ return f2fs_nc_effective_count(nr, sbi->blocksize);
+}
+
+unsigned long f2fs_nc_shrink_nodes(struct f2fs_sb_info *sbi,
+ unsigned long nr_to_scan)
+{
+ struct f2fs_cached_block_list *cache = NODE_CACHE(sbi);
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ unsigned long nr[F2FS_NC_NR_QUEUES];
+ unsigned long quota[F2FS_NC_NR_QUEUES];
+ unsigned long freed = 0;
+ unsigned int i;
+
+ if (!ctx)
+ return f2fs_shrink_cache_list(cache, &cache->lru_list,
+ ULONG_MAX, nr_to_scan, NULL);
+ f2fs_nc_population_snapshot(ctx, nr);
+ f2fs_nc_scan_quotas(nr, sbi->blocksize, nr_to_scan,
+ ctx->shrink_credit, quota);
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++) {
+ struct list_head *head;
+ unsigned long scanned = 0;
+ unsigned long queue_freed;
+
+ if (!quota[i])
+ continue;
+ head = f2fs_nc_queue_head(cache, i);
+ if (!head)
+ continue;
+ queue_freed = f2fs_shrink_cache_list(cache, head, nr[i], quota[i],
+ &scanned);
+ freed += queue_freed;
+ atomic64_add(scanned, &ctx->shrink_scanned[i]);
+ atomic64_add(queue_freed, &ctx->shrink_freed[i]);
+ }
+ return freed;
+}
+
void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
struct f2fs_nc_memory *memory)
{
diff --git a/fs/f2fs/node_cache_compress.h b/fs/f2fs/node_cache_compress.h
index e1a9b45d67e3..e7d6b07c1c98 100644
--- a/fs/f2fs/node_cache_compress.h
+++ b/fs/f2fs/node_cache_compress.h
@@ -34,6 +34,9 @@ struct list_head *f2fs_nc_queue_head(struct f2fs_cached_block_list *cache,
unsigned int f2fs_nc_entry_queue(const struct f2fs_cached_block *entry);
void f2fs_nc_entry_attached(struct f2fs_cached_block *entry);
void f2fs_nc_entry_detached(struct f2fs_cached_block *entry);
+unsigned long f2fs_nc_count_nodes(struct f2fs_sb_info *sbi);
+unsigned long f2fs_nc_shrink_nodes(struct f2fs_sb_info *sbi,
+ unsigned long nr_to_scan);
void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
struct f2fs_nc_memory *memory);
#else
@@ -66,6 +69,18 @@ f2fs_nc_entry_queue(const struct f2fs_cached_block *entry)
static inline void f2fs_nc_entry_attached(struct f2fs_cached_block *entry) { }
static inline void f2fs_nc_entry_detached(struct f2fs_cached_block *entry) { }
+static inline unsigned long f2fs_nc_count_nodes(struct f2fs_sb_info *sbi)
+{
+ return NODE_CACHE(sbi)->num_entries;
+}
+
+static inline unsigned long f2fs_nc_shrink_nodes(struct f2fs_sb_info *sbi,
+ unsigned long nr_to_scan)
+{
+ return f2fs_shrink_cache_list(NODE_CACHE(sbi),
+ &NODE_CACHE(sbi)->lru_list, ULONG_MAX, nr_to_scan, NULL);
+}
+
static inline void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
struct f2fs_nc_memory *memory)
{
diff --git a/fs/f2fs/node_cache_policy.c b/fs/f2fs/node_cache_policy.c
new file mode 100644
index 000000000000..3adcbf89d37d
--- /dev/null
+++ b/fs/f2fs/node_cache_policy.c
@@ -0,0 +1,163 @@
+// SPDX-License-Identifier: GPL-2.0
+#include <linux/math64.h>
+#include <linux/overflow.h>
+#include <linux/shrinker.h>
+#include <linux/string.h>
+
+#include "node_cache_policy.h"
+
+#define F2FS_NC_RECLAIM_SCALE_PCT 25U
+#define F2FS_NC_SCORE_HEADROOM 4U
+
+static void f2fs_nc_queue_weights(u32 blocksize,
+ u64 weight[F2FS_NC_NR_QUEUES])
+{
+ weight[F2FS_NC_RAW] = blocksize;
+ weight[F2FS_NC_256] = F2FS_NC_BUCKET_256_SIZE *
+ F2FS_NC_RECLAIM_SCALE_PCT / F2FS_NC_PERCENT_MAX;
+ weight[F2FS_NC_512] = F2FS_NC_BUCKET_512_SIZE *
+ F2FS_NC_RECLAIM_SCALE_PCT / F2FS_NC_PERCENT_MAX;
+ weight[F2FS_NC_1024] = F2FS_NC_BUCKET_1024_SIZE *
+ F2FS_NC_RECLAIM_SCALE_PCT / F2FS_NC_PERCENT_MAX;
+}
+
+static u64 f2fs_nc_queue_mass(unsigned long nr, u64 weight,
+ unsigned long budget)
+{
+ u64 limit;
+ u64 count;
+
+ if (!nr || !weight || !budget)
+ return 0;
+ /* Keep both the quota product and signed score arithmetic bounded. */
+ limit = min_t(u64, S64_MAX / F2FS_NC_SCORE_HEADROOM /
+ F2FS_NC_NR_QUEUES,
+ div64_u64(U64_MAX, budget));
+ if (limit < weight)
+ return 1;
+ count = min_t(u64, nr, div64_u64(limit, weight));
+ return count * weight;
+}
+
+u64 f2fs_nc_effective_count(const unsigned long nr[F2FS_NC_NR_QUEUES],
+ u32 blocksize)
+{
+ u64 weight[F2FS_NC_NR_QUEUES];
+ u64 cap = (u64)SHRINK_EMPTY - 1;
+ u64 count = 0;
+ u64 remainder = 0;
+ u64 carry;
+ u64 carry_remainder;
+ unsigned int i;
+
+ if (!blocksize)
+ return 0;
+ f2fs_nc_queue_weights(blocksize, weight);
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++) {
+ u64 quotient, count_remainder;
+
+ quotient = mul_u64_u64_div_u64(nr[i], weight[i], blocksize);
+ if (quotient > cap - count)
+ return cap;
+ count += quotient;
+ div64_u64_rem(nr[i], blocksize, &count_remainder);
+ remainder += count_remainder * weight[i] % blocksize;
+ }
+ carry = div64_u64_rem(remainder, blocksize, &carry_remainder);
+ if (carry_remainder)
+ carry++;
+ if (carry > cap - count)
+ return cap;
+ return count + carry;
+}
+
+void f2fs_nc_scan_quotas(const unsigned long nr[F2FS_NC_NR_QUEUES],
+ u32 blocksize, unsigned long requested,
+ s64 credit[F2FS_NC_NR_QUEUES],
+ unsigned long quota[F2FS_NC_NR_QUEUES])
+{
+ u64 weight[F2FS_NC_NR_QUEUES];
+ u64 mass[F2FS_NC_NR_QUEUES], total_mass;
+ s64 score[F2FS_NC_NR_QUEUES];
+ unsigned long population = 0, budget, assigned = 0;
+ unsigned int round, i;
+
+ memset(quota, 0, sizeof(*quota) * F2FS_NC_NR_QUEUES);
+ f2fs_nc_queue_weights(blocksize, weight);
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++) {
+ if (!nr[i] || !weight[i])
+ credit[i] = 0;
+ if (weight[i] &&
+ check_add_overflow(population, nr[i], &population))
+ population = ULONG_MAX;
+ }
+ budget = min(requested, population);
+ if (!budget || !blocksize)
+ return;
+ total_mass = 0;
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++) {
+ mass[i] = f2fs_nc_queue_mass(nr[i], weight[i], budget);
+ total_mass += mass[i];
+ }
+ if (!total_mass)
+ return;
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++)
+ score[i] = mass[i] ? clamp_t(s64, credit[i],
+ -(s64)total_mass,
+ (s64)total_mass) : 0;
+
+ for (round = 0; round < F2FS_NC_NR_QUEUES + 1 && assigned < budget;
+ round++) {
+ u64 active_mass = 0;
+ unsigned long before = assigned;
+ bool capped = false;
+
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++)
+ if (mass[i] && quota[i] < nr[i])
+ active_mass += mass[i];
+ if (!active_mass)
+ break;
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++) {
+ unsigned long capacity, share;
+ u64 product, quotient, remainder;
+
+ if (!mass[i] || quota[i] >= nr[i])
+ continue;
+ capacity = nr[i] - quota[i];
+ quotient = mul_u64_u64_div_u64(budget - assigned,
+ mass[i], active_mass);
+ share = min_t(u64, quotient, capacity);
+ if (share < quotient)
+ capped = true;
+ quota[i] += share;
+ before += share;
+ product = (budget - assigned) * mass[i];
+ remainder = product - quotient * active_mass;
+ score[i] = clamp_t(s64, score[i], -(s64)active_mass,
+ (s64)active_mass) + (s64)remainder;
+ }
+ assigned = before;
+ while (assigned < budget) {
+ unsigned int best = F2FS_NC_NR_QUEUES;
+
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++) {
+ if (!mass[i] || quota[i] >= nr[i])
+ continue;
+ if (best == F2FS_NC_NR_QUEUES ||
+ score[i] > score[best])
+ best = i;
+ }
+ if (best == F2FS_NC_NR_QUEUES)
+ break;
+ quota[best]++;
+ score[best] -= active_mass;
+ assigned++;
+ if (capped)
+ break;
+ }
+ if (!capped || assigned == budget)
+ break;
+ }
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++)
+ credit[i] = mass[i] ? score[i] : 0;
+}
diff --git a/fs/f2fs/node_cache_policy.h b/fs/f2fs/node_cache_policy.h
index efe906bfcc58..42af4ad1698f 100644
--- a/fs/f2fs/node_cache_policy.h
+++ b/fs/f2fs/node_cache_policy.h
@@ -8,6 +8,7 @@
#define F2FS_NC_BUCKET_512_SIZE 512U
#define F2FS_NC_BUCKET_1024_SIZE 1024U
#define F2FS_NC_MAX_OBJECT_SIZE F2FS_NC_BUCKET_1024_SIZE
+#define F2FS_NC_PERCENT_MAX 100U
/*
* Keep compressed queues in the same order as f2fs_nc_bucket_sizes[].
@@ -20,4 +21,12 @@ enum f2fs_nc_queue {
F2FS_NC_1024,
F2FS_NC_NR_QUEUES,
};
+
+u64 f2fs_nc_effective_count(const unsigned long nr[F2FS_NC_NR_QUEUES],
+ u32 blocksize);
+void f2fs_nc_scan_quotas(const unsigned long nr[F2FS_NC_NR_QUEUES],
+ u32 blocksize, unsigned long requested,
+ s64 credit[F2FS_NC_NR_QUEUES],
+ unsigned long quota[F2FS_NC_NR_QUEUES]);
+
#endif /* __F2FS_NODE_CACHE_POLICY_H__ */
diff --git a/fs/f2fs/shrinker.c b/fs/f2fs/shrinker.c
index 29f488531492..a4e9e61557cc 100644
--- a/fs/f2fs/shrinker.c
+++ b/fs/f2fs/shrinker.c
@@ -11,6 +11,7 @@
#include "f2fs.h"
#include "node.h"
+#include "node_cache_compress.h"
static LIST_HEAD(f2fs_list);
static DEFINE_SPINLOCK(f2fs_list_lock);
@@ -40,7 +41,7 @@ static unsigned long __count_extent_cache(struct f2fs_sb_info *sbi,
static unsigned long __count_cache(struct f2fs_sb_info *sbi)
{
return META_CACHE(sbi)->num_entries +
- NODE_CACHE(sbi)->num_entries +
+ f2fs_nc_count_nodes(sbi) +
COMPRESS_CACHE(sbi)->num_entries;
}
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH 4/6] f2fs: restore compressed node cache entries before access
2026-09-29 7:29 [RFC PATCH 0/6] f2fs: retain clean node blocks in a compressed cache Wenjie Qi
` (2 preceding siblings ...)
2026-09-29 7:29 ` [RFC PATCH 3/6] f2fs: bias node cache reclaim toward raw entries Wenjie Qi
@ 2026-09-29 7:29 ` Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 5/6] f2fs: compress clean node cache entries in background Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 6/6] f2fs: give referenced node cache entries a shared second chance Wenjie Qi
5 siblings, 0 replies; 7+ messages in thread
From: Wenjie Qi @ 2026-09-29 7:29 UTC (permalink / raw)
To: Jaegeuk Kim, Chao Yu
Cc: Barry Song, linux-f2fs-devel, linux-kernel, Wenjie Qi
A node-cache lookup may find a compressed entry, but existing callers
expect ordinary node bytes. Restore the raw block before returning the
entry to those callers.
Perform the restore under the entry lock. Validate the allocation size,
compressed length, LZ4 output size and raw-block CRC before publishing the
raw representation. Move a successful restore to the raw LRU tail, and
keep RESTORED set until the existing footer, inode-checksum and node-type
validation has completed.
A raw-buffer allocation failure leaves the compressed entry intact for a
later retry. Invalid compressed data is discarded and the lookup falls
back to the normal disk-read path. The entry never owns raw and compressed
storage at the same time.
Signed-off-by: Wenjie Qi <qiwenjie@xiaomi.com>
---
fs/f2fs/cache.c | 12 ++++
fs/f2fs/cache.h | 20 ++++++
fs/f2fs/data.c | 5 +-
fs/f2fs/f2fs.h | 10 +++
fs/f2fs/inode.c | 20 +++++-
fs/f2fs/node.c | 55 ++++++++++++---
fs/f2fs/node_cache_compress.c | 129 ++++++++++++++++++++++++++++++++++
fs/f2fs/node_cache_compress.h | 13 ++++
8 files changed, 251 insertions(+), 13 deletions(-)
diff --git a/fs/f2fs/cache.c b/fs/f2fs/cache.c
index 480412f6c729..9cb541dd5cea 100644
--- a/fs/f2fs/cache.c
+++ b/fs/f2fs/cache.c
@@ -64,6 +64,7 @@ bool f2fs_mark_cache_dirty(struct f2fs_cached_block *entry)
{
struct f2fs_cached_block_list *cache = entry->cache;
+ f2fs_nc_content_changed(entry);
f2fs_cache_set_uptodate(entry);
#ifdef CONFIG_F2FS_CHECK_FS
@@ -94,6 +95,8 @@ void f2fs_drop_cache_dirty(struct f2fs_cached_block *entry)
F2FS_DIRTY_META : F2FS_DIRTY_NODES;
f2fs_cache_clear_uptodate(entry);
+ if (!f2fs_cache_test_compressed(entry))
+ f2fs_nc_content_changed(entry);
if (!f2fs_cache_test_and_clear_dirty(entry))
return;
@@ -286,7 +289,16 @@ struct f2fs_cached_block *f2fs_grab_cache(
entry = f2fs_insert_cache(cache, index, new);
found:
if (lock) {
+ int ret;
+
f2fs_lock_cache(entry);
+ if (entry->cache == cache) {
+ ret = f2fs_nc_restore(entry);
+ if (ret) {
+ f2fs_put_cache(entry, true);
+ return ERR_PTR(ret);
+ }
+ }
/* has been truncated */
if (entry->cache != cache) {
f2fs_put_cache(entry, true);
diff --git a/fs/f2fs/cache.h b/fs/f2fs/cache.h
index 8a1eab71b1ae..603c5f7203b7 100644
--- a/fs/f2fs/cache.h
+++ b/fs/f2fs/cache.h
@@ -66,6 +66,7 @@ enum f2fs_cached_state {
F2FS_BLOCK_INLINE_DATA, /* indicate inline data */
#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
F2FS_BLOCK_COMPRESSED,
+ F2FS_BLOCK_RESTORED,
#endif
};
@@ -151,6 +152,25 @@ F2FS_CACHE_FLAG_TEST_AND_CLEAR_FUNC(referenced, REFERENCED);
F2FS_CACHE_FLAG_TEST_FUNC(compressed, COMPRESSED);
F2FS_CACHE_FLAG_SET_FUNC(compressed, COMPRESSED);
F2FS_CACHE_FLAG_CLEAR_FUNC(compressed, COMPRESSED);
+F2FS_CACHE_FLAG_TEST_FUNC(restored, RESTORED);
+F2FS_CACHE_FLAG_SET_FUNC(restored, RESTORED);
+F2FS_CACHE_FLAG_CLEAR_FUNC(restored, RESTORED);
+#else
+static inline bool f2fs_cache_test_compressed(const struct f2fs_cached_block *entry)
+{
+ return false;
+}
+
+static inline void f2fs_cache_set_compressed(struct f2fs_cached_block *entry) { }
+static inline void f2fs_cache_clear_compressed(struct f2fs_cached_block *entry) { }
+
+static inline bool f2fs_cache_test_restored(const struct f2fs_cached_block *entry)
+{
+ return false;
+}
+
+static inline void f2fs_cache_set_restored(struct f2fs_cached_block *entry) { }
+static inline void f2fs_cache_clear_restored(struct f2fs_cached_block *entry) { }
#endif
static inline void *cache_address(const struct f2fs_cached_block *entry)
diff --git a/fs/f2fs/data.c b/fs/f2fs/data.c
index 6d4ba5e77906..3c07fdecbd32 100644
--- a/fs/f2fs/data.c
+++ b/fs/f2fs/data.c
@@ -26,6 +26,7 @@
#include "node.h"
#include "segment.h"
#include "iostat.h"
+#include "node_cache_compress.h"
#include <trace/events/f2fs.h>
#define NUM_PREALLOC_POST_READ_CTXS 128
@@ -393,8 +394,10 @@ static void f2fs_cache_read_end_io(struct bio *bio)
entry->index, NODE_TYPE_REGULAR, true))
bio->bi_status = BLK_STS_IOERR;
- if (bio->bi_status == BLK_STS_OK)
+ if (bio->bi_status == BLK_STS_OK) {
+ f2fs_nc_content_changed(entry);
f2fs_cache_set_uptodate(entry);
+ }
dec_cache_count(sbi, io_type);
diff --git a/fs/f2fs/f2fs.h b/fs/f2fs/f2fs.h
index 16cc050124dd..a5d8ab041aa8 100644
--- a/fs/f2fs/f2fs.h
+++ b/fs/f2fs/f2fs.h
@@ -4001,6 +4001,16 @@ int f2fs_pin_file_control(struct inode *inode, bool inc);
*/
void f2fs_set_inode_flags(struct inode *inode);
bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry);
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+bool f2fs_inode_chksum_valid(struct f2fs_sb_info *sbi,
+ struct f2fs_cached_block *entry);
+#else
+static inline bool
+f2fs_inode_chksum_valid(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry)
+{
+ return true;
+}
+#endif
void f2fs_inode_chksum_set(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry);
struct inode *f2fs_iget(struct super_block *sb, unsigned long ino);
struct inode *f2fs_iget_retry(struct super_block *sb, unsigned long ino);
diff --git a/fs/f2fs/inode.c b/fs/f2fs/inode.c
index 9164a2b5d9f0..bd4a46dd7a9e 100644
--- a/fs/f2fs/inode.c
+++ b/fs/f2fs/inode.c
@@ -168,7 +168,9 @@ static __u32 f2fs_inode_chksum(struct f2fs_sb_info *sbi, struct f2fs_cached_bloc
return chksum;
}
-bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry)
+static bool __f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi,
+ struct f2fs_cached_block *entry,
+ bool report)
{
struct f2fs_inode *ri;
__u32 provided, calculated;
@@ -187,7 +189,7 @@ bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct f2fs_cached_block
provided = le32_to_cpu(ri->i_inode_checksum);
calculated = f2fs_inode_chksum(sbi, entry);
- if (provided != calculated)
+ if (provided != calculated && report)
f2fs_warn(sbi, "checksum invalid, nid = %lu, ino_of_node = %u, %x vs. %x",
entry->index, ino_of_node(sbi, entry),
provided, calculated);
@@ -195,6 +197,20 @@ bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct f2fs_cached_block
return provided == calculated;
}
+bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi,
+ struct f2fs_cached_block *entry)
+{
+ return __f2fs_inode_chksum_verify(sbi, entry, true);
+}
+
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+bool f2fs_inode_chksum_valid(struct f2fs_sb_info *sbi,
+ struct f2fs_cached_block *entry)
+{
+ return __f2fs_inode_chksum_verify(sbi, entry, false);
+}
+#endif
+
void f2fs_inode_chksum_set(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry)
{
struct f2fs_inode *ri = F2FS_INODE(entry);
diff --git a/fs/f2fs/node.c b/fs/f2fs/node.c
index be14dfa31e53..2cf8ebf0c114 100644
--- a/fs/f2fs/node.c
+++ b/fs/f2fs/node.c
@@ -19,6 +19,7 @@
#include "segment.h"
#include "xattr.h"
#include "iostat.h"
+#include "node_cache_compress.h"
#include <trace/events/f2fs.h>
#define on_f2fs_build_free_nids(nm_i) mutex_is_locked(&(nm_i)->build_lock)
@@ -1467,7 +1468,15 @@ static int read_node_cache(struct f2fs_cached_block *entry, blk_opf_t op_flags)
};
int err;
+ if (WARN_ON_ONCE(f2fs_cache_test_compressed(entry) || !entry->data))
+ return -EFSCORRUPTED;
+
if (f2fs_cache_test_uptodate(entry)) {
+ if (f2fs_cache_test_restored(entry) &&
+ !f2fs_inode_chksum_valid(sbi, entry)) {
+ f2fs_nc_validation_failed(entry);
+ goto read_disk;
+ }
if (!f2fs_inode_chksum_verify(sbi, entry)) {
f2fs_cache_clear_uptodate(entry);
return -EFSBADCRC;
@@ -1475,6 +1484,7 @@ static int read_node_cache(struct f2fs_cached_block *entry, blk_opf_t op_flags)
return LOCKED_CACHE;
}
+read_disk:
err = f2fs_get_node_info(sbi, entry->index, &ni, false);
if (err)
return err;
@@ -1520,14 +1530,14 @@ void f2fs_ra_node_cache(struct f2fs_sb_info *sbi, nid_t nid)
f2fs_put_cache(entry, err ? true : false);
}
-int f2fs_sanity_check_node_footer(struct f2fs_sb_info *sbi,
- struct f2fs_cached_block *entry, pgoff_t nid,
- enum node_type ntype, bool in_irq)
+static bool f2fs_node_footer_valid(struct f2fs_sb_info *sbi,
+ struct f2fs_cached_block *entry, pgoff_t nid,
+ enum node_type ntype)
{
bool is_inode, is_xnode;
if (unlikely(nid != nid_of_node(sbi, entry)))
- goto out_err;
+ return false;
is_inode = IS_INODE(sbi, entry);
is_xnode = f2fs_has_xattr_block(ofs_of_node(sbi, entry));
@@ -1535,27 +1545,46 @@ int f2fs_sanity_check_node_footer(struct f2fs_sb_info *sbi,
switch (ntype) {
case NODE_TYPE_REGULAR:
if (is_inode && is_xnode)
- goto out_err;
+ return false;
break;
case NODE_TYPE_INODE:
if (!is_inode || is_xnode)
- goto out_err;
+ return false;
break;
case NODE_TYPE_XATTR:
if (is_inode || !is_xnode)
- goto out_err;
+ return false;
break;
case NODE_TYPE_NON_INODE:
if (is_inode)
- goto out_err;
+ return false;
break;
case NODE_TYPE_NON_IXNODE:
if (is_inode || is_xnode)
- goto out_err;
+ return false;
break;
default:
break;
}
+ return true;
+}
+
+static bool f2fs_restored_footer_valid(struct f2fs_sb_info *sbi,
+ struct f2fs_cached_block *entry,
+ pgoff_t nid, enum node_type ntype)
+{
+ if (f2fs_node_footer_valid(sbi, entry, nid, ntype))
+ return true;
+ f2fs_nc_validation_failed(entry);
+ return false;
+}
+
+int f2fs_sanity_check_node_footer(struct f2fs_sb_info *sbi,
+ struct f2fs_cached_block *entry, pgoff_t nid,
+ enum node_type ntype, bool in_irq)
+{
+ if (!f2fs_node_footer_valid(sbi, entry, nid, ntype))
+ goto out_err;
if (time_to_inject(sbi, FAULT_INCONSISTENT_FOOTER))
goto out_err;
return 0;
@@ -1588,6 +1617,7 @@ static struct f2fs_cached_block *__get_node_cache(struct f2fs_sb_info *sbi, pgof
if (IS_ERR(entry))
return entry;
+read_cache:
err = read_node_cache(entry, 0);
if (err < 0)
goto out_put_err;
@@ -1614,9 +1644,14 @@ static struct f2fs_cached_block *__get_node_cache(struct f2fs_sb_info *sbi, pgof
goto out_err;
}
entry_hit:
+ if (f2fs_cache_test_restored(entry) &&
+ !f2fs_restored_footer_valid(sbi, entry, nid, ntype))
+ goto read_cache;
err = f2fs_sanity_check_node_footer(sbi, entry, nid, ntype, false);
- if (!err)
+ if (!err) {
+ f2fs_nc_validation_succeeded(entry);
return entry;
+ }
out_err:
f2fs_drop_cache_dirty(entry);
out_put_err:
diff --git a/fs/f2fs/node_cache_compress.c b/fs/f2fs/node_cache_compress.c
index 50e9c820eb81..9c2d19e5e821 100644
--- a/fs/f2fs/node_cache_compress.c
+++ b/fs/f2fs/node_cache_compress.c
@@ -1,6 +1,7 @@
// SPDX-License-Identifier: GPL-2.0
#include <linux/atomic.h>
#include <linux/f2fs_fs.h>
+#include <linux/lz4.h>
#include <linux/refcount.h>
#include <linux/slab.h>
@@ -99,6 +100,10 @@ struct f2fs_nc_ctx {
atomic64_t shrink_scanned[F2FS_NC_NR_QUEUES];
atomic64_t shrink_freed[F2FS_NC_NR_QUEUES];
atomic64_t detached[F2FS_NC_NR_QUEUES];
+ /* Cumulative restore statistics since mount. */
+ atomic64_t compressed_to_raw;
+ atomic64_t restored_hits;
+ atomic64_t restore_failures;
/* Signed quota history carried between shrinker calls. */
s64 shrink_credit[F2FS_NC_NR_QUEUES];
/* Mount reference plus references held by compressed objects. */
@@ -280,6 +285,130 @@ void f2fs_nc_free_data(struct f2fs_cached_block *entry)
f2fs_nc_ctx_put(ctx);
}
+static int
+f2fs_nc_publish_raw(struct f2fs_cached_block *entry,
+ struct f2fs_node_cached_block *node,
+ struct f2fs_nc_ctx *ctx, void *raw, bool restored)
+{
+ struct f2fs_cached_block_list *cache = entry->cache;
+ void *object = entry->data;
+ u32 len = node->compressed_len;
+ u32 alloc_size = node->compressed_alloc_size;
+ unsigned int queue = f2fs_nc_entry_queue(entry);
+
+ spin_lock(&cache->list_lock);
+ entry->data = raw;
+ f2fs_cache_clear_compressed(entry);
+ node->owner = NULL;
+ node->compressed_len = 0;
+ node->compressed_alloc_size = 0;
+ node->compressed_crc = 0;
+ if (restored) {
+ f2fs_cache_set_restored(entry);
+ f2fs_cache_set_uptodate(entry);
+ } else {
+ f2fs_cache_clear_restored(entry);
+ f2fs_cache_clear_uptodate(entry);
+ }
+ list_move_tail(&entry->list, &cache->lru_list);
+ f2fs_nc_account_del(ctx, queue, len, alloc_size);
+ f2fs_nc_account_add(ctx, F2FS_NC_RAW, 0, 0);
+ atomic64_inc(&ctx->compressed_to_raw);
+ spin_unlock(&cache->list_lock);
+
+ f2fs_nc_store_free(ctx->store, object, alloc_size);
+ f2fs_nc_ctx_put(ctx);
+ return 0;
+}
+
+static int
+f2fs_nc_restore_fallback(struct f2fs_cached_block *entry,
+ struct f2fs_node_cached_block *node,
+ struct f2fs_nc_ctx *ctx)
+{
+ void *raw;
+
+ raw = f2fs_kmalloc(ctx->sbi, ctx->sbi->blocksize, GFP_NOFS);
+ if (!raw)
+ return -ENOMEM;
+ atomic64_inc(&ctx->restore_failures);
+ return f2fs_nc_publish_raw(entry, node, ctx, raw, false);
+}
+
+int f2fs_nc_restore(struct f2fs_cached_block *entry)
+{
+ struct f2fs_node_cached_block *node;
+ struct f2fs_nc_ctx *ctx;
+ struct f2fs_sb_info *sbi;
+ void *raw;
+ int ret;
+
+ if (!entry || !entry->cache || !IS_NODE_CACHE(entry->cache) ||
+ !f2fs_cache_test_compressed(entry))
+ return 0;
+ if (WARN_ON_ONCE(!f2fs_cache_test_locked(entry)))
+ return -EFSCORRUPTED;
+ if (WARN_ON_ONCE(f2fs_cache_test_dirty(entry) ||
+ f2fs_cache_test_writeback(entry)))
+ return -EFSCORRUPTED;
+
+ node = f2fs_nc_node_entry(entry);
+ ctx = node->owner;
+ if (!ctx || ctx != entry->cache->sbi->node_compress || !entry->data)
+ return -EFSCORRUPTED;
+ sbi = ctx->sbi;
+ if (f2fs_nc_store_bucket_from_size(node->compressed_alloc_size) < 0)
+ return -EFSCORRUPTED;
+ if (!node->compressed_len ||
+ node->compressed_len > node->compressed_alloc_size)
+ return f2fs_nc_restore_fallback(entry, node, ctx);
+
+ raw = f2fs_kmalloc(sbi, sbi->blocksize, GFP_NOFS);
+ if (!raw)
+ return -ENOMEM;
+ ret = LZ4_decompress_safe(entry->data, raw, node->compressed_len,
+ sbi->blocksize);
+ if (ret != sbi->blocksize ||
+ f2fs_crc32(raw, sbi->blocksize) != node->compressed_crc) {
+ atomic64_inc(&ctx->restore_failures);
+ return f2fs_nc_publish_raw(entry, node, ctx, raw, false);
+ }
+ return f2fs_nc_publish_raw(entry, node, ctx, raw, true);
+}
+
+void f2fs_nc_validation_failed(struct f2fs_cached_block *entry)
+{
+ struct f2fs_nc_ctx *ctx;
+
+ if (f2fs_cache_test_restored(entry) && entry->cache) {
+ ctx = entry->cache->sbi->node_compress;
+ if (ctx)
+ atomic64_inc(&ctx->restore_failures);
+ }
+ f2fs_cache_clear_restored(entry);
+ f2fs_cache_clear_uptodate(entry);
+}
+
+void f2fs_nc_validation_succeeded(struct f2fs_cached_block *entry)
+{
+ struct f2fs_nc_ctx *ctx;
+
+ if (f2fs_cache_test_restored(entry) && entry->cache) {
+ ctx = entry->cache->sbi->node_compress;
+ if (ctx)
+ atomic64_inc(&ctx->restored_hits);
+ }
+ f2fs_cache_clear_restored(entry);
+}
+
+void f2fs_nc_content_changed(struct f2fs_cached_block *entry)
+{
+ if (!entry || !entry->cache || !IS_NODE_CACHE(entry->cache))
+ return;
+ WARN_ON_ONCE(f2fs_cache_test_compressed(entry));
+ f2fs_cache_clear_restored(entry);
+}
+
static void f2fs_nc_population_snapshot(struct f2fs_nc_ctx *ctx,
unsigned long nr[F2FS_NC_NR_QUEUES])
{
diff --git a/fs/f2fs/node_cache_compress.h b/fs/f2fs/node_cache_compress.h
index e7d6b07c1c98..faeb342b9cc3 100644
--- a/fs/f2fs/node_cache_compress.h
+++ b/fs/f2fs/node_cache_compress.h
@@ -29,6 +29,10 @@ size_t f2fs_nc_entry_alloc_size(struct f2fs_cached_block_list *cache);
void f2fs_nc_init(struct f2fs_sb_info *sbi);
void f2fs_nc_destroy(struct f2fs_sb_info *sbi);
void f2fs_nc_free_data(struct f2fs_cached_block *entry);
+int f2fs_nc_restore(struct f2fs_cached_block *entry);
+void f2fs_nc_validation_failed(struct f2fs_cached_block *entry);
+void f2fs_nc_validation_succeeded(struct f2fs_cached_block *entry);
+void f2fs_nc_content_changed(struct f2fs_cached_block *entry);
struct list_head *f2fs_nc_queue_head(struct f2fs_cached_block_list *cache,
unsigned int queue);
unsigned int f2fs_nc_entry_queue(const struct f2fs_cached_block *entry);
@@ -54,6 +58,15 @@ static inline void f2fs_nc_free_data(struct f2fs_cached_block *entry)
kfree(entry->data);
}
+static inline int f2fs_nc_restore(struct f2fs_cached_block *entry)
+{
+ return 0;
+}
+
+static inline void f2fs_nc_validation_failed(struct f2fs_cached_block *entry) { }
+static inline void f2fs_nc_validation_succeeded(struct f2fs_cached_block *entry) { }
+static inline void f2fs_nc_content_changed(struct f2fs_cached_block *entry) { }
+
static inline struct list_head *f2fs_nc_queue_head(struct f2fs_cached_block_list *cache,
unsigned int queue)
{
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH 5/6] f2fs: compress clean node cache entries in background
2026-09-29 7:29 [RFC PATCH 0/6] f2fs: retain clean node blocks in a compressed cache Wenjie Qi
` (3 preceding siblings ...)
2026-09-29 7:29 ` [RFC PATCH 4/6] f2fs: restore compressed node cache entries before access Wenjie Qi
@ 2026-09-29 7:29 ` Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 6/6] f2fs: give referenced node cache entries a shared second chance Wenjie Qi
5 siblings, 0 replies; 7+ messages in thread
From: Wenjie Qi @ 2026-09-29 7:29 UTC (permalink / raw)
To: Jaegeuk Kim, Chao Yu
Cc: Barry Song, linux-f2fs-devel, linux-kernel, Wenjie Qi
With the compressed representation, restore path and reclaim policy in
place, add background conversion of cold clean node-cache entries.
Reuse the per-superblock cache thread. Writeback and compression keep
separate due times but execute serially. Each compression pass scans at
most 1024 raw entries, retains at most 256 candidates and reschedules after
every 32 candidates.
Compress outside the cache list and radix-tree locks, then revalidate the
entry before replacing its raw block. Expose only the compression threshold
and worker interval through sysfs.
Signed-off-by: Wenjie Qi <qiwenjie@xiaomi.com>
---
Documentation/ABI/testing/sysfs-fs-f2fs | 22 ++
fs/f2fs/cache.c | 61 +++-
fs/f2fs/cache.h | 15 +
fs/f2fs/debug.c | 37 +++
fs/f2fs/node_cache_compress.c | 407 +++++++++++++++++++++++-
fs/f2fs/node_cache_compress.h | 37 +++
fs/f2fs/node_cache_policy.c | 48 +++
fs/f2fs/node_cache_policy.h | 24 +-
fs/f2fs/sysfs.c | 59 ++++
9 files changed, 691 insertions(+), 19 deletions(-)
diff --git a/Documentation/ABI/testing/sysfs-fs-f2fs b/Documentation/ABI/testing/sysfs-fs-f2fs
index c4746c416ac2..485d4231855f 100644
--- a/Documentation/ABI/testing/sysfs-fs-f2fs
+++ b/Documentation/ABI/testing/sysfs-fs-f2fs
@@ -1027,3 +1027,25 @@ Contact: "Chao Yu" <chao@kernel.org>
Description: This is a writable entry to control writeback interval of
f2fs_writeback-x:y, the range is [100, 30000], by default the value
is 5000, unit is ms.
+
+What: /sys/fs/f2fs/<disk>/node_compress_interval
+Date: September 2026
+Contact: "Wenjie Qi" <qiwenjie@xiaomi.com>
+Description: This is a writable entry to control how often the background
+ node-cache compression worker runs. The range is [100, 30000],
+ default value is 1000, and the unit is ms. Updating the value
+ reschedules the next worker deadline but does not synchronously
+ cancel a cycle that is already running. This entry is present only
+ when CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION is enabled.
+
+What: /sys/fs/f2fs/<disk>/node_compress_threshold
+Date: September 2026
+Contact: "Wenjie Qi" <qiwenjie@xiaomi.com>
+Description: This is a writable entry to set the maximum compressed payload
+ length as a percentage of the filesystem block size. The range is
+ [0, 100], default value is 0, which disables background node-cache
+ compression. The effective payload limit is also capped at 1024 bytes.
+ Updating the value affects subsequent worker scheduling and cycles;
+ it does not synchronously cancel a cycle that is already running.
+ This entry is present only when
+ CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION is enabled.
diff --git a/fs/f2fs/cache.c b/fs/f2fs/cache.c
index 9cb541dd5cea..b3158f4ab45c 100644
--- a/fs/f2fs/cache.c
+++ b/fs/f2fs/cache.c
@@ -686,6 +686,19 @@ unsigned long f2fs_shrink_cache(struct f2fs_sb_info *sbi,
return freed;
}
+static void f2fs_try_write_caches(struct f2fs_sb_info *sbi)
+{
+ if (f2fs_readonly(sbi->sb) || f2fs_cp_error(sbi) ||
+ unlikely(freezing(current)))
+ return;
+ if (!sb_start_write_trylock(sbi->sb))
+ return;
+
+ f2fs_write_meta_caches(sbi);
+ f2fs_write_node_caches(sbi);
+ sb_end_write(sbi->sb);
+}
+
static int f2fs_cache_writeback_kthread(void *data)
{
struct f2fs_sb_info *sbi = data;
@@ -694,6 +707,32 @@ static int f2fs_cache_writeback_kthread(void *data)
set_freezable();
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+ if (sbi->node_compress) {
+ while (!kthread_should_stop()) {
+ u64 now = ktime_to_ms(ktime_get());
+ u64 due = min(cache_thread->next_wb_ms,
+ f2fs_nc_next_deadline(sbi));
+
+ wait_event_freezable_timeout(*wq,
+ kthread_should_stop(),
+ msecs_to_jiffies(due > now ? due - now : 0));
+ if (kthread_should_stop())
+ break;
+ now = ktime_to_ms(ktime_get());
+ if (now >= cache_thread->next_wb_ms) {
+ f2fs_try_write_caches(sbi);
+ WRITE_ONCE(cache_thread->next_wb_ms,
+ ktime_to_ms(ktime_get()) +
+ cache_thread->cache_wb_interval);
+ }
+ if (!freezing(current) &&
+ ktime_to_ms(ktime_get()) >= f2fs_nc_next_deadline(sbi))
+ f2fs_nc_run(sbi);
+ }
+ return 0;
+ }
+#endif
while (!kthread_should_stop()) {
unsigned int interval = cache_thread->cache_wb_interval;
@@ -704,22 +743,7 @@ static int f2fs_cache_writeback_kthread(void *data)
if (kthread_should_stop())
break;
- if (f2fs_readonly(sbi->sb))
- continue;
-
- if (f2fs_cp_error(sbi))
- continue;
-
- if (unlikely(freezing(current)))
- continue;
-
- if (!sb_start_write_trylock(sbi->sb))
- continue;
-
- f2fs_write_meta_caches(sbi);
- f2fs_write_node_caches(sbi);
-
- sb_end_write(sbi->sb);
+ f2fs_try_write_caches(sbi);
}
return 0;
}
@@ -736,6 +760,11 @@ int f2fs_start_cache_wb_thread(struct f2fs_sb_info *sbi)
init_waitqueue_head(&cache_thread->cache_wb_wq);
cache_thread->cache_wb_interval = DEF_DIRTY_CACHE_TIMEOUT;
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+ cache_thread->next_wb_ms = ktime_to_ms(ktime_get()) +
+ cache_thread->cache_wb_interval;
+ f2fs_nc_schedule_start(sbi, ktime_to_ms(ktime_get()));
+#endif
snprintf(name, sizeof(name), "f2fs_writeback-%u:%u",
MAJOR(dev), MINOR(dev));
diff --git a/fs/f2fs/cache.h b/fs/f2fs/cache.h
index 603c5f7203b7..3c1fdc7894b7 100644
--- a/fs/f2fs/cache.h
+++ b/fs/f2fs/cache.h
@@ -67,6 +67,7 @@ enum f2fs_cached_state {
#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
F2FS_BLOCK_COMPRESSED,
F2FS_BLOCK_RESTORED,
+ F2FS_BLOCK_INCOMPRESSIBLE,
#endif
};
@@ -155,6 +156,9 @@ F2FS_CACHE_FLAG_CLEAR_FUNC(compressed, COMPRESSED);
F2FS_CACHE_FLAG_TEST_FUNC(restored, RESTORED);
F2FS_CACHE_FLAG_SET_FUNC(restored, RESTORED);
F2FS_CACHE_FLAG_CLEAR_FUNC(restored, RESTORED);
+F2FS_CACHE_FLAG_TEST_FUNC(incompressible, INCOMPRESSIBLE);
+F2FS_CACHE_FLAG_SET_FUNC(incompressible, INCOMPRESSIBLE);
+F2FS_CACHE_FLAG_CLEAR_FUNC(incompressible, INCOMPRESSIBLE);
#else
static inline bool f2fs_cache_test_compressed(const struct f2fs_cached_block *entry)
{
@@ -171,6 +175,14 @@ static inline bool f2fs_cache_test_restored(const struct f2fs_cached_block *entr
static inline void f2fs_cache_set_restored(struct f2fs_cached_block *entry) { }
static inline void f2fs_cache_clear_restored(struct f2fs_cached_block *entry) { }
+static inline bool
+f2fs_cache_test_incompressible(const struct f2fs_cached_block *entry)
+{
+ return false;
+}
+
+static inline void f2fs_cache_set_incompressible(struct f2fs_cached_block *entry) { }
+static inline void f2fs_cache_clear_incompressible(struct f2fs_cached_block *entry) { }
#endif
static inline void *cache_address(const struct f2fs_cached_block *entry)
@@ -264,6 +276,9 @@ struct f2fs_cache_kthread {
struct task_struct *cache_wb_task;
wait_queue_head_t cache_wb_wq;
unsigned int cache_wb_interval;
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+ u64 next_wb_ms;
+#endif
};
int f2fs_start_cache_wb_thread(struct f2fs_sb_info *sbi);
diff --git a/fs/f2fs/debug.c b/fs/f2fs/debug.c
index a6096537b495..984fed727a48 100644
--- a/fs/f2fs/debug.c
+++ b/fs/f2fs/debug.c
@@ -713,6 +713,43 @@ static int stat_show(struct seq_file *s, void *v)
seq_printf(s, " - compress: %4d, hit:%8d\n", si->compress_pages, si->compress_page_hit);
seq_printf(s, " - nodes: %4d in %4d\n",
si->ndirty_node, si->node_caches);
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+ {
+ struct f2fs_nc_stats nc;
+
+ f2fs_nc_get_stats(sbi, &nc);
+ seq_printf(s,
+ "NodeCacheCompress: initialized=%u threshold=%u interval=%u\n",
+ nc.initialized, nc.threshold_pct, nc.interval_ms);
+ seq_printf(s,
+ "NCQueues: raw=%llu 256=%llu 512=%llu 1024=%llu payload=%llu/%llu/%llu slots=%llu/%llu/%llu\n",
+ nc.attached[F2FS_NC_RAW], nc.attached[F2FS_NC_256],
+ nc.attached[F2FS_NC_512], nc.attached[F2FS_NC_1024],
+ nc.payload[0], nc.payload[1], nc.payload[2],
+ nc.slots[0], nc.slots[1], nc.slots[2]);
+ seq_printf(s,
+ "NCWorker: cycles=%llu visited=%llu candidates=%llu attempts=%llu converted=%llu candidate_refs=%llu transient=%llu alloc_fail=%llu\n",
+ nc.cycles, nc.visited, nc.candidates, nc.attempts,
+ nc.converted, nc.candidate_refs,
+ nc.transient_extra_bytes, nc.allocation_failures);
+ seq_printf(s,
+ "NCRestore: compressed_to_raw=%llu restored_hits=%llu restore_failures=%llu\n",
+ nc.compressed_to_raw, nc.restored_hits,
+ nc.restore_failures);
+ seq_printf(s,
+ "NCReclaim: scanned=%llu/%llu/%llu/%llu freed=%llu/%llu/%llu/%llu detached=%llu/%llu/%llu/%llu\n",
+ nc.shrink_scanned[F2FS_NC_RAW],
+ nc.shrink_scanned[F2FS_NC_256],
+ nc.shrink_scanned[F2FS_NC_512],
+ nc.shrink_scanned[F2FS_NC_1024],
+ nc.shrink_freed[F2FS_NC_RAW],
+ nc.shrink_freed[F2FS_NC_256],
+ nc.shrink_freed[F2FS_NC_512],
+ nc.shrink_freed[F2FS_NC_1024],
+ nc.detached[F2FS_NC_RAW], nc.detached[F2FS_NC_256],
+ nc.detached[F2FS_NC_512], nc.detached[F2FS_NC_1024]);
+ }
+#endif
seq_printf(s, " - dents: %4d in dirs:%4d (%4d)\n",
si->ndirty_dent, si->ndirty_dirs, si->ndirty_all);
seq_printf(s, " - data: %4d in files:%4d\n",
diff --git a/fs/f2fs/node_cache_compress.c b/fs/f2fs/node_cache_compress.c
index 9c2d19e5e821..6e474fea7241 100644
--- a/fs/f2fs/node_cache_compress.c
+++ b/fs/f2fs/node_cache_compress.c
@@ -1,13 +1,18 @@
// SPDX-License-Identifier: GPL-2.0
#include <linux/atomic.h>
#include <linux/f2fs_fs.h>
+#include <linux/kthread.h>
#include <linux/lz4.h>
#include <linux/refcount.h>
+#include <linux/sched.h>
#include <linux/slab.h>
#include "f2fs.h"
#include "node_cache_compress.h"
+#define F2FS_NC_ALLOC_BACKOFF_MS 1000U
+#define F2FS_NC_PROCESS_BATCH 32U
+
static const u32 f2fs_nc_bucket_sizes[] = {
F2FS_NC_BUCKET_256_SIZE,
F2FS_NC_BUCKET_512_SIZE,
@@ -20,6 +25,18 @@ struct f2fs_nc_store {
atomic_long_t objects[ARRAY_SIZE(f2fs_nc_bucket_sizes)];
};
+static int f2fs_nc_store_bucket(u32 len)
+{
+ int i;
+
+ if (!len)
+ return -EINVAL;
+ for (i = 0; i < ARRAY_SIZE(f2fs_nc_bucket_sizes); i++)
+ if (len <= f2fs_nc_bucket_sizes[i])
+ return i;
+ return -E2BIG;
+}
+
static int f2fs_nc_store_bucket_from_size(u32 alloc_size)
{
int i;
@@ -72,6 +89,22 @@ static void f2fs_nc_store_destroy(struct f2fs_nc_store *store)
kfree(store);
}
+static void *f2fs_nc_store_alloc(struct f2fs_nc_store *store, u32 len,
+ gfp_t gfp, u32 *alloc_size)
+{
+ void *object;
+ int bucket = f2fs_nc_store_bucket(len);
+
+ if (!store || !alloc_size || bucket < 0)
+ return NULL;
+ object = kmem_cache_alloc(store->caches[bucket], gfp);
+ if (!object)
+ return NULL;
+ atomic_long_inc(&store->objects[bucket]);
+ *alloc_size = f2fs_nc_bucket_sizes[bucket];
+ return object;
+}
+
static void f2fs_nc_store_free(struct f2fs_nc_store *store, void *object,
u32 alloc_size)
{
@@ -90,20 +123,35 @@ static void f2fs_nc_store_free(struct f2fs_nc_store *store, void *object,
struct f2fs_nc_ctx {
struct f2fs_sb_info *sbi;
struct f2fs_nc_store *store;
+ /* Per-mount workspace used only by the cache thread. */
+ void *workmem;
+ void *scratch;
+ struct f2fs_cached_block **candidates;
+ spinlock_t config_lock; /* protect config and next deadline */
+ struct f2fs_nc_config config;
/* Compressed queues only; raw entries use NODE_CACHE()->lru_list. */
struct list_head queues[F2FS_NC_NR_QUEUES - 1];
/* Current population and compressed bytes by queue. */
atomic_long_t attached[F2FS_NC_NR_QUEUES];
atomic64_t attached_payload[F2FS_NC_NR_QUEUES - 1];
atomic64_t attached_slot_bytes[F2FS_NC_NR_QUEUES - 1];
- /* Cumulative reclaim and detach statistics since mount. */
+ /* Cumulative reclaim, detach, restore and worker statistics. */
atomic64_t shrink_scanned[F2FS_NC_NR_QUEUES];
atomic64_t shrink_freed[F2FS_NC_NR_QUEUES];
atomic64_t detached[F2FS_NC_NR_QUEUES];
- /* Cumulative restore statistics since mount. */
atomic64_t compressed_to_raw;
atomic64_t restored_hits;
atomic64_t restore_failures;
+ atomic64_t cycles;
+ atomic64_t visited;
+ atomic64_t candidate_count;
+ atomic64_t attempts;
+ atomic64_t converted;
+ atomic64_t candidate_refs;
+ atomic64_t transient_extra_bytes;
+ atomic64_t allocation_failures;
+ /* Absolute CLOCK_MONOTONIC deadline; U64_MAX disables scheduling. */
+ u64 next_compress_ms;
/* Signed quota history carried between shrinker calls. */
s64 shrink_credit[F2FS_NC_NR_QUEUES];
/* Mount reference plus references held by compressed objects. */
@@ -121,6 +169,11 @@ static void f2fs_nc_ctx_release(struct f2fs_nc_ctx *ctx)
for (i = 0; i < F2FS_NC_NR_QUEUES; i++)
WARN_ON_ONCE(atomic_long_read(&ctx->attached[i]));
+ WARN_ON_ONCE(atomic64_read(&ctx->candidate_refs));
+ WARN_ON_ONCE(atomic64_read(&ctx->transient_extra_bytes));
+ kfree(ctx->candidates);
+ kfree(ctx->scratch);
+ kfree(ctx->workmem);
f2fs_nc_store_destroy(ctx->store);
kfree(ctx);
}
@@ -131,6 +184,11 @@ static void f2fs_nc_ctx_put(struct f2fs_nc_ctx *ctx)
f2fs_nc_ctx_release(ctx);
}
+static void f2fs_nc_ctx_get(struct f2fs_nc_ctx *ctx)
+{
+ refcount_inc(&ctx->refs);
+}
+
static void f2fs_nc_account_add(struct f2fs_nc_ctx *ctx, unsigned int queue,
u32 len, u32 alloc_size)
{
@@ -174,6 +232,18 @@ void f2fs_nc_init(struct f2fs_sb_info *sbi)
ctx->store = f2fs_nc_store_create(sbi);
if (!ctx->store)
goto fail_open;
+ ctx->workmem = kmalloc(LZ4_MEM_COMPRESS, GFP_NOFS);
+ if (!ctx->workmem)
+ goto fail_open;
+ ctx->scratch = kmalloc(F2FS_NC_MAX_OBJECT_SIZE, GFP_NOFS);
+ if (!ctx->scratch)
+ goto fail_open;
+ ctx->candidates = kcalloc(F2FS_NC_MAX_CANDIDATES,
+ sizeof(*ctx->candidates), GFP_NOFS);
+ if (!ctx->candidates)
+ goto fail_open;
+ spin_lock_init(&ctx->config_lock);
+ f2fs_nc_config_defaults(&ctx->config);
for (i = 0; i < ARRAY_SIZE(ctx->queues); i++)
INIT_LIST_HEAD(&ctx->queues[i]);
refcount_set(&ctx->refs, 1);
@@ -197,6 +267,91 @@ void f2fs_nc_destroy(struct f2fs_sb_info *sbi)
f2fs_nc_ctx_put(ctx);
}
+static void f2fs_nc_config_snapshot(struct f2fs_nc_ctx *ctx,
+ struct f2fs_nc_config *cfg)
+{
+ unsigned long flags;
+
+ spin_lock_irqsave(&ctx->config_lock, flags);
+ *cfg = ctx->config;
+ spin_unlock_irqrestore(&ctx->config_lock, flags);
+}
+
+static u64 f2fs_nc_compress_deadline(const struct f2fs_nc_config *cfg,
+ u64 now_ms)
+{
+ if (!cfg->compression_threshold_pct)
+ return U64_MAX;
+ return now_ms + cfg->compression_interval_ms;
+}
+
+u64 f2fs_nc_config_value(struct f2fs_sb_info *sbi, enum f2fs_nc_param id)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ struct f2fs_nc_config cfg;
+
+ if (!ctx) {
+ f2fs_nc_config_defaults(&cfg);
+ return f2fs_nc_config_get(&cfg, id);
+ }
+ f2fs_nc_config_snapshot(ctx, &cfg);
+ return f2fs_nc_config_get(&cfg, id);
+}
+
+int f2fs_nc_config_update(struct f2fs_sb_info *sbi,
+ enum f2fs_nc_param id, u64 value)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ struct f2fs_nc_config old, new;
+ unsigned long flags;
+ int ret;
+
+ if (!ctx) {
+ f2fs_nc_config_defaults(&old);
+ ret = f2fs_nc_config_set(&new, &old, id, value);
+ if (ret)
+ return ret;
+ return f2fs_nc_config_get(&old, id) == value ? 0 : -EOPNOTSUPP;
+ }
+
+ spin_lock_irqsave(&ctx->config_lock, flags);
+ old = ctx->config;
+ ret = f2fs_nc_config_set(&new, &old, id, value);
+ if (ret)
+ goto out_unlock;
+ ctx->config = new;
+ ctx->next_compress_ms =
+ f2fs_nc_compress_deadline(&new, ktime_to_ms(ktime_get()));
+out_unlock:
+ spin_unlock_irqrestore(&ctx->config_lock, flags);
+ return ret;
+}
+
+u64 f2fs_nc_next_deadline(struct f2fs_sb_info *sbi)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ unsigned long flags;
+ u64 deadline;
+
+ if (!ctx)
+ return U64_MAX;
+ spin_lock_irqsave(&ctx->config_lock, flags);
+ deadline = ctx->next_compress_ms;
+ spin_unlock_irqrestore(&ctx->config_lock, flags);
+ return deadline;
+}
+
+void f2fs_nc_schedule_start(struct f2fs_sb_info *sbi, u64 now_ms)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ unsigned long flags;
+
+ if (!ctx)
+ return;
+ spin_lock_irqsave(&ctx->config_lock, flags);
+ ctx->next_compress_ms = f2fs_nc_compress_deadline(&ctx->config, now_ms);
+ spin_unlock_irqrestore(&ctx->config_lock, flags);
+}
struct list_head *f2fs_nc_queue_head(struct f2fs_cached_block_list *cache,
unsigned int queue)
{
@@ -407,6 +562,214 @@ void f2fs_nc_content_changed(struct f2fs_cached_block *entry)
return;
WARN_ON_ONCE(f2fs_cache_test_compressed(entry));
f2fs_cache_clear_restored(entry);
+ f2fs_cache_clear_incompressible(entry);
+}
+
+static bool f2fs_nc_candidate(struct f2fs_cached_block *entry,
+ struct f2fs_cached_block_list *cache, int refs)
+{
+ return entry->cache == cache && entry->data &&
+ !f2fs_cache_test_compressed(entry) &&
+ f2fs_cache_test_uptodate(entry) &&
+ !f2fs_cache_test_dirty(entry) &&
+ !f2fs_cache_test_writeback(entry) &&
+ !f2fs_cache_test_locked(entry) &&
+ !f2fs_cache_test_referenced(entry) &&
+ !f2fs_cache_test_incompressible(entry) &&
+ atomic_read(&entry->refcount) == refs;
+}
+
+enum f2fs_nc_compress_result {
+ F2FS_NC_COMPRESS_SKIPPED,
+ F2FS_NC_COMPRESS_CONVERTED,
+ F2FS_NC_COMPRESS_NO_MEMORY,
+};
+
+static enum f2fs_nc_compress_result
+f2fs_nc_compress(struct f2fs_cached_block *entry,
+ const struct f2fs_nc_config *cfg)
+{
+ struct f2fs_cached_block_list *cache = entry->cache;
+ struct f2fs_node_cached_block *node;
+ struct f2fs_nc_ctx *ctx;
+ struct f2fs_sb_info *sbi;
+ unsigned long flags;
+ unsigned int max_len;
+ void *object;
+ void *raw;
+ u32 alloc_size;
+ u32 crc;
+ int bucket;
+ int len;
+
+ if (!cache || !IS_NODE_CACHE(cache) || !cfg ||
+ f2fs_cache_test_compressed(entry))
+ return F2FS_NC_COMPRESS_SKIPPED;
+ if (WARN_ON_ONCE(!f2fs_cache_test_locked(entry)))
+ return F2FS_NC_COMPRESS_SKIPPED;
+ if (!f2fs_cache_test_uptodate(entry) || f2fs_cache_test_dirty(entry) ||
+ f2fs_cache_test_writeback(entry) ||
+ f2fs_cache_test_referenced(entry) ||
+ f2fs_cache_test_incompressible(entry) || !entry->data ||
+ atomic_read(&entry->refcount) != 2)
+ return F2FS_NC_COMPRESS_SKIPPED;
+
+ sbi = cache->sbi;
+ ctx = sbi->node_compress;
+ if (!ctx || !cfg->compression_threshold_pct)
+ return F2FS_NC_COMPRESS_SKIPPED;
+ max_len = min_t(u64,
+ (u64)sbi->blocksize * cfg->compression_threshold_pct /
+ F2FS_NC_PERCENT_MAX,
+ F2FS_NC_MAX_OBJECT_SIZE);
+ if (!max_len)
+ return F2FS_NC_COMPRESS_SKIPPED;
+
+ atomic64_inc(&ctx->attempts);
+ len = LZ4_compress_default(entry->data, ctx->scratch, sbi->blocksize,
+ max_len, ctx->workmem);
+ if (len <= 0) {
+ if (max_len == F2FS_NC_MAX_OBJECT_SIZE)
+ f2fs_cache_set_incompressible(entry);
+ return F2FS_NC_COMPRESS_SKIPPED;
+ }
+ object = f2fs_nc_store_alloc(ctx->store, len, GFP_NOFS, &alloc_size);
+ if (!object) {
+ atomic64_inc(&ctx->allocation_failures);
+ return F2FS_NC_COMPRESS_NO_MEMORY;
+ }
+ bucket = f2fs_nc_store_bucket_from_size(alloc_size);
+ if (WARN_ON_ONCE(bucket < 0)) {
+ f2fs_nc_store_free(ctx->store, object, alloc_size);
+ return F2FS_NC_COMPRESS_SKIPPED;
+ }
+ /*
+ * Before publication only the new slot is extra. After publication,
+ * the old raw block remains the extra allocation until it is freed.
+ */
+ atomic64_add(alloc_size, &ctx->transient_extra_bytes);
+ memcpy(object, ctx->scratch, len);
+ raw = entry->data;
+ crc = f2fs_crc32(raw, sbi->blocksize);
+
+ spin_lock(&cache->list_lock);
+ spin_lock_irqsave(&cache->tree_lock, flags);
+ if (entry->cache != cache || entry->data != raw ||
+ atomic_read(&entry->refcount) != 2 || f2fs_cp_error(sbi) ||
+ f2fs_cache_test_referenced(entry) ||
+ !f2fs_cache_test_uptodate(entry) || f2fs_cache_test_dirty(entry) ||
+ f2fs_cache_test_writeback(entry) ||
+ f2fs_cache_test_incompressible(entry)) {
+ spin_unlock_irqrestore(&cache->tree_lock, flags);
+ spin_unlock(&cache->list_lock);
+ f2fs_nc_store_free(ctx->store, object, alloc_size);
+ atomic64_sub(alloc_size, &ctx->transient_extra_bytes);
+ return F2FS_NC_COMPRESS_SKIPPED;
+ }
+ node = f2fs_nc_node_entry(entry);
+ f2fs_nc_ctx_get(ctx);
+ node->owner = ctx;
+ node->compressed_len = len;
+ node->compressed_alloc_size = alloc_size;
+ node->compressed_crc = crc;
+ entry->data = object;
+ f2fs_cache_set_compressed(entry);
+ list_move_tail(&entry->list, &ctx->queues[bucket]);
+ f2fs_nc_account_del(ctx, F2FS_NC_RAW, 0, 0);
+ f2fs_nc_account_add(ctx, bucket + 1, len, alloc_size);
+ atomic64_add((s64)sbi->blocksize - alloc_size,
+ &ctx->transient_extra_bytes);
+ spin_unlock_irqrestore(&cache->tree_lock, flags);
+ spin_unlock(&cache->list_lock);
+
+ kfree(raw);
+ atomic64_sub(sbi->blocksize, &ctx->transient_extra_bytes);
+ atomic64_inc(&ctx->converted);
+ return F2FS_NC_COMPRESS_CONVERTED;
+}
+
+void f2fs_nc_run(struct f2fs_sb_info *sbi)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ struct f2fs_cached_block_list *cache = NODE_CACHE(sbi);
+ struct f2fs_nc_config cfg;
+ struct f2fs_cached_block *entry;
+ unsigned long flags;
+ u64 now_ms;
+ unsigned long raw_population;
+ unsigned int goal, visited = 0, count = 0, i;
+ bool allocation_failed = false;
+
+ if (!ctx)
+ return;
+ f2fs_nc_config_snapshot(ctx, &cfg);
+ if (!cfg.compression_threshold_pct || f2fs_cp_error(sbi) ||
+ f2fs_readonly(sbi->sb))
+ goto out;
+ /* Count enabled worker passes, including passes with no candidates. */
+ atomic64_inc(&ctx->cycles);
+ raw_population = atomic_long_read(&ctx->attached[F2FS_NC_RAW]);
+ goal = min_t(unsigned long, raw_population, F2FS_NC_SCAN_MAX);
+ if (!goal)
+ goto out;
+
+ spin_lock(&cache->list_lock);
+ list_for_each_entry(entry, &cache->lru_list, list) {
+ if (visited == goal || count == F2FS_NC_MAX_CANDIDATES)
+ break;
+ visited++;
+ if (!f2fs_nc_candidate(entry, cache, 1))
+ continue;
+ f2fs_cache_get(entry);
+ ctx->candidates[count++] = entry;
+ atomic64_inc(&ctx->candidate_refs);
+ }
+ spin_unlock(&cache->list_lock);
+ atomic64_add(visited, &ctx->visited);
+ atomic64_add(count, &ctx->candidate_count);
+
+ for (i = 0; i < count; i++) {
+ enum f2fs_nc_compress_result compress_result =
+ F2FS_NC_COMPRESS_SKIPPED;
+
+ entry = ctx->candidates[i];
+ if (f2fs_trylock_cache(entry)) {
+ compress_result = f2fs_nc_compress(entry, &cfg);
+ f2fs_unlock_cache(entry);
+ }
+ f2fs_put_cache(entry, false);
+ atomic64_dec(&ctx->candidate_refs);
+ ctx->candidates[i] = NULL;
+ switch (compress_result) {
+ case F2FS_NC_COMPRESS_CONVERTED:
+ break;
+ case F2FS_NC_COMPRESS_NO_MEMORY:
+ allocation_failed = true;
+ i++;
+ goto put_remaining;
+ case F2FS_NC_COMPRESS_SKIPPED:
+ break;
+ }
+ if ((i + 1) % F2FS_NC_PROCESS_BATCH == 0)
+ cond_resched();
+ }
+put_remaining:
+ for (; i < count; i++) {
+ f2fs_put_cache(ctx->candidates[i], false);
+ atomic64_dec(&ctx->candidate_refs);
+ ctx->candidates[i] = NULL;
+ }
+out:
+ now_ms = ktime_to_ms(ktime_get());
+ spin_lock_irqsave(&ctx->config_lock, flags);
+ if (allocation_failed && ctx->config.compression_threshold_pct)
+ WRITE_ONCE(ctx->next_compress_ms, now_ms +
+ max_t(u32, ctx->config.compression_interval_ms,
+ F2FS_NC_ALLOC_BACKOFF_MS));
+ else
+ WRITE_ONCE(ctx->next_compress_ms,
+ f2fs_nc_compress_deadline(&ctx->config, now_ms));
+ spin_unlock_irqrestore(&ctx->config_lock, flags);
}
static void f2fs_nc_population_snapshot(struct f2fs_nc_ctx *ctx,
@@ -492,3 +855,43 @@ void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
memory->entry_bytes = entries * sizeof(struct f2fs_node_cached_block);
memory->data_bytes = data;
}
+
+void f2fs_nc_get_stats(struct f2fs_sb_info *sbi,
+ struct f2fs_nc_stats *stats)
+{
+ struct f2fs_nc_ctx *ctx = sbi->node_compress;
+ struct f2fs_nc_config cfg;
+ unsigned int i;
+
+ memset(stats, 0, sizeof(*stats));
+ if (!ctx)
+ return;
+ f2fs_nc_config_snapshot(ctx, &cfg);
+ stats->initialized = true;
+ stats->threshold_pct = cfg.compression_threshold_pct;
+ stats->interval_ms = cfg.compression_interval_ms;
+ for (i = 0; i < F2FS_NC_NR_QUEUES; i++) {
+ stats->attached[i] = atomic_long_read(&ctx->attached[i]);
+ stats->shrink_scanned[i] = atomic64_read(&ctx->shrink_scanned[i]);
+ stats->shrink_freed[i] = atomic64_read(&ctx->shrink_freed[i]);
+ stats->detached[i] = atomic64_read(&ctx->detached[i]);
+ if (i != F2FS_NC_RAW) {
+ stats->payload[i - 1] =
+ atomic64_read(&ctx->attached_payload[i - 1]);
+ stats->slots[i - 1] =
+ atomic64_read(&ctx->attached_slot_bytes[i - 1]);
+ }
+ }
+ stats->compressed_to_raw = atomic64_read(&ctx->compressed_to_raw);
+ stats->restored_hits = atomic64_read(&ctx->restored_hits);
+ stats->restore_failures = atomic64_read(&ctx->restore_failures);
+ stats->cycles = atomic64_read(&ctx->cycles);
+ stats->visited = atomic64_read(&ctx->visited);
+ stats->candidates = atomic64_read(&ctx->candidate_count);
+ stats->attempts = atomic64_read(&ctx->attempts);
+ stats->converted = atomic64_read(&ctx->converted);
+ stats->candidate_refs = atomic64_read(&ctx->candidate_refs);
+ stats->transient_extra_bytes =
+ atomic64_read(&ctx->transient_extra_bytes);
+ stats->allocation_failures = atomic64_read(&ctx->allocation_failures);
+}
diff --git a/fs/f2fs/node_cache_compress.h b/fs/f2fs/node_cache_compress.h
index faeb342b9cc3..efe2eceef352 100644
--- a/fs/f2fs/node_cache_compress.h
+++ b/fs/f2fs/node_cache_compress.h
@@ -13,6 +13,33 @@ struct f2fs_nc_memory {
u64 data_bytes; /* Raw buffers plus compressed slab slots. */
};
+/* Snapshot used by debugfs; queue arrays use enum f2fs_nc_queue. */
+struct f2fs_nc_stats {
+ /* Current queue state; payload and slots start at F2FS_NC_256. */
+ u64 attached[F2FS_NC_NR_QUEUES];
+ u64 payload[F2FS_NC_NR_QUEUES - 1];
+ u64 slots[F2FS_NC_NR_QUEUES - 1];
+ /* Cumulative statistics since mount. */
+ u64 shrink_scanned[F2FS_NC_NR_QUEUES];
+ u64 shrink_freed[F2FS_NC_NR_QUEUES];
+ u64 detached[F2FS_NC_NR_QUEUES];
+ u64 compressed_to_raw;
+ u64 restored_hits;
+ u64 restore_failures;
+ u64 cycles;
+ u64 visited;
+ u64 candidates;
+ u64 attempts;
+ u64 converted;
+ u64 candidate_refs; /* Worker-held references right now. */
+ u64 transient_extra_bytes; /* Temporary raw/compressed overlap. */
+ u64 allocation_failures;
+ /* Current configuration snapshot. */
+ u32 threshold_pct;
+ u32 interval_ms;
+ bool initialized; /* A compression context is available. */
+};
+
#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
/* Extended NODE_CACHE entry; base must remain the first member. */
struct f2fs_node_cached_block {
@@ -28,6 +55,13 @@ static_assert(offsetof(struct f2fs_node_cached_block, base) == 0);
size_t f2fs_nc_entry_alloc_size(struct f2fs_cached_block_list *cache);
void f2fs_nc_init(struct f2fs_sb_info *sbi);
void f2fs_nc_destroy(struct f2fs_sb_info *sbi);
+u64 f2fs_nc_next_deadline(struct f2fs_sb_info *sbi);
+void f2fs_nc_schedule_start(struct f2fs_sb_info *sbi, u64 now_ms);
+void f2fs_nc_run(struct f2fs_sb_info *sbi);
+u64 f2fs_nc_config_value(struct f2fs_sb_info *sbi,
+ enum f2fs_nc_param id);
+int f2fs_nc_config_update(struct f2fs_sb_info *sbi,
+ enum f2fs_nc_param id, u64 value);
void f2fs_nc_free_data(struct f2fs_cached_block *entry);
int f2fs_nc_restore(struct f2fs_cached_block *entry);
void f2fs_nc_validation_failed(struct f2fs_cached_block *entry);
@@ -43,6 +77,8 @@ unsigned long f2fs_nc_shrink_nodes(struct f2fs_sb_info *sbi,
unsigned long nr_to_scan);
void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
struct f2fs_nc_memory *memory);
+void f2fs_nc_get_stats(struct f2fs_sb_info *sbi,
+ struct f2fs_nc_stats *stats);
#else
static inline size_t f2fs_nc_entry_alloc_size(struct f2fs_cached_block_list *cache)
{
@@ -101,6 +137,7 @@ static inline void f2fs_nc_memory_usage(struct f2fs_sb_info *sbi,
sizeof(struct f2fs_cached_block);
memory->data_bytes = (u64)NODE_CACHE(sbi)->num_entries * sbi->blocksize;
}
+
#endif
#endif /* __F2FS_NODE_CACHE_COMPRESS_H__ */
diff --git a/fs/f2fs/node_cache_policy.c b/fs/f2fs/node_cache_policy.c
index 3adcbf89d37d..7c82557c05ff 100644
--- a/fs/f2fs/node_cache_policy.c
+++ b/fs/f2fs/node_cache_policy.c
@@ -9,6 +9,54 @@
#define F2FS_NC_RECLAIM_SCALE_PCT 25U
#define F2FS_NC_SCORE_HEADROOM 4U
+void f2fs_nc_config_defaults(struct f2fs_nc_config *cfg)
+{
+ *cfg = (struct f2fs_nc_config) {
+ .compression_threshold_pct = F2FS_NC_DEFAULT_THRESHOLD_PCT,
+ .compression_interval_ms = F2FS_NC_DEFAULT_INTERVAL_MS,
+ };
+}
+
+int f2fs_nc_config_set(struct f2fs_nc_config *out,
+ const struct f2fs_nc_config *old, enum f2fs_nc_param id, u64 value)
+{
+ struct f2fs_nc_config new;
+
+ if (!out || !old)
+ return -EINVAL;
+ new = *old;
+ switch (id) {
+ case F2FS_NC_PARAM_THRESHOLD:
+ if (value > F2FS_NC_PERCENT_MAX)
+ return -ERANGE;
+ new.compression_threshold_pct = value;
+ break;
+ case F2FS_NC_PARAM_INTERVAL:
+ if (value < F2FS_NC_MIN_INTERVAL_MS ||
+ value > F2FS_NC_MAX_INTERVAL_MS)
+ return -ERANGE;
+ new.compression_interval_ms = value;
+ break;
+ default:
+ return -EINVAL;
+ }
+ *out = new;
+ return 0;
+}
+
+u64 f2fs_nc_config_get(const struct f2fs_nc_config *cfg,
+ enum f2fs_nc_param id)
+{
+ switch (id) {
+ case F2FS_NC_PARAM_THRESHOLD:
+ return cfg->compression_threshold_pct;
+ case F2FS_NC_PARAM_INTERVAL:
+ return cfg->compression_interval_ms;
+ default:
+ return 0;
+ }
+}
+
static void f2fs_nc_queue_weights(u32 blocksize,
u64 weight[F2FS_NC_NR_QUEUES])
{
diff --git a/fs/f2fs/node_cache_policy.h b/fs/f2fs/node_cache_policy.h
index 42af4ad1698f..fb305413aad6 100644
--- a/fs/f2fs/node_cache_policy.h
+++ b/fs/f2fs/node_cache_policy.h
@@ -4,11 +4,17 @@
#include <linux/types.h>
+#define F2FS_NC_PERCENT_MAX 100U
+#define F2FS_NC_DEFAULT_THRESHOLD_PCT 0U
+#define F2FS_NC_DEFAULT_INTERVAL_MS 1000U
+#define F2FS_NC_MIN_INTERVAL_MS 100U
+#define F2FS_NC_MAX_INTERVAL_MS 30000U
+#define F2FS_NC_SCAN_MAX 1024U
+#define F2FS_NC_MAX_CANDIDATES 256U
#define F2FS_NC_BUCKET_256_SIZE 256U
#define F2FS_NC_BUCKET_512_SIZE 512U
#define F2FS_NC_BUCKET_1024_SIZE 1024U
#define F2FS_NC_MAX_OBJECT_SIZE F2FS_NC_BUCKET_1024_SIZE
-#define F2FS_NC_PERCENT_MAX 100U
/*
* Keep compressed queues in the same order as f2fs_nc_bucket_sizes[].
@@ -22,6 +28,22 @@ enum f2fs_nc_queue {
F2FS_NC_NR_QUEUES,
};
+enum f2fs_nc_param {
+ F2FS_NC_PARAM_THRESHOLD,
+ F2FS_NC_PARAM_INTERVAL,
+};
+
+struct f2fs_nc_config {
+ u32 compression_threshold_pct; /* Max payload percentage; 0 disables. */
+ u32 compression_interval_ms; /* Delay between worker passes, in ms. */
+};
+
+void f2fs_nc_config_defaults(struct f2fs_nc_config *cfg);
+int f2fs_nc_config_set(struct f2fs_nc_config *out,
+ const struct f2fs_nc_config *old, enum f2fs_nc_param id, u64 value);
+u64 f2fs_nc_config_get(const struct f2fs_nc_config *cfg,
+ enum f2fs_nc_param id);
+
u64 f2fs_nc_effective_count(const unsigned long nr[F2FS_NC_NR_QUEUES],
u32 blocksize);
void f2fs_nc_scan_quotas(const unsigned long nr[F2FS_NC_NR_QUEUES],
diff --git a/fs/f2fs/sysfs.c b/fs/f2fs/sysfs.c
index 9749da70089a..ea657b6e7419 100644
--- a/fs/f2fs/sysfs.c
+++ b/fs/f2fs/sysfs.c
@@ -18,6 +18,7 @@
#include "segment.h"
#include "gc.h"
#include "iostat.h"
+#include "node_cache_compress.h"
#include <trace/events/f2fs.h>
static struct proc_dir_entry *f2fs_proc_root;
@@ -312,6 +313,52 @@ static ssize_t mounted_time_sec_show(struct f2fs_attr *a,
return sysfs_emit(buf, "%llu\n", SIT_I(sbi)->mounted_time);
}
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+static ssize_t node_compress_interval_show(struct f2fs_attr *a,
+ struct f2fs_sb_info *sbi, char *buf)
+{
+ return sysfs_emit(buf, "%llu\n",
+ f2fs_nc_config_value(sbi, F2FS_NC_PARAM_INTERVAL));
+}
+
+static ssize_t node_compress_threshold_show(struct f2fs_attr *a,
+ struct f2fs_sb_info *sbi, char *buf)
+{
+ return sysfs_emit(buf, "%llu\n",
+ f2fs_nc_config_value(sbi, F2FS_NC_PARAM_THRESHOLD));
+}
+
+static ssize_t f2fs_nc_config_store(struct f2fs_sb_info *sbi,
+ enum f2fs_nc_param id, const char *buf,
+ size_t count)
+{
+ u64 value;
+ int ret;
+
+ if (kstrtou64(skip_spaces(buf), 0, &value))
+ return -EINVAL;
+ if (!down_read_trylock(&sbi->sb->s_umount))
+ return -EAGAIN;
+ ret = f2fs_nc_config_update(sbi, id, value);
+ up_read(&sbi->sb->s_umount);
+ return ret ? ret : count;
+}
+
+static ssize_t node_compress_interval_store(struct f2fs_attr *a,
+ struct f2fs_sb_info *sbi,
+ const char *buf, size_t count)
+{
+ return f2fs_nc_config_store(sbi, F2FS_NC_PARAM_INTERVAL, buf, count);
+}
+
+static ssize_t node_compress_threshold_store(struct f2fs_attr *a,
+ struct f2fs_sb_info *sbi,
+ const char *buf, size_t count)
+{
+ return f2fs_nc_config_store(sbi, F2FS_NC_PARAM_THRESHOLD, buf, count);
+}
+#endif
+
#ifdef CONFIG_F2FS_STAT_FS
static ssize_t moved_blocks_foreground_show(struct f2fs_attr *a,
struct f2fs_sb_info *sbi, char *buf)
@@ -1363,6 +1410,14 @@ ATGC_INFO_RW_ATTR(atgc_age_threshold, age_threshold);
/* WB_THREAD ATTR */
WB_THREAD_RW_ATTR(cache_wb_interval, cache_wb_interval);
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+static struct f2fs_attr f2fs_attr_node_compress_interval =
+ __ATTR(node_compress_interval, 0644, node_compress_interval_show,
+ node_compress_interval_store);
+static struct f2fs_attr f2fs_attr_node_compress_threshold =
+ __ATTR(node_compress_threshold, 0644, node_compress_threshold_show,
+ node_compress_threshold_store);
+#endif
F2FS_GENERAL_RO_ATTR(dirty_segments);
F2FS_GENERAL_RO_ATTR(free_segments);
@@ -1551,6 +1606,10 @@ static struct attribute *f2fs_attrs[] = {
ATTR_LIST(adjust_lock_priority),
ATTR_LIST(critical_task_priority),
ATTR_LIST(cache_wb_interval),
+#ifdef CONFIG_F2FS_FS_NODE_CACHE_COMPRESSION
+ ATTR_LIST(node_compress_interval),
+ ATTR_LIST(node_compress_threshold),
+#endif
NULL,
};
ATTRIBUTE_GROUPS(f2fs);
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
* [RFC PATCH 6/6] f2fs: give referenced node cache entries a shared second chance
2026-09-29 7:29 [RFC PATCH 0/6] f2fs: retain clean node blocks in a compressed cache Wenjie Qi
` (4 preceding siblings ...)
2026-09-29 7:29 ` [RFC PATCH 5/6] f2fs: compress clean node cache entries in background Wenjie Qi
@ 2026-09-29 7:29 ` Wenjie Qi
5 siblings, 0 replies; 7+ messages in thread
From: Wenjie Qi @ 2026-09-29 7:29 UTC (permalink / raw)
To: Jaegeuk Kim, Chao Yu
Cc: Barry Song, linux-f2fs-devel, linux-kernel, Wenjie Qi
The shrinker already gives a referenced cache entry a second chance by
clearing REFERENCED and moving it to the queue tail. The compression worker
only skips such entries, so it can revisit the same referenced prefix on
every pass.
Use the same clear-and-move rule in both paths. The worker records the
original raw-list tail so entries moved during a pass are not visited again
in that pass.
If an access races with compression after candidate collection, discard
the temporary compressed object and give the raw entry the same second
chance.
Signed-off-by: Wenjie Qi <qiwenjie@xiaomi.com>
---
fs/f2fs/cache.c | 15 ++++++++++++---
fs/f2fs/cache.h | 3 +++
fs/f2fs/debug.c | 4 ++--
fs/f2fs/node_cache_compress.c | 36 ++++++++++++++++++++++++++---------
fs/f2fs/node_cache_compress.h | 1 +
5 files changed, 45 insertions(+), 14 deletions(-)
diff --git a/fs/f2fs/cache.c b/fs/f2fs/cache.c
index b3158f4ab45c..2b0a6eccd3af 100644
--- a/fs/f2fs/cache.c
+++ b/fs/f2fs/cache.c
@@ -60,6 +60,17 @@ void f2fs_cache_update_tag(struct f2fs_cached_block *entry,
spin_unlock_irqrestore(&cache->tree_lock, flags);
}
+bool f2fs_cache_clear_referenced_and_move(struct f2fs_cached_block_list *cache,
+ struct list_head *head,
+ struct f2fs_cached_block *entry)
+{
+ lockdep_assert_held(&cache->list_lock);
+ if (!f2fs_cache_test_and_clear_referenced(entry))
+ return false;
+ list_move_tail(&entry->list, head);
+ return true;
+}
+
bool f2fs_mark_cache_dirty(struct f2fs_cached_block *entry)
{
struct f2fs_cached_block_list *cache = entry->cache;
@@ -594,10 +605,8 @@ f2fs_shrink_cache_list(struct f2fs_cached_block_list *cache,
scanned++;
/* If accessed, give it a second chance to rotate to tail */
- if (f2fs_cache_test_and_clear_referenced(entry)) {
- list_move_tail(&entry->list, head);
+ if (f2fs_cache_clear_referenced_and_move(cache, head, entry))
continue;
- }
if (f2fs_cache_test_dirty(entry) ||
f2fs_cache_test_writeback(entry) ||
diff --git a/fs/f2fs/cache.h b/fs/f2fs/cache.h
index 3c1fdc7894b7..8acbbbaf897d 100644
--- a/fs/f2fs/cache.h
+++ b/fs/f2fs/cache.h
@@ -232,6 +232,9 @@ void f2fs_cache_wait_writeback_cond(struct f2fs_cached_block *entry,
void f2fs_cache_wait_writeback(struct f2fs_cached_block *entry);
void f2fs_cache_update_tag(struct f2fs_cached_block *entry,
unsigned int clear_from, unsigned int set_to);
+bool f2fs_cache_clear_referenced_and_move(struct f2fs_cached_block_list *cache,
+ struct list_head *head,
+ struct f2fs_cached_block *entry);
struct f2fs_cached_block *f2fs_grab_cache(struct f2fs_cached_block_list *cache,
unsigned long index, int flags);
void f2fs_truncate_locked_cache(struct f2fs_cached_block *entry,
diff --git a/fs/f2fs/debug.c b/fs/f2fs/debug.c
index 984fed727a48..8625dccc2131 100644
--- a/fs/f2fs/debug.c
+++ b/fs/f2fs/debug.c
@@ -728,9 +728,9 @@ static int stat_show(struct seq_file *s, void *v)
nc.payload[0], nc.payload[1], nc.payload[2],
nc.slots[0], nc.slots[1], nc.slots[2]);
seq_printf(s,
- "NCWorker: cycles=%llu visited=%llu candidates=%llu attempts=%llu converted=%llu candidate_refs=%llu transient=%llu alloc_fail=%llu\n",
+ "NCWorker: cycles=%llu visited=%llu candidates=%llu attempts=%llu converted=%llu referenced_moved=%llu candidate_refs=%llu transient=%llu alloc_fail=%llu\n",
nc.cycles, nc.visited, nc.candidates, nc.attempts,
- nc.converted, nc.candidate_refs,
+ nc.converted, nc.referenced_moved, nc.candidate_refs,
nc.transient_extra_bytes, nc.allocation_failures);
seq_printf(s,
"NCRestore: compressed_to_raw=%llu restored_hits=%llu restore_failures=%llu\n",
diff --git a/fs/f2fs/node_cache_compress.c b/fs/f2fs/node_cache_compress.c
index 6e474fea7241..cf13110b2961 100644
--- a/fs/f2fs/node_cache_compress.c
+++ b/fs/f2fs/node_cache_compress.c
@@ -147,6 +147,7 @@ struct f2fs_nc_ctx {
atomic64_t candidate_count;
atomic64_t attempts;
atomic64_t converted;
+ atomic64_t referenced_moved;
atomic64_t candidate_refs;
atomic64_t transient_extra_bytes;
atomic64_t allocation_failures;
@@ -574,7 +575,6 @@ static bool f2fs_nc_candidate(struct f2fs_cached_block *entry,
!f2fs_cache_test_dirty(entry) &&
!f2fs_cache_test_writeback(entry) &&
!f2fs_cache_test_locked(entry) &&
- !f2fs_cache_test_referenced(entry) &&
!f2fs_cache_test_incompressible(entry) &&
atomic_read(&entry->refcount) == refs;
}
@@ -601,6 +601,7 @@ f2fs_nc_compress(struct f2fs_cached_block *entry,
u32 crc;
int bucket;
int len;
+ bool referenced;
if (!cache || !IS_NODE_CACHE(cache) || !cfg ||
f2fs_cache_test_compressed(entry))
@@ -654,9 +655,14 @@ f2fs_nc_compress(struct f2fs_cached_block *entry,
spin_lock(&cache->list_lock);
spin_lock_irqsave(&cache->tree_lock, flags);
+ referenced = f2fs_cache_test_referenced(entry);
+ /* A concurrent access wins; rotate the entry and discard this result. */
+ if (referenced && entry->cache == cache &&
+ f2fs_cache_clear_referenced_and_move(cache, &cache->lru_list, entry))
+ atomic64_inc(&ctx->referenced_moved);
if (entry->cache != cache || entry->data != raw ||
atomic_read(&entry->refcount) != 2 || f2fs_cp_error(sbi) ||
- f2fs_cache_test_referenced(entry) ||
+ referenced ||
!f2fs_cache_test_uptodate(entry) || f2fs_cache_test_dirty(entry) ||
f2fs_cache_test_writeback(entry) ||
f2fs_cache_test_incompressible(entry)) {
@@ -693,7 +699,7 @@ void f2fs_nc_run(struct f2fs_sb_info *sbi)
struct f2fs_nc_ctx *ctx = sbi->node_compress;
struct f2fs_cached_block_list *cache = NODE_CACHE(sbi);
struct f2fs_nc_config cfg;
- struct f2fs_cached_block *entry;
+ struct f2fs_cached_block *entry, *next, *scan_tail;
unsigned long flags;
u64 now_ms;
unsigned long raw_population;
@@ -714,15 +720,26 @@ void f2fs_nc_run(struct f2fs_sb_info *sbi)
goto out;
spin_lock(&cache->list_lock);
- list_for_each_entry(entry, &cache->lru_list, list) {
+ /* Do not revisit entries rotated to the tail during this pass. */
+ scan_tail = list_empty(&cache->lru_list) ? NULL :
+ list_last_entry(&cache->lru_list, struct f2fs_cached_block, list);
+ list_for_each_entry_safe(entry, next, &cache->lru_list, list) {
if (visited == goal || count == F2FS_NC_MAX_CANDIDATES)
break;
visited++;
- if (!f2fs_nc_candidate(entry, cache, 1))
- continue;
- f2fs_cache_get(entry);
- ctx->candidates[count++] = entry;
- atomic64_inc(&ctx->candidate_refs);
+ if (f2fs_nc_candidate(entry, cache, 1)) {
+ if (f2fs_cache_clear_referenced_and_move(cache,
+ &cache->lru_list,
+ entry)) {
+ atomic64_inc(&ctx->referenced_moved);
+ } else {
+ f2fs_cache_get(entry);
+ ctx->candidates[count++] = entry;
+ atomic64_inc(&ctx->candidate_refs);
+ }
+ }
+ if (entry == scan_tail)
+ break;
}
spin_unlock(&cache->list_lock);
atomic64_add(visited, &ctx->visited);
@@ -890,6 +907,7 @@ void f2fs_nc_get_stats(struct f2fs_sb_info *sbi,
stats->candidates = atomic64_read(&ctx->candidate_count);
stats->attempts = atomic64_read(&ctx->attempts);
stats->converted = atomic64_read(&ctx->converted);
+ stats->referenced_moved = atomic64_read(&ctx->referenced_moved);
stats->candidate_refs = atomic64_read(&ctx->candidate_refs);
stats->transient_extra_bytes =
atomic64_read(&ctx->transient_extra_bytes);
diff --git a/fs/f2fs/node_cache_compress.h b/fs/f2fs/node_cache_compress.h
index efe2eceef352..542093f2eca7 100644
--- a/fs/f2fs/node_cache_compress.h
+++ b/fs/f2fs/node_cache_compress.h
@@ -31,6 +31,7 @@ struct f2fs_nc_stats {
u64 candidates;
u64 attempts;
u64 converted;
+ u64 referenced_moved;
u64 candidate_refs; /* Worker-held references right now. */
u64 transient_extra_bytes; /* Temporary raw/compressed overlap. */
u64 allocation_failures;
--
2.43.0
^ permalink raw reply [flat|nested] 7+ messages in thread
end of thread, other threads:[~2026-09-29 7:29 UTC | newest]
Thread overview: 7+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-29 7:29 [RFC PATCH 0/6] f2fs: retain clean node blocks in a compressed cache Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 1/6] f2fs: generalize metadata cache shrinking to explicit lists Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 2/6] f2fs: add compressed clean node cache representation Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 3/6] f2fs: bias node cache reclaim toward raw entries Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 4/6] f2fs: restore compressed node cache entries before access Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 5/6] f2fs: compress clean node cache entries in background Wenjie Qi
2026-09-29 7:29 ` [RFC PATCH 6/6] f2fs: give referenced node cache entries a shared second chance Wenjie Qi
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®