From: Partha Satapathy <partha.satapathy@oracle.com>
To: Jan Kara <jack@suse.cz>
Cc: amir73il@gmail.com, viro@zeniv.linux.org.uk, brauner@kernel.org,
linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [External] : Re: [PATCH 0/1] fsnotify: Check parent watches without child dentry flags
Date: Fri, 9 Oct 2026 17:18:17 +0530 [thread overview]
Message-ID: <9499fb1a-85f2-41da-88b7-5e7de6fdff0a@oracle.com> (raw)
In-Reply-To: <qlw6x3hyky5yoievlkd6ylz3ov432hdjlzkcyboswqktnk43il@fvccpqnil5sp>
On 08-10-2026 15:12, Jan Kara wrote:
> On Wed 07-10-26 15:20:50, Partha Sarathi Satapathy wrote:
>> This patch removes the cached-child dentry walk performed when a directory
>> first starts watching child events. Instead, the event path takes a reference
>> to the parent dentry and checks its child-watch mask. It also removes the
>> DCACHE_FSNOTIFY_PARENT_WATCHED flag and its update sites.
>
> If we can get rid of it, it would be a nice simplification but I have
> doubts this is feasible.
>
>> On the test host, 12 million files were populated and their positive dentries
>> were warmed before each run. With eight concurrent unlink/recreate workers
>> and 100 sole-watch add/close cycles, inotify_add_watch() measured:
>>
>> Before After
>> run 1 average 107.420 ms 0.004 ms
>> run 2 average 125.433 ms 0.004 ms
>> run 1 maximum 122.811 ms 0.113 ms
>> run 2 maximum 141.198 ms 0.095 ms
>>
>> The displayed 0.004 ms averages are rounded to three decimal places.
>> Maximum unlink/recreate latency was 777.852 and 770.569 ms before, versus
>> 8.525 and 15.624 ms after. Those mutation maxima are supporting observations:
>> workers ran only during the watch-add loops, so the pre-change and post-change
>> mutation measurement windows had very different lengths.
>
> This looks good and is kind of expected. We don't have to traverse the huge
> children list.
>
>> One 100,000-iteration, single-worker open/close event-path run per mode gave:
>>
>> Before After
>> none: p50/p99 ns 9564 / 15703 9684 / 15023
>> throughput 102221 ops/s 101432 ops/s
>> other: p50/p99 ns 9554 / 15573 9684 / 14962
>> throughput 99442 ops/s 101497 ops/s
>> parent: p50/p99 ns 12529 / 19179 12499 / 17716
>> throughput 78196 ops/s 78655 ops/s
>>
>> The 'other' mode watches a different directory on the same filesystem, so it
>> exercises the superblock watcher path without target-file event delivery.
>> These are single runs; they do not establish a small event-path cost or its
>> absence. No open/close regression is apparent at this measurement resolution.
>> The test used worker CPU 2 and listener CPU 4 on kernels
>> 6.19.0-rc8.V_fsn0.el9.omm0.x86_64 and 6.19.0-rc8.fsn3.el9.omm3.x86_64,
>> respectively. The test host reported XFS for its working directory.
>>
>> The inotify correctness test and broader fsnotify functional test passed on
>> both kernels. The fanotify permission subtest skipped with EPERM in both runs,
>> so FAN_OPEN_PERM remains unvalidated by these results.
>
> This isn't a load for which the DCACHE_FSNOTIFY_PARENT_WATCHED optimization
> is that interesting. Try the following:
>
> Create a directory on tmpfs with say 1024 files, each 4k large.
> Place IN_MODIFY inotify watch on the directory.
> Start X clients (where X can be 1,2,4,8,...,1024 to see the scaling)
> The client will open the file based on its number (so different clients
> work on different file)
> Do 1000000 writes of 1 byte at offset 0 to the file => measure time for
> this
> Close the file
>
> I would bet you would see the difference already at 1 client and as the
> number of clients grows and the cacheline contention on the parent's
> refcount increases, it will get progressively worse. In particular if you
> try on a NUMA machine where the cacheline will ping-pong across NUMA nodes.
>
> We've got performance regression reports for similar loads already when we
> added one cacheline load to fsnotify_parent(). Doing the "grab parent
> refcount" dance is much more expensive than that.
>
> Honza
>
Hi Honza,
Thanks for suggesting this workload.
The setup and before/after timing comparisons are below.
Architecture: x86_64
CPU(s): 512
On-line CPU(s) list: 0-511
NUMA:
NUMA node(s): 2
NUMA node0 CPU(s): 0-127,256-383
NUMA node1 CPU(s): 128-255,384-511
Testcase :
Using /dev/shm tmpfs to create private test directory.
Create 1,024 files each fully populated with 4kb.
Add IN_MODIFY inotify watch to the directory.
Start a listener thread to read notifications and count queue overflows.
Start X(1,2,4..1024) client threads; each opens a different file once.
Wait until every client is ready, then release them together.
Each client starts its timer and performs 1000000
one-byte pwrite() at offset0 on its own file.
Each client stops its timer and closes its file.
Wait for all clients to finish, drain remaining notifications,
and stop the listener.
Print write times, throughput, notification count, and queue-overflow count.
Close inotify and delete the test files and directory.
Note : File creation, open, close, and cleanup are outside the timed
write loops.
Run 1:
-------
| Clients | Before average time | After average time | Time change |
|---:|---:|---:|---:|
| 1 | 1.392 s | 1.484 s | +6.6% |
| 2 | 2.205 s | 2.402 s | +8.9% |
| 4 | 3.969 s | 3.644 s | −8.2% |
| 8 | 5.851 s | 5.804 s | −0.8% |
| 16 | 10.607 s | 11.538 s | +8.8% |
| 32 | 21.093 s | 22.492 s | +6.6% |
| 64 | 48.698 s | 49.576 s | +1.8% |
| 128 | 100.208 s | 100.467 s | +0.3% |
| 256 | 198.773 s | 197.669 s | −0.6% |
| 512 | 400.056 s | 399.550 s | −0.1% |
Run2 :
--------
| Clients | Before average time | After average time | Time change |
|---:|---:|---:|---:|
| 1 | 1.491 s | 1.484 s | −0.4% |
| 2 | 2.391 s | 2.402 s | +0.5% |
| 4 | 4.050 s | 3.644 s | −10.0% |
| 8 | 6.630 s | 5.804 s | −12.5% |
| 16 | 11.931 s | 11.538 s | −3.3% |
| 32 | 22.513 s | 22.492 s | −0.1% |
| 64 | 50.231 s | 49.576 s | −1.3% |
| 128 | 100.699 s | 100.467 s | −0.2% |
| 256 | 198.771 s | 197.669 s | −0.6% |
| 512 | 401.860 s | 399.550 s | −0.6% |
| 1,024 | 803.812 s | 806.593 s | +0.3% |
Positive means After Fix took longer; negative means faster.
The measured write performance is broadly comparable before and after
the fix, with no consistent substantial regression or progressively
worsening slowdown as concurrency increases.
The differences vary between baseline batches: the increases at 16
and 32 clients in the first comparison disappear or reverse in the
second. At 128–1024 clients, average write times differ by approximately
0.6% or less. These shifts are consistent with run-to-run variation;
the results do not show a persistent substantial slowdown from the fix.
Thanks,
Partha
>>
>> Partha Sarathi Satapathy (1):
>> fsnotify: Check parent watches without child dentry flags
>>
>> fs/dcache.c | 4 --
>> fs/notify/fsnotify.c | 99 ++++++++------------------------
>> fs/notify/fsnotify.h | 6 --
>> fs/notify/mark.c | 35 -----------
>> include/linux/dcache.h | 1 -
>> include/linux/fsnotify.h | 7 +--
>> include/linux/fsnotify_backend.h | 29 ++--------
>> 7 files changed, 29 insertions(+), 152 deletions(-)
> --
> Jan Kara <jack@suse.com>
> SUSE Labs, CR
>
prev parent reply other threads:[~2026-10-09 11:48 UTC|newest]
Thread overview: 4+ messages / expand[flat|nested] mbox.gz Atom feed top
2026-10-07 15:20 Partha Sarathi Satapathy
2026-10-07 15:20 ` [PATCH 1/1] " Partha Sarathi Satapathy
2026-10-08 9:42 ` [PATCH 0/1] " Jan Kara
2026-10-09 11:48 ` Partha Satapathy [this message]
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=9499fb1a-85f2-41da-88b7-5e7de6fdff0a@oracle.com \
--to=partha.satapathy@oracle.com \
--cc=amir73il@gmail.com \
--cc=brauner@kernel.org \
--cc=jack@suse.cz \
--cc=linux-fsdevel@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=viro@zeniv.linux.org.uk \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®