From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-wm1-f46.google.com (mail-wm1-f46.google.com [209.85.128.46]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 68183404BE8 for ; Mon, 3 Aug 2026 12:51:54 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=209.85.128.46 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785761516; cv=none; b=cGujxM/JiEgCM6nNOQ9SHmW5aDHOrFdwVzTEVoebsKlV7lP1qCKgWzsuolxs8ftF3IPGRDYnbq2fmpNdmv/PIYKz0qRlo9nwoQswxAij42M8p0xtvspUZ6l0cyuZdnpGImriqkA1lo/Nldd2+RC2Erc/EKjZTYnED3uGZozKvqc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1785761516; c=relaxed/simple; bh=qJDZyJaPrTUKbQ3b/EdHrd0SnnQCDvFSy6slc3wqj2U=; h=From:To:Cc:Subject:Date:Message-ID:MIME-Version; b=Uc5PBRsdimz9X66rmCbzUFryeU5/5E/kTNuUC76DY+nJQI0BJq6AkLQrbYrgD8TN3xw1i/Lp1BZXLKIy2+oDWlLOGbXOJUQUDtjGb6QN9N07dJjUDlxNNUZyEwpR5KoGJe26nWMF2prBLECkqCT0EnNLHgKAKed4d247J4GiM50= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com; spf=pass smtp.mailfrom=gmail.com; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b=OQu+4v8c; arc=none smtp.client-ip=209.85.128.46 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=gmail.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=gmail.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=gmail.com header.i=@gmail.com header.b="OQu+4v8c" Received: by mail-wm1-f46.google.com with SMTP id 5b1f17b1804b1-4955de8797cso12148795e9.3 for ; Mon, 03 Aug 2026 05:51:54 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20251104; t=1785761512; x=1786366312; darn=vger.kernel.org; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:from:to:cc:subject:date:message-id:reply-to:content-type; bh=8UBY5JJk6ucDrVNZKtjRD505n3yq4bKSlPUs5EcW+mg=; b=OQu+4v8cFP68s8D6pdo+Zb8WgzxEVQB2Ntvw/QHkYAKS3KQ+xRUcbrVlI6WCo2oFIN 2T/rGtNyuByHmIwWDoG2/DBDjHXJt5MB2GmpRjSHMk0ZPWTqLojRAJQUYjlO4+/ZYk29 REsQAWBXBH68HP8I3WLiURiJdOW1Nt//yJ2mFUaQx+IaWL/cuxM8eGZcgpSIYQzDK0zc QAk73uI+Yy/osoh11KvAMBM/wiUvuYPsMpNLbaYDIBQRDGLSpsJYAkXJE30bymiO8hwX eNbkDzuUGbWQLkKf4RzFzqk4T732jUMtsqoTe/8iMHrYjmjrQEeo3WmAWeesrxyvMQx2 Qw0A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1785761512; x=1786366312; h=content-transfer-encoding:mime-version:message-id:date:subject:cc :to:from:x-gm-gg:x-gm-message-state:from:to:cc:subject:date :message-id:reply-to:content-type; bh=8UBY5JJk6ucDrVNZKtjRD505n3yq4bKSlPUs5EcW+mg=; b=ZawExev1T8LNB6jZUZIuUYJhS0NUnz8leJE0M1mH53IwcjCoMEPSlcLEI4m3Fs34q0 6aOCJw9WHXBu6V/RFOXCYtVbGRYvZT4zC/8vFU4V/bDzhVVc2hwE/7m4i+XLL3ZtJ3g3 2Ea2ID9i9jEt4ld/ITWb1BqT8AL3JX7AoemCiX3uE4fVxhurdI5WZpZdUS9naa2wlpB8 nuxoYHpf+1DpmzXo0WYxVsFel2NwxOzA8AVhZstjhcBfMD+SYK7hWarBsD3XB/TZUTqH dtijI+nLA10Z+LTwIjr3QJ6qmd2U5H1aIfX4SufS9svcIgBQI8/3ffrJWcXZ3ahPXZkM 2wXw== X-Forwarded-Encrypted: i=1; AHgh+RrPmKNPJng8kON8WvkrR1QhwetvUgYUbRAmx+SCmA9+UsYtZf4JqZJwN/yT0xHTA76ZjW3wbu1NvzRdt+0=@vger.kernel.org X-Gm-Message-State: AOJu0Yw6cEql1II4/0gWmdUyy+zvpwpwbeiDd60GtE7rGMZI1cWlvI+J uMr7HiRzI4fQT4ltUb/mnogg+QO3Umj35GtT5xROj9PBYCsdS5leltMP X-Gm-Gg: AR+sD11GwZZ15BFb7wwDWd5rITCzMj7qtseW8P8d4ea64njPXLh9o0oEhBKiScUWWuD t5QsHkn8OOhJBuI2II6ss+GCo3Mrt7VBc4zV6mMsPfEqoajMHWQh+SNvMgHlqan/n4UcZ2VAJQr gUljbKeSmr/qQcz5ouoipDOZXCiE16H/kHtvXDZQqADbQJjHDlFRq8mIScSByt38yItGCyeNy/R kv7A4fOQpUpXYocQvIIKMgel/PQ1afNOoYklyI3SmqqUGnBIHebJmGYO7H/w56Isv/Gj68/zeJy 2xKHlha8Ch1kUSZ+VBJkmocU6bLxF8WSnRNqFP09tcDzYOMLc3g5SEVkskHuy2ZaX04/ylYUX7N hpdtGhiJ1YIYkT7sYHPMDzGxXt8DVuG9IgtRpMWSh9BQEyEdWxUWMpupPfa20JIsXkZ2JYOOQQj sZdHaMaAnPvOrRuacBQfJ1E2uq+qUpCiSKlrVmzcF7Gyh7Y6cTaZ+O42JShPpbIxYsL1k+LUqEy U9BmzaHqt+j5D4tiz2NIopKAhhT/rWuzYSydM4= X-Received: by 2002:a05:600c:8489:b0:495:495b:9248 with SMTP id 5b1f17b1804b1-4980c64b73emr247171055e9.4.1785761512390; Mon, 03 Aug 2026 05:51:52 -0700 (PDT) Received: from f.. (cst-prg-94-167.cust.vodafone.cz. [46.135.94.167]) by smtp.gmail.com with ESMTPSA id 5b1f17b1804b1-49807b916e1sm196922935e9.9.2026.08.03.05.51.50 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 03 Aug 2026 05:51:51 -0700 (PDT) From: Mateusz Guzik To: brauner@kernel.org Cc: viro@zeniv.linux.org.uk, jack@suse.cz, linux-kernel@vger.kernel.org, linux-fsdevel@vger.kernel.org, Mateusz Guzik Subject: [PATCH v5] fs: avoid spurious dentry ref/unref cycle on open Date: Mon, 3 Aug 2026 14:51:38 +0200 Message-ID: <20260803125138.1937674-1-mjguzik@gmail.com> X-Mailer: git-send-email 2.43.0 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Opening a file grabs a reference on the terminal dentry in __legitimize_path(), then another one in do_dentry_open() and finally drops the initial reference in terminate_walk(). That's 2 modifications which don't need to be there -- do_dentry_open() can consume the already held reference instead. When benchmarking on a 20-core vm using will-it-scale to open the same file read-only, the results are (ops/s): before: 4043375 after: 5629378 (+39%) Signed-off-by: Mateusz Guzik --- The spurious ref cycle remains an issue and it is trivially avoidable, for the most common case anyway. Al Viro had a more involved patchset which got stalled, see: https://lore.kernel.org/linux-fsdevel/20240822003359.GO504335@ZenIV/ I already pointed this out over a year ago when sending v3. Given lack of traffic on the more involved variant, the nice win from my simple patch and its overall triviality, I think it should go in. Worst case, if the more involved work ever gets off the ground it can be trivially reverted later. bench is: $ cat tests/openro3.c #include #include #include #include #include #include static char tmpfile[] = "/tmp/willitscale.XXXXXX"; char *testcase_description = "Same file open/close read-only"; void testcase_prepare(unsigned long nr_tasks) { int fd = mkstemp(tmpfile); assert(fd >= 0); close(fd); } void testcase(unsigned long long *iterations, unsigned long nr) { while (1) { int fd = open(tmpfile, O_RDONLY); assert(fd >= 0); close(fd); (*iterations)++; } } void testcase_cleanup(void) { unlink(tmpfile); } v5: - the extra ref is of course needed, i blame the heatwave for thinking it is not this time around v4: - rebase - don't grab the extra ref on mnt for truncate - bench opening things r/o. note perf improved from last year thanks to other changes fs/internal.h | 1 + fs/namei.c | 15 ++++++++++++--- fs/open.c | 27 ++++++++++++++++++++++++++- 3 files changed, 39 insertions(+), 4 deletions(-) diff --git a/fs/internal.h b/fs/internal.h index c658c8a5ebd5..9632239036ac 100644 --- a/fs/internal.h +++ b/fs/internal.h @@ -205,6 +205,7 @@ int do_fchownat(int dfd, const char __user *filename, uid_t user, gid_t group, int flag); int chown_common(const struct path *path, uid_t user, gid_t group); extern int vfs_open(const struct path *, struct file *); +int vfs_open_consume(struct path *, struct file *); /* * inode.c diff --git a/fs/namei.c b/fs/namei.c index 3f9bf103ba12..f71481b9bf8f 100644 --- a/fs/namei.c +++ b/fs/namei.c @@ -4789,6 +4789,7 @@ static const char *open_last_lookups(struct nameidata *nd, static int do_open(struct nameidata *nd, struct file *file, const struct open_flags *op) { + struct vfsmount *mnt; struct mnt_idmap *idmap; int open_flag = op->open_flag; bool do_truncate; @@ -4830,11 +4831,17 @@ static int do_open(struct nameidata *nd, error = mnt_want_write(nd->path.mnt); if (error) return error; + /* + * A dedicated reference is needed because after the call to + * vfs_open_consume() we no longer own the reference in nd->path.mnt + * while we need to undo write acess below. + */ + mnt = mntget(nd->path.mnt); do_truncate = true; } error = may_open(idmap, &nd->path, acc_mode, open_flag); if (!error && !(file->f_mode & FMODE_OPENED)) - error = vfs_open(&nd->path, file); + error = vfs_open_consume(&nd->path, file); if (!error) error = security_file_post_open(file, op->acc_mode); if (!error && do_truncate) @@ -4843,8 +4850,10 @@ static int do_open(struct nameidata *nd, WARN_ON(1); error = -EINVAL; } - if (do_truncate) - mnt_drop_write(nd->path.mnt); + if (do_truncate) { + mnt_drop_write(mnt); + mntput(mnt); + } return error; } diff --git a/fs/open.c b/fs/open.c index 6b1c14e684a9..2a7697cee00b 100644 --- a/fs/open.c +++ b/fs/open.c @@ -931,6 +931,11 @@ static inline int file_get_write_access(struct file *f) return error; } +/* + * Populate struct file + * + * NOTE: it assumes f_path is populated and consumes the caller's reference. + */ static int do_dentry_open(struct file *f, int (*open)(struct inode *, struct file *)) { @@ -938,7 +943,6 @@ static int do_dentry_open(struct file *f, struct inode *inode = f->f_path.dentry->d_inode; int error; - path_get(&f->f_path); f->f_inode = inode; f->f_mapping = inode->i_mapping; f->f_wb_err = filemap_sample_wb_err(f->f_mapping); @@ -1055,6 +1059,7 @@ int finish_open(struct file *file, struct dentry *dentry, BUG_ON(file->f_mode & FMODE_OPENED); /* once it's opened, it's opened */ file->__f_path.dentry = dentry; + path_get(&file->f_path); return do_dentry_open(file, open); } EXPORT_SYMBOL(finish_open); @@ -1098,6 +1103,7 @@ int vfs_open(const struct path *path, struct file *file) int ret; file->__f_path = *path; + path_get(&file->f_path); ret = do_dentry_open(file, NULL); if (!ret) { /* @@ -1110,6 +1116,25 @@ int vfs_open(const struct path *path, struct file *file) return ret; } +/** + * vfs_open_consume - open the file at the given path and consume the reference + * @path: path to open + * @file: newly allocated file with f_flag initialized + */ +int vfs_open_consume(struct path *path, struct file *file) +{ + int ret; + + file->__f_path = *path; + path->mnt = NULL; + path->dentry = NULL; + ret = do_dentry_open(file, NULL); + if (!ret) { + fsnotify_open(file); + } + return ret; +} + struct file *dentry_open(const struct path *path, int flags, const struct cred *cred) { -- 2.53.0