mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* [RFC PATCH 0/1] close(): stop exposing non-retryable EINTR
@ 2026-09-13 19:38 Mikko Rantalainen
  2026-09-13 19:38 ` [RFC PATCH 1/1] fs: don't return EINTR from close() Mikko Rantalainen
                   ` (2 more replies)
  0 siblings, 3 replies; 11+ messages in thread
From: Mikko Rantalainen @ 2026-09-13 19:38 UTC (permalink / raw)
  To: linux-fsdevel
  Cc: linux-api, linux-kernel, brauner, viro, jack, alx, dalias,
	Mikko Rantalainen

This is an RFC because it deliberately changes a long-established raw
syscall ABI. Jan Kara raised userspace-regression concerns when this was
discussed in 2025:

    https://lore.kernel.org/linux-fsdevel/ddqmhjc2rpzk2jjvunbt3l3eukcn4xzkocqzdg3j4msihdhzko@fizekvxndg2d/

while musl and Android bionic already normalize this result to success
in libc.

The Linux implementation of close() will always close the file descriptor
given as argument, except for the invalid file descriptor which will
return -EBADF.

However, currently Linux kernel will return EINTR in some cases for
close(). There is no good way for caller to recover from this case using
the original fd. Whatever action was actually interrupted cannot be
resumed or retried through this fd, because the fd has already been
consumed. Even worse, EINTR conventionally invites retrying an operation,
but retrying close() is unsafe: the same file descriptor number may
already refer to another file opened by another thread by the time
close() returns EINTR.

In addition, POSIX.1-2024 requires that if close() reports EINTR, the
descriptor must remain open. It also explicitly permits an interrupted
close() to return success after closing the descriptor.

This patch is about implementing the second option to be compatible with
both POSIX.1-2024 and real-world applications.

Since close() on Linux first relinquishes the file descriptor and only
then performs ->flush() work, interruption of that later work cannot be
recovered through the original fd. No matter how important that
close-time work was, ownership of the fd has already been irrevocably
relinquished.

For regular files, applications requiring durability already need an
explicit synchronization operation such as fsync() or fdatasync() before
relinquishing the fd. This proposal does not suppress meaningful
delayed-I/O errors such as EIO, ENOSPC, or EDQUOT; it only changes
interruption results whose conventional recovery action (retrying the
operation) is unsafe for close().

Historical discussions about this subject:

- https://inbox.sourceware.org/libc-alpha/efaffc5a404cf104f225c26dbc96e0001cede8f9.1747399542.git.alx@kernel.org/T/

- https://sourceware.org/pipermail/libc-alpha/2025-May/166675.html

- https://lkml.rescloud.iu.edu/hypermail/linux/kernel/2205.3/06731.html

- https://lwn.net/Articles/576478/

- https://yarchive.net/comp/linux/must_check.html

- https://sourceware.org/pipermail/libc-alpha/2025-May/166907.html

- https://sourceware.org/pipermail/libc-alpha/2025-May/166722.html


POSIX.1-2024 also allows EINPROGRESS after the descriptor has been closed.
I considered using that result, but it appears less useful than success
for Linux. It would preserve diagnostic information about interrupted
close-time work, but there is no operation the caller can perform on the
original fd to resume or complete that work. It would therefore turn an
irrevocably completed ownership transfer into an apparent failure without
providing a recovery path. Returning success avoids that ambiguity and
still leaves genuinely useful delayed-I/O errors such as EIO, ENOSPC, and
EDQUOT untouched.

This also matches the direction taken by musl, which initially used
EINPROGRESS for this case and later changed to success because existing
applications were prone to interpret EINPROGRESS as a failure and could
incorrectly infer that the fd was still open.

Automatically replacing EINTR with success does change the *observable*
raw syscall ABI for applications that distinguish EINTR from successful
close(). For applications that already treat the descriptor as consumed,
this changes control flow to the normal successful-close path. I would be
particularly interested in concrete examples where distinguishing EINTR
provides useful recovery semantics, given that the original fd has
already been consumed and cannot be used to resume the interrupted
close-time work.

Mikko Rantalainen (1):
  fs: don't return EINTR from close()

 fs/open.c | 11 ++++++++---
 1 file changed, 8 insertions(+), 3 deletions(-)

-- 
2.43.0


^ permalink raw reply	[flat|nested] 11+ messages in thread

end of thread, other threads:[~2026-09-14 12:40 UTC | newest]

Thread overview: 11+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-09-13 19:38 [RFC PATCH 0/1] close(): stop exposing non-retryable EINTR Mikko Rantalainen
2026-09-13 19:38 ` [RFC PATCH 1/1] fs: don't return EINTR from close() Mikko Rantalainen
2026-09-13 20:53 ` [RFC PATCH 0/1] close(): stop exposing non-retryable EINTR Rich Felker
2026-09-13 21:27   ` Alejandro Colomar
2026-09-14  6:36   ` Mikko Rantalainen
2026-09-14  9:44     ` Alejandro Colomar
2026-09-13 22:42 ` Matthew Wilcox
2026-09-13 23:51   ` Rich Felker
2026-09-14  9:38     ` Mikko Rantalainen
2026-09-14  9:07   ` Mikko Rantalainen
2026-09-14 12:40     ` Mikko Rantalainen

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®