mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Mikko Rantalainen <mikko.rantalainen@peda.net>
To: Matthew Wilcox <willy@infradead.org>
Cc: linux-fsdevel@vger.kernel.org, linux-api@vger.kernel.org,
	linux-kernel@vger.kernel.org, brauner@kernel.org,
	viro@zeniv.linux.org.uk, jack@suse.cz, alx@kernel.org,
	dalias@libc.org
Subject: Re: [RFC PATCH 0/1] close(): stop exposing non-retryable EINTR
Date: Mon, 14 Sep 2026 12:07:41 +0300	[thread overview]
Message-ID: <bcdb8f85-d53a-4824-8e34-8d190ec7f4b2@peda.net> (raw)
In-Reply-To: <aqcmz5uI7PgN5CT7@casper.infradead.org>

Matthew Wilcox (2026-09-14 01:42 Europe/Helsinki):
> On Sun, Sep 13, 2026 at 10:38:14PM +0300, Mikko Rantalainen wrote:
>> However, currently Linux kernel will return EINTR in some cases for
>> close(). There is no good way for caller to recover from this case using
>> the original fd. Whatever action was actually interrupted cannot be
>> resumed or retried through this fd, because the fd has already been
>> consumed. Even worse, EINTR conventionally invites retrying an operation,
>> but retrying close() is unsafe: the same file descriptor number may
>> already refer to another file opened by another thread by the time
>> close() returns EINTR.
>>
>> In addition, POSIX.1-2024 requires that if close() reports EINTR, the
>> descriptor must remain open. It also explicitly permits an interrupted
>> close() to return success after closing the descriptor.
> 
> I think any filesystem / device driver / ... which returns -EINTR from
> close() is broken.  There is one exception though -- if the signal is
> fatal.  It's like read()/write() being killable; if the signal is fatal,
> the task dies before it gets to see the errno.  So it doesn't matter.
> 
> So that's my preferred solution; track down the bad kernel code that's
> doing things in close() that are "interruptible" and convert them to
> "killable".  We don't want SIGWINCH or SIGALRM interrupting close();
> that's just dumb.

Am I reading this correctly as a proposed VFS invariant: ->flush() must
not return an interruption result which could become visible as EINTR
to a surviving userspace caller of close()? A fatal signal is the
exception because the task will die before observing the return value.

That seems like a reasonable invariant, but I couldn't find it
documented anywhere. Documentation/filesystems/vfs.rst currently says
only:

    flush
        called by the close(2) system call to flush a file

without specifying allowed return values or signal semantics.

I had originally approached this from the userspace contract. close(2)
guarantees relinquishing the descriptor, but does not provide a general
synchronization guarantee. In particular, successful close() does not
mean regular-file data has reached storage. Applications which require
that guarantee need an explicit synchronization operation such as
fsync() or fdatasync() before close().

The close(2) documentation does say that later close-time operations,
including flushing data to a filesystem or device, *can* report errors.
But I don't read it as guaranteeing that all such work completes or that
all outstanding errors are discovered before close() returns.

That's why I was thinking EINTR should simply be converted to success
by close() implementation.

So I think the important distinction is:

1. userspace is not generally promised completion of arbitrary
   close-time work; but

2. if a subsystem deliberately performs synchronous work in ->flush(),
   the kernel may nevertheless require that work to complete for the
   subsystem's own semantics.

If (2) is the intended VFS rule, then I agree that converting EINTR to
success in close() would hide a bug rather than fix it. The bug would
be an ->flush() implementation allowing an ordinary signal to abandon
work which it intended to perform synchronously.

Would it make sense to document that invariant explicitly, e.g. that
->flush() must not return -EINTR or -ERESTART* due to an ordinary
non-fatal signal?

If those results should be considered implementation bugs, perhaps a
useful diagnostic would also be something along these lines:

---
retval = filp_flush(file, current->files);

WARN_ONCE(retval == -EINTR ||
          retval == -ERESTARTSYS ||
          retval == -ERESTARTNOINTR ||
          retval == -ERESTARTNOHAND ||
          retval == -ERESTART_RESTARTBLOCK,
          "close: ->flush %ps returned interrupt error %d\n",
          file->f_op->flush, retval);
---

That would leave the existing userspace ABI unchanged while making
remaining offending implementations easier to find and fix.

I also considered retrying filp_flush() inside close(), but I don't
think that can be done generically. ->flush() is not documented as
safe to restart from the beginning after partial execution, and
an interruptible wait could immediately encounter the same
still-pending signal again. So fixing the interruptibility at the
offending wait seems safer if the above invariant is indeed
the intended one.

-- 
Mikko


  parent reply	other threads:[~2026-09-14  9:07 UTC|newest]

Thread overview: 11+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-09-13 19:38 Mikko Rantalainen
2026-09-13 19:38 ` [RFC PATCH 1/1] fs: don't return EINTR from close() Mikko Rantalainen
2026-09-13 20:53 ` [RFC PATCH 0/1] close(): stop exposing non-retryable EINTR Rich Felker
2026-09-13 21:27   ` Alejandro Colomar
2026-09-14  6:36   ` Mikko Rantalainen
2026-09-14  9:44     ` Alejandro Colomar
2026-09-13 22:42 ` Matthew Wilcox
2026-09-13 23:51   ` Rich Felker
2026-09-14  9:38     ` Mikko Rantalainen
2026-09-14  9:07   ` Mikko Rantalainen [this message]
2026-09-14 12:40     ` Mikko Rantalainen

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=bcdb8f85-d53a-4824-8e34-8d190ec7f4b2@peda.net \
    --to=mikko.rantalainen@peda.net \
    --cc=alx@kernel.org \
    --cc=brauner@kernel.org \
    --cc=dalias@libc.org \
    --cc=jack@suse.cz \
    --cc=linux-api@vger.kernel.org \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=viro@zeniv.linux.org.uk \
    --cc=willy@infradead.org \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®