mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Linus Torvalds <torvalds@linux-foundation.org>
To: Al Viro <viro@ZenIV.linux.org.uk>
Cc: OGAWA Hirofumi <hirofumi@mail.parknet.co.jp>,
	linux-fsdevel@vger.kernel.org, linux-kernel@vger.kernel.org
Subject: Re: [RFC] readdir mess
Date: Tue, 12 Aug 2008 18:51:42 -0700 (PDT)	[thread overview]
Message-ID: <alpine.LFD.1.10.0808121841030.3462@nehalem.linux-foundation.org> (raw)
In-Reply-To: <20080813011909.GA28946@ZenIV.linux.org.uk>



On Wed, 13 Aug 2008, Al Viro wrote:
> 
> What _can_ a common helper do, anyway, when we are busy parsing an arseload of
> possibly corrupt data in whatever weird format fs insists upon?

Well, the parsing has to be done by the low-level filesystem code, yes.

However, the whole thing with races with "f_pos" and all the locking - 
that's only because we see the filesystem "readdir" code as being the 
primary source of data.

Quite frankly, if we had a "readdir page cache", the low-level filesystem 
would still have to parse the insane low-level data with corruption 
issues, but we could make it totally independent of f_pos (because we 
would never use in the _real_ file->f_pos - we would just populate the 
cache), and the locking issues would be only a cold-cache issue, with the 
hot-cache hopefully needing little locking at all.

For an exmple of that: you did a good job with all the "seq_file" helpers, 
which meant that the low-level "filesystem" ops didn't need to know 
_anything_ about partial results etc, and it automatically did the right 
thing wrt f_pos updates and lseek etc.

I'm not saying that readdir() would use the _same_ model, but I do suspect 
that a common format in between the disk format and the eventual readdir() 
output, that also could be cached, might mitigate a lot of the problems.

As to the issues with lookup() - yes, a lookup would need to get the lock 
for writing, but only for the last entry, and only if O_CREAT is set. 
There's nothing wrogn with concurrent read-only lookups, I think (apart 
from having to protect the dentries from being duplicated, of course, but 
that would be a per-dentry lock flag, not a directory lock, methinks).

I dunno.

That said, I think you are right that we could also just improve on the 
current non-caching version with soem higher-level semantics. Including 
flags like "yes, we've seen the end", so that we don't need to always call 
into the low-level filesystem one extra time to see that final zero 
return.

So yes, instead of separate "filldir_t" and "void *data" things, having a 
"struct filldir_t" with several fields in common might be worth it.

			Linus

  reply	other threads:[~2008-08-13  1:52 UTC|newest]

Thread overview: 38+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2008-08-12  6:22 Al Viro
2008-08-12 17:02 ` OGAWA Hirofumi
2008-08-12 17:18   ` Linus Torvalds
2008-08-12 18:10     ` Al Viro
2008-08-12 18:22       ` Al Viro
2008-08-12 18:37         ` Al Viro
2008-08-12 19:24           ` Al Viro
2008-08-12 20:02       ` Linus Torvalds
2008-08-12 20:21       ` Linus Torvalds
2008-08-12 20:38         ` Al Viro
2008-08-12 21:04           ` Linus Torvalds
2008-08-13  0:04             ` Al Viro
2008-08-13  0:28               ` Linus Torvalds
2008-08-13  1:19                 ` Al Viro
2008-08-13  1:51                   ` Linus Torvalds [this message]
2008-08-13  8:36               ` Brad Boyer
2008-08-13 16:19                 ` Al Viro
2008-08-15  5:06               ` Jan Harkes
2008-08-15  5:34                 ` Al Viro
2008-08-15 16:58                 ` Linus Torvalds
2008-08-24 10:10                   ` Al Viro
2008-08-24 11:03                     ` Al Viro
2008-08-25 16:16                       ` J. Bruce Fields
2008-08-24 17:20                     ` Linus Torvalds
2008-08-24 19:59                       ` Al Viro
2008-08-24 23:51                         ` Linus Torvalds
2008-08-25  1:33                           ` Al Viro
2008-08-25  1:44                             ` Al Viro
2008-08-12 19:45     ` OGAWA Hirofumi
2008-08-12 20:05       ` Linus Torvalds
2008-08-12 20:59         ` Al Viro
2008-08-12 21:24           ` Linus Torvalds
2008-08-12 21:54             ` Al Viro
2008-08-12 22:04               ` Linus Torvalds
2008-08-13 16:20                 ` J. Bruce Fields
2008-08-12 21:47         ` Alan Cox
2008-08-12 22:20           ` Linus Torvalds
2008-08-12 22:10             ` Alan Cox

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=alpine.LFD.1.10.0808121841030.3462@nehalem.linux-foundation.org \
    --to=torvalds@linux-foundation.org \
    --cc=hirofumi@mail.parknet.co.jp \
    --cc=linux-fsdevel@vger.kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=viro@ZenIV.linux.org.uk \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®