mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Trond Myklebust <trond.myklebust@fys.uio.no>
To: Chris Wedgwood <cw@f00f.org>
Cc: Craig Soules <soules@happyplace.pdl.cmu.edu>,
	jrs@world.std.com, linux-kernel@vger.kernel.org
Subject: Re: NFS Client patch
Date: Wed, 11 Jul 2001 10:14:34 +0200	[thread overview]
Message-ID: <15180.2794.472409.969703@charged.uio.no> (raw)
In-Reply-To: <20010711013805.C31799@weta.f00f.org>
In-Reply-To: <15178.3722.86802.671534@charged.uio.no> <Pine.LNX.3.96L.1010709175623.16113S-100000@happyplace.pdl.cmu.edu> <15178.47928.328862.678031@charged.uio.no> <20010711013805.C31799@weta.f00f.org>

>>>>> " " == Chris Wedgwood <cw@f00f.org> writes:

     > On Tue, Jul 10, 2001 at 10:22:16AM +0200, Trond Myklebust
     > wrote:
     >     Imagine if somebody gives you a 1Gb directory. Would it or
     >     would it not piss you off if your file pointer got reset to
     >     0 every time somebody created a file?
    
     >     The current semantics are scalable. Anything which resets
     >     the file pointer upon change of a file/directory/whatever
     >     isn't...

     > Anyone using a 1GB directory deserves for it not to scale.  I
     > think this is a very poor example.

It's an extreme case, but it illustrates something the kernel should
be able to cope with.

     > No that I disagree with you, the largest directories I have on
     > my system here are 2.6MB (freedb, lots of hashed flat-files in
     > one directory), here I do agree that you should not have to
     > reset the counter everytime.

Right: this is what most people expect. The reason why the readdir
code went through several quite different incarnations in the 2.3.x
series was that duplicate directory entries were not acceptable to
people.

readdir() is not an atomic operation. You can't lock a directory while
doing a series of readdir calls either on local filesystems nor over
NFS. As such, the idea of volatile cookies doesn't really make sense,
nor is it supported in rfc1094 (NFSv2):

   Each "entry" contains a "fileid" which consists of a unique number
   to identify the file within a filesystem, the "name" of the file,
   and a "cookie" which is an opaque pointer to the next entry in the
   directory.  The cookie is used in the next READDIR call to get more
   entries starting at a given point in the directory.

Nothing there states that the cookie can be invalidated, nor is there
even an error to tell you that this is the case.


In rfc1813 (NFSv3), they recognized that NFSv2 couldn't cope with
stale cookies (yes: this fact is explicitly written down on pages 77
and 78), and hence they introduced the cookie verifier and the
NFS3ERR_BAD_COOKIE error, that can be used by the server to declare a
cookie as being stale. In this case, some extra recovery action might
make sense. I can see 3 possible solutions:

  1) The behaviour in this case is undefined. Leave it up to the
     user to reopen the directory, reset the file pointer, or whatever.
     This is what we do now.

  2) implement some extra caching info to allow an improved recovery
     of the last file position. This would likely have to involve
     storing the fileid + filename of the last entry somewhere in the
     struct file.

     This scheme means that lseek() breaks, and can undermine
     glibc. The latter has a lousy getdents algorithm in which it
     reads a number n of entries into a temporary buffer, the copy <=
     n entries to the user (because their struct dirent is larger than
     the kernel struct dirent), and then use lseek() to jump back.

  3) Implement something like Craig suggests whereby you reset the
     file pointer.

     This gives unexpected results as far as the user is concerned as
     it causes duplicate entries to pop up without any warning. It too
     breaks lseek(). It's a policy decision on behalf of the user.

If someone can persuade the glibc people to implement a sane algorithm
for getdents() that precalculates the upper limit on how much padding
is needed, and drops the use of lseek(), then (2) might possibly be
worth doing for 2.5.x (I believe I heard that Solaris does something
along these lines).

If not, given a choice between (1) and (3), I choose (1).



Finally, any scheme that assumes cookie staleness in all cases where
the directory mtime changes will not be passed on to Linus.

Cheers,
  Trond

  parent reply	other threads:[~2001-07-11  8:15 UTC|newest]

Thread overview: 36+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2001-07-09 17:28 Craig Soules
2001-07-09 18:59 ` Trond Myklebust
2001-07-09 19:45   ` Craig Soules
2001-07-09 19:53     ` Charles Cazabon
2001-07-09 21:46     ` J. Richard Sladkey
2001-07-10 15:06       ` Craig Soules
2001-07-09 20:05   ` Trond Myklebust
2001-07-09 22:09     ` Craig Soules
2001-07-10  8:22     ` Trond Myklebust
2001-07-10 13:38       ` Chris Wedgwood
2001-07-11  8:14       ` Trond Myklebust [this message]
     [not found] <Pine.LNX.3.96L.1010709131315.16113O-200000@happyplace.pdl.cmu.edu.suse.lists.linux.kernel>
2001-07-09 18:33 ` Andi Kleen
2001-07-10 13:33   ` Chris Wedgwood
2001-07-10 13:41     ` Andi Kleen
2001-07-10 16:48       ` Craig Soules
2001-07-10 17:06         ` Andi Kleen
2001-07-10 18:04           ` Chris Wedgwood
2001-07-12 20:57             ` Alan Cox
2001-07-13 11:26               ` Chris Wedgwood
2001-07-17 22:02   ` Hans Reiser
2001-07-17 22:14     ` Craig Soules
2001-07-17 22:21       ` Hans Reiser
2001-07-18 13:30         ` Daniel Phillips
2001-07-18 14:46           ` Hans Reiser
2001-07-18 14:00         ` Jan Harkes
2001-07-18 14:46           ` Hans Reiser
2001-07-19 18:24             ` Pavel Machek
2001-07-22 15:15               ` Rob Landley
2001-07-23  2:02                 ` Horst von Brand
2001-07-23  9:57                   ` Rob Landley
2001-07-18 13:57     ` Chris Mason
2001-07-19 11:35       ` Trond Myklebust
2001-07-19 18:02         ` Hans Reiser
2001-07-20  8:50         ` Trond Myklebust
2001-07-20 11:30           ` Hans Reiser
2001-07-20 14:07           ` Chris Mason

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=15180.2794.472409.969703@charged.uio.no \
    --to=trond.myklebust@fys.uio.no \
    --cc=cw@f00f.org \
    --cc=jrs@world.std.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=soules@happyplace.pdl.cmu.edu \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®