mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Daniel Phillips <phillips@bonn-fries.net>
To: Linus Torvalds <torvalds@transmeta.com>
Cc: <dmccr@us.ibm.com>, Kernel Mailing List <linux-kernel@vger.kernel.org>
Subject: Re: [RFC] Page table sharing
Date: Sat, 16 Feb 2002 22:08:05 +0100	[thread overview]
Message-ID: <E16cC3z-0004Kf-00@starship.berlin> (raw)
In-Reply-To: <Pine.LNX.4.33.0202161212200.24268-100000@home.transmeta.com>
In-Reply-To: <Pine.LNX.4.33.0202161212200.24268-100000@home.transmeta.com>

On February 16, 2002 09:21 pm, Linus Torvalds wrote:
> On Sat, 16 Feb 2002, Daniel Phillips wrote:
> >
> > I think this patch is ready to look at now.  It's been pretty stable, 
> > though I haven't gone as far as booting with it - page table sharing is 
> > still restricted to uid 9999.  I'm running it on a 2 way under moderate 
> > load without apparent problems.  The speedup on forking from a parent 
> > with large vm is *way* more than I expected.
> 
> I'd really like to hear what happens when you enable it unconditionally,
> and then run various real loads along with things like "lmbench".

I'm running on it unconditionally now.  I'm still up though I haven't run 
really heavy stress tests.  I've had one bug report, curiously with the 
version keying on uid 9999 while *not* running as uid 9999.  This makes me 
think that there are some uninitialized page use counts on page tables 
somewhere in the system.

> Also, when you test forking over a parent, do you test just the fork, or
> do you test the "fork+wait" combination that waits for the child to exit
> too? The latter is the only really meaningful thing to test.

Will do.  Does this do the trick:

    wait();
    gettimeofday(&etime, NULL);

If so, it doesn't affect the timings at all, in other words I haven't just 
pushed the work into the child's exit.

> Anyway, the patch certainly looks pretty simple and small. Great.
> 
> > I haven't fully analyzed the locking yet, but I'm beginning to suspect it
> > just works as is, i.e., I haven't exposed any new critical regions.  I'd
> > be happy to be corrected on that though.
> 
> What's the protection against two different MM's doing a
> "zap_page_range()" concurrently, both thinking that they can just drop the
> page table directory entry, and neither actually freeing it? I don't see
> any such logic there..

Nothing prevents that, duh.

> I suspect that the only _good_ way to handle it is to do
> 
> 	pmd_page = ..
> 
> 	if (put_page_testzero(pmd_page)) {
> 		.. free the actual page table entries ..
> 		__free_pages_ok(pmd_page, 0);
> 	}
> 
> instead of using the free_page() logic. Maybe you do that already, I
> didn't go through the patches _that_ closely.

I do something similar in clear_page_tables->free_one_pmd, after the entries
are all gone.  I have to do something different in zap_page_range - it wants
to free the pmd only if the count is *greater* than one, and can't tolerate 
two mms thinking that at the same time.  I think I'd better lock the pmd page 
there.

> Ie you'd do a "two-phase" page free - first do the count handling, and if
> that indicates you should really free the pmd, you free the lower page
> tables before you physically free the pmd page (ie the page is "live" even
> though it has a count of zero).

Actually, I'm not freeing the pmd I'm freeing *pmd, a page table.  So, if the
count is > 1 it can be freed without doing anything to the ptes on it.  This
is the entire source of the speedup, by the way.

-- 
Daniel

  reply	other threads:[~2002-02-16 21:03 UTC|newest]

Thread overview: 53+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2002-02-16 18:07 Daniel Phillips
2002-02-16 20:21 ` Linus Torvalds
2002-02-16 21:08   ` Daniel Phillips [this message]
2002-02-17  6:23     ` Linus Torvalds
2002-02-17 19:39       ` Daniel Phillips
2002-02-17 20:16         ` Daniel Phillips
2002-02-17 22:16         ` Hugh Dickins
2002-02-18  1:35           ` Daniel Phillips
2002-02-18  8:09             ` Hugh Dickins
2002-02-18  9:41               ` Daniel Phillips
2002-02-18 11:32                 ` Daniel Phillips
2002-02-19  0:01                   ` Daniel Phillips
2002-02-18 19:04                 ` Hugh Dickins
2002-02-18 23:37                   ` Daniel Phillips
2002-02-19  0:56                     ` Linus Torvalds
2002-02-19  1:22                       ` Rik van Riel
2002-02-19  1:29                         ` Daniel Phillips
2002-02-19  1:48                         ` Linus Torvalds
2002-02-19  1:53                           ` Rik van Riel
2002-02-19  2:05                             ` Linus Torvalds
2002-02-19  2:22                               ` Daniel Phillips
2002-02-19  2:35                                 ` Linus Torvalds
2002-02-19  2:55                                   ` Daniel Phillips
2002-02-19  3:11                                   ` Daniel Phillips
2002-02-19  3:22                                   ` Linus Torvalds
2002-02-19  3:45                                     ` Daniel Phillips
2002-02-19 17:29                                       ` Linus Torvalds
2002-02-19 18:11                                         ` Hugh Dickins
2002-02-20 14:18                                           ` Daniel Phillips
2002-02-20 15:30                                             ` Hugh Dickins
2002-02-20 14:10                                         ` Daniel Phillips
2002-02-20 14:38                                           ` Hugh Dickins
2002-02-20 14:57                                             ` Daniel Phillips
2002-02-19 11:39                                     ` Daniel Phillips
2002-02-19 12:22                                       ` Hugh Dickins
2002-02-19 12:43                                         ` Daniel Phillips
2002-02-19 10:02                                   ` Roman Zippel
2002-02-22  5:29                               ` Daniel Phillips
2002-02-22  6:32                                 ` Daniel Phillips
2002-02-22  9:21                                   ` [RFC] Page table sharing, leak gone Daniel Phillips
2002-02-19  1:57                           ` [RFC] Page table sharing Daniel Phillips
2002-02-19  1:23                       ` Daniel Phillips
2002-02-19  1:50                       ` Daniel Phillips
2002-02-19  1:53                         ` Linus Torvalds
2002-02-19  2:12                           ` Daniel Phillips
2002-02-18 23:48                   ` Daniel Phillips
2002-02-18 23:59                   ` Daniel Phillips
2002-02-19  0:03                     ` Hugh Dickins
2002-02-19  0:27                       ` Daniel Phillips
2002-02-19  4:27                         ` Eric W. Biederman
2002-02-19 17:30                           ` Linus Torvalds
     [not found] <Pine.LNX.4.33.0202182000320.5124-100000@coffee.psychology.mcmaster.ca>
2002-02-19  1:11 ` Daniel Phillips
2002-02-19 18:18 Qing Huang

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=E16cC3z-0004Kf-00@starship.berlin \
    --to=phillips@bonn-fries.net \
    --cc=dmccr@us.ibm.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=torvalds@transmeta.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®