mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
* Multi-Threaded fork() correctness on Linux 2.4 & 2.6
@ 2005-09-19 16:24 Smarduch Mario-CMS063
  2005-09-19 18:42 ` Hugh Dickins
  0 siblings, 1 reply; 9+ messages in thread
From: Smarduch Mario-CMS063 @ 2005-09-19 16:24 UTC (permalink / raw)
  To: linux-kernel

MMU/Kernel experts,
    
    recently I've been involved in debugging MT forks on IA-64 the issues
found were related to the way IA64 does its TLB invalidation. But there
is still one  issue that appears to be Linux in general related, and it
has to do with MT forks. 
 
For example a process has 3 threads T1, T2, T3.
 
1 - T1 issues a fork()
    - under page_table_lock write bit is reset in src and dest pte's
2 - T2 may have a TLB mis, the new write protected pte is inserted
    (hw walker or sw tlb hdlr)
3 - T2 winds up in do_wp_page() the page is copied to a new one
4 - In the mean time T3 may be working off the same page, the
    TLB invalidation (flush_tlb_mm()) has no occurred yet.
5 - Eventually TLB is globally flushed so threads will see the new pte.
6 - As a result the MT task may experience inconsisitent state.
    - During 3 & 4. For example locks may be acquired by both
      threads depending on the timing of the copy.

The TLB refill (hw/sw) work independently of the page_table_lock,
it seems all threads should be forced to see the new pte
before the new page is copied over by forcing the pte to 0 
and allowing page_table_lock to synchronize the threads.

I come from an SVR4 background and relatively new to
Linux, any insights or corrections would be greatly appreciated.
Please copy my email id as its to miss email on this list.

	- mario.
  

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: Multi-Threaded fork() correctness on Linux 2.4 & 2.6
  2005-09-19 16:24 Multi-Threaded fork() correctness on Linux 2.4 & 2.6 Smarduch Mario-CMS063
@ 2005-09-19 18:42 ` Hugh Dickins
  2005-09-19 19:24   ` Linus Torvalds
  0 siblings, 1 reply; 9+ messages in thread
From: Hugh Dickins @ 2005-09-19 18:42 UTC (permalink / raw)
  To: Smarduch Mario-CMS063; +Cc: Linus Torvalds, linux-kernel

On Mon, 19 Sep 2005, Smarduch Mario-CMS063 wrote:

> MMU/Kernel experts,
>     
>     recently I've been involved in debugging MT forks on IA-64 the issues
> found were related to the way IA64 does its TLB invalidation. But there
> is still one  issue that appears to be Linux in general related, and it
> has to do with MT forks. 
>  
> For example a process has 3 threads T1, T2, T3.
>  
> 1 - T1 issues a fork()
>     - under page_table_lock write bit is reset in src and dest pte's
> 2 - T2 may have a TLB mis, the new write protected pte is inserted
>     (hw walker or sw tlb hdlr)
> 3 - T2 winds up in do_wp_page() the page is copied to a new one
> 4 - In the mean time T3 may be working off the same page, the
>     TLB invalidation (flush_tlb_mm()) has no occurred yet.
> 5 - Eventually TLB is globally flushed so threads will see the new pte.
> 6 - As a result the MT task may experience inconsisitent state.
>     - During 3 & 4. For example locks may be acquired by both
>       threads depending on the timing of the copy.

I do think you've hit upon something interesting here.  Though it's
not quite as you describe.  We don't have to wait for the flush_tlb_mm
to sort it out: the flush_tlb_page in ptep_establish in break_cow in
do_wp_page resolves the discrepancy much sooner.  But your point is,
that's already too late: T2 inserts a "smudged" copy of the page which
T3 was working on, it does not contain all the data T3 had written there.

> The TLB refill (hw/sw) work independently of the page_table_lock,
> it seems all threads should be forced to see the new pte
> before the new page is copied over by forcing the pte to 0 
> and allowing page_table_lock to synchronize the threads.

I don't get your pte to 0 suggestion.  What we seem to need is to
flush the TLB sooner (as well as in ptep_establish).  A first guess
is that do_wp_page needs (in these circumstances) to flush_tlb_page
before copying the page.

But this stuff is subtle, and TLB flushes shouldn't be added lightly.
I'm certainly not going to rush to propose the fix (and I could easily
be wrong in seeing the problem).  But perhaps others will be surer.

> I come from an SVR4 background and relatively new to
> Linux, any insights or corrections would be greatly appreciated.
> Please copy my email id as its to miss email on this list.

Welcome!

Hugh

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: Multi-Threaded fork() correctness on Linux 2.4 & 2.6
  2005-09-19 18:42 ` Hugh Dickins
@ 2005-09-19 19:24   ` Linus Torvalds
  2005-09-19 20:14     ` Hugh Dickins
  0 siblings, 1 reply; 9+ messages in thread
From: Linus Torvalds @ 2005-09-19 19:24 UTC (permalink / raw)
  To: Hugh Dickins; +Cc: Smarduch Mario-CMS063, linux-kernel



On Mon, 19 Sep 2005, Hugh Dickins wrote:
> On Mon, 19 Sep 2005, Smarduch Mario-CMS063 wrote:
> >     
> >     recently I've been involved in debugging MT forks on IA-64 the issues
> > found were related to the way IA64 does its TLB invalidation. But there
> > is still one  issue that appears to be Linux in general related, and it
> > has to do with MT forks. 
> >  
> > For example a process has 3 threads T1, T2, T3.
> >  
> > 1 - T1 issues a fork()
> >     - under page_table_lock write bit is reset in src and dest pte's
> > 2 - T2 may have a TLB mis, the new write protected pte is inserted
> >     (hw walker or sw tlb hdlr)
> > 3 - T2 winds up in do_wp_page() the page is copied to a new one
> > 4 - In the mean time T3 may be working off the same page, the
> >     TLB invalidation (flush_tlb_mm()) has no occurred yet.
> > 5 - Eventually TLB is globally flushed so threads will see the new pte.
> > 6 - As a result the MT task may experience inconsisitent state.
> >     - During 3 & 4. For example locks may be acquired by both
> >       threads depending on the timing of the copy.
> 
> I do think you've hit upon something interesting here.  Though it's
> not quite as you describe.  We don't have to wait for the flush_tlb_mm
> to sort it out: the flush_tlb_page in ptep_establish in break_cow in
> do_wp_page resolves the discrepancy much sooner.  But your point is,
> that's already too late: T2 inserts a "smudged" copy of the page which
> T3 was working on, it does not contain all the data T3 had written there.

Hmm. 

We hold the page_table_lock when doing the fork(), so T2 can't actually be 
copying the page until we've done the TLB flush, no? And once the TLB 
flush is done, all the writes by T3 should be in the page, so we copy the 
right thing at that point, and there is no consistency problems?

No? What am I missing?

		Linus

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: Multi-Threaded fork() correctness on Linux 2.4 & 2.6
  2005-09-19 19:24   ` Linus Torvalds
@ 2005-09-19 20:14     ` Hugh Dickins
  2005-09-19 20:39       ` Linus Torvalds
  0 siblings, 1 reply; 9+ messages in thread
From: Hugh Dickins @ 2005-09-19 20:14 UTC (permalink / raw)
  To: Linus Torvalds; +Cc: Smarduch Mario-CMS063, linux-kernel

On Mon, 19 Sep 2005, Linus Torvalds wrote:
> 
> We hold the page_table_lock when doing the fork(), so T2 can't actually be 
> copying the page until we've done the TLB flush, no? And once the TLB 
> flush is done, all the writes by T3 should be in the page, so we copy the 
> right thing at that point, and there is no consistency problems?

I was totally overlooking the page_table_lock during the fork.

But no matter, it's not good enough: src_mm->page_table_lock is acquired
and dropped at the inner level, in copy_pte_range (looking at latest 2.6):
it cannot be held across allocating page tables for dst_mm.

So each time T1 drops it, there's a window for the T2 vs. T3 problem.
Yet we don't much want to flush TLB each time we leave copy_pte_range.

Hugh

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: Multi-Threaded fork() correctness on Linux 2.4 & 2.6
  2005-09-19 20:14     ` Hugh Dickins
@ 2005-09-19 20:39       ` Linus Torvalds
  2005-09-19 21:30         ` Hugh Dickins
  0 siblings, 1 reply; 9+ messages in thread
From: Linus Torvalds @ 2005-09-19 20:39 UTC (permalink / raw)
  To: Hugh Dickins; +Cc: Smarduch Mario-CMS063, linux-kernel



On Mon, 19 Sep 2005, Hugh Dickins wrote:
> 
> I was totally overlooking the page_table_lock during the fork.
> 
> But no matter, it's not good enough: src_mm->page_table_lock is acquired
> and dropped at the inner level, in copy_pte_range (looking at latest 2.6):
> it cannot be held across allocating page tables for dst_mm.
> 
> So each time T1 drops it, there's a window for the T2 vs. T3 problem.
> Yet we don't much want to flush TLB each time we leave copy_pte_range.

Hmm. But we do hold the mmap_sem for writing, and we flush before we 
release it, so it should still be ok. The page fault case needs to get it 
for reading anyway.

Yeah, the page_table_lock might make more sense, but I think the mmap_sem 
thing works equally well.

			Linus

^ permalink raw reply	[flat|nested] 9+ messages in thread

* Re: Multi-Threaded fork() correctness on Linux 2.4 & 2.6
  2005-09-19 20:39       ` Linus Torvalds
@ 2005-09-19 21:30         ` Hugh Dickins
  0 siblings, 0 replies; 9+ messages in thread
From: Hugh Dickins @ 2005-09-19 21:30 UTC (permalink / raw)
  To: Linus Torvalds; +Cc: Smarduch Mario-CMS063, linux-kernel

On Mon, 19 Sep 2005, Linus Torvalds wrote:
> 
> Hmm. But we do hold the mmap_sem for writing, and we flush before we 
> release it, so it should still be ok. The page fault case needs to get it 
> for reading anyway.
> 
> Yeah, the page_table_lock might make more sense, but I think the mmap_sem 
> thing works equally well.

Yes, you're both ahead of me: thank you for working that out: no issue.

Hugh

^ permalink raw reply	[flat|nested] 9+ messages in thread

* RE: Multi-Threaded fork() correctness on Linux 2.4 & 2.6
  2005-09-19 20:21 Smarduch Mario-CMS063
@ 2005-09-19 20:47 ` Linus Torvalds
  0 siblings, 0 replies; 9+ messages in thread
From: Linus Torvalds @ 2005-09-19 20:47 UTC (permalink / raw)
  To: Smarduch Mario-CMS063; +Cc: Hugh Dickins, linux-kernel



On Mon, 19 Sep 2005, Smarduch Mario-CMS063 wrote:
>
> But I don't see the samething in 2.4.31 (which what we're using now).

I'm pretty sure 2.4.x should have the mmap_sem too.

But I think there it's aquired by the caller of dup_mmap() (ie it's 
"copy_mm()" that gets the mmap_sem and holds it for the duration of 
dup_mmap).

In fact, I'm sure: the mmap_sem was changed from a regular semaphore to a
read-write semaphore in 2.4.2.5, and we definitely took it around
dup_mmap() at that point. 

Of course, maybe some broken change removed it later in 2.4.x, I don't 
keep a current 2.4.x tree around.

		Linus

^ permalink raw reply	[flat|nested] 9+ messages in thread

* RE: Multi-Threaded fork() correctness on Linux 2.4 & 2.6
@ 2005-09-19 20:34 Smarduch Mario-CMS063
  0 siblings, 0 replies; 9+ messages in thread
From: Smarduch Mario-CMS063 @ 2005-09-19 20:34 UTC (permalink / raw)
  To: Hugh Dickins, Linus Torvalds; +Cc: linux-kernel

mmap_sem is also acquired in 2.4 properly. 
It seemed a little bit too obvious.
Thanks for your help!

	- mario

-----Original Message-----
From: Hugh Dickins [mailto:hugh@veritas.com] 
Sent: Monday, September 19, 2005 3:14 PM
To: Linus Torvalds
Cc: Smarduch Mario-CMS063; linux-kernel@vger.kernel.org
Subject: Re: Multi-Threaded fork() correctness on Linux 2.4 & 2.6

On Mon, 19 Sep 2005, Linus Torvalds wrote:
> 
> We hold the page_table_lock when doing the fork(), so T2 can't 
> actually be copying the page until we've done the TLB flush, no? And 
> once the TLB flush is done, all the writes by T3 should be in the 
> page, so we copy the right thing at that point, and there is no consistency problems?

I was totally overlooking the page_table_lock during the fork.

But no matter, it's not good enough: src_mm->page_table_lock is acquired and dropped at the inner level, in copy_pte_range (looking at latest 2.6):
it cannot be held across allocating page tables for dst_mm.

So each time T1 drops it, there's a window for the T2 vs. T3 problem.
Yet we don't much want to flush TLB each time we leave copy_pte_range.

Hugh

^ permalink raw reply	[flat|nested] 9+ messages in thread

* RE: Multi-Threaded fork() correctness on Linux 2.4 & 2.6
@ 2005-09-19 20:21 Smarduch Mario-CMS063
  2005-09-19 20:47 ` Linus Torvalds
  0 siblings, 1 reply; 9+ messages in thread
From: Smarduch Mario-CMS063 @ 2005-09-19 20:21 UTC (permalink / raw)
  To: Linus Torvalds, Hugh Dickins; +Cc: linux-kernel

On second look 'dup_mmap()' does a 'down_write()
on oldmm->mmap_sem and doesn't release it until
flush_tlb_mm() runs. Intermitent faults while
copy_pte_range() (taking releasing src page_table_lock)
runs should block in various platforms do_page_fault()
until oldmm->mmap_sem is released.
The eventual flush_tlb_mm() gets them to see
the same pte. So at no point can a threaded
task work off 2 different mappings.So I think this seems 
ok in 2.6. But I don't see the samething in 2.4.31 
(which what we're using now).

- Mario



-----Original Message-----
From: Linus Torvalds [mailto:torvalds@osdl.org] 
Sent: Monday, September 19, 2005 2:25 PM
To: Hugh Dickins
Cc: Smarduch Mario-CMS063; linux-kernel@vger.kernel.org
Subject: Re: Multi-Threaded fork() correctness on Linux 2.4 & 2.6



On Mon, 19 Sep 2005, Hugh Dickins wrote:
> On Mon, 19 Sep 2005, Smarduch Mario-CMS063 wrote:
> >     
> >     recently I've been involved in debugging MT forks on IA-64 the 
> > issues found were related to the way IA64 does its TLB invalidation. 
> > But there is still one  issue that appears to be Linux in general 
> > related, and it has to do with MT forks.
> >  
> > For example a process has 3 threads T1, T2, T3.
> >  
> > 1 - T1 issues a fork()
> >     - under page_table_lock write bit is reset in src and dest pte's
> > 2 - T2 may have a TLB mis, the new write protected pte is inserted
> >     (hw walker or sw tlb hdlr)
> > 3 - T2 winds up in do_wp_page() the page is copied to a new one
> > 4 - In the mean time T3 may be working off the same page, the
> >     TLB invalidation (flush_tlb_mm()) has no occurred yet.
> > 5 - Eventually TLB is globally flushed so threads will see the new pte.
> > 6 - As a result the MT task may experience inconsisitent state.
> >     - During 3 & 4. For example locks may be acquired by both
> >       threads depending on the timing of the copy.
> 
> I do think you've hit upon something interesting here.  Though it's 
> not quite as you describe.  We don't have to wait for the flush_tlb_mm 
> to sort it out: the flush_tlb_page in ptep_establish in break_cow in 
> do_wp_page resolves the discrepancy much sooner.  But your point is, 
> that's already too late: T2 inserts a "smudged" copy of the page which
> T3 was working on, it does not contain all the data T3 had written there.

Hmm. 

We hold the page_table_lock when doing the fork(), so T2 can't actually be copying the page until we've done the TLB flush, no? And once the TLB flush is done, all the writes by T3 should be in the page, so we copy the right thing at that point, and there is no consistency problems?

No? What am I missing?

		Linus

^ permalink raw reply	[flat|nested] 9+ messages in thread

end of thread, other threads:[~2005-09-19 21:31 UTC | newest]

Thread overview: 9+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2005-09-19 16:24 Multi-Threaded fork() correctness on Linux 2.4 & 2.6 Smarduch Mario-CMS063
2005-09-19 18:42 ` Hugh Dickins
2005-09-19 19:24   ` Linus Torvalds
2005-09-19 20:14     ` Hugh Dickins
2005-09-19 20:39       ` Linus Torvalds
2005-09-19 21:30         ` Hugh Dickins
2005-09-19 20:21 Smarduch Mario-CMS063
2005-09-19 20:47 ` Linus Torvalds
2005-09-19 20:34 Smarduch Mario-CMS063

This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®