mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Linus Torvalds <torvalds@linux-foundation.org>
To: Satyam Sharma <ssatyam@cse.iitk.ac.in>
Cc: Nick Piggin <nickpiggin@yahoo.com.au>,
	Linux Kernel Mailing List <linux-kernel@vger.kernel.org>,
	David Howells <dhowells@redhat.com>, Andi Kleen <ak@suse.de>,
	Andrew Morton <akpm@linux-foundation.org>
Subject: Re: [PATCH 8/8] i386: bitops: smp_mb__{before, after}_clear_bit() definitions
Date: Tue, 24 Jul 2007 10:12:50 -0700 (PDT)	[thread overview]
Message-ID: <alpine.LFD.0.999.0707240959331.3607@woody.linux-foundation.org> (raw)
In-Reply-To: <Pine.LNX.4.64.0707241708350.1433@cselinux1.cse.iitk.ac.in>



On Tue, 24 Jul 2007, Satyam Sharma wrote:
> 
> Looks like when you said "CPU memory barrier extends to all memory
> references" you were probably referring to a _given_ CPU ... yes,
> that statement is correct in that case.

No. CPU memory barriers extend to all CPU's. End of discussion.

It's not about "that cacheline". The whole *point* of a CPU memory barrier 
is that it's about independent memory accesses.

Yes, for a memory barrier to be effective, all CPU's involved in the 
transaction have to have the barriers - the same way a lock needs to be 
taken by everybody in order for it to make sense - but the point is, CPU 
barriers are about *global* behaviour, not local ones.

So there's a *huge* difference between

	clear_bit(x,y);

and

	clear_bit(x,y);
	smp_mb__before_after_clear_bit();

and it has absolutely nothing to do with the particular cacheline that "y" 
is in, it's about the *global* memory ordering.

Any write you do after that "smp_mb__before_after_clear_bit()" will be 
guaranteed to be visible to _other_ CPU's *after* they have seen the bit 
being cleared. Yes, those other CPU's need to have a read barrier between 
reading the bit and reading some other thign, but the point is, this hass 
*nothing* to do with cache coherency, and the particular cache line that 
"y" is in.

And no, "smp_mb__before/after_clear_bit()" must *not* be just an empty "do 
{} while (0)". It needs to be a compiler barrier even when it has no 
actual CPU meaning, unless clear_bit() itself is guaranteed to be a 
compiler barrier (which it isn't, although the "volatile" on the asm in 
practice makes it something *close* to that).

Why? Think of the sequence like this:

	clear_bit(x,y);
	smp_mb__after_clear_bit();
	other_variable = 10;

the whole *point* of this sequence is that if another CPU does

	x = other_variable;
	smp_rmb();
	bit = test_bit(x,y)

then if it sees "x" being 10, then the bit *has* to be clear.

And this is why the compiler barrier in "smp_mb__after_clear_bit()" needs 
to be a compiler barrier:

 - it doesn't matter for the action of the "clear_bit()" itself: that one 
   is locked, and on x86 it thus also happens to be a serializing 
   instruction, and the cache coherency and lock obviously means that the 
   bit clearing *itself* is safe!

 - but it *does* matter for the compiler scheduling. If the compiler were 
   to decide that "y" and "other_variable" are totally independent, it 
   might otherwise decide to move the "other_variable = 10" assignment to 
   *before* the clear_bit(), which would make the whole code pointless!

See? We have two totally independent issues:

 - the CPU itself can re-order the visibility of accesses. x86 doesn't do 
   this very much, and doesn't do it at all across a locked instruction, 
   but it's still a real issue, even if it tends to be much easier to see 
   on other architectures.

 - the compiler doesn't care about rules of "locked instruction" at all, 
   because it has no clue. It has *different* rules about how it can 
   re-order instructions and accesses, and maybe the "asm volatile" will 
   guarantee that the compiler won't re-order things around the 
   clear_bit(), and maybe it won't. But making it a compiler barrier (by 
   using the "memory clobber" thing, *guarantees* that gcc cannot reorder 
   memory writes or reads.

See? Two different - and _totally_ independent - levels of ordering, and 
we need to make sure that both are valid.

		Linus

  parent reply	other threads:[~2007-07-24 17:13 UTC|newest]

Thread overview: 92+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2007-07-23 16:05 [PATCH 0/8] i386: bitops: Cleanup, sanitize, optimize Satyam Sharma
2007-07-23 16:05 ` [PATCH 1/8] i386: bitops: Update/correct comments Satyam Sharma
2007-07-23 16:05 ` [PATCH 2/8] i386: bitops: Rectify bogus "Ir" constraints Satyam Sharma
2007-07-23 16:10   ` Andi Kleen
2007-07-23 16:21     ` Satyam Sharma
2007-07-23 16:30       ` Andi Kleen
2007-07-23 16:36         ` Jan Hubicka
2007-07-23 18:05         ` H. Peter Anvin
2007-07-23 18:28           ` H. Peter Anvin
2007-07-23 17:57   ` Linus Torvalds
2007-07-23 18:14     ` Satyam Sharma
2007-07-23 18:32       ` Andi Kleen
2007-07-23 18:39     ` H. Peter Anvin
2007-07-23 18:52       ` Satyam Sharma
2007-07-23 16:05 ` [PATCH 3/8] i386: bitops: Rectify bogus "+m" constraints Satyam Sharma
2007-07-23 16:37   ` Andi Kleen
2007-07-23 17:15     ` Satyam Sharma
2007-07-23 17:46   ` Linus Torvalds
2007-07-24  9:22   ` David Howells
2007-07-23 16:05 ` [PATCH 4/8] i386: bitops: Kill volatile-casting of memory addresses Satyam Sharma
2007-07-23 17:52   ` Linus Torvalds
2007-07-24  4:19     ` Nick Piggin
2007-07-24  6:23       ` Satyam Sharma
2007-07-24  7:16         ` Nick Piggin
2007-07-24  9:49     ` Benjamin Herrenschmidt
2007-07-24 17:20       ` Linus Torvalds
2007-07-24 17:39         ` Jeff Garzik
2007-07-25  4:54         ` Nick Piggin
2007-07-23 16:05 ` [PATCH 5/8] i386: bitops: Contain warnings fallout from the death of volatiles Satyam Sharma
2007-07-23 16:05 ` [PATCH 6/8] i386: bitops: Don't mark memory as clobbered unnecessarily Satyam Sharma
2007-07-23 16:13   ` Andi Kleen
2007-07-23 16:26     ` Satyam Sharma
2007-07-23 16:33       ` Andi Kleen
2007-07-23 17:12         ` Satyam Sharma
2007-07-23 17:49           ` Jeremy Fitzhardinge
2007-07-23 17:55   ` Linus Torvalds
2007-07-24  9:52     ` Benjamin Herrenschmidt
2007-07-24 17:24       ` Linus Torvalds
2007-07-24 17:42         ` Trond Myklebust
2007-07-24 18:13           ` Linus Torvalds
2007-07-24 18:28             ` Trond Myklebust
2007-07-24 21:37             ` Benjamin Herrenschmidt
2007-07-24 21:55               ` Trond Myklebust
2007-07-24 22:32                 ` Benjamin Herrenschmidt
2007-07-25  4:10                   ` Nick Piggin
2007-07-24 21:36         ` Benjamin Herrenschmidt
2007-07-24  3:57   ` Nick Piggin
2007-07-24  6:38     ` Satyam Sharma
2007-07-24  7:24       ` Nick Piggin
2007-07-24  8:29         ` Satyam Sharma
2007-07-24  8:39           ` Nick Piggin
2007-07-24  8:38         ` Trent Piepho
2007-07-24 19:39           ` Linus Torvalds
2007-07-24 20:37             ` Andi Kleen
2007-07-24 20:08               ` Linus Torvalds
2007-07-24 21:31                 ` Jeremy Fitzhardinge
2007-07-24 21:46                   ` Linus Torvalds
2007-07-26  1:07             ` Trent Piepho
2007-07-26  1:18               ` Linus Torvalds
2007-07-26  1:22                 ` Linus Torvalds
2007-07-24  9:44     ` David Howells
2007-07-24 10:02       ` Satyam Sharma
2007-07-23 16:06 ` [PATCH 7/8] i386: bitops: Kill needless usage of __asm__ __volatile__ Satyam Sharma
2007-07-23 16:18   ` Andi Kleen
2007-07-23 16:22     ` [PATCH 7/8] i386: bitops: Kill needless usage of __asm__ __volatile__ II Andi Kleen
2007-07-23 16:32     ` [PATCH 7/8] i386: bitops: Kill needless usage of __asm__ __volatile__ Satyam Sharma
2007-07-23 16:23   ` Jeremy Fitzhardinge
2007-07-23 16:43     ` Satyam Sharma
2007-07-23 17:39       ` Jeremy Fitzhardinge
2007-07-23 18:07         ` Satyam Sharma
2007-07-23 18:28           ` Jeremy Fitzhardinge
2007-07-23 20:29             ` Trent Piepho
2007-07-23 20:40               ` Jeremy Fitzhardinge
2007-07-23 21:06                 ` Trent Piepho
2007-07-23 21:30               ` Andi Kleen
2007-07-23 21:48                 ` Nicholas Miell
2007-07-23 16:06 ` [PATCH 8/8] i386: bitops: smp_mb__{before, after}_clear_bit() definitions Satyam Sharma
2007-07-24  3:53   ` Nick Piggin
2007-07-24  7:34     ` Satyam Sharma
2007-07-24  7:48       ` Jeremy Fitzhardinge
2007-07-24  8:31         ` Nick Piggin
2007-07-24  8:20       ` Nick Piggin
2007-07-24  9:21         ` Satyam Sharma
2007-07-24 10:25           ` Nick Piggin
2007-07-24 11:10             ` Satyam Sharma
2007-07-24 11:32               ` Nick Piggin
2007-07-24 11:45                 ` Satyam Sharma
2007-07-24 12:01                   ` Nick Piggin
2007-07-24 17:12                   ` Linus Torvalds [this message]
2007-07-24 19:01                     ` Satyam Sharma
2007-07-30 17:57 ` [PATCH 0/8] i386: bitops: Cleanup, sanitize, optimize Denis Vlasenko
2007-07-31  1:07   ` Satyam Sharma

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=alpine.LFD.0.999.0707240959331.3607@woody.linux-foundation.org \
    --to=torvalds@linux-foundation.org \
    --cc=ak@suse.de \
    --cc=akpm@linux-foundation.org \
    --cc=dhowells@redhat.com \
    --cc=linux-kernel@vger.kernel.org \
    --cc=nickpiggin@yahoo.com.au \
    --cc=ssatyam@cse.iitk.ac.in \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®