mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Nicolas Pitre <nico@cam.org>
To: Nick Piggin <nickpiggin@yahoo.com.au>
Cc: Ingo Molnar <mingo@elte.hu>,
	David Woodhouse <dwmw2@infradead.org>,
	Zwane Mwaikambo <zwane@arm.linux.org.uk>,
	lkml <linux-kernel@vger.kernel.org>,
	Linus Torvalds <torvalds@osdl.org>, Andrew Morton <akpm@osdl.org>,
	Arjan van de Ven <arjanv@infradead.org>,
	Steven Rostedt <rostedt@goodmis.org>,
	Alan Cox <alan@lxorguk.ukuu.org.uk>,
	Christoph Hellwig <hch@infradead.org>, Andi Kleen <ak@suse.de>,
	David Howells <dhowells@redhat.com>,
	Alexander Viro <viro@parcelfarce.linux.theplanet.co.uk>,
	Oleg Nesterov <oleg@tv-sign.ru>, Paul Jackson <pj@sgi.com>
Subject: Re: [patch 04/15] Generic Mutex Subsystem, add-atomic-call-func-x86_64.patch
Date: Tue, 20 Dec 2005 11:35:22 -0500 (EST)	[thread overview]
Message-ID: <Pine.LNX.4.64.0512201049310.26663@localhost.localdomain> (raw)
In-Reply-To: <43A81DD4.30906@yahoo.com.au>

On Wed, 21 Dec 2005, Nick Piggin wrote:

> Nicolas Pitre wrote:
> > On Wed, 21 Dec 2005, Nick Piggin wrote:
> > 
> > 
> > > Nicolas Pitre wrote:
> > > 
> > > > On Tue, 20 Dec 2005, Nick Piggin wrote:
> > > 
> > > > > Considering that on UP, the arm should not need to disable interrupts
> > > > > for this function (or has someone refuted Linus?), how about:
> > > > 
> > > > 
> > > > Kernel preemption.
> > > > 
> > > 
> > > preempt_disable() ?
> > 
> > 
> > Sure, and we're now more costly than the current implementation with irq
> > disabling.
> > 
> 
> Why? It is just a plain increment of a location that will almost certainly
> be in cache. I can't see how it would be more than half the cost of the
> irq implementation (based on looking at your measurements). How do you
> figure?

ARM is a load/store architecture.  So the preempt_count has to be 
loaded, incremented, and stored back (3 instructions just there).  Then 
the preempt_count itself isn't straight there, it is reached with the 
thread_info pointer which requires 2 additional instructions to 
generate.  So we're already up to 5 instructions while the disabling of 
interrupts takes 3.

OK now we want to decrement the semaphore count.  One load, one 
decrement, one store, i.e. 3 more instructions.  We're up to 8.

Oh and we're still supposed to be in a fast path.  And we even aren't 
ready to decide if we acquired the semaphore just yet.

Now we have to call preempt_enable().  Given that preempt_count() will 
use 2 instructions again to get to the thread_info pointer, let's say 
for the demonstration that we instead open coded our own 
preempt_enable() and preempt_disable() so to be able to cache the 
thread_info pointer amongst those two.  So let's forget about those 2 
additional instructions.

Let's also suppose that we preserved the original preempt count so that 
we don't have to re-load and decrement it which would have used 2 other 
additional instructions.  Good!  But register pressure is increasing, 
and btw we're using completely non generic kernel infrastructure at this 
point.

We nevertheless still have to store the original preempt count back.  
One instruction i.e. we're up to 9.

And with 9 instructions we're already worse than the implementation with 
interrupt disabling.  Maybe not in terms of cycles but hey we're not 
done yet!

Any preempt_enable() equivalent look at the TIF_NEED_RESCHED flag 
(luckily we still have the thread_info pointer handy since we're not 
using test_thread_flag() but loading the flag directly i.e. one 
instruction), testing it (one instruction), conditionally branching to 
preempt_schedule (one instruction).  Now up to 12 instructions.  Amd to 
make things even worse: since the flag has to be tested right after it 
has been loaded you must account for result delays (more cycles) for the 
load.

OK, did we acquire that damn semaphore or not?  Oh since we clobbered 
the processor's condition flag while testing the TIF_NEED_RESCHED flag 
we are now forced to test that semaphore count again to see if it went 
negative: one instruction.  And finally branch to the contention routine 
if the semaphore was already locked: one instruction.

So 14 instructions total with preemption disabling, and that's with the 
best implementation possible by open coding everything instead of 
relying on generic functions/macros.

Compare that to my 4 (or 3 when gcc can cse a constant) 
instructions needed for the mutex equivalent.


Nicolas

  reply	other threads:[~2005-12-20 16:36 UTC|newest]

Thread overview: 34+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2005-12-19  1:35 Ingo Molnar
2005-12-19 17:49 ` Zwane Mwaikambo
2005-12-19 20:58   ` David Woodhouse
2005-12-20  4:31     ` Ingo Molnar
2005-12-20  6:30       ` Nicolas Pitre
2005-12-20  8:12         ` Nick Piggin
2005-12-20 14:10           ` Nicolas Pitre
2005-12-20 14:12             ` Nick Piggin
2005-12-20 14:35               ` Nicolas Pitre
2005-12-20 15:05                 ` Nick Piggin
2005-12-20 16:35                   ` Nicolas Pitre [this message]
2005-12-20 19:20                     ` Russell King
2005-12-20 19:32                       ` Arjan van de Ven
2005-12-20 19:37                         ` Russell King
2005-12-20 19:59                         ` Nicolas Pitre
2005-12-20 19:43                       ` Steven Rostedt
2005-12-20 19:57                         ` Russell King
2005-12-20 20:35                         ` Horst von Brand
2005-12-20 19:55                       ` Nicolas Pitre
2005-12-20 18:27                 ` Linus Torvalds
2005-12-20 19:34                   ` Russell King
2005-12-20 20:10                     ` Linus Torvalds
2005-12-20 21:18                       ` Nicolas Pitre
2005-12-20 22:04                         ` Linus Torvalds
2005-12-20 22:19                           ` Steven Rostedt
2005-12-20 22:43                           ` Nicolas Pitre
2005-12-21  6:00                           ` Ingo Molnar
2005-12-21  2:12                         ` Nick Piggin
2005-12-20 19:49                   ` Nicolas Pitre
2005-12-20 18:25               ` Linus Torvalds
2005-12-21  6:20               ` Ingo Molnar
2005-12-21  4:16         ` Nicolas Pitre
2005-12-21  6:26           ` Ingo Molnar
2005-12-20  4:29   ` Ingo Molnar

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=Pine.LNX.4.64.0512201049310.26663@localhost.localdomain \
    --to=nico@cam.org \
    --cc=ak@suse.de \
    --cc=akpm@osdl.org \
    --cc=alan@lxorguk.ukuu.org.uk \
    --cc=arjanv@infradead.org \
    --cc=dhowells@redhat.com \
    --cc=dwmw2@infradead.org \
    --cc=hch@infradead.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=mingo@elte.hu \
    --cc=nickpiggin@yahoo.com.au \
    --cc=oleg@tv-sign.ru \
    --cc=pj@sgi.com \
    --cc=rostedt@goodmis.org \
    --cc=torvalds@osdl.org \
    --cc=viro@parcelfarce.linux.theplanet.co.uk \
    --cc=zwane@arm.linux.org.uk \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®