On Fri, Apr 12, 2013 at 09:55:56AM +0200, Peter Zijlstra wrote: > > The above is totally untested, but each step is pretty damn simple and > > fairly cheap. Sure, it's a loop, but it's bounded to 32 (cheap) > > iterations, and the normal case is that it's not done at all, or done > > only a few times. > > Right it gets gradually heavier the bigger the numbers get; which is > more and more unlikely. > > > And the advantage is that the end result is always that simple > > 32x32/32 case that we started out with as the common case. > > > > I dunno. Maybe I'm overlooking something, and the above is horrible, > > but the above seems reasonably efficient if not optimal, and > > *understandable*. > > I suppose that entirely matters on what one is used to ;-) I had to > stare rather hard at it for a little while. > > But yes, you take it one step further and are willing to ditch rtime > bits too and I suppose that's fine. > > Should work,.. Stanislaw could you stick this into your userspace > thingy and verify the numbers are sane enough? It works fine - gives relative error less than 0.1% for very big numbers. For the record I'm attaching test program and script. Thanks Stanislaw