From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752840AbbJ3Mk3 (ORCPT ); Fri, 30 Oct 2015 08:40:29 -0400 Received: from unicorn.mansr.com ([81.2.72.234]:33870 "EHLO unicorn.mansr.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1752319AbbJ3Mk2 convert rfc822-to-8bit (ORCPT ); Fri, 30 Oct 2015 08:40:28 -0400 From: =?iso-8859-1?Q?M=E5ns_Rullg=E5rd?= To: Nicolas Pitre Cc: Alexey Brodkin , linux-kernel@vger.kernel.org, linux-snps-arc@lists.infradead.org, Vineet Gupta , Ingo Molnar , Stephen Hemminger , "David S. Miller" , Russell King Subject: Re: [PATCH] __div64_32: implement division by multiplication for 32-bit arches References: <1446072455-16074-1-git-send-email-abrodkin@synopsys.com> Date: Fri, 30 Oct 2015 12:40:19 +0000 In-Reply-To: (Nicolas Pitre's message of "Thu, 29 Oct 2015 21:26:25 -0400 (EDT)") Message-ID: User-Agent: Gnus/5.13 (Gnus v5.13) Emacs/24.5 (gnu/linux) MIME-Version: 1.0 Content-Type: text/plain; charset=iso-8859-1 Content-Transfer-Encoding: 8BIT Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org Nicolas Pitre writes: > On Wed, 28 Oct 2015, Nicolas Pitre wrote: > >> On Thu, 29 Oct 2015, Alexey Brodkin wrote: >> >> > Fortunately we already have much better __div64_32() for 32-bit ARM. >> > There in case of division by constant preprocessor calculates so-called >> > "magic number" which is later used in multiplications instead of divisions. >> >> It's not magic, it is science. :-) >> >> > It's really nice and very optimal but obviously works only for ARM >> > because ARM assembly is involved. >> > >> > Now why don't we extend the same approach to all other 32-bit arches >> > with multiplication part implemented in pure C. With good compiler >> > resulting assembly will be quite close to manually written assembly. > > Well... not as close at least on ARM. Maybe 2x to 3x more costly than > the one with assembly. Still better than 100x or so without this > optimization. That's more or less what I found on MIPS too. >> > But there's at least 1 problem which I don't know how to solve. >> > Preprocessor magic only happens if __div64_32() is inlined (that's >> > obvious - preprocessor has to know if divider is constant or not). >> > >> > But __div64_32() is already marked as weak function (which in its turn >> > is required to allow some architectures to provide its own optimal >> > implementations). I.e. addition of "inline" for __div64_32() is not an >> > option. >> >> You can't inline __div64_32(). It should remain as is and used only for >> the slow path. >> >> For the constant based optimization to work, you need to modify do_div() >> in include/asm-generic/div64.h directly. > > OK... I was intrigued, so I adapted my ARM code to the generic case, > including the overflow avoidance optimizations. Please have look and > tell me how this works for you. > > If this patch is accepted upstream, then it could be possible to > abstract only the actual multiplication part with some architecture > specific assembly. Good idea. -- Måns Rullgård mans@mansr.com