From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: X-Spam-Checker-Version: SpamAssassin 3.4.0 (2014-02-07) on aws-us-west-2-korg-lkml-1.web.codeaurora.org Received: from vger.kernel.org (vger.kernel.org [23.128.96.18]) by smtp.lore.kernel.org (Postfix) with ESMTP id 8345EC71153 for ; Sun, 10 Sep 2023 21:20:53 +0000 (UTC) Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S233417AbjIJVUz convert rfc822-to-8bit (ORCPT ); Sun, 10 Sep 2023 17:20:55 -0400 Received: from lindbergh.monkeyblade.net ([23.128.96.19]:57414 "EHLO lindbergh.monkeyblade.net" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S230254AbjIJVUy (ORCPT ); Sun, 10 Sep 2023 17:20:54 -0400 Received: from eu-smtp-delivery-151.mimecast.com (eu-smtp-delivery-151.mimecast.com [185.58.85.151]) by lindbergh.monkeyblade.net (Postfix) with ESMTPS id D3916188 for ; Sun, 10 Sep 2023 14:20:49 -0700 (PDT) Received: from AcuMS.aculab.com (156.67.243.121 [156.67.243.121]) by relay.mimecast.com with ESMTP with both STARTTLS and AUTH (version=TLSv1.2, cipher=TLS_ECDHE_RSA_WITH_AES_256_CBC_SHA384) id uk-mta-146-jIjh0yxkNzCFrfh3QFoU7A-1; Sun, 10 Sep 2023 22:20:42 +0100 X-MC-Unique: jIjh0yxkNzCFrfh3QFoU7A-1 Received: from AcuMS.Aculab.com (10.202.163.4) by AcuMS.aculab.com (10.202.163.4) with Microsoft SMTP Server (TLS) id 15.0.1497.48; Sun, 10 Sep 2023 22:20:33 +0100 Received: from AcuMS.Aculab.com ([::1]) by AcuMS.aculab.com ([::1]) with mapi id 15.00.1497.048; Sun, 10 Sep 2023 22:20:33 +0100 From: David Laight To: 'Charlie Jenkins' , Conor Dooley CC: Palmer Dabbelt , Samuel Holland , "linux-riscv@lists.infradead.org" , "linux-kernel@vger.kernel.org" , Paul Walmsley , Albert Ou Subject: RE: [PATCH v2 1/5] riscv: Checksum header Thread-Topic: [PATCH v2 1/5] riscv: Checksum header Thread-Index: AQHZ4bMJWeMEu2xwd0uaIDwvnyVCgbAUlGVw Date: Sun, 10 Sep 2023 21:20:33 +0000 Message-ID: References: <20230905-optimize_checksum-v2-0-ccd658db743b@rivosinc.com> <20230905-optimize_checksum-v2-1-ccd658db743b@rivosinc.com> <20230907-f8c8993dbeb24d5ea5310ec7@fedora> In-Reply-To: Accept-Language: en-GB, en-US X-MS-Has-Attach: X-MS-TNEF-Correlator: x-ms-exchange-transport-fromentityheader: Hosted x-originating-ip: [10.202.205.107] MIME-Version: 1.0 X-Mimecast-Spam-Score: 0 X-Mimecast-Originator: aculab.com Content-Language: en-US Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8BIT Precedence: bulk List-ID: X-Mailing-List: linux-kernel@vger.kernel.org ... > > > +/* > > > + * Fold a partial checksum without adding pseudo headers > > > + */ > > > +static inline __sum16 csum_fold(__wsum sum) > > > +{ > > > + sum += (sum >> 16) | (sum << 16); > > > + return (__force __sum16)(~(sum >> 16)); > > > +} I'm intrigued, gcc normally compiler that quite well. The very similar (from arch/arc): return (~sum - rol32(sum, 16)) >> 16; is slightly better on most architectures. (Especially if the ~sum and rol() can be executed together.) The only odd archs I saw were sparc32 (carry flag bug no rotate) and arm (barrel shifter on all instructions). It is better than the current asm for a lot of archs including x64. David - Registered Address Lakeside, Bramley Road, Mount Farm, Milton Keynes, MK1 1PT, UK Registration No: 1397386 (Wales)