From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1752446AbcAGKSs (ORCPT ); Thu, 7 Jan 2016 05:18:48 -0500 Received: from mail.fireflyinternet.com ([87.106.93.118]:60985 "EHLO fireflyinternet.com" rhost-flags-OK-OK-OK-FAIL) by vger.kernel.org with ESMTP id S1750809AbcAGKSo (ORCPT ); Thu, 7 Jan 2016 05:18:44 -0500 X-Default-Received-SPF: pass (skip=forwardok (res=PASS)) x-ip-name=78.156.65.138; Date: Thu, 7 Jan 2016 10:16:52 +0000 From: Chris Wilson To: linux-kernel@vger.kernel.org Cc: Ross Zwisler , "H . Peter Anvin" , Andy Lutomirski , Borislav Petkov , Brian Gerst , Denys Vlasenko , "H. Peter Anvin" , Linus Torvalds , Thomas Gleixner , Imre Deak , Daniel Vetter , dri-devel@lists.freedesktop.org Subject: Re: [PATCH] x86: Add an explicit barrier() to clflushopt() Message-ID: <20160107101652.GF652@nuc-i3427.alporthouse.com> Mail-Followup-To: Chris Wilson , linux-kernel@vger.kernel.org, Ross Zwisler , "H . Peter Anvin" , Andy Lutomirski , Borislav Petkov , Brian Gerst , Denys Vlasenko , "H. Peter Anvin" , Linus Torvalds , Thomas Gleixner , Imre Deak , Daniel Vetter , dri-devel@lists.freedesktop.org References: <1445248735-11915-1-git-send-email-chris@chris-wilson.co.uk> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <1445248735-11915-1-git-send-email-chris@chris-wilson.co.uk> User-Agent: Mutt/1.5.21 (2010-09-15) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Mon, Oct 19, 2015 at 10:58:55AM +0100, Chris Wilson wrote: > During testing we observed that the last cacheline was not being flushed > from a > > mb() > for (addr = addr & -clflush_size; addr < end; addr += clflush_size) > clflushopt(); > mb() > > loop (where the initial addr and end were not cacheline aligned). > > Changing the loop from addr < end to addr <= end, or replacing the > clflushopt() with clflush() both fixed the testcase. Hinting that GCC > was miscompling the assembly within the loop and specifically the > alternative within clflushopt() was confusing the loop optimizer. > > Adding a barrier() into clflushopt() is enough for GCC to dtrt, but > solving why GCC is not seeing the constraints from the alternative_io() > would be smarter... > > Bugzilla: https://bugs.freedesktop.org/show_bug.cgi?id=92501 > Testcase: gem_tiled_partial_pwrite_pread/read > Signed-off-by: Chris Wilson > Cc: Ross Zwisler > Cc: H. Peter Anvin > Cc: Imre Deak > Cc: Daniel Vetter > Cc: dri-devel@lists.freedesktop.org > --- > arch/x86/include/asm/special_insns.h | 5 +++++ > 1 file changed, 5 insertions(+) > > diff --git a/arch/x86/include/asm/special_insns.h b/arch/x86/include/asm/special_insns.h > index 2270e41b32fd..0c7aedbf8930 100644 > --- a/arch/x86/include/asm/special_insns.h > +++ b/arch/x86/include/asm/special_insns.h > @@ -199,6 +199,11 @@ static inline void clflushopt(volatile void *__p) > ".byte 0x66; clflush %P0", > X86_FEATURE_CLFLUSHOPT, > "+m" (*(volatile char __force *)__p)); > + /* GCC (4.9.1 and 5.2.1 at least) appears to be very confused when > + * meeting this alternative() and demonstrably miscompiles loops > + * iterating over clflushopts. > + */ > + barrier(); > } Or an alternative: +#define alternative_output(oldinstr, newinstr, feature, output) \ + asm volatile (ALTERNATIVE(oldinstr, newinstr, feature) \ + : output : "i" (0) : "memory") I would really appreciate some knowledgeable folks taking a look at the asm for clflushopt() as it still affects today's kernel and gcc. Fwiw, I have confirmed that arch/x86/mm/pageattr.c clflush_cache_range() is similarly affected. -Chris -- Chris Wilson, Intel Open Source Technology Centre