From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1754168AbbJPNfR (ORCPT ); Fri, 16 Oct 2015 09:35:17 -0400 Received: from mga03.intel.com ([134.134.136.65]:47852 "EHLO mga03.intel.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S1751509AbbJPNfP (ORCPT ); Fri, 16 Oct 2015 09:35:15 -0400 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.17,689,1437462000"; d="scan'208";a="812593033" Date: Fri, 16 Oct 2015 06:35:14 -0700 From: Andi Kleen To: Peter Zijlstra Cc: Andi Kleen , linux-kernel@vger.kernel.org, Ingo Molnar , Thomas Gleixner Subject: Re: [PATCH 1/4] x86, perf: Use a new PMU ack sequence on Skylake Message-ID: <20151016133514.GB15102@tassilo.jf.intel.com> References: <1444952280-24184-1-git-send-email-andi@firstfloor.org> <1444952280-24184-2-git-send-email-andi@firstfloor.org> <20151016115107.GV3816@twins.programming.kicks-ass.net> MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20151016115107.GV3816@twins.programming.kicks-ass.net> User-Agent: Mutt/1.5.24 (2015-08-30) Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org > What would happen with a 'stuck' event in the new scheme? The usual reason events are stuck is that perf programs a too low period in frequency mode after fork. We have two safety mechanisms: - The perf overflow code checks against the max number of overflows per second, and forces a lower period if needed. - The NMI duration limiter checks the total length of NMI With the low period problem it will bump into the first limit. The second limit also makes sure not too much CPU time is lost. It wouldn't handle a case where the interrupt handler gets re-issued without overflowing, but I don't think that can really happen (except for the check point case, which is also covered) > > In principle the sequence should work on other CPUs too, but > > since I only tested on Skylake it is only enabled there. > > I would very much like a reduction of the ack states. You introduced the > late thing, which should also work for everyone, and now you introduce > yet another variant. Ingo suggested to do it this way. Originally I thought it wasn't needed, but I think now that late-ack made some of the races that eventually caused Skylake LBR to fall over worse. So in hindsight it was a good idea to not use it everywhere. > I would very much prefer a single ack scheme if at all possible. Could enable it everywhere, but then users would need to test it on most types of CPUs, as I can't. -Andi -- ak@linux.intel.com -- Speaking for myself only