From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S966492AbbDXIoK (ORCPT ); Fri, 24 Apr 2015 04:44:10 -0400 Received: from www.linutronix.de ([62.245.132.108]:40388 "EHLO Galois.linutronix.de" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S966254AbbDXIoG (ORCPT ); Fri, 24 Apr 2015 04:44:06 -0400 Date: Fri, 24 Apr 2015 10:44:05 +0200 (CEST) From: Thomas Gleixner To: Andi Kleen cc: Alexander Shishkin , Ingo Molnar , "H. Peter Anvin" , x86@kernel.org, Don Zickus , Frederic Weisbecker , Adrian Hunter , Anton Blanchard , Michael Ellerman , linux-kernel@vger.kernel.org, stable@vger.kernel.org Subject: Re: [PATCH v1] watchdog: Use a reference cycle counter to avoid scaling issues In-Reply-To: <20150424005133.GK13605@tassilo.jf.intel.com> Message-ID: References: <1429801408-11309-1-git-send-email-alexander.shishkin@linux.intel.com> <20150423203236.GJ13605@tassilo.jf.intel.com> <20150424005133.GK13605@tassilo.jf.intel.com> User-Agent: Alpine 2.11 (DEB 23 2013-08-11) MIME-Version: 1.0 Content-Type: TEXT/PLAIN; charset=US-ASCII X-Linutronix-Spam-Score: -1.0 X-Linutronix-Spam-Level: - X-Linutronix-Spam-Status: No , -1.0 points, 5.0 required, ALL_TRUSTED=-1,SHORTCIRCUIT=-0.0001 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Thu, 23 Apr 2015, Andi Kleen wrote: > > We can just detect the deviation in the callback itself: > > > > u64 now = ktime_get_mono_fast_ns(); > > > > if (now - __this_cpu_read(nmi_timestamp) < period) > > return; > > > > __this_cpu_write(nmi_timestamp, now); > > > > It's that simple. > > It's a simple short term hac^wsolution. Yes, and way simpler and less complex for pushing into stable. > But if we had a (hypothetical) system with let's say 10*TSC max you > may end up with quite a few false ticks, as in unnecessary > interrupts. With 100*TSC it would be really bad. And hypothetical systems with 100*TSC justify all that? > There were systems in the past that ran TSC at a much slower frequency, > such as the early AMD Barcelona systems. > > So the problem may eventually come back if not solved properly. There are better ways to do that than using heuristics. We have to deal with 3 variants of the reference counter: 1) Core and Atom: counts bus cycles and we know that frequency already from the local apic calibration 2) Nehalem, Westmere: Same as TSC 3) Sandybridge and later: XCLK which is 100MHz No magic calibration, just use the information which we have on our hands already. Thanks, tglx