From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S1947786AbcBRUQZ (ORCPT ); Thu, 18 Feb 2016 15:16:25 -0500 Received: from www.linutronix.de ([62.245.132.108]:51712 "EHLO Galois.linutronix.de" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S965344AbcBRUQY (ORCPT ); Thu, 18 Feb 2016 15:16:24 -0500 Date: Thu, 18 Feb 2016 21:15:05 +0100 (CET) From: Thomas Gleixner To: Vikas Shivappa cc: LKML , Peter Zijlstra , x86@kernel.org, Stephane Eranian , Matt Fleming Subject: Re: [PATCH] x86/perf/intel/cqm: Get rid of the silly for_each_cpu lookups In-Reply-To: Message-ID: References: User-Agent: Alpine 2.11 (DEB 23 2013-08-11) MIME-Version: 1.0 Content-Type: TEXT/PLAIN; charset=US-ASCII X-Linutronix-Spam-Score: -1.0 X-Linutronix-Spam-Level: - X-Linutronix-Spam-Status: No , -1.0 points, 5.0 required, ALL_TRUSTED=-1,SHORTCIRCUIT=-0.0001 Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Wed, 17 Feb 2016, Thomas Gleixner wrote: > On Wed, 17 Feb 2016, Vikas Shivappa wrote: > > Please stop top posting, finally! > > > But we have an extra static - static to avoid having it in the stack.. > > It's not about the cpu mask on the stack. The reason was that with cpumask off > stack cpumask_and_mask() requires an allocation, which then can't be used in > the starting/dying callbacks. > > Darn, you are right to remind me. > > Now, the proper solution for this stuff is to provide a library function as we > need that for several drivers. No point to duplicate that functionality. I'll > cook something up and repost the uncore/cqm set tomorrow. Second thoughts on that. cpumask_any_but() is fine as is, if we feed it topology_core_cpumask(cpu). The worst case search is two bitmap_find_next() if the first search returned cpu. Now cpumask_any_and() does a search as well, but the number of bitmap_find_next() invocations is limited to the number of sockets if we feed the cqm_cpu_mask as first argument. So for 4 or 8 sockets that's still a reasonable limit. If the people with insane large machines care, we can revisit that topic. It's still faster than for_each_online_cpu() :) Thanks, tglx