From mboxrd@z Thu Jan 1 00:00:00 1970 Return-Path: DMARC-Filter: OpenDMARC Filter v1.3.2 smtp.codeaurora.org 068C7601D2 Authentication-Results: pdx-caf-mail.web.codeaurora.org; dmarc=none (p=none dis=none) header.from=arm.com Authentication-Results: pdx-caf-mail.web.codeaurora.org; spf=none smtp.mailfrom=linux-kernel-owner@vger.kernel.org Received: (majordomo@vger.kernel.org) by vger.kernel.org via listexpand id S932378AbeFFJNv (ORCPT + 25 others); Wed, 6 Jun 2018 05:13:51 -0400 Received: from foss.arm.com ([217.140.101.70]:38226 "EHLO foss.arm.com" rhost-flags-OK-OK-OK-OK) by vger.kernel.org with ESMTP id S932265AbeFFJNt (ORCPT ); Wed, 6 Jun 2018 05:13:49 -0400 Date: Wed, 06 Jun 2018 10:13:44 +0100 Message-ID: <86a7s89t13.wl-marc.zyngier@arm.com> From: Marc Zyngier To: Yang Yingliang Cc: , , Subject: Re: [PATCH v2] irqchip/gic-v3-its: fix ITS queue timeout In-Reply-To: <1528252824-15144-1-git-send-email-yangyingliang@huawei.com> References: <1528252824-15144-1-git-send-email-yangyingliang@huawei.com> User-Agent: Wanderlust/2.15.9 (Almost Unreal) SEMI-EPG/1.14.7 (Harue) FLIM/1.14.9 (=?UTF-8?B?R29qxY0=?=) APEL/10.8 EasyPG/1.0.0 Emacs/25.1 (aarch64-unknown-linux-gnu) MULE/6.0 (HANACHIRUSATO) Organization: ARM Ltd MIME-Version: 1.0 (generated by SEMI-EPG 1.14.7 - "Harue") Content-Type: text/plain; charset=US-ASCII Sender: linux-kernel-owner@vger.kernel.org List-ID: X-Mailing-List: linux-kernel@vger.kernel.org On Wed, 06 Jun 2018 03:40:24 +0100, Yang Yingliang wrote: [I'm travelling, so please do not expect any quick answer...] > > When the kernel booted with maxcpus=x, 'x' is smaller > than actual cpu numbers, the TAs of offline cpus won't TA? Target Address? Target Affinity? Timing Advance? Terrible Acronym? > be set to its->collection. > > If LPI is bind to offline cpu, sync cmd will use zero TA, > it leads to ITS queue timeout. Fix this by choosing a > online cpu, if there is no online cpu in cpu_mask. So instead of fixing the emission of a sync command on a non-mapped collection, you hack set_affinity? It doesn't feel like the right thing to do. It is also worth noticing that mapping an LPI to a collection that is not mapped yet is perfectly legal. > Signed-off-by: Yang Yingliang > --- > drivers/irqchip/irq-gic-v3-its.c | 9 +++++++-- > 1 file changed, 7 insertions(+), 2 deletions(-) > > diff --git a/drivers/irqchip/irq-gic-v3-its.c b/drivers/irqchip/irq-gic-v3-its.c > index 5416f2b..d8b9539 100644 > --- a/drivers/irqchip/irq-gic-v3-its.c > +++ b/drivers/irqchip/irq-gic-v3-its.c > @@ -2309,7 +2309,9 @@ static int its_irq_domain_activate(struct irq_domain *domain, > cpu_mask = cpumask_of_node(its_dev->its->numa_node); > > /* Bind the LPI to the first possible CPU */ > - cpu = cpumask_first(cpu_mask); > + cpu = cpumask_first_and(cpu_mask, cpu_online_mask); > + if (cpu >= nr_cpu_ids) > + cpu = cpumask_first(cpu_online_mask); Now you're completely ignoring cpu_mask which constraints the NUMA affinity. On some systems, this ends up with a deadlock (Cavium TX1, if I remember well). Wouldn't it be better to just return that the affinity setting request is impossible to satisfy? And more to the point, how comes we end-up in such a case? Thanks, M. -- Jazz is not dead, it just smell funny.