From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from smtp.kernel.org (aws-us-west-2-korg-mail-1.web.codeaurora.org [10.30.226.201]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 78C6680B for ; Fri, 1 Aug 2025 12:01:43 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=10.30.226.201 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1754049703; cv=none; b=FXEFdXKQ0lGkeo+kUOoYmQyQfUGrVuG86k4a1eB3MI4I+fuwLzW+kbsnAooCQD6mX0o3Yna+34cB4chfp5ew8ihZeyCu6BJEPaKq1cg5vIz0OYcv6zRqeG9CJ9z+3mPoCQF+NtuA6lrg4n0JhgFgJUV5M0UuTas1hZuOuNsuPiM= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1754049703; c=relaxed/simple; bh=uOoLORZbVz/2DjHrvQwdE/xKoSWkkYOsM2fYvuYSxFs=; h=Date:Message-ID:From:To:Cc:Subject:In-Reply-To:References: MIME-Version:Content-Type; b=o8x9Jge56O+O7oop20bC1JLjU7dNL0MXKBHxjoHgPg8RRn2wT7TEDDn+axCbO1DkL5+TeonCTiXkCDUx0I9rg1IJVUECAnI6Dqg4uMdG0rx7VZrzNhCi0+Lsz0kRrwojSVtLiQlJip3Z1fcl1Nq8a8RGIibwq1foHO6aGADWqxs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b=h+/JIR9n; arc=none smtp.client-ip=10.30.226.201 Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=kernel.org header.i=@kernel.org header.b="h+/JIR9n" Received: by smtp.kernel.org (Postfix) with ESMTPSA id EF2F4C4CEE7; Fri, 1 Aug 2025 12:01:42 +0000 (UTC) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=kernel.org; s=k20201202; t=1754049703; bh=uOoLORZbVz/2DjHrvQwdE/xKoSWkkYOsM2fYvuYSxFs=; h=Date:From:To:Cc:Subject:In-Reply-To:References:From; b=h+/JIR9nibIna+akEcLThAyq6TgJbtwZzNpTbQWAatG3VAMzBqM3DKsQp6hY2ycFo 81JXKMc/UEQQVZOUC2uzJ+XuHJjO/2Ia9P4fsBUKScsKVtGNLcoBA5txTHyTc65CmY bIA7Xf9OkC5Xr/bYvGx2kRj+W1jdWOqyhBqa3zNGbDlXuUFGTx2JgCEQUrsJj8Rj/1 lqjMt5lHDaCERJFddCo1ocjpq2aHU9TklTLEAmxwzNmCRKAycow3Gn7WjFbeuMA0My H64+VHCaL/v9r75YkjCVZMBn1hZ/QnfuUAccGGDJ3ZZHE+NIaYPL64bFcLTq+AD7RR VkKizrKCltwZQ== Received: from sofa.misterjones.org ([185.219.108.64] helo=goblin-girl.misterjones.org) by disco-boy.misterjones.org with esmtpsa (TLS1.3) tls TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384 (Exim 4.95) (envelope-from ) id 1uhoSB-003Bli-P7; Fri, 01 Aug 2025 13:01:40 +0100 Date: Fri, 01 Aug 2025 13:01:38 +0100 Message-ID: <86y0s36yjh.wl-maz@kernel.org> From: Marc Zyngier To: wangwudi Cc: Thomas Gleixner , , , , , Zenghui Yu Subject: Re: Question on the scheduling of timer interrupt and FIO interrupt In-Reply-To: <8c6eb963-0a3a-8b75-8ab4-a0b2e10f3d40@hisilicon.com> References: <8c6eb963-0a3a-8b75-8ab4-a0b2e10f3d40@hisilicon.com> User-Agent: Wanderlust/2.15.9 (Almost Unreal) SEMI-EPG/1.14.7 (Harue) FLIM-LB/1.14.9 (=?UTF-8?B?R29qxY0=?=) APEL-LB/10.8 EasyPG/1.0.0 Emacs/30.1 (aarch64-unknown-linux-gnu) MULE/6.0 (HANACHIRUSATO) Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 (generated by SEMI-EPG 1.14.7 - "Harue") Content-Type: text/plain; charset=US-ASCII X-SA-Exim-Connect-IP: 185.219.108.64 X-SA-Exim-Rcpt-To: wangwudi@hisilicon.com, tglx@linutronix.de, linux-arm-kernel@lists.infradead.org, linux-kernel@vger.kernel.org, yangwei24@huawei.com, yaohongshi@hisilicon.com, zenghui.yu@linux.dev X-SA-Exim-Mail-From: maz@kernel.org X-SA-Exim-Scanned: No (on disco-boy.misterjones.org); SAEximRunCond expanded to false + Zenghui, in case he has seen this before. On Fri, 01 Aug 2025 07:26:20 +0100, wangwudi wrote: > > Hi, all > When running some FIO tests on ARM64 server(Kunpeng), frequent NVMe interrupts occupy the > CPU, and the CPU's hardirq load is 100%. The watchdog feed interrupt arch_timer cannot be > responded, triggering the hardlockup. I am extremely surprised that even with a screaming NVMe (or even several of them), you end up in a situation where you don't have the resource to take the timer interrupt. > > GIC driver uses GICV3_PRIO_IRQ to set the same priority for arch_timer interrupt and NVMe > interrupt. In GIC spec, "If, on a particular CPU interface, multiple pending interrupts > have the same priority, and have sufficient priority for the interface to signal them to > the PE, it is IMPLEMENTATION DEFINED how the interface selects which interrupt to signal." > Shell we consider setting a higher priority for the arch_timer interrupt to fix this case? Linux only deals with two priorities: the normal interrupt priority, and NMI, where the NMI can preempt any other interrupt. obviously, we don't want to make the timer an NMI, as it would break a lot of things. Which means that even if you were to give the timer a higher priority, it should not be allowed to preempt any other interrupt. Which means that you'd need to set the binary point so that both the NVMe and timer priorities fall into the same preemption bucket. But it also means that you now are eating into the few bits of priority that we have, and that will cause problems with the NMI priority. Also, how to you decide what interrupts should be of a higher priority? I find it surprising that your GIC doesn't have some form of round-robin scheme to pick the next HPPI, because that's clearly a fairness problem, and punting that on SW is pretty ugly. Thanks, M. -- Without deviation from the norm, progress is not possible.