From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id CEAAA37BE7D for ; Thu, 4 Jun 2026 10:45:09 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780569911; cv=none; b=N9bEuy5sCSIvy5X0ebJ09M3IOq0g4i2QBuMfXvkry21EijZ8SV5InGtnN2voP24wfK4RB8R3dHQBBIwAGrIUSJqxaMnycmrzvXoX8qvdMa2iRXv/pu5rHJPFarfnvj3Tc/S+CX9Bm77YGhsiM5ZqZF3mgIsbMjF3msHG9GIBwjE= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1780569911; c=relaxed/simple; bh=pXOC3ieQ42cIu2RuSkFAP20upsmn2zTEUHKrLTllEys=; h=Message-ID:Date:MIME-Version:Subject:To:Cc:References:From: In-Reply-To:Content-Type; b=LGSyC81cOkmk51/Te4DXcr1S9i9Sz8owNCXnEzPGrowhG9peF0l9lXRmSg1BYKjdZC70WPaYdHi5m2euMOuN0dxFz7hnMzPga9SuQ0N/uqO9lWz3oEYD3Uj8RxGcaIyrUJBdAl97eL98mZ+alH1Wp0OyAjnIe4Pb1Q8CdApJZPs= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=B82TNq9C; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b=RYDH67KF; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="B82TNq9C"; dkim=pass (2048-bit key) header.d=redhat.com header.i=@redhat.com header.b="RYDH67KF" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1780569909; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=wbsaCU5Vs3pFhS9fw5gs64Im6dVDjZtVePKSfJVpfmE=; b=B82TNq9C7CU9wajSDKTsX9f6HT6Sd8VD3QhviMzJjQhy4tEOM3lgvatP8/cxpvkgzBpB5e I5b5KGewjnC7zpBVEIy54vg8QIP3yNZStLzqKk3UZYZtjHJzs83kWvwvY8mSaVBGMjm4Dp UKdrTLnsblnNee2D/oobRt9y0kIsK1M= Received: from mail-wr1-f70.google.com (mail-wr1-f70.google.com [209.85.221.70]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-342-ZvDNbBEpPeyPUFlxUirrdg-1; Thu, 04 Jun 2026 06:45:07 -0400 X-MC-Unique: ZvDNbBEpPeyPUFlxUirrdg-1 X-Mimecast-MFC-AGG-ID: ZvDNbBEpPeyPUFlxUirrdg_1780569907 Received: by mail-wr1-f70.google.com with SMTP id ffacd0b85a97d-45ef6b407b4so257968f8f.1 for ; Thu, 04 Jun 2026 03:45:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=google; t=1780569906; x=1781174706; darn=vger.kernel.org; h=content-transfer-encoding:in-reply-to:content-language:from :references:cc:to:subject:user-agent:mime-version:date:message-id :from:to:cc:subject:date:message-id:reply-to; bh=wbsaCU5Vs3pFhS9fw5gs64Im6dVDjZtVePKSfJVpfmE=; b=RYDH67KFRkeI/Rx5xKDrJSt6mX+ruPR8jRHHwDxrvbsv3zwHhNVS1KF7m5RuupNqqO IBl2YLRfYKYZwMXgy5vKESkAQKOnYVxn46BvsrmkdIIjGykX1vCcXvibSsmLRLcoOIav KfLNzZcQ9q/cb0zc86dOwXtVNseqyiiaHxviFmIv74nZS2MSXdcECSm7kSDD290Oa2d2 tuKsuqJUkRC68OOTxZN+QC8Lr2rWzXWvDWyo/ML/Emn+3dX1itBjC/y26ODHNY1gWdiU 0rNtx81VP8HILA7XMF8m0MODyPRhdkNfE038/3BofRYgHArZ3rJgyy4mfW/Qp60v7pZW G05A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20251104; t=1780569906; x=1781174706; h=content-transfer-encoding:in-reply-to:content-language:from :references:cc:to:subject:user-agent:mime-version:date:message-id :x-gm-gg:x-gm-message-state:from:to:cc:subject:date:message-id :reply-to; bh=wbsaCU5Vs3pFhS9fw5gs64Im6dVDjZtVePKSfJVpfmE=; b=eAiQKqbNBh4mCSc4AxJejLlOG7ATNWPHnEawAzrgn6FKfhS13KogM51qjsl/mjoquS ADQOzS+4vXsvHbV7TAqTdVuLkG7Ayut6lui6PBRjIcPlOhyGEszfjMuwEDqu+7PdIa7t TSZ8uCEGaxZzIX1lxu98u8csHM9+pdbbqosEg9AFTdu9e2W8WBkQbjkrNtAEyWCS360L 9vnxpfGZvO3kHSj8+T8sXYSGik6tCJy9c/i4XNxTMgf0keUjm1dpymsoB0VkigH613xl t92n5xGaFvwPz/vFYhV+wtwS4gUlb4SJ7Ew6h/xI0bPY8dmSVdYD1/VCvN5fc1QEiBqV pKTQ== X-Forwarded-Encrypted: i=1; AFNElJ/f3S03MJVUBn0TmQzbIiGYBFMQ5y9szeyzCPFI6RVkpryursLYXmtgDj0/T617JFJeueHO7HSfXPHgh7w=@vger.kernel.org X-Gm-Message-State: AOJu0Yyk1+2n3SJFytteIg/jVT2Gf7LC4TohzQV0vkqhYrtACZH0tJ7j kw7Z5d5EdlP2QeBF+8DOkGLd1BSjP1blzYRgWAw1unt9+XI/yeydwXAwaqedoX2IaY/G4EUGUWX 7xh8HntgjQY5vl5CWrQpCikvduxWLe5UugosHlzT3Qf+9C4PeOtxPfX8xZjHEOjh3cQ== X-Gm-Gg: Acq92OGdOW2GwTDxMLyMvsz5DArnKVMFQZyR3iGIg71MCRH63KhQHsJWy8n1wkMq6JB mDNqpKC57FvKDmgT+zOgumFh6+nra83BvWh1VpvfibnRXKyfna3Ul7/Y4BwcWnQFUiW73AlEyh4 N44tMrJYzg1xr9GoN4+RUKcFbv9URB7cso8gKyk9vpWG27OOWFkH85jm2bCB7p0Bi4/Z3jl3PWD Bnk1eXgDzSwtusJN1iqGE6FCmb7c/xoMYYYPc3ekhgnarpr6exk6hd3k/RotaqE8TmY6g9B0uxx t5bSPhdmY9cE4D2zFo/G1iF6bZr1XY64gOu7iVAMTLAyQ8/FtlIpK1x8ETnFZ0Y9ZQmDN89jezS u41RwWZEzUxZk3ywOzclJLjSmTxVmouMT11oKDFR/tbkDkoKoUVqvR5Ur+ztpauRU5VQ= X-Received: by 2002:a5d:630a:0:b0:45e:e44b:3147 with SMTP id ffacd0b85a97d-4602178318cmr7176629f8f.5.1780569906581; Thu, 04 Jun 2026 03:45:06 -0700 (PDT) X-Received: by 2002:a5d:630a:0:b0:45e:e44b:3147 with SMTP id ffacd0b85a97d-4602178318cmr7176572f8f.5.1780569906004; Thu, 04 Jun 2026 03:45:06 -0700 (PDT) Received: from [192.168.88.32] ([212.105.155.59]) by smtp.gmail.com with ESMTPSA id ffacd0b85a97d-4601f2dcad5sm16588064f8f.5.2026.06.04.03.45.04 (version=TLS1_3 cipher=TLS_AES_128_GCM_SHA256 bits=128/128); Thu, 04 Jun 2026 03:45:05 -0700 (PDT) Message-ID: <6d1fa9d9-73c2-48b5-95a1-51710d81b3ed@redhat.com> Date: Thu, 4 Jun 2026 12:45:03 +0200 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 User-Agent: Mozilla Thunderbird Subject: Re: [PATCH net v3] net: mana: Optimize irq affinity for low vcpu configs To: Shradha Gupta , Dexuan Cui , Wei Liu , Haiyang Zhang , "K. Y. Srinivasan" , Andrew Lunn , "David S. Miller" , Eric Dumazet , Jakub Kicinski , Konstantin Taranov , Simon Horman , Erni Sri Satya Vennela , Dipayaan Roy , Shiraz Saleem , Michael Kelley , Long Li , Yury Norov Cc: linux-hyperv@vger.kernel.org, linux-kernel@vger.kernel.org, netdev@vger.kernel.org, Paul Rosswurm , Shradha Gupta , Saurabh Singh Sengar , stable@vger.kernel.org References: <20260601102749.1768304-1-shradhagupta@linux.microsoft.com> From: Paolo Abeni Content-Language: en-US In-Reply-To: <20260601102749.1768304-1-shradhagupta@linux.microsoft.com> Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 7bit On 6/1/26 12:27 PM, Shradha Gupta wrote: > In mana driver, the number of IRQs allocated is capped by the > min(num_cpu + 1, queue count). In cases, where the IRQ count is greater > than the vcpu count, we want to utilize all the vCPUs, irrespective of > their NUMA/core bindings. > > This is important, especially in the envs where number of vCPUs are so > few that the softIRQ handling overhead on two IRQs on the same vCPU is > much more than their overheads if they were spread across sibling vCPUs. > > This behaviour is more evident with dynamic IRQ allocation. Since MANA > IRQs are assigned at a later stage compared to static allocation, other > device IRQs may already be affinitized to the vCPUs. As a result, IRQ > weights become imbalanced, causing multiple MANA IRQs to land on the > same vCPU, while some vCPUs have none. > > In such cases when many parallel TCP connections are tested, the > throughput drops significantly. > > Test envs: > ======================================================= > Case 1: without this patch > ======================================================= > 4 vcpu(2 cores), 5 MANA IRQs (1 HWC + 4 Queue) > > TYPE effective vCPU aff > ======================================================= > IRQ0: HWC 0 > IRQ1: mana_q1 0 > IRQ2: mana_q2 2 > IRQ3: mana_q3 0 > IRQ4: mana_q4 3 > > %soft on each vCPU(mpstat -P ALL 1) on receiver > vCPU 0 1 2 3 > ======================================================= > pass 1: 38.85 0.03 24.89 24.65 > pass 2: 39.15 0.03 24.57 25.28 > pass 3: 40.36 0.03 23.20 23.17 > > ======================================================= > Case 2: with this patch > ======================================================= > 4 vcpu(2 cores), 5 MANA IRQs (1 HWC + 4 Queue) > > TYPE effective vCPU aff > ======================================================= > IRQ0: HWC 0 > IRQ1: mana_q1 0 > IRQ2: mana_q2 1 > IRQ3: mana_q3 2 > IRQ4: mana_q4 3 > > %soft on each vCPU(mpstat -P ALL 1) on receiver > vCPU 0 1 2 3 > ======================================================= > pass 1: 15.42 15.85 14.99 14.51 > pass 2: 15.53 15.94 15.81 15.93 > pass 3: 16.41 16.35 16.40 16.36 > > ======================================================= > Throughput Impact(in Gbps, same env) > ======================================================= > TCP conn with patch w/o patch > 20480 15.65 7.73 > 10240 15.63 8.93 > 8192 15.64 9.69 > 6144 15.64 13.16 > 4096 15.69 15.75 > 2048 15.69 15.83 > 1024 15.71 15.28 > > Fixes: 755391121038 ("net: mana: Allocate MSI-X vectors dynamically") > Cc: stable@vger.kernel.org > Co-developed-by: Erni Sri Satya Vennela > Signed-off-by: Erni Sri Satya Vennela > Signed-off-by: Shradha Gupta > Reviewed-by: Haiyang Zhang > Reviewed-by: Simon Horman Why do you consider this patch a fix? To me is a configuration improvement and should land on net-next. > @@ -1717,11 +1719,24 @@ static int irq_setup(unsigned int *irqs, unsigned int len, int node, > return 0; > } > > +/* should be called with cpus_read_lock() held */ Minor nit: s/should/must/ or just drop the comment, as `for_each_online_cpu()` usage implies that. > +static void irq_setup_linear(unsigned int *irqs, unsigned int len) > +{ > + int cpu; > + > + for_each_online_cpu(cpu) { > + if (len == 0) > + break; > + > + irq_set_affinity_and_hint(*irqs++, cpumask_of(cpu)); > + len--; > + } As this is another heuristic regarding irq spreading, why don't you implement that inside irq_setup()? > @@ -1767,13 +1784,42 @@ static int mana_gd_setup_dyn_irqs(struct pci_dev *pdev, int nvec) > * first CPU sibling group since they are already affinitized to HWC IRQ > */ > cpus_read_lock(); > - if (gc->num_msix_usable <= num_online_cpus()) > - skip_first_cpu = true; > + if (gc->num_msix_usable <= num_online_cpus()) { > + err = irq_setup(irqs, nvec, gc->numa_node, true); > + if (err) { > + cpus_read_unlock(); > + goto free_irq; > + } > + } else { > + /* > + * When num_msix_usable are more than num_online_cpus, our > + * queue IRQs should be equal to num of online vCPUs. > + * We try to make sure queue IRQs spread across all vCPUs. > + * In such a case NUMA or CPU core affinity does not matter. > + * Note: in this case the total mana IRQ should always be > + * num_online_cpus + 1. The first HWC IRQ is already handled > + * in HWC setup calls > + * However, if CPUs went offline since num_msix_usable was > + * computed, queue IRQs will be more than num_online_cpus(). > + * In such cases remaining extra IRQs will retain their default > + * affinity. > + */ > + int first_unassigned = num_online_cpus(); > + if (nvec > first_unassigned) { An empty line is needed between the variable declaration and the code. /P