From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mgamail.intel.com (mgamail.intel.com [198.175.65.20]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E74A3230BE9 for ; Thu, 6 Aug 2026 21:18:22 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=198.175.65.20 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786051104; cv=none; b=Xr4aF8FQ14hvcx577oUQZP3AXVI/SawdW7DjglLpFGMI2UOKid4Jd3hApfZVxQ9EVKIJButox4ewK57Gvr18CWOe4oLQ6si8pQjfdhE8dP6KoSXKg2B0C5EV996RHCaBFjHobe3wxtl9iHUiiBspwVzsYZuHQk4tqRfGtvMx73k= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1786051104; c=relaxed/simple; bh=B+3CZsZUdtppbI5i1SbA+c8iZ3CR4rp2wPMwmZUxf/4=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=dGcEjNX9moDdQp56Wgouw7EaujTfp/ybKtjz9dtl4bmvK+vNF+rGwGqnMmJGQH3aeN9r5dN58xsd3tZO9zzF9KSCMBi3sxTHob9cUvF5vVPVL3UPCknGwlE+pn2GfSeYkBXNPkQop/H25/lWohmPMb29E91cQKUhJr17Fmx5ioM= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com; spf=pass smtp.mailfrom=intel.com; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b=ECZEYp1T; arc=none smtp.client-ip=198.175.65.20 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=intel.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=intel.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=intel.com header.i=@intel.com header.b="ECZEYp1T" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/simple; d=intel.com; i=@intel.com; q=dns/txt; s=Intel; t=1786051103; x=1817587103; h=date:from:to:cc:subject:message-id:references: mime-version:content-transfer-encoding:in-reply-to; bh=B+3CZsZUdtppbI5i1SbA+c8iZ3CR4rp2wPMwmZUxf/4=; b=ECZEYp1TbCEe4st3Ukg427bdf66dC1ZcelnXLAmbPqxq/kY0wQ2t3xR5 MhuaKA+XJg7RydSnhHcVldNGE5haSh297oXy8RwXX2w6ola63imrkgnnH M6pvRCoKqhWBcH/WOK08NCnVRpQYy3rvMLMiu3vBsx63o1qjVZ6f9iyMo xnUnwf/LlamG18CJgZ6C3kibtpq5h7/nVkW8PaqmKQoWCsTu69mwJm0k5 dUzIG1gTcpE3B80yYdPPQIee912EamLdbSnIdJFZOqz1EZ3qsXpZEQdBE LuLomKuXuUqTbEFyP2L1oXQtxG1TVbtGVKuIJq+QN2dGUdOpTjJ+kzz2S Q==; X-CSE-ConnectionGUID: bc01cUgzQGKx0eGhGn0zQQ== X-CSE-MsgGUID: TmLEnW3IRJ+Ie3uUzAH8Fw== X-IronPort-AV: E=McAfee;i="6800,10657,11867"; a="86425323" X-IronPort-AV: E=Sophos;i="6.25,209,1779174000"; d="scan'208";a="86425323" Received: from orviesa008.jf.intel.com ([10.64.159.148]) by orvoesa112.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 06 Aug 2026 14:18:22 -0700 X-CSE-ConnectionGUID: EQQ1tZQaRjOGm1+ipb2BEA== X-CSE-MsgGUID: EAhyCzXqQv+Tl6JaEY0Q6w== X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="6.25,209,1779174000"; d="scan'208";a="261697084" Received: from ettammin-mobl3.ger.corp.intel.com (HELO localhost) ([10.245.245.50]) by orviesa008-auth.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 06 Aug 2026 14:18:19 -0700 Date: Fri, 7 Aug 2026 00:18:16 +0300 From: Andy Shevchenko To: Fengnan Chang Cc: linux-nvme@lists.infradead.org, Keith Busch , Jens Axboe , Christoph Hellwig , Sagi Grimberg , Bart Van Assche , Thomas Gleixner , Jun Zeng , Gang Cao , Jun I Jin , Liang A Fang , Yong Hu , linux-kernel@vger.kernel.org, Guzebing Subject: Re: [RFC PATCH] nvme-pci: adaptively poll completions on busy queues Message-ID: References: <20260723120556.58670-1-changfengnan@bytedance.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Disposition: inline Content-Transfer-Encoding: 8bit In-Reply-To: <20260723120556.58670-1-changfengnan@bytedance.com> Organization: Intel Finland Oy - BIC 0357606-4 - c/o Alberga Business Park, 6 krs, Bertel Jungin Aukio 5, 02600 Espoo On Thu, Jul 23, 2026 at 08:05:56PM +0800, Fengnan Chang wrote: > In high-IOPS scenarios, relying on interrupts to handle I/O operations > can limit performance. This issue becomes particularly pronounced in > multi-disk environments, where performance is constrained by the CPU’s > interrupt-handling capacity. > > Each Solidigm SB5PH27X038T device used for testing can deliver about 3.2M > 4 KiB random-read IOPS. Four devices therefore have about 12.8M IOPS of > aggregate capability, but interrupt-driven completion topped out at 7.81M > IOPS, or about 61% of that capability. > > Add optional adaptive polling for busy interrupt-driven I/O queues. After > an interrupt drains the CQ, use the ring distance between last_sq_tail and > cq_head as an approximate host-side inflight count. If the count reaches > a per-controller high watermark, let the threaded handler keep draining > completions instead of returning to interrupt-driven completion right away. > > The thread always checks for pending CQEs first. When the CQ is empty, it > compares the inflight count with a per-controller low watermark and sleeps > for a configurable interval before checking again. It returns to > interrupt-driven completion after three consecutive empty-CQ checks below > the low watermark or when the loop limit is reached. > > The mode is disabled by default. It can be enabled at module load time. > The high and low watermarks, empty-CQ interval, and loop limit are > available through sysfs. The defaults are 32 commands, 3 commands, > 30 us, and 1000 iterations. > > Tested with the default settings and 4 KiB random reads on four Solidigm > SB5PH27X038T devices. With adaptive polling off and on: > > - t/io_uring, depth 1024 and four workers per device: > 7.810M and 12.780M aggregate IOPS (+63.64%). > - fio/libaio, iodepth 1024, numjobs 4 and completion batches of 32: > 7.699M and 12.785M aggregate IOPS (+66.05%). > > In a paired libaio and io_uring boundary sweep, QD8, QD16, QD32, and > QD16/numjobs=4 changed by -0.15% to +0.48%. QD64 improved by 28.61% to > 59.75%. ... > +What: /sys/class/nvme/nvmeX/adaptive_poll_inflight_high > +Date: July 2026 Can't be July for v7.3. See crystal ball predictor for the dates (rc1 or release). > +KernelVersion: 7.3 ... > +static void nvme_adaptive_unmask_irq(struct nvme_queue *nvmeq, int irq) > +{ > + struct nvme_dev *dev = nvmeq->dev; > + struct pci_dev *pdev = to_pci_dev(dev->dev); > + > + clear_bit(NVMEQ_ADAPTIVE_POLLING, &nvmeq->flags); > + if (pdev->msi_enabled) { Wondering if MSI-X also should be considered here. With that in mind, can we use pci_dev_msi_enabled()? > + writel(BIT(nvmeq->cq_vector), dev->bar + NVME_REG_INTMC); > + readl(dev->bar + NVME_REG_INTMS); > + } else { > + enable_irq(irq); > + } > +} ... > +static ssize_t adaptive_poll_interval_us_store(struct device *dev, > + struct device_attribute *attr, > + const char *buf, size_t count) > +{ > + struct nvme_dev *ndev = to_nvme_dev(dev_get_drvdata(dev)); > + unsigned int value; > + int ret; > + > + ret = kstrtouint(buf, 10, &value); > + if (ret) > + return ret; > + if (!value) > + return -EINVAL; ERANGE? EDOM? Ditto for other similar cases. > + WRITE_ONCE(ndev->adaptive_poll_interval_us, value); > + return count; > +} -- With Best Regards, Andy Shevchenko