From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from mail-pz2-f12.google.com (mail-pz2-f12.google.com [74.125.228.12]) (using TLSv1.2 with cipher ECDHE-RSA-AES128-GCM-SHA256 (128/128 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 0278438B7DD for ; Mon, 21 Sep 2026 20:36:07 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=74.125.228.12 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790022969; cv=none; b=s/wnxYxvslXT8NMW0mGs73qjO4DlFYk0wfnUqP56kVCyVrRJqbuUIn75HVHKjYJnaXEwHwLSj9lsmyFnT/m3JRyX05ecyWI0Im9hMcH4FLnaJYQtvIz02Hb4WRs0yhU89FiIG13uV0VEvd+zukIe0QKdyOoispn3WM5TVGKq8MI= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1790022969; c=relaxed/simple; bh=TBnKjZCAAvcwm8HIgliKTp3sPJIMdyVfOm0dVU9jd1A=; h=Date:From:To:Cc:Subject:Message-ID:References:MIME-Version: Content-Type:Content-Disposition:In-Reply-To; b=pJBfxho6716pqSQpsiSaU1dlYIlolol/YIU4KKljic83llKTMPhYa7I7tZgRZAxji4NXkM8l4aWUprCpX5XRmPXaWRs7MTr06yeX/QUIL6JH/9I1JyGMBIPrzB1JghqRnjOZG/Voc5NG3W5BxvdnPwxHtR5dBZ14z6mjMJ/1Z18= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com; spf=pass smtp.mailfrom=google.com; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b=q/iQ6tpZ; arc=none smtp.client-ip=74.125.228.12 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=google.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=google.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=google.com header.i=@google.com header.b="q/iQ6tpZ" Received: by mail-pz2-f12.google.com with SMTP id d2e1a72fcca58-8674704dab1so4696908b3a.2 for ; Mon, 21 Sep 2026 13:36:07 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20251104; t=1790022967; x=1790627767; darn=vger.kernel.org; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:from:to:cc:subject :date:message-id:reply-to:content-type; bh=groqaSpL753eNcDTGAY1Cn/tLSUz1r+DZjsuraXcs+A=; b=q/iQ6tpZ2FiuBXS4Blc2q7Kh2GRYBvLQv5QV1uh5B0sG6X3tMKIkM51tksz5kaXBms ZvkNKPs2FfTevrz29o8VGhVuymU4YQsmKyQu+ho/gIm/rXATdba8XgvRHejAcgxk7S// AJUtS1qeaLYR30yDaC5QS0+SkNY5etKh2EOQSyZqVJi1EhgtSBJmLJKTytnNjZ7qDs9F 8DpScHfLvwz013jF/L7KgJUZra924D0W/mJHgijyAvYkJve7Vf+1j7ti5pyysk+cPEWg Ps7L4wfP4rSOnIMn0m681ir/2k7BoDxJQJzirUwl7+LqQIvg6jx5Gvb3MXdvpyTXyBY4 LPqw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20260707; t=1790022967; x=1790627767; h=in-reply-to:content-disposition:content-type:mime-version :references:message-id:subject:cc:to:from:date:x-gm-gg :x-gm-message-state:from:to:cc:subject:date:message-id:reply-to :content-type; bh=groqaSpL753eNcDTGAY1Cn/tLSUz1r+DZjsuraXcs+A=; b=pmIvAlV4Iy21ehEg0yXQz7zAEkLHuNbNhNuHdUIiJWnDuuBEujtbcoHGHeLIreRrOh 5qtf5KXwkYd6EYYow0s4/pJvap9TxOaEVk8OIbkN5kQAYseZd0Xv1WKxdN2TpD3kABQu wph64d3phar7g+RYlWUjgknwY5vHiZjNJw8SJlLeg4wuVHEWi7h7rMokRTQKVqmDTS4q VdUdBFoI6AFyPIYhNYnvkBEh+njNcjPYZdzfAdnQCUDPe1c+UStAuG4simhBkwJpZ4NO ZT5SZalvfyWM2c8+A0tnM1qp15ysWdY+92L6HuASBe/yFxilpT4cc30snsGL8A2HeYV9 Oe/Q== X-Forwarded-Encrypted: i=1; AKwUvBxK3/ujx2ezRKRV3ZV3dK0eWxb6PIWlM6RsJeOtWsS7XoSb7hU6rWCE4XglxOYwF2WbWfq1ekuF/ZI989k=@vger.kernel.org X-Gm-Message-State: AFuF++k3YzakiSp/EqdK/UOLiSFclQqywtvqB2PF93UXwyXcFDiFogRR 8xz+W6QWK+c0jlTC4G2dcai9S+H4IDVSh8xK7AKGkVEZy+B3zJC+84EFqN4Uk+/OjQ== X-Gm-Gg: AYBFou0OYLXIYW3P4/koKNDTz7L8Kfy0qUIO/f+ZN0uN9VQN9jc+TyIYXE9FE1ixdnv IZUC3OP3xtfQg2jZppSyeemTbykxjoKD2dki7UWhaRB4YS1DlB9n49WqTAnJ4CWXa0nTSiVvG24 NeGcvoS+x5IpsvE4wr7dogZzD6ysVVqY9G76fCfiFT6HAWy/OeEiGTEeh8Ofviqf1TRG+NovotI jWP9/NgHIEExhVkLjUdZI51RzAYJpet1eXXnwaxRA1xZ07RgWalnLlyRG8Rd7XSe0AWQIyW021P euYM7nEeiAdQjXxMUdv3fjZBj8pgNiOjkDhhLMj0MiKUsMGaaXDC50v7Tnd+Jj0NoZjyrytdVsR 2uMYXZOeZPSjhCFyj2N6/4F2uHR4PUFd/QF0hbuxcJfgdSWEmVT9VsIYQpcVcktO/Yy2r28yrqc +Udr378/owuYqL4dwqxrRhuuAlO1pytr4YIi9kRWXrQe8LukcsW17cG8WiKSo04o5p5GbAv4G+X FvrIFW3fr8HRZlrNyscqQpSX8f/jVJqNRWek1np X-Received: by 2002:a05:6a00:2e89:b0:874:708d:b63a with SMTP id d2e1a72fcca58-874dea03034mr17531732b3a.27.1790022966708; Mon, 21 Sep 2026 13:36:06 -0700 (PDT) Received: from google.com (192.150.203.35.bc.googleusercontent.com. [35.203.150.192]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-87bf5785276sm45140b3a.8.2026.09.21.13.36.05 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Mon, 21 Sep 2026 13:36:05 -0700 (PDT) Date: Mon, 21 Sep 2026 20:36:01 +0000 From: David Matlack To: Alex Williamson Cc: kexec@lists.infradead.org, linux-doc@vger.kernel.org, linux-kernel@vger.kernel.org, linux-mm@kvack.org, linux-pci@vger.kernel.org, Adithya Jayachandran , Alexander Graf , Bjorn Helgaas , Chris Li , David Rientjes , Jacob Pan , Jason Gunthorpe , Jonathan Corbet , Josh Hilke , Leon Romanovsky , Lukas Wunner , Mike Rapoport , Parav Pandit , Pasha Tatashin , Pranjal Shrivastava , Pratyush Yadav , Randy Dunlap , Saeed Mahameed , Samiullah Khawaja , Shuah Khan , Vipin Sharma , William Tu , Yi Liu Subject: Re: [PATCH v9 08/13] PCI: Save and restore the ACS Control register Message-ID: References: <20260918200640.887030-1-dmatlack@google.com> <20260918200640.887030-9-dmatlack@google.com> <20260918191846.2f68b23b@shazbot.org> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=us-ascii Content-Disposition: inline In-Reply-To: <20260918191846.2f68b23b@shazbot.org> On 2026-09-18 07:18 PM, Alex Williamson wrote: > On Fri, 18 Sep 2026 20:06:34 +0000 > David Matlack wrote: > > > Save the ACS Control register in pci_save_state() and write it back > > in pci_restore_state(), instead of recomputing the ACS controls from > > scratch with pci_enable_acs(). > > > > This makes ACS symmetric with the rest of a device's saved state. Today > > pci_save_state() ignores ACS entirely and pci_restore_state() re-enables > > the ACS controls from the kernel's current ACS policy. As a result, a > > device can come out of a reset with different ACS controls than it went > > in with, e.g. any controls programmed outside of pci_enable_acs() are > > silently dropped. > > Controls, yes, but if we're restoring controls that might have EC set > now, should the Egress Control Vector also be part of the save state? > > Currently EC always gets cleared on reset and won't be restored by > pci_enable_acs(), so we can lose both EC and the EC vector. With this, > I think we restore EC but still lose the EC vector. Thanks, Hi Alex, The kernel does not enable Egress Control today or program the vector. Would this be to cover the case where firmware or userspace enabled it? Here is an updated patch to save/restore the vector, but I don't have any devices that support Egress Control on my normal testing system so I haven't been able to really test it yet. From: David Matlack Date: Fri, 11 Sep 2026 20:05:04 +0000 Subject: [PATCH] PCI: Save and restore the ACS Control register and Egress Control Vector Save the ACS Control register and the ACS Egress Control Vector in pci_save_state() and write them back in pci_restore_state(), instead of recomputing the ACS controls from scratch with pci_enable_acs(). This makes ACS symmetric with the rest of a device's saved state. Today pci_save_state() ignores ACS entirely and pci_restore_state() re-enables the ACS controls from the kernel's current ACS policy. As a result, a device can come out of a reset with different ACS controls than it went in with, e.g. any controls programmed outside of pci_enable_acs() are silently dropped. Notably, this prepares the kernel to be able to adopt the ACS controls established by a previous kernel across a Live Update rather than assigning new ones through pci_enable_acs(). Save and restore the Egress Control Vector as well to keep it in sync with the now-properly-restored Egress Control Enable bit in the ACS control register. The kernel never enables Egress Controls or programs the vector, but the firmware could have and they need to be kept in sync to avoid changing how P2P traffic is rounted. The size of the Egress Control Vector is not known when pci_allocate_cap_save_buffers() runs, as the ACS Capability register is only read later, in pci_acs_init(). Rather than move the allocation, size the save buffer for the largest vector a device can implement, which costs at most 32 bytes per ACS-capable device. pci_enable_acs() runs when a driver binds to a device (pci_dma_configure()), i.e. after pci_bus_add_device() has already saved the device's state. Refresh the saved ACS state there as well, otherwise a subsequent reset would revert ACS back to the configuration left behind by firmware. Devices that rely on device-specific quirks to enable an ACS equivalent keep that configuration outside of the ACS Control register, so keep configuring ACS from scratch for them. Do the same for devices that have no saved ACS state at all. Assisted-by: Claude:claude-opus-5 Signed-off-by: David Matlack --- drivers/pci/pci.c | 117 +++++++++++++++++++++++++++++++++- drivers/pci/pci.h | 5 ++ drivers/pci/quirks.c | 7 ++ include/uapi/linux/pci_regs.h | 1 + 4 files changed, 129 insertions(+), 1 deletion(-) diff --git a/drivers/pci/pci.c b/drivers/pci/pci.c index b2879a6be5f8..e1c05866e2ee 100644 --- a/drivers/pci/pci.c +++ b/drivers/pci/pci.c @@ -1021,6 +1021,106 @@ static void pci_std_enable_acs(struct pci_dev *dev, struct pci_acs *caps) caps->ctrl |= (dev->acs_capabilities & PCI_ACS_TB); } +/* + * Layout of the ACS save buffer. @ecv holds the ACS Egress Control Vector and + * is sized for the largest vector a device can implement. + */ +struct pci_acs_saved_state { + u16 ctrl; + u32 ecv[8]; +}; + +/* + * Return the size in bytes of the ACS Egress Control Vector, or 0 if the + * device does not implement P2P Egress Control. + */ +static unsigned int pci_acs_ecv_size(struct pci_dev *dev) +{ + unsigned int bits; + + if (!dev->acs_cap || !(dev->acs_capabilities & PCI_ACS_EC)) + return 0; + + /* An Egress Control Vector Size of 0 means 256 bits */ + bits = FIELD_GET(PCI_ACS_EGRESS_BITS_MASK, dev->acs_capabilities); + if (!bits) + bits = 256; + + /* The vector is implemented as a series of DWORD registers */ + return round_up(bits, 32) / 8; +} + +/** + * pci_save_acs_state - save the ACS Control register and Egress Control Vector + * @dev: the PCI device + * + * Record the ACS configuration currently programmed in hardware so that + * pci_restore_acs_state() can reapply it after a reset. + */ +static void pci_save_acs_state(struct pci_dev *dev) +{ + struct pci_cap_saved_state *save_state; + struct pci_acs_saved_state *acs; + unsigned int i, dwords; + + if (!dev->acs_cap) + return; + + save_state = pci_find_saved_ext_cap(dev, PCI_EXT_CAP_ID_ACS); + if (!save_state) + return; + + acs = (struct pci_acs_saved_state *)save_state->cap.data; + + pci_read_config_word(dev, dev->acs_cap + PCI_ACS_CTRL, &acs->ctrl); + + dwords = pci_acs_ecv_size(dev) / sizeof(u32); + for (i = 0; i < dwords; i++) + pci_read_config_dword(dev, dev->acs_cap + PCI_ACS_EGRESS_CTL_V + + i * sizeof(u32), &acs->ecv[i]); +} + +/** + * pci_restore_acs_state - restore the ACS Control register and Egress Control + * Vector + * @dev: the PCI device + */ +static void pci_restore_acs_state(struct pci_dev *dev) +{ + struct pci_cap_saved_state *save_state = NULL; + struct pci_acs_saved_state *acs; + unsigned int i, dwords; + + if (dev->acs_cap && !pci_need_dev_specific_enable_acs(dev)) + save_state = pci_find_saved_ext_cap(dev, PCI_EXT_CAP_ID_ACS); + + /* + * Devices that rely on device-specific quirks to enable an ACS + * equivalent keep that configuration outside of the ACS Control + * register, so there is nothing useful to restore for them. Configure + * ACS from scratch instead, which also covers devices that have no + * saved ACS state at all. + */ + if (!save_state) { + pci_enable_acs(dev); + return; + } + + acs = (struct pci_acs_saved_state *)save_state->cap.data; + + /* + * Restore the Egress Control Vector before the ACS Control register. + * The vector resets to zero, so enabling P2P Egress Control first + * would briefly apply the reset vector instead of the saved one. + */ + dwords = pci_acs_ecv_size(dev) / sizeof(u32); + for (i = 0; i < dwords; i++) + pci_write_config_dword(dev, dev->acs_cap + PCI_ACS_EGRESS_CTL_V + + i * sizeof(u32), acs->ecv[i]); + + pci_write_config_word(dev, dev->acs_cap + PCI_ACS_CTRL, acs->ctrl); +} + /** * pci_enable_acs - enable ACS if hardware support it * @dev: the PCI device @@ -1057,6 +1157,15 @@ void pci_enable_acs(struct pci_dev *dev) __pci_config_acs(dev, &caps, config_acs_param, 0, 0); pci_write_config_word(dev, pos + PCI_ACS_CTRL, caps.ctrl); + + /* + * pci_enable_acs() runs when a driver binds to the device, i.e. after + * pci_bus_add_device() has already saved the device's state. Refresh + * the saved ACS state so that a subsequent reset restores the + * configuration programmed here rather than the one left behind by + * firmware. + */ + pci_save_acs_state(dev); } /** @@ -1800,6 +1909,7 @@ int pci_save_state(struct pci_dev *dev) pci_save_aer_state(dev); pci_save_ptm_state(dev); pci_save_tph_state(dev); + pci_save_acs_state(dev); return pci_save_vc_state(dev); } EXPORT_SYMBOL(pci_save_state); @@ -1877,7 +1987,7 @@ void pci_restore_state(struct pci_dev *dev) pci_restore_msi_state(dev); /* Restore ACS and IOV configuration state */ - pci_enable_acs(dev); + pci_restore_acs_state(dev); pci_restore_iov_state(dev); dev->state_saved = false; @@ -3532,6 +3642,11 @@ void pci_allocate_cap_save_buffers(struct pci_dev *dev) if (error) pci_err(dev, "unable to allocate suspend buffer for LTR\n"); + error = pci_add_ext_cap_save_buffer(dev, PCI_EXT_CAP_ID_ACS, + sizeof(struct pci_acs_saved_state)); + if (error) + pci_err(dev, "unable to allocate suspend buffer for ACS\n"); + pci_allocate_vc_save_buffers(dev); } diff --git a/drivers/pci/pci.h b/drivers/pci/pci.h index ba3c3fddddc2..037c1674f164 100644 --- a/drivers/pci/pci.h +++ b/drivers/pci/pci.h @@ -1095,6 +1095,7 @@ void pci_acs_init(struct pci_dev *dev); void pci_enable_acs(struct pci_dev *dev); #ifdef CONFIG_PCI_QUIRKS int pci_dev_specific_acs_enabled(struct pci_dev *dev, u16 acs_flags); +bool pci_need_dev_specific_enable_acs(struct pci_dev *dev); int pci_dev_specific_enable_acs(struct pci_dev *dev); int pci_dev_specific_disable_acs_redir(struct pci_dev *dev); void pci_disable_broken_acs_cap(struct pci_dev *pdev); @@ -1105,6 +1106,10 @@ static inline int pci_dev_specific_acs_enabled(struct pci_dev *dev, { return -ENOTTY; } +static inline bool pci_need_dev_specific_enable_acs(struct pci_dev *dev) +{ + return false; +} static inline int pci_dev_specific_enable_acs(struct pci_dev *dev) { return -ENOTTY; diff --git a/drivers/pci/quirks.c b/drivers/pci/quirks.c index 7aee30734303..e500c202d2ec 100644 --- a/drivers/pci/quirks.c +++ b/drivers/pci/quirks.c @@ -5476,6 +5476,13 @@ static const struct pci_dev_acs_ops *pci_dev_acs_ops_get(struct pci_dev *dev) return NULL; } +bool pci_need_dev_specific_enable_acs(struct pci_dev *dev) +{ + const struct pci_dev_acs_ops *p = pci_dev_acs_ops_get(dev); + + return p && p->enable_acs; +} + int pci_dev_specific_enable_acs(struct pci_dev *dev) { const struct pci_dev_acs_ops *p = pci_dev_acs_ops_get(dev); diff --git a/include/uapi/linux/pci_regs.h b/include/uapi/linux/pci_regs.h index facaa324bd86..66359cb94f0d 100644 --- a/include/uapi/linux/pci_regs.h +++ b/include/uapi/linux/pci_regs.h @@ -1023,6 +1023,7 @@ #define PCI_ACS_UF 0x0010 /* Upstream Forwarding */ #define PCI_ACS_EC 0x0020 /* P2P Egress Control */ #define PCI_ACS_DT 0x0040 /* Direct Translated P2P */ +#define PCI_ACS_EGRESS_BITS_MASK 0xff00 /* Egress Control Vector Size */ #define PCI_ACS_EGRESS_BITS 0x05 /* ACS Egress Control Vector Size */ #define PCI_ACS_CTRL 0x06 /* ACS Control Register */ #define PCI_ACS_EGRESS_CTL_V 0x08 /* ACS Egress Control Vector */