From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from CY7PR03CU001.outbound.protection.outlook.com (mail-westcentralusazon11010056.outbound.protection.outlook.com [40.93.198.56]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id A8C64370AD7 for ; Thu, 2 Jul 2026 19:28:58 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.93.198.56 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783020540; cv=fail; b=pXfqFI7lfV7P1PI+cj1Zrw3sjepftGd+rUiT0RsrI7/ZB3gx2Uf9IghFD7lsoKZchrWJpxH3NqwHZ6PkZVL1W9N4nLWjLB24S4qgq787ifbf4OJHqccKHY4+DvUe14ri3m9gzj8Qj8m3/5/1zWBPD3S5KdGO4SFr8jvPQ8/hX/c= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783020540; c=relaxed/simple; bh=VdNBvHm+8D5lFkuRv5CCh5zUoWr56sz7yZ3BbXx9Foo=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=SATtIdAbuaV6lG293pv+NBBXz0I4XcWQCELyDOCyapkjPIa7tz9vsRwYejYpW4rSlo8iH0L2pfymKkEwV8jA3ExzjjJRpEbWxTLNfEyGNrOUT/ERu96RI5LNNSOQjzPBdhLKEW2kbbhiH7LKbk8Nu66Su2UFRiVB9uhAWCyJdJI= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com; spf=fail smtp.mailfrom=nvidia.com; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b=lst6xbVT; arc=fail smtp.client-ip=40.93.198.56 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=nvidia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b="lst6xbVT" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=qeWA4x41hOZzsmYHEEMd2t3N0+4PXBDqDbyLfFS/M5D7qOqhMMSWuukKqN/EAchuHJ5NS8YX2zAM39U/7cNN9NTjUL3iLgE/YqXcxBxUhfJOuqgBtIelqVHzP3bdcWuVYqSJX8r2omsZHvpMY+VYLFLaAlQZaVPvn3SZif5ETNhtILZIKQxGACQIigFOpS6DtR46yKUSKVqRailszr/SyKe0kqTYg72OcV6svBlsj1yfoih793iVqI2BqpEGq5WCAJDpmVQ8J5CPHs93CbD8nmyTAQ/BkWgJowIDpBWpUMRggZvK3wWwYt2Hl2vi8Ei2cPwYjvy/8eXFSpBastk1Pg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=Y/sftroTW9MxmeowOr/XNmGyPDAzONhvdM4DerF9H1w=; b=o/g7GWDpoiMpHmDed1f8l6hcGguVLNFwDrYQ0dDAYibE3wxVqHj8dueFfJ2DgZZjrEozHigRI/TUdJPHFaNRHuJcwxf7nhngmJDT6sF621BwCXiwoKCA6ElFJSvfHR8talMEpybRvVtyUqsDGFOfSUdoRGovLnFhFzKt1s8ANYk/JN4CUBWSfR20n7Egfv2QPtruJbvHsXW6IB5OdD6YbduyIor61l4A/xL87XLo7E6XHV0WN1p7FoTOkhjw7BIDJ63gvihvTA44Pnli/FFRbWIqmL/rrNOJEYPzmMbm2m3F5t9657hExJhiZF7d+lSMXAEKq/bAthhKE1NkC2Aa7A== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.117.160) smtp.rcpttodomain=linuxfoundation.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=Y/sftroTW9MxmeowOr/XNmGyPDAzONhvdM4DerF9H1w=; b=lst6xbVTFGcduhH90NZO7Fg6mTzq2i3AHUGCcrk1H8JGg2zfpnWQMejRKYtyW6nREtVuAjHyAdgXtTlND+F61C5jrp42/8xEIORqDrV4Cq/ne0HrpZcE/3eE0BpKMsbKP2HuYdz2cOLRst4QRJhLVjTH3PjavNjMjQV3U/s7qB2Hp/SCOtlF5O97lb/ANuQ42hY6bGlHV9y+OHTzlCwlMoh82jgSbhIRdoAPVr0BxQqH48CuB8fXanNscpeb1XL3CtGuttT5ngHLlHAJ5cUFzlCUkMs0J2+dp/kmocL+NJa2pkjlWJliliNZWuL5aKsEyQRH92lIWeIrmV+l/hRdIA== Received: from MW4PR04CA0228.namprd04.prod.outlook.com (2603:10b6:303:87::23) by CY1PR12MB9627.namprd12.prod.outlook.com (2603:10b6:930:104::9) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.8; Thu, 2 Jul 2026 19:26:00 +0000 Received: from SJ1PEPF000026C9.namprd04.prod.outlook.com (2603:10b6:303:87:cafe::58) by MW4PR04CA0228.outlook.office365.com (2603:10b6:303:87::23) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.181.10 via Frontend Transport; Thu, 2 Jul 2026 19:26:00 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 216.228.117.160) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.117.160 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.117.160; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.117.160) by SJ1PEPF000026C9.mail.protection.outlook.com (10.167.244.106) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.6 via Frontend Transport; Thu, 2 Jul 2026 19:26:00 +0000 Received: from rnnvmail202.nvidia.com (10.129.68.7) by mail.nvidia.com (10.129.200.66) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Thu, 2 Jul 2026 12:25:36 -0700 Received: from rnnvmail203.nvidia.com (10.129.68.9) by rnnvmail202.nvidia.com (10.129.68.7) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Thu, 2 Jul 2026 12:25:35 -0700 Received: from localhost.nvidia.com (10.127.8.12) by mail.nvidia.com (10.129.68.9) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20 via Frontend Transport; Thu, 2 Jul 2026 12:25:35 -0700 From: Ankit Agrawal To: , , CC: , , , , , , , , Subject: [PATCH 4/4] platform/nvidia: Register EGM PFNMAP range with memory_failure Date: Thu, 2 Jul 2026 19:25:32 +0000 Message-ID: <20260702192532.455400-5-ankita@nvidia.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260702192532.455400-1-ankita@nvidia.com> References: <20260702192532.455400-1-ankita@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Content-Type: text/plain X-NV-OnPremToCloud: ExternallySecured X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SJ1PEPF000026C9:EE_|CY1PR12MB9627:EE_ X-MS-Office365-Filtering-Correlation-Id: 8166fe90-8c93-4bb4-8c17-08ded86fbfbb X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|36860700016|1800799024|376014|82310400026|23010399003|18002099003|5023799004|22082099003|11063799006|56012099006|6133799003; X-Microsoft-Antispam-Message-Info: GWOp0/xqGzErFQkwi4Cc0s9ODx96mp7Ci3dm00d5pKpUvBAWwynhqdG7nmcvHoacO2CuEva1dxgUzjq4qZ6V+bm6+KEFAb9xwWxNi+939rihB8G1Ct/ejXdfOLODrEao9Zj6B65zg+JRr9zqxdYCcj/HvWfojvjANm9OVc/2eeRKEEdr0MKN7nItLCCEWL9Y19Unx/h1+1W1L4JABuX+5BzNVHmapShuSqx6dPg+cl3eCvytl2kmz8ZbzVRhcdSoAzLGD3A8ugrv5U/3z9nRXx93M5LZ/CDlV6i9ResaXXEB3ILS+KNkgo8MaIqwv5uDamH0neRwMp3y8Y9dLhte944+57bHS1RpZDfX5wnZuXifHj4DXcrqId5a6x/IP5RH/0CXSULqodjRTbexW5IN7DHGMGvlxsU0hYry2c+NXyHvy1whKnuhEywESgaofcYu+Lq12CSMuqXwMIqukW+5NjTkjMPWQ0BHbKfkAvp0l+4p5vv2YLq6Y0f+HuHi+PlBlgvSS3d+/6xIRR+3UPglubUhrbXpBpKOsFiiKSUcckVt42wuexZ45CbB/65HnoNtgFu46JVl+ZTMxNLA6bk+6miOh8BPxB5fmAfap39CA10Mx5G4/zYHNvaVpEofvZvdcGVvI/5S5hemMHJJamhRghUZK2BqbVkBdulCN+QphZeLZCp0gHNpkk3Qh7Ic6VhCSHAm54sH5Bp+TKeCv1EdHg== X-Forefront-Antispam-Report: CIP:216.228.117.160;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:mail.nvidia.com;PTR:dc6edge1.nvidia.com;CAT:NONE;SFS:(13230040)(36860700016)(1800799024)(376014)(82310400026)(23010399003)(18002099003)(5023799004)(22082099003)(11063799006)(56012099006)(6133799003);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: cRVYffwpxyAyy23IT3857hG0+O+/VdXGxD7qSL5QVh4faWK7jLM0wDdcwxRyZ1CeL9L65yO5OumU5jhTKHDhlqih711LwBXF+nDDd59nDqjuXkWPN4r0WR50TH7ntrQmsWXYiMaqfNBFXcgF9MvyC9CRFl0wEY6meV4rYIvhwTk/U1vvvmdCyPY0eUVy9Q2zPq8HT3ensUzVSQgKQbZQ5JN29L1VLJ3KV0IMRZOH/yQ17BbIcSzRLtlN4uNq1WP8/jfR6zeE7u0YxdoO12novn2b9pPa8SxoV2+DiOlWkUy53DLVJ2mnnVwpnpQZmtsAz2KpkSOcMwckXQCg3KhJlr5Ir29P+uJ8ELqQ30ay0TeJeFrxXBDC1svehFXom3zkGIpWSOqrxbYqnuNJiNkko4hI/RITrqN7IACZH58gClIgM3jSx9S4dvJhzLBgVhOh X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 02 Jul 2026 19:26:00.1542 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: 8166fe90-8c93-4bb4-8c17-08ded86fbfbb X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a;Ip=[216.228.117.160];Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: SJ1PEPF000026C9.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: CY1PR12MB9627 EGM carveout memory is mapped directly into userspace (QEMU) and is not added to the kernel. It is not managed by the kernel page allocator and has no struct pages. The module can thus utilize the Linux memory manager's memory_failure mechanism for regions with no struct pages. The Linux MM code exposes register/unregister APIs allowing modules to register such memory regions for memory_failure handling. On the occurrence of the memory_failure, Linux MM calls the registered function if the PFN is within the range. These dynamically occurring memory_failure pages are tracked alongside the statically populated list by the system firmware in the xarray. So the userspace can get both the information through the common EGM_RETIRED_PAGES_LIST ioctl. Register the EGM PFN range with the memory_failure infrastructure on first open() and unregister it on the last close(). Provide a pfn_to_vma_pgoff callback that: - Validates that the reported PFN falls within the EGM region and the current VMA. - Converts the PFN to a file offset within the EGM region. - Records the poisoned offset in the existing retired-pages hashtable so it is reported to userspace via EGM_RETIRED_PAGES_LIST alongside statically configured retired pages from SBIOS. Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Ankit Agrawal --- drivers/platform/nvidia/egm.c | 123 +++++++++++++++++++++++++++++++--- 1 file changed, 115 insertions(+), 8 deletions(-) diff --git a/drivers/platform/nvidia/egm.c b/drivers/platform/nvidia/egm.c index df963caedaa8..39c4be768f75 100644 --- a/drivers/platform/nvidia/egm.c +++ b/drivers/platform/nvidia/egm.c @@ -3,6 +3,7 @@ * Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved */ +#include #include #include #include @@ -40,7 +41,7 @@ struct gpu_node { struct nvgrace_egm_dev { struct device device; struct cdev cdev; - /* serialises the first-open scrub */ + /* serialises the first-open scrub + PFN registration */ struct mutex open_lock; /* protected by open_lock */ unsigned int open_count; @@ -49,6 +50,7 @@ struct nvgrace_egm_dev { * EGM region and keyed by page index. */ struct xarray retired_pages; + struct pfn_address_space pfn_address_space; phys_addr_t egmphys; size_t egmlength; phys_addr_t retiredpagesphys; @@ -71,6 +73,104 @@ static LIST_HEAD(egm_chardevs); static void cleanup_retired_pages(struct nvgrace_egm_dev *egm_dev); static int nvgrace_egm_fetch_retired_pages(struct nvgrace_egm_dev *egm_dev); +static int pfn_memregion_offset(struct nvgrace_egm_dev *egm_dev, + unsigned long pfn, + pgoff_t *pfn_offset_in_region) +{ + unsigned long start_pfn, num_pages; + + start_pfn = PHYS_PFN(egm_dev->egmphys); + num_pages = egm_dev->egmlength >> PAGE_SHIFT; + + if (pfn < start_pfn || pfn >= start_pfn + num_pages) + return -EFAULT; + + *pfn_offset_in_region = pfn - start_pfn; + + return 0; +} + +static int track_ecc_offset(struct nvgrace_egm_dev *egm_dev, + unsigned long mem_offset) +{ + void *old; + + /* + * This runs from the pfn_to_vma_pgoff callback, which the + * memory_failure path invokes under rcu_read_lock() (see + * collect_procs_pfn()). Sleeping is not allowed in that atomic + * context and so use GFP_ATOMIC rather than GFP_KERNEL. + * + * Storing by page index makes a repeated report at the same offset a + * harmless overwrite, so no explicit de-duplication is needed. + */ + old = xa_store(&egm_dev->retired_pages, mem_offset >> PAGE_SHIFT, + EGM_RETIRED_MARK, GFP_ATOMIC); + if (xa_is_err(old)) + return xa_err(old); + + return 0; +} + +static int nvgrace_egm_pfn_to_vma_pgoff(struct vm_area_struct *vma, + unsigned long pfn, + pgoff_t *pgoff) +{ + struct nvgrace_egm_dev *egm_dev = vma->vm_file->private_data; + pgoff_t vma_offset_in_region = vma->vm_pgoff; + pgoff_t pfn_offset_in_region; + int ret; + + ret = pfn_memregion_offset(egm_dev, pfn, &pfn_offset_in_region); + if (ret) + return ret; + + /* Ensure PFN is not before VMA's start within the region */ + if (pfn_offset_in_region < vma_offset_in_region) + return -EFAULT; + + /* Ensure PFN is not past the VMA's end within the region */ + if (pfn_offset_in_region >= vma_offset_in_region + vma_pages(vma)) + return -EFAULT; + + /* + * The VMA's vm_pgoff is the page offset of its start within the EGM + * region, so the region offset of the pfn is already the pgoff to + * report. + */ + *pgoff = pfn_offset_in_region; + + /* + * Record the poisoned offset for reporting via + * EGM_RETIRED_PAGES_LIST. This bookkeeping is best-effort: its + * failure must not be reported as "vma does not map pfn", or the + * owning process would not be signalled for the poisoned page. + */ + track_ecc_offset(egm_dev, *pgoff << PAGE_SHIFT); + + return 0; +} + +static int +nvgrace_egm_register_pfn_range(struct inode *inode, + struct nvgrace_egm_dev *egm_dev) +{ + unsigned long pfn, nr_pages; + int ret; + + pfn = PHYS_PFN(egm_dev->egmphys); + nr_pages = egm_dev->egmlength >> PAGE_SHIFT; + + egm_dev->pfn_address_space.node.start = pfn; + egm_dev->pfn_address_space.node.last = pfn + nr_pages - 1; + egm_dev->pfn_address_space.mapping = inode->i_mapping; + egm_dev->pfn_address_space.pfn_to_vma_pgoff = nvgrace_egm_pfn_to_vma_pgoff; + + ret = register_pfn_address_space(&egm_dev->pfn_address_space); + + return ret; +} + static int nvgrace_egm_open(struct inode *inode, struct file *file) { struct nvgrace_egm_dev *egm_dev = @@ -84,9 +184,9 @@ static int nvgrace_egm_open(struct inode *inode, struct file *file) /* * The EGM region is a physical carveout handed to a single VM at a - * time and must be scrubbed before each handover. Take open_lock - * killably around the scrub which can run for long time, so that - * the waiters stay killable. + * time and must be scrubbed and registered with memory_failure before + * each handover. Take open_lock killably around the scrub which can + * run for long time, so that the waiters stay killable. */ if (mutex_lock_killable(&egm_dev->open_lock)) return -EINTR; @@ -142,10 +242,15 @@ static int nvgrace_egm_open(struct inode *inode, struct file *file) remaining -= chunk_size; } + ret = nvgrace_egm_register_pfn_range(inode, egm_dev); + if (ret && ret != -EOPNOTSUPP) + goto unlock; + /* - * Mark the device open only after the scrub completes so that a - * concurrent opener cannot observe a non-zero count and proceed - * before the region has been cleared. + * Only mark the device open once the scrub and registration have + * succeeded. On failure open_count stays 0, so this opener returns + * an error (release() never runs for it) and the next opener + * retries the scrub and registration from a clean state. */ egm_dev->open_count = 1; ret = 0; @@ -162,8 +267,10 @@ static int nvgrace_egm_release(struct inode *inode, struct file *file) guard(mutex)(&egm_dev->open_lock); - if (!--egm_dev->open_count) + if (!--egm_dev->open_count) { + unregister_pfn_address_space(&egm_dev->pfn_address_space); file->private_data = NULL; + } return 0; } -- 2.34.1