From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from SA9PR02CU001.outbound.protection.outlook.com (mail-southcentralusazon11013011.outbound.protection.outlook.com [40.93.196.11]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id BD6B51FF1DA for ; Thu, 2 Jul 2026 19:26:05 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=40.93.196.11 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783020369; cv=fail; b=B4RJE0PRSmRxBmBMm6ufe1FPHwqX5f4fPqL8bQq5p2Eq0dbDt/bEPsu3pNiY/CSGBcFSWp5g20R6KVDvSgftTTYTMYpyfH5DvXRiQT9gezhL758VetKXH2ClfJQMAA4opryJwAsqdED4iWkNb4x1COBHeaSH/Ns508RJ/9WzubM= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783020369; c=relaxed/simple; bh=eXVg2nqffn/Q1h9X1M6/J9MNvEp1OhELvqbFhnn9BPM=; h=From:To:CC:Subject:Date:Message-ID:MIME-Version:Content-Type; b=TnTnaCgjTnuzVAQ/onRZcqCMkukQwcC1IQI3V4xo8OAGj2fIruoMUNHmcSAge8ND8Avl4zwPpO898NuXwSiNmrguIZ+XI/PP+P8xF4n8Kxtfil8EVwlSROd0TFdVzxvKkivzxHAP10ziWrY3wytEh2bgZfLjj+C19sy+W3hmh6Y= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com; spf=fail smtp.mailfrom=nvidia.com; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b=EAAQDw/g; arc=fail smtp.client-ip=40.93.196.11 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=nvidia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b="EAAQDw/g" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=AHX+R2QCK5KpWZY208BTRfsqOEuFv4zv0ATxnbS6pUb+tF1Z4a036uG829ApIqK2WEa0ZNzTMpeNyLgTMO+Pn9JFLnI5Fao5wf7cr1SWlwV3Em1o+TyZfPttr8UJZhUT/rXvFydCUt8Z9PTtjoteaaU4Zfm6UBkeHgCw24M9g/ARtWe8VVonA0Wah5AjgTRFXDO7K1AVq4U054eOuBM2uIkbsFd4+SRDfFQtxDlfpVUVTl25CqjYUdBLXdzYtzGyT7PMzHSFbQMYt5f0lw9PxZZ21YJTnnoa9ybIVfhJkBzK+A5iDaTizjrGY9CaeQAEIppNXxHf+aWoPxF+dhKrZg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=9xoH9K1G1t7nckMKdQFrNUfr8SIMcWgvKKjZ7VLCROg=; b=vd3W+3SQ0hIXwZ20+k88fk9xhf4+1nXIqKIBMVmFKrEUXLVcNTgGi7Vfmpgz5anzbgfWhUBqBUmWMpIPWWHHY18oNK89Fxz+EzrtpNQCLeIHfRlHDkssK9oU+PHDbNKl2abRVY2+yX8n3nMwk5Mt+grJR/8HQVgQCeviBe3UoB2rCtsTiWf8Rd+9fOTIXKGyTxNsXH1iRcNHyjDefK+02P5PIK7Bsd8kNACKqVfjjVmL1WgbzKdSSTo2is2RlXaeBSBlFa6TtoVNeTh/PzFUSYvOjYA2CQH8nwTkdJdji0VelaNROESh71EgA5Bxp8Y8WiZ1JeSwp13WBPseLA8xSA== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.117.161) smtp.rcpttodomain=linuxfoundation.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=9xoH9K1G1t7nckMKdQFrNUfr8SIMcWgvKKjZ7VLCROg=; b=EAAQDw/gTJvUnw/nxQAPBoCuNb3aCs1d9fekOCC1481yx3RlO+ssv/t0QWEEEeqoiiBTwCXSI48rM5P/oapa8IweKPbgWoDdr9Qwynlb8GODqxf8MSwmCcaVdSDq4ikH4R+qP6OgpGWovkL+D28IW4jyBvlkLkSySEwjNEaqS7TGFZ2/6p0QmgYKhk5GZ6Dl7LARIHKT1upWbTPtndXyDV/a3sah1sG5MblbahJMDen9hHl2arLYD8fId974qgnlwjd6TaY3L+XBvtX48Mx7NWcNSkC+xMovieRAewAF7p1AgR2SAJ44Jow078diQ0/cyegFc49D9xGTp7RCS2hdJg== Received: from SA1PR04CA0018.namprd04.prod.outlook.com (2603:10b6:806:2ce::23) by SA1PR12MB8987.namprd12.prod.outlook.com (2603:10b6:806:386::22) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.8; Thu, 2 Jul 2026 19:25:58 +0000 Received: from SN1PEPF0002529F.namprd05.prod.outlook.com (2603:10b6:806:2ce:cafe::21) by SA1PR04CA0018.outlook.office365.com (2603:10b6:806:2ce::23) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.181.11 via Frontend Transport; Thu, 2 Jul 2026 19:25:58 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 216.228.117.161) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.117.161 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.117.161; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.117.161) by SN1PEPF0002529F.mail.protection.outlook.com (10.167.242.6) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.6 via Frontend Transport; Thu, 2 Jul 2026 19:25:57 +0000 Received: from rnnvmail205.nvidia.com (10.129.68.10) by mail.nvidia.com (10.129.200.67) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Thu, 2 Jul 2026 12:25:33 -0700 Received: from rnnvmail203.nvidia.com (10.129.68.9) by rnnvmail205.nvidia.com (10.129.68.10) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Thu, 2 Jul 2026 12:25:33 -0700 Received: from localhost.nvidia.com (10.127.8.12) by mail.nvidia.com (10.129.68.9) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20 via Frontend Transport; Thu, 2 Jul 2026 12:25:33 -0700 From: Ankit Agrawal To: , , CC: , , , , , , , , Subject: [PATCH 0/4] Introduce nvgrace-egm driver for Extended GPU Memory Date: Thu, 2 Jul 2026 19:25:28 +0000 Message-ID: <20260702192532.455400-1-ankita@nvidia.com> X-Mailer: git-send-email 2.34.1 Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Content-Type: text/plain X-NV-OnPremToCloud: ExternallySecured X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SN1PEPF0002529F:EE_|SA1PR12MB8987:EE_ X-MS-Office365-Filtering-Correlation-Id: da3cfbf8-3af8-41d6-95cb-08ded86fbe69 X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|82310400026|376014|23010399003|36860700016|1800799024|13003099007|6133799003|11063799006|56012099006|18002099003|3023799007; X-Microsoft-Antispam-Message-Info: uleQPQ7mcjmWKkjz16bccj+aKQE9k91QZ+CWpRBO0kuT/AA1QaAuqLo/T5TMnVIqgjuh/wr2+IQ4LqvzB0mEyzsaE00N3tO//nDwA+7UmnYqwKW5BaU2WGkM70HSxQ+R9543CgL3q0g2cMRQihoT+ZMz1WJbd8cXCepQuIryWrCVXYUjaGVRDn5YzJTKJGd/64CLsNO2EDpHUmOa04X3sJ0d4dgJj2CarU3G2dlw3SxhWz8acl2QV1mCIKhdPqO++oGedmd3A0VExORaqDzgMHcgCQR3Hld2VtgOLvIahzvPwtBbq5jgR1BYmMod7Ld9qvOY+kYs+uk9yO66PlpD8VJQZz8toVY6gAYT75FpncaZrx8uzTbRQYa48xzdPACnRv8h6rVOvRb0Zb7tmrLyzU1xkR3CRDklipBtTKOpoyOMQJ52KpYuksCims5N4RQLyBew7jfPRow67VYZsUnXXtFVK0aa8gwGqPQnK/AkcDGpTkij/y0g3x1nNMxFkZEELeARg8GLUrTKA3r5SEvMekGyhLzJaHpAROfzFigxDvnW7eFZAknST2VgM9N4V3q+LFcxldo0N6j0jqyUYU8kTmHQOh9iKY8ZyGcqxp0oq+7joOxqlaSJh2UF5w5aVtAy9HZetcUgEAJ4c3E4HT/Wqi0N3r5MsuTL7XGNO667kndNurPCfCvEJHyDehtCE7DQ3VZItpMz0ZdzLlLsPdSuvQ== X-Forefront-Antispam-Report: CIP:216.228.117.161;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:mail.nvidia.com;PTR:dc6edge2.nvidia.com;CAT:NONE;SFS:(13230040)(82310400026)(376014)(23010399003)(36860700016)(1800799024)(13003099007)(6133799003)(11063799006)(56012099006)(18002099003)(3023799007);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: KYsULKPKOXfK7HPae7zD9M0ca9VKmqsxZMukQjpmpMHJyCt5+eWv6fSZ4O8Cg9FUMJzzINfIKqNbrvoJP0h5s2FHPYroXeg8v1gDxeiZ0/csNrGXdg/xGU3fcoSYVuINriQMNV7FzSWmkULXHZ2KhrmUa58dW6f/IBYCH+fVnIvKn7EyTcbqmFQCitSxQjJ+1GuX3hWVCeV7Ha781e8J4xZKNnSVOXgCwJTdF1sBJfKB6Eu0RVbe8bcpDv3ZhCabPBvE+SM1L3zomZ9P44eQrZC94TWNjFwSPEfEEd2xZvvv5g7oPbgsg6H2D98Ec9z5uKZEaIu5l1WH4+a7NnUQXITiJZaCiHSj3IQHmE8LJM67QNQsX7kkIA9jx/R2X6h3sPtvNpYcAZMVyKO7hmlx4Rmh7AYrkAnh12csgJnrMebnkBrw5NRkI+zFzY2uEDFm X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 02 Jul 2026 19:25:57.9089 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: da3cfbf8-3af8-41d6-95cb-08ded86fbe69 X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a;Ip=[216.228.117.161];Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: SN1PEPF0002529F.namprd05.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: SA1PR12MB8987 Background ========== NVIDIA Grace-based systems (Grace Hopper, Grace Blackwell Superchip) have a special bios-configurable mode called EGM (Extended GPU Memory) that allows a carveout to be created from the system memory. This carveout is a contiguous region of special range of system DRAM reserved by the UEFI boot firmware. This range does not appear in the EFI memory map and is therefore invisible to the host kernel. i.e. there are no struct pages for it and it cannot be managed by the normal page allocator. This range is described in ACPI as vendor DSD properties on the GPU PCI device node (see ACPI Design below). It is intended to be used by a VMM to back an entire virtual machine's physical memory. The VMM maps the range into its address space via mmap() on a char device and passes it to KVM as the backing store for guest RAM similar to memfd or hugetlbfs. KVM builds Stage-2 page tables mapping the guest's GPA range linearly to the EGM SPAs i.e. a guest offset X maps to SPA (egm_base + X). The range is special because the GPU on these systems has direct high-speed access to it that bypasses the host system IOMMU. Normal PCI DMA is gated by the IOMMU and requires an explicit IOMMU mapping. The GPU's path to this memory goes via NVLink/NVSwitch directly to the memory controller and does not involve the host IOMMU. Since KVM maps guest GPA to EGM SPA linearly, the GPU can reach guest memory at those same SPAs without needing any additional host IOMMU mapping. Normal system memory that the kernel sees is protected by the IOMMU in the usual way and this range is deliberately kept separate. To make this work a driver is needed to parse the ACPI, capture the physical memory range, expose it as a char device and allow it to be mmap'd for KVM. A VMM like qemu can open this char dev and then use it much like a hugetlbfs or memfd to back the VM. Due to the security difference the memory has to be strictly separated from the rest of the hypervisor memory. Exposing the range as a DMABUF for iommufd is planned as future work and is not part of this series. This hardware feature allows the system memory to be access across nodes on a multi-node system such as Grace Blackwell NVL72. It can also provide throughput improvements because the GPU can access VM system memory via its high-bandwidth fabric rather than going through the IOMMU-gated PCIe path. Moreover skipping IOMMU translations could also improve performance for any use cases targeted towards high performance. Due to the security difference, the EGM range is strictly isolated from the rest of the hypervisor memory (which is covered through IOMMU). A driver is needed to: - Parse ACPI to locate the physical range (base SPA, size, proximity domain) - Expose the range as a char device (/dev/egmX, X = NUMA proximity domain) that a VMM such as QEMU can open and mmap to back guest RAM - Zero the region on first open to isolate successive VM uses - Track and expose retired ECC pages to the VMM via ioctl - Register the PFN range with the memory_failure infrastructure for runtime uncorrectable error handling ACPI Design =========== The EGM region properties are exposed as NVIDIA vendor DSD properties on the GPU's PCI ACPI device node: nvidia,egm-pxm NUMA proximity domain of the region nvidia,egm-base-pa Physical base address (SPA) of the region nvidia,egm-size Length of the region in bytes nvidia,egm-retired-pages-data-base Physical address of the firmware- provided retired ECC pages table This ACPI construct helps with describing and discovering this physically present memory range that is intentionally absent from the EFI memory map and is tied to a specific device for its access properties. On a multi-socket system there is one EGM region per socket. Multiple GPUs on the same socket advertise identical region properties (same proximity domain, base address and size). The driver deduplicates by proximity domain: the first GPU probed for a given domain creates the char device and subsequent GPUs on the same socket are linked to it via sysfs. Implementation ============== Patch 1 introduces the driver skeleton: Kconfig/Makefile/MAINTAINERS, the egm device class, the char device minor number range (MAX_EGM_NODES=4 matching the largest 4-socket Grace configuration), ACPI property walking (nvidia,egm-pxm / nvidia,egm-base-pa / nvidia,egm-size), per-region state tracking with PXM-based deduplication, char device creation (/dev/egmX), GPU-to-EGM sysfs symlinks and the egm_size sysfs attribute. Patch 2 implements the file operations: memory scrubbing on the first open() using memremap()/memset() in 1 GiB chunks with cond_resched() to avoid RCU stalls and mmap() via remap_pfn_range(). Patch 3 adds retired ECC page handling: reads the firmware-provided retired-pages table from the ACPI DSD property nvidia,egm-retired-pages-data-base at module init, stores the page offsets in a per-device xarray and exposes them to userspace via the EGM_RETIRED_PAGES_LIST ioctl (UAPI header include/uapi/linux/egm.h). Each firmware entry covers a 64 KiB retired region. The driver inserts one xarray entry per kernel page so a lookup at the running page size hits the right retired range on both 4K and 64K kernels. Patch 4 registers the EGM PFN range with the memory_failure infrastructure on the first open() and unregisters it on the last close(). The pfn_to_vma_pgoff callback validates that a reported PFN falls within the EGM region, converts it to an in-region file offset, and inserts it into the retired-pages xarray so runtime uncorrectable errors are surfaced alongside ACPI-reported retired pages via the same ioctl. Enablement ========== EGM mode is enabled through a flag in the platform firmware (SBIOS). The size of the host-visible partition is separately configurable; all remaining system DRAM is reserved as the EGM region invisible to the hypervisor. RFC Posting =========== Link: https://lore.kernel.org/all/20260223155514.152435-1-ankita@nvidia.com/ Verification ============ Tested on a 2-socket (4 GPU) Grace Blackwell platform by booting a VM with QEMU [1] and kernel branch [2] and confirming EGM capability is visible inside the guest and basic cuda workloads. Note that this has dependency on the generic dmabuf importer/exporter work in progress [3]. Until then, a bypass PFNMAP workaround is used [4] (also part of the kernel branch [2]. Link: https://github.com/NVIDIA/QEMU/tree/nvidia_stable-10.1 [1] Link: https://github.com/ankita-nv/linux/tree/v7.2-egm-01072026 [2] Link: https://lore.kernel.org/all/0-v1-b5cab63049c0+191af-dmabuf_map_type_jgg@nvidia.com/ [3] Link: https://github.com/ankita-nv/linux/commit/f2a915d00916c5ef5e25f7bc2ad5c10efd36fb4b [4] Suggested-by: Alex Williamson Suggested-by: Jason Gunthorpe Suggested-by: Aniket Agashe Suggested-by: Vikram Sethi Suggested-by: Matthew R. Ochs Ankit Agrawal (4): platform/nvidia: Introduce nvgrace-egm driver and enumerate EGM regions platform/nvidia: Implement mmap and memory scrubbing for EGM chardev platform/nvidia: Handle retired ECC pages and expose via ioctl platform/nvidia: Register EGM PFNMAP range with memory_failure .../userspace-api/ioctl/ioctl-number.rst | 1 + MAINTAINERS | 6 + drivers/platform/Kconfig | 1 + drivers/platform/Makefile | 1 + drivers/platform/nvidia/Kconfig | 10 + drivers/platform/nvidia/Makefile | 3 + drivers/platform/nvidia/egm.c | 874 ++++++++++++++++++ include/uapi/linux/egm.h | 28 + 8 files changed, 924 insertions(+) create mode 100644 drivers/platform/nvidia/Kconfig create mode 100644 drivers/platform/nvidia/Makefile create mode 100644 drivers/platform/nvidia/egm.c create mode 100644 include/uapi/linux/egm.h -- 2.34.1