From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from BL0PR03CU003.outbound.protection.outlook.com (mail-eastusazon11012035.outbound.protection.outlook.com [52.101.53.35]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 7D9EF38AC8C for ; Thu, 2 Jul 2026 19:26:23 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=fail smtp.client-ip=52.101.53.35 ARC-Seal:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783020384; cv=fail; b=BkJ7ys090LR0eijp7l9fb2uZD2SokdP0SxZPIRBV+t4YlwyGZwLwyj1AjPg3dHqzIJw8RllmE1WJ66G4R7czYrS31VX8O9hLXyNfYiatEHST19A5G+AyTRbVLKGT0UzCLda2sdlZpKYiB55HnQyLQ2d8psspEnxo746k1+4UU+o= ARC-Message-Signature:i=2; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1783020384; c=relaxed/simple; bh=lrJ7OfOQcA9rP8xa6dvGdBlApwZC+J7Ii54rC2y6jFg=; h=From:To:CC:Subject:Date:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=HdJp+ZVl5MxoC1We2ry+vXGhbHzTSTSOVpMCc41do83Txui57ddnGdMGz6f3gzuR9Z7LiyxReXUou1oXnzFFmXZVK6OnoqnPOFJblIXLvfuzehO0/rWRjVZyq554nN6ssuTfZ3n4gTpI7gsN8hrUQeUgp0vUFOoowJttJHgjLWA= ARC-Authentication-Results:i=2; smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com; spf=fail smtp.mailfrom=nvidia.com; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b=Z2k/D5T6; arc=fail smtp.client-ip=52.101.53.35 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=reject dis=none) header.from=nvidia.com Authentication-Results: smtp.subspace.kernel.org; spf=fail smtp.mailfrom=nvidia.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (2048-bit key) header.d=Nvidia.com header.i=@Nvidia.com header.b="Z2k/D5T6" ARC-Seal: i=1; a=rsa-sha256; s=arcselector10001; d=microsoft.com; cv=none; b=eUmP+Ppyth9gquKa3Uyu5/y3wYT1pNtU4m0pRK33NITS6wXrgfr4JnJvi6URu5PGiGTRsrvSqTu+EkLr/mmMs3Ao3fGgegLONQyb4pAWwz7MiqKE03HJdgYNYKKyv4uChvJNeVNQR5GWeo0WYkE+1quQ46DWpf6WWbWb/Q2WX2rCiTj1buZwJOqdLgO84aSOD6QU7CjUD/ga9sG5Tc9zi8nDKcwMda+VkxKHEgSwrLTwunQOVuhGaoqenkMmap5vz5vNRP/wo1Qc6YYUuoPAOJJRAFP79YfPY9Qg+S0fQawoVK6+rYqvfMNX7iYacE/4HZxoBSMuNNGBGF+WtpGG9w== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=microsoft.com; s=arcselector10001; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-AntiSpam-MessageData-ChunkCount:X-MS-Exchange-AntiSpam-MessageData-0:X-MS-Exchange-AntiSpam-MessageData-1; bh=AbTWrD7o4oTBUr3/HYKEtlpjNGsEnBrDp0PZhxCmqwI=; b=fMyTFAauUKDXoMb1Q6nJhR3UJ67aP2Y+xloz8TWT/82OVFReRs9G2g13Lx1ZgaL3ctV3WpBjDIv4FPUonGIONh11/grLfdCmK1BpM5ToNx9m5+cGqhwkf2t9zHAfLTTRuOF6DFzSo32mlOAFIMsqOBm25nvGz8N5bbgmFKZxCpgSq3Yo+HphbJhQXRwnn+RPXCM7R9zcJjsrAaFYozhEDyB/BJiBSFnLReQRzqP82Qz2Wt9M5+wpK9TCBtMuktpObCzzDEhoPo2rpetT4XH4HUw04sXojzy0844XRpkPbMoCCfMivR3NPevOU7ZGptldwTZnvV0vdJnadL3slt1yPw== ARC-Authentication-Results: i=1; mx.microsoft.com 1; spf=pass (sender ip is 216.228.117.160) smtp.rcpttodomain=linuxfoundation.org smtp.mailfrom=nvidia.com; dmarc=pass (p=reject sp=reject pct=100) action=none header.from=nvidia.com; dkim=none (message not signed); arc=none (0) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=Nvidia.com; s=selector2; h=From:Date:Subject:Message-ID:Content-Type:MIME-Version:X-MS-Exchange-SenderADCheck; bh=AbTWrD7o4oTBUr3/HYKEtlpjNGsEnBrDp0PZhxCmqwI=; b=Z2k/D5T6LNcMmMY/rNN5UmohQ7130jIl5+euxEEuxjR2zMfm5J1/WQ65ApIPKv4xeGjQnqSeFbAmw6EKDZnOGax+4wKRlMteOpEu/SM0MaagbJ3lk3Evmy4FgrB8JVf9actQVjU0mKcdJp+KDd1cFkH7xc2AZ2ijbmxYfhq6v9rRjCFnjFngLBjeugOYEZoOK3VBkt7duVuyyQNY1/EAZlQS8FL2BLJER18nBC5bM6CE/9ukb5/8SVlZVfHVhBUWXs2LYAjOSNbtz5fODbWLHmCdt8k4XDgI9K62jcmdAYloxUoXrmvxj5XrliseFxZdRGFR8ORR3uU5Hi+T53ziEA== Received: from MW4PR04CA0224.namprd04.prod.outlook.com (2603:10b6:303:87::19) by DS5PPFEC0C6BDA1.namprd12.prod.outlook.com (2603:10b6:f:fc00::668) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.8; Thu, 2 Jul 2026 19:25:59 +0000 Received: from SJ1PEPF000026C9.namprd04.prod.outlook.com (2603:10b6:303:87:cafe::d) by MW4PR04CA0224.outlook.office365.com (2603:10b6:303:87::19) with Microsoft SMTP Server (version=TLS1_3, cipher=TLS_AES_256_GCM_SHA384) id 15.21.181.9 via Frontend Transport; Thu, 2 Jul 2026 19:25:59 +0000 X-MS-Exchange-Authentication-Results: spf=pass (sender IP is 216.228.117.160) smtp.mailfrom=nvidia.com; dkim=none (message not signed) header.d=none;dmarc=pass action=none header.from=nvidia.com; Received-SPF: Pass (protection.outlook.com: domain of nvidia.com designates 216.228.117.160 as permitted sender) receiver=protection.outlook.com; client-ip=216.228.117.160; helo=mail.nvidia.com; pr=C Received: from mail.nvidia.com (216.228.117.160) by SJ1PEPF000026C9.mail.protection.outlook.com (10.167.244.106) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.21.181.6 via Frontend Transport; Thu, 2 Jul 2026 19:25:58 +0000 Received: from rnnvmail202.nvidia.com (10.129.68.7) by mail.nvidia.com (10.129.200.66) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Thu, 2 Jul 2026 12:25:35 -0700 Received: from rnnvmail203.nvidia.com (10.129.68.9) by rnnvmail202.nvidia.com (10.129.68.7) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20; Thu, 2 Jul 2026 12:25:34 -0700 Received: from localhost.nvidia.com (10.127.8.12) by mail.nvidia.com (10.129.68.9) with Microsoft SMTP Server (version=TLS1_2, cipher=TLS_ECDHE_RSA_WITH_AES_256_GCM_SHA384) id 15.2.2562.20 via Frontend Transport; Thu, 2 Jul 2026 12:25:34 -0700 From: Ankit Agrawal To: , , CC: , , , , , , , , Subject: [PATCH 2/4] platform/nvidia: Implement mmap and memory scrubbing for EGM chardev Date: Thu, 2 Jul 2026 19:25:30 +0000 Message-ID: <20260702192532.455400-3-ankita@nvidia.com> X-Mailer: git-send-email 2.34.1 In-Reply-To: <20260702192532.455400-1-ankita@nvidia.com> References: <20260702192532.455400-1-ankita@nvidia.com> Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Transfer-Encoding: 8bit Content-Type: text/plain X-NV-OnPremToCloud: ExternallySecured X-EOPAttributedMessage: 0 X-MS-PublicTrafficType: Email X-MS-TrafficTypeDiagnostic: SJ1PEPF000026C9:EE_|DS5PPFEC0C6BDA1:EE_ X-MS-Office365-Filtering-Correlation-Id: cff2eda7-119e-418d-6bb0-08ded86fbeeb X-MS-Exchange-SenderADCheck: 1 X-MS-Exchange-AntiSpam-Relay: 0 X-Microsoft-Antispam: BCL:0;ARA:13230040|23010399003|1800799024|376014|82310400026|36860700016|11063799006|56012099006|18002099003|22082099003; X-Microsoft-Antispam-Message-Info: hK4ZHis5usXDG029OIPeAGiNoUyjBE+RIt21v39nKrjx+cTORGmCSdJ7EZtTSJLFvGUayshrMybd4n7DkL4wrWvJmVaFggl2y5w7ixUwx5JH7YK+HeJ4HkpjQIBvIxhjBN9WxEma6UtPttzS0xITxeGTdAQdOajDtozwOMP9CKOA2CkJL5WBDvl3BG3Iu++dXKekEx9sqyT80n94z9Nrk8UFCckbx1JyJH5Pd94ZZzh4djxVbgWkeUSRR9xa5JLQbRIL6urb/Wc2K4ysbJAqSyBOGI5KgoDT9LEF1R4AAj4YLcL2zZB8nOzbs6DM+8SE/ZZcePG/jrZqMv46XTPys3l2O8Uo2KKegY0TB+62HwWm7/c4qqOnYeKqcQp4n8PghWfreHmegQvC8VFTzht0m4w6g8XbxM0EOwemTM8QG00SOWXwVSzY+0fxLR9FgLagTJoVqX7gc5horO/ml2pMTXo5bCzUxS4N9wj5ulOtPs4LuxdqiyUlcpxtW0Od2K9ydJ079MQNbfvCV6cKKjY3cu2aCx9X8AKbKZuW+N+JSuPoz0iqkxoc9S7aTC49O53dkwaOZea1voLFDbSU2cd/b3qDG51OsMhm5ibRuw152GilXLx8CA19JBgzqN1r0P4bzpNyo36JIPCZ0/RqSj9nh8BBmUl3NQ/RMQ48wFIrHgG0ioAbCF/sc2ZVRYb6dCrK6vOl3zURIIt97Kq9f2/a+w== X-Forefront-Antispam-Report: CIP:216.228.117.160;CTRY:US;LANG:en;SCL:1;SRV:;IPV:NLI;SFV:NSPM;H:mail.nvidia.com;PTR:dc6edge1.nvidia.com;CAT:NONE;SFS:(13230040)(23010399003)(1800799024)(376014)(82310400026)(36860700016)(11063799006)(56012099006)(18002099003)(22082099003);DIR:OUT;SFP:1101; X-MS-Exchange-AntiSpam-MessageData-ChunkCount: 1 X-MS-Exchange-AntiSpam-MessageData-0: ED+6Pgh/7f5tScR2iPHuI8osV8Zju7Y0Gr1GBjncCsVfqYb3UkF5cVEnZ1rEQ+vvjsaIiUFIw/XqDjwx26mERBxHw2BWCdtbkC909YIS6h91xvuvwwKTcKGBLFLSQzNFThqGUXknVS5r2FTSLhW/QPAh5ASjxw1XzRrPv4rOZctGc2QZ/OZTcNuLjjULbPOfTyqsXiD2QAr3DoB9cfc/OB1uSdJk+gNW0v8X8fFrKKJ+f60osQMDyFKxzhqPtRFgj0FvDixcawqTxaHOWzcP2lacbDVX6Zj+I7iciNa6lo2/FLEmpEDTf/sT4pfuunP/JXcpH7Jl8S0rfS1Cku4YII0p546soqIUIfsmCJfX4eb8FTB3dXxhPYybTppRB4TYHQs92QBVunqv0C+Y1uo11ErKPwiwpceoBqwjHzlfTOvbvaL7AOWHu3EAuJlfdkK0 X-OriginatorOrg: Nvidia.com X-MS-Exchange-CrossTenant-OriginalArrivalTime: 02 Jul 2026 19:25:58.7895 (UTC) X-MS-Exchange-CrossTenant-Network-Message-Id: cff2eda7-119e-418d-6bb0-08ded86fbeeb X-MS-Exchange-CrossTenant-Id: 43083d15-7273-40c1-b7db-39efd9ccc17a X-MS-Exchange-CrossTenant-OriginalAttributedTenantConnectingIp: TenantId=43083d15-7273-40c1-b7db-39efd9ccc17a;Ip=[216.228.117.160];Helo=[mail.nvidia.com] X-MS-Exchange-CrossTenant-AuthSource: SJ1PEPF000026C9.namprd04.prod.outlook.com X-MS-Exchange-CrossTenant-AuthAs: Anonymous X-MS-Exchange-CrossTenant-FromEntityHeader: HybridOnPrem X-MS-Exchange-Transport-CrossTenantHeadersStamped: DS5PPFEC0C6BDA1 Setup the file operations (open, close and mmap) for the EGM char device. Memory scrubbing: - The EGM region is invisible to the host Linux kernel and is not managed by the page allocator. So the driver is responsible for zeroing it before handing it to a VM. - Clear the entire EGM region on the first open() call using memremap() and memset() in 1 GiB chunks with cond_resched() between iterations to avoid RCU stalls on large regions. - The region is handed to a single VM at a time and must be scrubbed before each handover. Serialise openers with a mutex guarding open_count and refuse a second opener with -EBUSY while the region is still held. mmap: - Implement nvgrace_egm_mmap() using remap_pfn_range(). The EGM region is not part of the host EFI map and has no struct pages making remap_pfn_range() the correct interface for mapping it directly into QEMU's virtual address space. - Validate that the requested range fits within the EGM region. Reject a vm_pgoff that already lies outside the region before it is shifted into a byte count so that the page offset cannot overflow and slip past the range check. Suggested-by: Aniket Agashe Suggested-by: Vikram Sethi Assisted-by: Claude:claude-opus-4.8 Signed-off-by: Ankit Agrawal --- drivers/platform/nvidia/egm.c | 125 ++++++++++++++++++++++++++++++++-- 1 file changed, 121 insertions(+), 4 deletions(-) diff --git a/drivers/platform/nvidia/egm.c b/drivers/platform/nvidia/egm.c index a340c4fcbc7a..09ab00213de2 100644 --- a/drivers/platform/nvidia/egm.c +++ b/drivers/platform/nvidia/egm.c @@ -3,6 +3,9 @@ * Copyright (c) 2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved */ +#include +#include +#include #include #define NVGRACE_EGM_DEV_NAME "egm" @@ -19,6 +22,10 @@ struct gpu_node { struct nvgrace_egm_dev { struct device device; struct cdev cdev; + /* serialises the first-open scrub */ + struct mutex open_lock; + /* protected by open_lock */ + unsigned int open_count; phys_addr_t egmphys; size_t egmlength; u64 egmpxm; @@ -39,21 +46,129 @@ static LIST_HEAD(egm_chardevs); static int nvgrace_egm_open(struct inode *inode, struct file *file) { - return 0; + struct nvgrace_egm_dev *egm_dev = + container_of(inode->i_cdev, struct nvgrace_egm_dev, cdev); + phys_addr_t phys; + size_t remaining, chunk_size; + void *chunk_addr; + int ret; + + file->private_data = egm_dev; + + /* + * The EGM region is a physical carveout handed to a single VM at a + * time and must be scrubbed before each handover. Take open_lock + * killably around the scrub which can run for long time, so that + * the waiters stay killable. + */ + if (mutex_lock_killable(&egm_dev->open_lock)) + return -EINTR; + + /* + * Refuse a second opener while the region is still held. release() + * runs only when the last reference to the struct file drops but an + * mmap keeps that reference (and hence open_count) alive past + * close for the lifetime of the VMA. So open_count returns to zero + * only when every mapping is gone. A concurrent opener cannot be + * handed the region: it cannot be re-scrubbed without destroying the + * current consumer's live data. + */ + if (egm_dev->open_count) { + ret = -EBUSY; + goto unlock; + } + + /* + * nvgrace-egm module is responsible to manage the EGM memory as + * the host kernel has no knowledge of it. Clear the region before + * handing over to userspace. + * + * The EGM region can be very large (hundreds of GiB). So Map and zero + * one chunk at a time rather than mapping the whole region at once. + */ + phys = egm_dev->egmphys; + remaining = egm_dev->egmlength; + + while (remaining > 0) { + /* + * The scrub holds the lock for a long time; let it be + * SIGKILL'd. + */ + if (fatal_signal_pending(current)) { + ret = -EINTR; + goto unlock; + } + + chunk_size = min(remaining, SZ_1G); + + chunk_addr = memremap(phys, chunk_size, MEMREMAP_WB); + if (!chunk_addr) { + ret = -ENOMEM; + goto unlock; + } + + memset(chunk_addr, 0, chunk_size); + memunmap(chunk_addr); + cond_resched(); + + phys += chunk_size; + remaining -= chunk_size; + } + + /* + * Mark the device open only after the scrub completes so that a + * concurrent opener cannot observe a non-zero count and proceed + * before the region has been cleared. + */ + egm_dev->open_count = 1; + ret = 0; + +unlock: + mutex_unlock(&egm_dev->open_lock); + return ret; } static int nvgrace_egm_release(struct inode *inode, struct file *file) { + struct nvgrace_egm_dev *egm_dev = + container_of(inode->i_cdev, struct nvgrace_egm_dev, cdev); + + guard(mutex)(&egm_dev->open_lock); + + if (!--egm_dev->open_count) + file->private_data = NULL; + return 0; } static int nvgrace_egm_mmap(struct file *file, struct vm_area_struct *vma) { + struct nvgrace_egm_dev *egm_dev = file->private_data; + u64 req_len, pgoff, end; + unsigned long start_pfn, num_pages; + + pgoff = vma->vm_pgoff; + num_pages = egm_dev->egmlength >> PAGE_SHIFT; + + /* Reject a page offset that already lies outside the EGM region. */ + if (pgoff >= num_pages) + return -EINVAL; + + if (check_sub_overflow(vma->vm_end, vma->vm_start, &req_len) || + check_add_overflow(PHYS_PFN(egm_dev->egmphys), pgoff, &start_pfn) || + check_add_overflow(PFN_PHYS(pgoff), req_len, &end)) + return -EOVERFLOW; + + if (end > egm_dev->egmlength) + return -EINVAL; + /* - * Mapping the EGM region into userspace is implemented by a later - * patch. Until then refuse the mmap. + * EGM memory is invisible to the host kernel and is not managed + * by it. Map the usermode VMA to the EGM region. */ - return -EOPNOTSUPP; + return remap_pfn_range(vma, vma->vm_start, + start_pfn, req_len, + vma->vm_page_prot); } static const struct file_operations file_ops = { @@ -122,6 +237,7 @@ static void egm_chardev_release(struct device *dev) struct nvgrace_egm_dev *egm_chardev = container_of(dev, struct nvgrace_egm_dev, device); remove_gpus(egm_chardev); + mutex_destroy(&egm_chardev->open_lock); kfree(egm_chardev); } @@ -155,6 +271,7 @@ static struct nvgrace_egm_dev *setup_egm_chardev(u64 egmphys, u64 egmlength, egm_chardev->egmphys = egmphys; egm_chardev->egmlength = egmlength; egm_chardev->egmpxm = egmpxm; + mutex_init(&egm_chardev->open_lock); INIT_LIST_HEAD(&egm_chardev->gpus); egm_chardev->device.devt = MKDEV(MAJOR(dev), egm_chardev->egmpxm); -- 2.34.1