From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.129.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id E5B6F191499 for ; Thu, 9 Jan 2025 19:21:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.129.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736450496; cv=none; b=SCjqX5VDU84JxXv+51OZbTHbtUrnI2MxFsv+ymw8QvqmDXv41GNY0dBl4kTc+K7iYkjpJLJ3aZRP81BPV3BxtZvt9g9hO8ix5z70R7SGALw720Hk4VoYhgy3chY7AjKmODhwKqgwO6RL31fXAlhaN9Tlb1iuUAUZuPd/LA0IkCc= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1736450496; c=relaxed/simple; bh=1VXcu8J4zDHhTrL5/OvuJJqPYBknxtMOtfvGWdeV7/I=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=tIHgaBoODV/7EJLxLYzLOzr+1Rd2xwAbcTkgvwoWroh0MPVp20pZ+qAF5ifCmoKP33galOLJ9kCZjkLqEzAL0Iv0XGZ1uoBFyFbp+1I4wSTtNpGka6XREY3cy7iuwz8/1W76TAepqGVwQLSgr46F5M/4JeZ7BkFroSUcNLMY13I= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=hACrHxoH; arc=none smtp.client-ip=170.10.129.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="hACrHxoH" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1736450493; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=iAyjAb7PBWRyAuz+1PdB6LhTCq5OIkw/AzaQ/w/A76E=; b=hACrHxoHV+/BlSgQP90oMv4y4q3g544fuSp8VNoS8KdqonH+YQtnLQ33zj3F2Y2ewpO47p nP7MChDocSF5YSbk9ceNtnVijqcbPG0pCAJrz4f3IQ9CTHs3enHj5jGOLFUDHh3GgDJ0U+ 1+wt+swA2o0IHcehMGfu+4sJLldTEbg= Received: from mail-io1-f69.google.com (mail-io1-f69.google.com [209.85.166.69]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-169-faTACZCCMCOZiYuBqh9heA-1; Thu, 09 Jan 2025 14:21:32 -0500 X-MC-Unique: faTACZCCMCOZiYuBqh9heA-1 X-Mimecast-MFC-AGG-ID: faTACZCCMCOZiYuBqh9heA Received: by mail-io1-f69.google.com with SMTP id ca18e2360f4ac-84a3a4ec598so21371239f.2 for ; Thu, 09 Jan 2025 11:21:32 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1736450492; x=1737055292; h=content-transfer-encoding:mime-version:organization:references :in-reply-to:message-id:subject:cc:to:from:date:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=iAyjAb7PBWRyAuz+1PdB6LhTCq5OIkw/AzaQ/w/A76E=; b=CEU/iqprtiAsgO4ok0oSGqtCZXs/pQKS6QVTvHlkt48r9rpp/oMuEgn4R4bZwxEAfR 4hjkYprSxAzqOsF6pZnM5aYxNDFdk8kzvTMQhqAl6n2M2WqLoaNR0mG/JwKES9byLsOA jqILCOGTUHla403l0A4zW8GUgHI/4cf/hq5SeK5m+nFi8ZNqb4UxvpbFO33Cg3iDU6In XKEhS39a2IUub//siWzjmnbwc6ZO8qZYTF5tI0NUdL8Yly3FhPhx3AJlFdGcD/V/RlDW dWx3Rg0tDVS981PMDn6lfd6iDsnjvUKzs8zOz/YQEXQBno9cUhlYkoaY+LjK5OHe93Ua pzVw== X-Forwarded-Encrypted: i=1; AJvYcCULQjbp1SAa1w3NoZuYa9uA1Z7ZovK55FLPa1qvOQ964BCOuenB77ElpKtowY0Jy4t1E8IaPDtfyliDKsA=@vger.kernel.org X-Gm-Message-State: AOJu0YzGUCR9GLR6YN756rxuRXA0HIRoERp3lfZqRapRLOAjePVMgH4X GrYJddEZQ5L96aiSfkHK2G/1vQE6i3jEQUrQjZmwxQtcJFGyOB5E25v/BzqazWySvfc8p8XHg4l 31DwZqFbMAZ4OQ792/eRGhEy/zBbV1sUKqf2Pw+0qLas0WzMtMbohZaGSCvzkCg== X-Gm-Gg: ASbGnctayhcPFHqmfuWbi7WVeNcrjmV5VxPOx0UjU/4ng28v72N88rTyvJe+KsMjC3G 0MTodZFAtoXs8aDWIdrkuy7t+5riY4eYrO3Ia2GsfErEA9L/BuDHKo1PVYWKml2NAwhb5G5i8n8 DggoitqR1qbxC110bdnSbIVBBPMlGmHMBG4iHddxnestkESbRymGJp98q7y90kXgY2rBS68+EwS ipRu/qi3PmAYO6/vzyaoFbS89HnEoVJCxlX8yDNb7VJbyFNOiPKtkUTdqyn X-Received: by 2002:a05:6e02:16c7:b0:3a7:bfc6:be with SMTP id e9e14a558f8ab-3ce3a8f134dmr17530685ab.5.1736450491726; Thu, 09 Jan 2025 11:21:31 -0800 (PST) X-Google-Smtp-Source: AGHT+IEndpSc5anNrpCccoAV3Ich4D4eit45O/0ZYPOo019ot95MvdlNfVAdN2PfeTL0Ef9iJsbseA== X-Received: by 2002:a05:6e02:16c7:b0:3a7:bfc6:be with SMTP id e9e14a558f8ab-3ce3a8f134dmr17530555ab.5.1736450491344; Thu, 09 Jan 2025 11:21:31 -0800 (PST) Received: from redhat.com ([38.15.36.11]) by smtp.gmail.com with ESMTPSA id e9e14a558f8ab-3ce4afc329dsm5221535ab.65.2025.01.09.11.21.29 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Thu, 09 Jan 2025 11:21:30 -0800 (PST) Date: Thu, 9 Jan 2025 14:21:23 -0500 From: Alex Williamson To: Cc: , , , , , , , , , , , , , , , , , Subject: Re: [PATCH v2 2/3] vfio/nvgrace-gpu: Expose the blackwell device PF BAR1 to the VM Message-ID: <20250109142123.3537519a.alex.williamson@redhat.com> In-Reply-To: <20250105173615.28481-3-ankita@nvidia.com> References: <20250105173615.28481-1-ankita@nvidia.com> <20250105173615.28481-3-ankita@nvidia.com> Organization: Red Hat Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Sun, 5 Jan 2025 17:36:14 +0000 wrote: > From: Ankit Agrawal > > There is a HW defect on Grace Hopper (GH) to support the > Multi-Instance GPU (MIG) feature [1] that necessiated the presence > of a 1G region carved out from the device memory and mapped as > uncached. The 1G region is shown as a fake BAR (comprising region 2 and 3) > to workaround the issue. > > The Grace Blackwell systems (GB) differ from GH systems in the following > aspects: > 1. The aforementioned HW defect is fixed on GB systems. > 2. There is a usable BAR1 (region 2 and 3) on GB systems for the > GPUdirect RDMA feature [2]. > > This patch accommodate those GB changes by showing the 64b physical > device BAR1 (region2 and 3) to the VM instead of the fake one. This > takes care of both the differences. > > Moreover, the entire device memory is exposed on GB as cacheable to > the VM as there is no carveout required. > > Link: https://www.nvidia.com/en-in/technologies/multi-instance-gpu/ [1] > Link: https://docs.nvidia.com/cuda/gpudirect-rdma/ [2] > > Signed-off-by: Ankit Agrawal > --- > drivers/vfio/pci/nvgrace-gpu/main.c | 32 +++++++++++++++++++++-------- > 1 file changed, 24 insertions(+), 8 deletions(-) > > diff --git a/drivers/vfio/pci/nvgrace-gpu/main.c b/drivers/vfio/pci/nvgrace-gpu/main.c > index 85eacafaffdf..44a276c886e1 100644 > --- a/drivers/vfio/pci/nvgrace-gpu/main.c > +++ b/drivers/vfio/pci/nvgrace-gpu/main.c > @@ -72,7 +72,7 @@ nvgrace_gpu_memregion(int index, > if (index == USEMEM_REGION_INDEX) > return &nvdev->usemem; > > - if (index == RESMEM_REGION_INDEX) > + if (!nvdev->has_mig_hw_bug_fix && index == RESMEM_REGION_INDEX) > return &nvdev->resmem; > > return NULL; > @@ -715,6 +715,16 @@ static const struct vfio_device_ops nvgrace_gpu_pci_core_ops = { > .detach_ioas = vfio_iommufd_physical_detach_ioas, > }; > > +static void > +nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, > + struct nvgrace_gpu_pci_core_device *nvdev, > + u64 memphys, u64 memlength) > +{ > + nvdev->usemem.memphys = memphys; > + nvdev->usemem.memlength = memlength; > + nvdev->usemem.bar_size = roundup_pow_of_two(nvdev->usemem.memlength); > +} > + > static int > nvgrace_gpu_fetch_memory_property(struct pci_dev *pdev, > u64 *pmemphys, u64 *pmemlength) > @@ -752,9 +762,9 @@ nvgrace_gpu_fetch_memory_property(struct pci_dev *pdev, > } > > static int > -nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, > - struct nvgrace_gpu_pci_core_device *nvdev, > - u64 memphys, u64 memlength) > +nvgrace_gpu_nvdev_struct_workaround(struct pci_dev *pdev, > + struct nvgrace_gpu_pci_core_device *nvdev, > + u64 memphys, u64 memlength) > { > int ret = 0; > > @@ -864,10 +874,16 @@ static int nvgrace_gpu_probe(struct pci_dev *pdev, > * Device memory properties are identified in the host ACPI > * table. Set the nvgrace_gpu_pci_core_device structure. > */ > - ret = nvgrace_gpu_init_nvdev_struct(pdev, nvdev, > - memphys, memlength); > - if (ret) > - goto out_put_vdev; > + if (nvdev->has_mig_hw_bug_fix) { > + nvgrace_gpu_init_nvdev_struct(pdev, nvdev, > + memphys, memlength); > + } else { > + ret = nvgrace_gpu_nvdev_struct_workaround(pdev, nvdev, > + memphys, > + memlength); > + if (ret) > + goto out_put_vdev; > + } Doesn't this work out much more naturally if we just do something like: diff --git a/drivers/vfio/pci/nvgrace-gpu/main.c b/drivers/vfio/pci/nvgrace-gpu/main.c index 85eacafaffdf..43a9457442ff 100644 --- a/drivers/vfio/pci/nvgrace-gpu/main.c +++ b/drivers/vfio/pci/nvgrace-gpu/main.c @@ -17,9 +17,6 @@ #define RESMEM_REGION_INDEX VFIO_PCI_BAR2_REGION_INDEX #define USEMEM_REGION_INDEX VFIO_PCI_BAR4_REGION_INDEX -/* Memory size expected as non cached and reserved by the VM driver */ -#define RESMEM_SIZE SZ_1G - /* A hardwired and constant ABI value between the GPU FW and VFIO driver. */ #define MEMBLK_SIZE SZ_512M @@ -72,7 +69,7 @@ nvgrace_gpu_memregion(int index, if (index == USEMEM_REGION_INDEX) return &nvdev->usemem; - if (index == RESMEM_REGION_INDEX) + if (nvdev->resmem.memlength && index == RESMEM_REGION_INDEX) return &nvdev->resmem; return NULL; @@ -757,6 +754,13 @@ nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, u64 memphys, u64 memlength) { int ret = 0; + u64 resmem_size = 0; + + /* + * Comment about the GH bug that requires this and fix in GB + */ + if (!nvdev->has_mig_hw_bug_fix) + resmem_size = SZ_1G; /* * The VM GPU device driver needs a non-cacheable region to support @@ -780,7 +784,7 @@ nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, * memory (usemem) is added to the kernel for usage by the VM * workloads. Make the usable memory size memblock aligned. */ - if (check_sub_overflow(memlength, RESMEM_SIZE, + if (check_sub_overflow(memlength, resmem_size, &nvdev->usemem.memlength)) { ret = -EOVERFLOW; goto done; @@ -813,7 +817,9 @@ nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, * the BAR size for them. */ nvdev->usemem.bar_size = roundup_pow_of_two(nvdev->usemem.memlength); - nvdev->resmem.bar_size = roundup_pow_of_two(nvdev->resmem.memlength); + if (nvdev->resmem.memlength) + nvdev->resmem.bar_size = + roundup_pow_of_two(nvdev->resmem.memlength); done: return ret; } Thanks, Alex