From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id DA72F189F5C for ; Fri, 24 Jan 2025 16:24:59 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737735901; cv=none; b=n4Xu5r/r8YwBSk4KMpYcHr3PuLz+obcOj4juR3iiQ8YtGnEhoEWGN7YBlVCLTLHiCrILIgT818+s0UnSqwTqN2E1QumdhMRuP1TWljC/B9WStn1+4McU8sMNvLYRIuetEFfZNSVCQoQtsXjfOxQuXtevrEFR9gLo5CpXnuveNqU= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1737735901; c=relaxed/simple; bh=CWfK3tPR2GmT2twHa8FOPccArDnfK3LdlDUjKo2vxxk=; h=Date:From:To:Cc:Subject:Message-ID:In-Reply-To:References: MIME-Version:Content-Type; b=H/X4CzntPWMgPH561b6RTzmRG1bbidhTZrYCrniCsZ91FMTYLr845OY0U1Gy1fOth4qDy3bJIPJfVkOVRQE9QqeHrWPwKQAv3t3ONF7EHFh1rPgChYRMuzQ3XY4QihzUkAC9xJ8L7JvSe019H3xQ2NMsaXAH2bj4tBwus/3Vfxw= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=RGvUqAP6; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=none dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="RGvUqAP6" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1737735898; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: content-transfer-encoding:content-transfer-encoding: in-reply-to:in-reply-to:references:references; bh=Cn0+KzXyVl0N+dq6Z+l3EIZmrDN/yiykev2yHytr19o=; b=RGvUqAP6y8kUb6ub/dHGkOlzMxnhwLBmkdpmBEcBEa6VYxdTNtLiTsDbQ9qH7ae3yMbM/K CoKcyIyxS430MhdMgEAGTR9Fpv/qnsaxmPlg3sgmDPae2/ZgI9ltbV2gantOuIQhuOl1rr EFDWpW5J/a3aOPqwI2xn4nVZSBhEfUs= Received: from mail-il1-f200.google.com (mail-il1-f200.google.com [209.85.166.200]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-180-1baKD2y8OIG8iDaFEbrluA-1; Fri, 24 Jan 2025 11:24:57 -0500 X-MC-Unique: 1baKD2y8OIG8iDaFEbrluA-1 X-Mimecast-MFC-AGG-ID: 1baKD2y8OIG8iDaFEbrluA Received: by mail-il1-f200.google.com with SMTP id e9e14a558f8ab-3ce8c06be57so1613075ab.1 for ; Fri, 24 Jan 2025 08:24:57 -0800 (PST) X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1737735896; x=1738340696; h=content-transfer-encoding:mime-version:organization:references :in-reply-to:message-id:subject:cc:to:from:date:x-gm-message-state :from:to:cc:subject:date:message-id:reply-to; bh=Cn0+KzXyVl0N+dq6Z+l3EIZmrDN/yiykev2yHytr19o=; b=L2JVETzsT3IIqVLdikMLbPJklrCrwATLahCCyhGiagF39AkMBZCiWnweo3EbjHv7Qt OLOpHnQXZBmWLTf0TV6fsvaKRzmd2rqaEFwHoUbn8FNd7foA1QHGCzJ4IDKUzpkxISdL pg8NGd12YEiV1avmdfnwpQtn/ywKdS6poF/nlk99d2UzdHiOlkTxPzbtt3XrWa0W7khV jMEHFFcTd4W5nfFk9kcAnUjByF6VJHokLrpU2ApJhBOZEbh8VD1+qzeFSeuLIuEtPbba lTLtgQQALi7L7jqcS4bDyeXTIZH5hBF7e1SPy6cVV68i0lpuNNHHtevzgpxmoYPf/9kf gUMw== X-Forwarded-Encrypted: i=1; AJvYcCU8HDTP2r5bwgjl+U9UQGGlH+pb5PXoJL6LEpu8o32UtpRWm+MC5JmdAnTlr4vE1f//R+0K1Pe85Lhv82U=@vger.kernel.org X-Gm-Message-State: AOJu0YxbbB+m4xLxcQTbKbcrc2EM7udZYaCgFzBPS6tXqeyNniPz5HD1 i0yUzrnLnKPGcrNcrdxx5nP8offn2HpiN7LLOaHKg0+NIz+9vkgQCapZMbaCHF0SgR6sgBGwM/f OHqHyseLel5BdTXHMmKrk8NdklvEyMJjBVTV50VZnxruGDYbW0M2YTEwOFI4IuQ== X-Gm-Gg: ASbGncvsRy2pZIRIJ+cjm6hcUsXR+DjWk6fBlQRpns3jWhYvSRn+jec2MqprDylO8DV rXSjM2/HQo/7xYPH09I1jlvYtKiqHu6DVSyefmKY//lJWo3C61YFI1P0UQzWu9z0XK0Ot766+Vg +CWsUzotxa18MQ0MxRe3RFRqT0ZWOeP1UGOrURBuwZHVuqz/GRBZeJDT4xL9pBwV/CUUxRBWWfM LWsUA8vyXLrCjarLIKPVUrPiYw49eD8J25IasXiTg59yUrFvJnT8+Wrxt5l+RrtcNLNrZ2kOQ== X-Received: by 2002:a05:6e02:1707:b0:3cf:bc24:2336 with SMTP id e9e14a558f8ab-3cfbc2423f8mr20614635ab.5.1737735896506; Fri, 24 Jan 2025 08:24:56 -0800 (PST) X-Google-Smtp-Source: AGHT+IGs8kE7HGWUswtGXw81IR02XMeOqo9fhlNsHZw+bYVVSIj6GrGv/AhGfHW9ODHGR+Ewrm8cLA== X-Received: by 2002:a05:6e02:1707:b0:3cf:bc24:2336 with SMTP id e9e14a558f8ab-3cfbc2423f8mr20614535ab.5.1737735896110; Fri, 24 Jan 2025 08:24:56 -0800 (PST) Received: from redhat.com ([38.15.36.11]) by smtp.gmail.com with ESMTPSA id 8926c6da1cb9f-4ec1dbaf035sm694919173.134.2025.01.24.08.24.54 (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Fri, 24 Jan 2025 08:24:55 -0800 (PST) Date: Fri, 24 Jan 2025 09:24:53 -0700 From: Alex Williamson To: Cc: , , , , , , , , , , , , , , , , , , Subject: Re: [PATCH v5 1/3] vfio/nvgrace-gpu: Read dvsec register to determine need for uncached resmem Message-ID: <20250124092453.7d3df3d6.alex.williamson@redhat.com> In-Reply-To: <20250123174854.3338-2-ankita@nvidia.com> References: <20250123174854.3338-1-ankita@nvidia.com> <20250123174854.3338-2-ankita@nvidia.com> Organization: Red Hat Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset=US-ASCII Content-Transfer-Encoding: 7bit On Thu, 23 Jan 2025 17:48:52 +0000 wrote: > From: Ankit Agrawal > > NVIDIA's recently introduced Grace Blackwell (GB) Superchip is a > continuation with the Grace Hopper (GH) superchip that provides a > cache coherent access to CPU and GPU to each other's memory with > an internal proprietary chip-to-chip cache coherent interconnect. > > There is a HW defect on GH systems to support the Multi-Instance > GPU (MIG) feature [1] that necessiated the presence of a 1G region > with uncached mapping carved out from the device memory. The 1G > region is shown as a fake BAR (comprising region 2 and 3) to > workaround the issue. This is fixed on the GB systems. > > The presence of the fix for the HW defect is communicated by the > device firmware through the DVSEC PCI config register with ID 3. > The module reads this to take a different codepath on GB vs GH. > > Scan through the DVSEC registers to identify the correct one and use > it to determine the presence of the fix. Save the value in the device's > nvgrace_gpu_pci_core_device structure. > > Link: https://www.nvidia.com/en-in/technologies/multi-instance-gpu/ [1] > > CC: Jason Gunthorpe > CC: Kevin Tian > Signed-off-by: Ankit Agrawal > --- > drivers/vfio/pci/nvgrace-gpu/main.c | 30 +++++++++++++++++++++++++++++ > 1 file changed, 30 insertions(+) > > diff --git a/drivers/vfio/pci/nvgrace-gpu/main.c b/drivers/vfio/pci/nvgrace-gpu/main.c > index a467085038f0..dde2daa597f8 100644 > --- a/drivers/vfio/pci/nvgrace-gpu/main.c > +++ b/drivers/vfio/pci/nvgrace-gpu/main.c > @@ -23,6 +23,11 @@ > /* A hardwired and constant ABI value between the GPU FW and VFIO driver. */ > #define MEMBLK_SIZE SZ_512M > > +#define DVSEC_BITMAP_OFFSET 0xA > +#define MIG_SUPPORTED_WITH_CACHED_RESMEM BIT(0) > + > +#define GPU_CAP_DVSEC_REGISTER 3 > + > /* > * The state of the two device memory region - resmem and usemem - is > * saved as struct mem_region. > @@ -46,6 +51,7 @@ struct nvgrace_gpu_pci_core_device { > struct mem_region resmem; > /* Lock to control device memory kernel mapping */ > struct mutex remap_lock; > + bool has_mig_hw_bug; > }; > > static void nvgrace_gpu_init_fake_bar_emu_regs(struct vfio_device *core_vdev) > @@ -812,6 +818,26 @@ nvgrace_gpu_init_nvdev_struct(struct pci_dev *pdev, > return ret; > } > > +static bool nvgrace_gpu_has_mig_hw_bug(struct pci_dev *pdev) > +{ > + int pcie_dvsec; > + u16 dvsec_ctrl16; > + > + pcie_dvsec = pci_find_dvsec_capability(pdev, PCI_VENDOR_ID_NVIDIA, > + GPU_CAP_DVSEC_REGISTER); > + > + if (pcie_dvsec) { > + pci_read_config_word(pdev, > + pcie_dvsec + DVSEC_BITMAP_OFFSET, > + &dvsec_ctrl16); > + > + if (dvsec_ctrl16 & MIG_SUPPORTED_WITH_CACHED_RESMEM) > + return false; > + } > + > + return true; > +} > + > static int nvgrace_gpu_probe(struct pci_dev *pdev, > const struct pci_device_id *id) > { > @@ -832,6 +858,8 @@ static int nvgrace_gpu_probe(struct pci_dev *pdev, > dev_set_drvdata(&pdev->dev, &nvdev->core_device); > > if (ops == &nvgrace_gpu_pci_ops) { > + nvdev->has_mig_hw_bug = nvgrace_gpu_has_mig_hw_bug(pdev); > + > /* > * Device memory properties are identified in the host ACPI > * table. Set the nvgrace_gpu_pci_core_device structure. > @@ -868,6 +896,8 @@ static const struct pci_device_id nvgrace_gpu_vfio_pci_table[] = { > { PCI_DRIVER_OVERRIDE_DEVICE_VFIO(PCI_VENDOR_ID_NVIDIA, 0x2345) }, > /* GH200 SKU */ > { PCI_DRIVER_OVERRIDE_DEVICE_VFIO(PCI_VENDOR_ID_NVIDIA, 0x2348) }, > + /* GB200 SKU */ > + { PCI_DRIVER_OVERRIDE_DEVICE_VFIO(PCI_VENDOR_ID_NVIDIA, 0x2941) }, > {} > }; > GB support isn't really complete until patch 3, so shouldn't we hold off on adding the ID to the table until a trivial patch 4, adding only the chunk above? Thanks, Alex