From: David Hildenbrand <david@redhat.com>
To: lizhe.67@bytedance.com, alex.williamson@redhat.com
Cc: kvm@vger.kernel.org, linux-kernel@vger.kernel.org,
muchun.song@linux.dev, peterx@redhat.com
Subject: Re: [PATCH v4] vfio/type1: optimize vfio_pin_pages_remote() for large folio
Date: Thu, 22 May 2025 09:22:50 +0200 [thread overview]
Message-ID: <81d73c4c-28c4-4fa0-bc71-aef6429e2c31@redhat.com> (raw)
In-Reply-To: <20250522034956.56617-1-lizhe.67@bytedance.com>
On 22.05.25 05:49, lizhe.67@bytedance.com wrote:
> On Wed, 21 May 2025 13:17:11 -0600, alex.williamson@redhat.com wrote:
>
>>> From: Li Zhe <lizhe.67@bytedance.com>
>>>
>>> When vfio_pin_pages_remote() is called with a range of addresses that
>>> includes large folios, the function currently performs individual
>>> statistics counting operations for each page. This can lead to significant
>>> performance overheads, especially when dealing with large ranges of pages.
>>>
>>> This patch optimize this process by batching the statistics counting
>>> operations.
>>>
>>> The performance test results for completing the 8G VFIO IOMMU DMA mapping,
>>> obtained through trace-cmd, are as follows. In this case, the 8G virtual
>>> address space has been mapped to physical memory using hugetlbfs with
>>> pagesize=2M.
>>>
>>> Before this patch:
>>> funcgraph_entry: # 33813.703 us | vfio_pin_map_dma();
>>>
>>> After this patch:
>>> funcgraph_entry: # 16071.378 us | vfio_pin_map_dma();
>>>
>>> Signed-off-by: Li Zhe <lizhe.67@bytedance.com>
>>> Co-developed-by: Alex Williamson <alex.williamson@redhat.com>
>>> Signed-off-by: Alex Williamson <alex.williamson@redhat.com>
>>> ---
>>
>> Given the discussion on v3, this is currently a Nak. Follow-up in that
>> thread if there are further ideas how to salvage this. Thanks,
>
> How about considering the solution David mentioned to check whether the
> pages or PFNs are actually consecutive?
>
> I have conducted a preliminary attempt, and the performance testing
> revealed that the time consumption is approximately 18,000 microseconds.
> Compared to the previous 33,000 microseconds, this also represents a
> significant improvement.
>
> The modification is quite straightforward. The code below reflects the
> changes I have made based on this patch.
>
> diff --git a/drivers/vfio/vfio_iommu_type1.c b/drivers/vfio/vfio_iommu_type1.c
> index bd46ed9361fe..1cc1f76d4020 100644
> --- a/drivers/vfio/vfio_iommu_type1.c
> +++ b/drivers/vfio/vfio_iommu_type1.c
> @@ -627,6 +627,19 @@ static long vaddr_get_pfns(struct mm_struct *mm, unsigned long vaddr,
> return ret;
> }
>
> +static inline long continuous_page_num(struct vfio_batch *batch, long npage)
> +{
> + long i;
> + unsigned long next_pfn = page_to_pfn(batch->pages[batch->offset]) + 1;
> +
> + for (i = 1; i < npage; ++i) {
> + if (page_to_pfn(batch->pages[batch->offset + i]) != next_pfn)
> + break;
> + next_pfn++;
> + }
> + return i;
> +}
What might be faster is obtaining the folio, and then calculating the
next expected page pointer, comparing whether the page pointers match.
Essentially, using folio_page() to calculate the expected next page.
nth_page() is a simple pointer arithmetic with CONFIG_SPARSEMEM_VMEMMAP,
so that might be rather fast.
So we'd obtain
start_idx = folio_idx(folio, batch->pages[batch->offset]);
and then check for
batch->pages[batch->offset + i] == folio_page(folio, start_idx + i)
--
Cheers,
David / dhildenb
next prev parent reply other threads:[~2025-05-22 7:22 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-05-21 4:25 [PATCH v4] vfio/type1: optimize vfio_pin_pages_remote() for large folios lizhe.67
2025-05-21 6:36 ` David Hildenbrand
2025-05-21 19:17 ` Alex Williamson
2025-05-22 3:49 ` [PATCH v4] vfio/type1: optimize vfio_pin_pages_remote() for large folio lizhe.67
2025-05-22 7:22 ` David Hildenbrand [this message]
2025-05-22 8:25 ` lizhe.67
2025-05-22 20:52 ` Alex Williamson
2025-05-23 3:42 ` lizhe.67
2025-05-23 14:54 ` Alex Williamson
2025-05-26 3:37 ` lizhe.67
2025-05-27 19:14 ` Alex Williamson
2025-05-28 4:21 ` lizhe.67
2025-05-28 20:11 ` Alex Williamson
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=81d73c4c-28c4-4fa0-bc71-aef6429e2c31@redhat.com \
--to=david@redhat.com \
--cc=alex.williamson@redhat.com \
--cc=kvm@vger.kernel.org \
--cc=linux-kernel@vger.kernel.org \
--cc=lizhe.67@bytedance.com \
--cc=muchun.song@linux.dev \
--cc=peterx@redhat.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®