From: Sheng Zhao <sheng.zhao@bytedance.com>
To: Jason Wang <jasowang@redhat.com>
Cc: mst@redhat.com, xuanzhuo@linux.alibaba.com, eperezma@redhat.com,
virtualization@lists.linux.dev, linux-kernel@vger.kernel.org,
xieyongji@bytedance.com
Subject: Re: Re: [PATCH] vduse: Use fixed 4KB bounce pages for arm64 64KB page size
Date: Wed, 24 Sep 2025 12:05:36 +0800 [thread overview]
Message-ID: <aed938d0-e70a-4af6-9950-d4d0b7d6a93f@bytedance.com> (raw)
In-Reply-To: <CACGkMEv3pUBF3Uv2s3MM0Qn--fP3mwN92SqE9NX4gNuMALBTUg@mail.gmail.com>
On 2025/9/24 08:57, Jason Wang wrote:
> On Tue, Sep 23, 2025 at 8:37 PM Sheng Zhao <sheng.zhao@bytedance.com> wrote:
>>
>>
>>
>> On 2025/9/17 16:16, Jason Wang wrote:
>>> On Mon, Sep 15, 2025 at 3:34 PM <sheng.zhao@bytedance.com> wrote:
>>>>
>>>> From: Sheng Zhao <sheng.zhao@bytedance.com>
>>>>
>>>> The allocation granularity of bounce pages is PAGE_SIZE. This may cause
>>>> even small IO requests to occupy an entire bounce page exclusively. The
>>>> kind of memory waste will be more significant on arm64 with 64KB pages.
>>>
>>> Let's tweak the title as there are archs that are using non 4KB pages
>>> other than arm.
>>>
>>
>> Got it. I will modify this in v2.
>>
>>>>
>>>> So, optimize it by using fixed 4KB bounce pages.
>>>>
>>>> Signed-off-by: Sheng Zhao <sheng.zhao@bytedance.com>
>>>> ---
>>>> drivers/vdpa/vdpa_user/iova_domain.c | 120 +++++++++++++++++----------
>>>> drivers/vdpa/vdpa_user/iova_domain.h | 5 ++
>>>> 2 files changed, 83 insertions(+), 42 deletions(-)
>>>>
>>>> diff --git a/drivers/vdpa/vdpa_user/iova_domain.c b/drivers/vdpa/vdpa_user/iova_domain.c
>>>> index 58116f89d8da..768313c80b62 100644
>>>> --- a/drivers/vdpa/vdpa_user/iova_domain.c
>>>> +++ b/drivers/vdpa/vdpa_user/iova_domain.c
>>>> @@ -103,19 +103,26 @@ void vduse_domain_clear_map(struct vduse_iova_domain *domain,
>>>> static int vduse_domain_map_bounce_page(struct vduse_iova_domain *domain,
>>>> u64 iova, u64 size, u64 paddr)
>>>> {
>>>> - struct vduse_bounce_map *map;
>>>> + struct vduse_bounce_map *map, *head_map;
>>>> + struct page *tmp_page;
>>>> u64 last = iova + size - 1;
>>>>
>>>> while (iova <= last) {
>>>> - map = &domain->bounce_maps[iova >> PAGE_SHIFT];
>>>> + map = &domain->bounce_maps[iova >> BOUNCE_PAGE_SHIFT];
>>>
>>> BOUNCE_PAGE_SIZE is kind of confusing as it's not the size of any page
>>> at all when PAGE_SIZE is not 4K.
>>>
>>
>> How about BOUNCE_MAP_SIZE?
>
> Fine with me.
>
>>
>>>> if (!map->bounce_page) {
>>>> - map->bounce_page = alloc_page(GFP_ATOMIC);
>>>> - if (!map->bounce_page)
>>>> - return -ENOMEM;
>>>> + head_map = &domain->bounce_maps[(iova & PAGE_MASK) >> BOUNCE_PAGE_SHIFT];
>>>> + if (!head_map->bounce_page) {
>>>> + tmp_page = alloc_page(GFP_ATOMIC);
>>>> + if (!tmp_page)
>>>> + return -ENOMEM;
>>>> + if (cmpxchg(&head_map->bounce_page, NULL, tmp_page))
>>>> + __free_page(tmp_page);
>>>
>>> I don't understand why we need cmpxchg() logic.
>>>
>>> Btw, it looks like you want to make multiple bounce_map to point to
>>> the same 64KB page? I wonder what's the advantages of doing this. Can
>>> we simply keep the 64KB page in bounce_map?
>>>
>>> Thanks
>>>
>>
>> That's correct. We use fixed 4KB-sized bounce pages, and there will be a
>> many-to-one relationship between these 4KB bounce pages and the 64KB
>> memory pages.
>>
>> Bounce pages are allocated on demand. As a result, it may occur that
>> multiple bounce pages corresponding to the same 64KB memory page attempt
>> to allocate memory simultaneously, so we use cmpxchg to handle this
>> concurrency.
>>
>> In the current implementation, the bounce_map structure requires no
>> modification. However, if we keep the 64KB page into a single bounce_map
>> while still wanting to implement a similar logic, we may need an
>> additional array to store multiple orig_phys values in order to
>> accommodate the many-to-one relationship.
>
> Or simply having a bitmap is sufficient per bounce_map?
>
Yes, using a bitmap can mark the usage status of each 4KB, but it may
not simplify things overall.
- we will inevitably need to add an additional array per bounce_map to
store the orig_phys corresponding to each 4KB for subsequent copying
(vduse_domain_bounce).
- compared to the current commit, this modification may only be a
structural change and fail to reduce the amount of changes to the code
logic. For instance, cmpxchg is still required.
Thanks
> Thanks
>
>>
>> Thanks
>>
>
next prev parent reply other threads:[~2025-09-24 4:05 UTC|newest]
Thread overview: 13+ messages / expand[flat|nested] mbox.gz Atom feed top
2025-09-15 7:34 sheng.zhao
2025-09-15 8:21 ` Jason Wang
2025-09-15 11:07 ` [External] " Sheng Zhao
2025-09-16 7:34 ` Jason Wang
2025-09-16 13:54 ` Sheng Zhao
2025-09-17 8:16 ` Jason Wang
2025-09-23 12:37 ` Sheng Zhao
2025-09-24 0:57 ` Jason Wang
2025-09-24 4:05 ` Sheng Zhao [this message]
2025-09-24 4:15 ` Jason Wang
2025-09-24 6:38 ` Sheng Zhao
2025-09-25 0:20 ` Jason Wang
2025-09-25 3:13 ` Sheng Zhao
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=aed938d0-e70a-4af6-9950-d4d0b7d6a93f@bytedance.com \
--to=sheng.zhao@bytedance.com \
--cc=eperezma@redhat.com \
--cc=jasowang@redhat.com \
--cc=linux-kernel@vger.kernel.org \
--cc=mst@redhat.com \
--cc=virtualization@lists.linux.dev \
--cc=xieyongji@bytedance.com \
--cc=xuanzhuo@linux.alibaba.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®