From: "Grant Grundler" <grundler@google.com>
To: "FUJITA Tomonori" <fujita.tomonori@lab.ntt.co.jp>
Cc: linux-kernel@vger.kernel.org, mgross@linux.intel.com,
linux-scsi@vger.kernel.org
Subject: Re: Intel IOMMU (and IOMMU for Virtualization) performances
Date: Thu, 5 Jun 2008 11:34:56 -0700 [thread overview]
Message-ID: <da824cf30806051134u4fb8d419p2ba2dafcb1ba33a7@mail.gmail.com> (raw)
In-Reply-To: <20080605235322L.fujita.tomonori@lab.ntt.co.jp>
On Thu, Jun 5, 2008 at 7:49 AM, FUJITA Tomonori
<fujita.tomonori@lab.ntt.co.jp> wrote:
...
>> You can easily emulate SSD drives by doing sequential 4K reads
>> from a normal SATA HD. That should result in ~7-8K IOPS since the disk
>> will recognize the sequential stream and read ahead. SAS/SCSI/FC will
>> probably work the same way with different IOP rates.
>
> Yeah, probabaly right. I thought that 10GbE give the IOMMU more
> workloads than SSD does and tried to emulate something like that.
10GbE might exercise a different code path. NICs typically use map_single
and storage devices typically use map_sg. But they both exercise the same
underlying resource management code since it's the same IOMMU they poke at.
...
>> Sorry, I didn't see a replacement for the deferred_flush_tables.
>> Mark Gross and I agree this substantially helps with unmap performance.
>> See http://lkml.org/lkml/2008/3/3/373
>
> Yeah, I can add a nice trick in parisc sba_iommu uses. I'll try next
> time.
>
> But it probably gives the bitmap method less gain than the RB tree
> since clear the bitmap takes less time than changing the tree.
>
> The deferred_flush_tables also batches flushing TLB. The patch flushes
> TLB only when it reaches the end of the bitmap (it's a trick that some
> IOMMUs like SPARC does).
The batching of the TLB flushes is the key thing. I was being paranoid
by not marking the resource free until after the TLB was flushed. If we
know the allocation is going to be circular through the bitmap, flushing
the TLB once per iteration through the bitmap should be sufficient since
we can guarantee the IO Pdir resource won't get re-used until a full
cycle through the bitmap has been completed.
I expect this will work for parisc too and I can test that. Funny that didn't
"click" with me when I original wrote the parisc code. DaveM had even told
me the SPARC code was only flushing the IOTLB once per iteration.
...
> Agreed. VT-d can handle DMA virtual address space larger than 32 bits
> but it means that we need more memory for the bitmap. I think that the
> majority of systems don't need DMA virtual address space larger than
> 32 bits. Making it as a kernel parameter is a reasonable approach, I
> think.
Agreed. It needs a resonable default and a way to change it at runtime
for odd cases.
...
>> "32-PAGE_SHIFT_4K" expression is used in several places but I didn't see
>> an explanation of why 32. Can you add one someplace?
>
> OK, I'll do next time. Most of them are about 4GB virtual address
> space that the patch uses.
thanks! The comment should then explain why 4GB is "reasonable" (vs
1GB for example).
...
> Thanks a lot! I didn't expect this patch to be reviewed. I really
> appreciate it.
very welcome,
grant
next prev parent reply other threads:[~2008-06-05 18:35 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2008-06-04 14:47 FUJITA Tomonori
2008-06-04 16:56 ` Andi Kleen
2008-06-05 14:49 ` FUJITA Tomonori
2008-06-04 18:06 ` Grant Grundler
2008-06-05 14:49 ` FUJITA Tomonori
2008-06-05 18:34 ` Grant Grundler [this message]
2008-06-05 19:01 ` James Bottomley
2008-06-06 4:44 ` FUJITA Tomonori
2008-06-06 5:48 ` Grant Grundler
2008-06-09 9:36 ` FUJITA Tomonori
2008-06-06 20:23 ` Muli Ben-Yehuda
2008-06-06 20:21 ` Muli Ben-Yehuda
2008-06-06 21:28 ` Grant Grundler
2008-06-06 21:36 ` Muli Ben-Yehuda
2008-06-06 21:51 ` Grant Grundler
2008-06-09 8:17 ` Andi Kleen
2008-06-09 9:36 ` FUJITA Tomonori
2008-06-09 10:20 ` Andi Kleen
2008-06-05 22:02 ` mark gross
2008-06-06 4:44 ` FUJITA Tomonori
2008-06-23 17:54 ` mark gross
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=da824cf30806051134u4fb8d419p2ba2dafcb1ba33a7@mail.gmail.com \
--to=grundler@google.com \
--cc=fujita.tomonori@lab.ntt.co.jp \
--cc=linux-kernel@vger.kernel.org \
--cc=linux-scsi@vger.kernel.org \
--cc=mgross@linux.intel.com \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®