* dma_sync_sg_for_cpu applied to a single scatterlist element
@ 2010-03-15 21:30 Alan Stern
2010-03-15 22:59 ` James Bottomley
2010-03-15 23:16 ` FUJITA Tomonori
0 siblings, 2 replies; 8+ messages in thread
From: Alan Stern @ 2010-03-15 21:30 UTC (permalink / raw)
To: James Bottomley; +Cc: Kernel development list
This is addressed to James Bottomley as he is the author of
Documentation/DMA-API.txt, but anyone else who can contribute is
invited to do so.
Suppose a scatter-gather transfer with multiple scatterlist elements
has been mapped via dma_map_sg(). Is it then valid to call
dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the
mapped scatterlist elements (not necessarily the first one) and the
"nelems" argument set to 1?
Thanks,
Alan Stern
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element
2010-03-15 21:30 dma_sync_sg_for_cpu applied to a single scatterlist element Alan Stern
@ 2010-03-15 22:59 ` James Bottomley
2010-03-16 14:30 ` Alan Stern
2010-03-15 23:16 ` FUJITA Tomonori
1 sibling, 1 reply; 8+ messages in thread
From: James Bottomley @ 2010-03-15 22:59 UTC (permalink / raw)
To: Alan Stern; +Cc: Kernel development list
On Mon, 2010-03-15 at 17:30 -0400, Alan Stern wrote:
> This is addressed to James Bottomley as he is the author of
> Documentation/DMA-API.txt, but anyone else who can contribute is
> invited to do so.
>
> Suppose a scatter-gather transfer with multiple scatterlist elements
> has been mapped via dma_map_sg(). Is it then valid to call
> dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the
> mapped scatterlist elements (not necessarily the first one) and the
> "nelems" argument set to 1?
It's not the design of the API, but I'm guessing, given the way the API
works on most arch's that it will work. However, if you just want a
single element sync'd, won't dma_sync_single_for_cpu do that
transparently (as in just feed in the address and length from the sg
list), without mucking with the sg API?
James
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element
2010-03-15 21:30 dma_sync_sg_for_cpu applied to a single scatterlist element Alan Stern
2010-03-15 22:59 ` James Bottomley
@ 2010-03-15 23:16 ` FUJITA Tomonori
1 sibling, 0 replies; 8+ messages in thread
From: FUJITA Tomonori @ 2010-03-15 23:16 UTC (permalink / raw)
To: stern; +Cc: James.Bottomley, linux-kernel
On Mon, 15 Mar 2010 17:30:19 -0400 (EDT)
Alan Stern <stern@rowland.harvard.edu> wrote:
> This is addressed to James Bottomley as he is the author of
> Documentation/DMA-API.txt, but anyone else who can contribute is
> invited to do so.
>
> Suppose a scatter-gather transfer with multiple scatterlist elements
> has been mapped via dma_map_sg(). Is it then valid to call
> dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the
> mapped scatterlist elements (not necessarily the first one) and the
> "nelems" argument set to 1?
As James said, probably it works. As long as passed scatterlist
elements points to mapped regions, it works.
However, I think that the latest DMA-API.txt makes it clear that only
dma_sync_single_for_{cpu|device} supports a partial sync. So it's not
recommended, I guess.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element
2010-03-15 22:59 ` James Bottomley
@ 2010-03-16 14:30 ` Alan Stern
2010-03-16 22:28 ` FUJITA Tomonori
0 siblings, 1 reply; 8+ messages in thread
From: Alan Stern @ 2010-03-16 14:30 UTC (permalink / raw)
To: James Bottomley, FUJITA Tomonori; +Cc: Kernel development list
On Mon, 15 Mar 2010, James Bottomley wrote:
> On Mon, 2010-03-15 at 17:30 -0400, Alan Stern wrote:
> > This is addressed to James Bottomley as he is the author of
> > Documentation/DMA-API.txt, but anyone else who can contribute is
> > invited to do so.
> >
> > Suppose a scatter-gather transfer with multiple scatterlist elements
> > has been mapped via dma_map_sg(). Is it then valid to call
> > dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the
> > mapped scatterlist elements (not necessarily the first one) and the
> > "nelems" argument set to 1?
>
> It's not the design of the API, but I'm guessing, given the way the API
> works on most arch's that it will work. However, if you just want a
> single element sync'd, won't dma_sync_single_for_cpu do that
> transparently (as in just feed in the address and length from the sg
> list), without mucking with the sg API?
On Tue, 16 Mar 2010, FUJITA Tomonori wrote:
> As James said, probably it works. As long as passed scatterlist
> elements points to mapped regions, it works.
>
> However, I think that the latest DMA-API.txt makes it clear that only
> dma_sync_single_for_{cpu|device} supports a partial sync. So it's not
> recommended, I guess.
What the documentation actually says about the dma_sync_* functions is:
All the parameters must be the same as those passed into the
single mapping API.
So it isn't clear that dma_sync_sg_for_cpu(dev, sg, 1, dir) can be used
on a mapping created by dma_map_sg(dev, sg, n, dir), and it isn't
clear that dma_sync_single_for_cpu() can be used on a mapping created
by dma_map_sg().
But if you guys say it will work, I'll go ahead and use
dma_sync_single_for_cpu().
Thanks,
Alan Stern
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element
2010-03-16 14:30 ` Alan Stern
@ 2010-03-16 22:28 ` FUJITA Tomonori
2010-03-17 15:22 ` Alan Stern
0 siblings, 1 reply; 8+ messages in thread
From: FUJITA Tomonori @ 2010-03-16 22:28 UTC (permalink / raw)
To: stern; +Cc: James.Bottomley, fujita.tomonori, linux-kernel
On Tue, 16 Mar 2010 10:30:47 -0400 (EDT)
Alan Stern <stern@rowland.harvard.edu> wrote:
> On Mon, 15 Mar 2010, James Bottomley wrote:
>
> > On Mon, 2010-03-15 at 17:30 -0400, Alan Stern wrote:
> > > This is addressed to James Bottomley as he is the author of
> > > Documentation/DMA-API.txt, but anyone else who can contribute is
> > > invited to do so.
> > >
> > > Suppose a scatter-gather transfer with multiple scatterlist elements
> > > has been mapped via dma_map_sg(). Is it then valid to call
> > > dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the
> > > mapped scatterlist elements (not necessarily the first one) and the
> > > "nelems" argument set to 1?
> >
> > It's not the design of the API, but I'm guessing, given the way the API
> > works on most arch's that it will work. However, if you just want a
> > single element sync'd, won't dma_sync_single_for_cpu do that
> > transparently (as in just feed in the address and length from the sg
> > list), without mucking with the sg API?
>
> On Tue, 16 Mar 2010, FUJITA Tomonori wrote:
>
> > As James said, probably it works. As long as passed scatterlist
> > elements points to mapped regions, it works.
> >
> > However, I think that the latest DMA-API.txt makes it clear that only
> > dma_sync_single_for_{cpu|device} supports a partial sync. So it's not
> > recommended, I guess.
>
> What the documentation actually says about the dma_sync_* functions is:
>
> All the parameters must be the same as those passed into the
> single mapping API.
It's true to dma_sync_sg_for_*. You can see there:
With the sync_single API, you can use dma_handle and size parameters
that aren't identical to those passed into the single mapping API to
do a partial sync.
dma_sync_single_for_* can do a partial sync but dma_sync_sg_for_*
doesn't support a partial sync.
> So it isn't clear that dma_sync_sg_for_cpu(dev, sg, 1, dir) can be used
> on a mapping created by dma_map_sg(dev, sg, n, dir),
You should not do (though it might work).
> and it isn't
> clear that dma_sync_single_for_cpu() can be used on a mapping created
> by dma_map_sg().
You should not do (though it might work).
> But if you guys say it will work, I'll go ahead and use
> dma_sync_single_for_cpu().
>
Well, it's undocumented. It might work but might not.
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element
2010-03-16 22:28 ` FUJITA Tomonori
@ 2010-03-17 15:22 ` Alan Stern
2010-03-18 1:46 ` FUJITA Tomonori
0 siblings, 1 reply; 8+ messages in thread
From: Alan Stern @ 2010-03-17 15:22 UTC (permalink / raw)
To: FUJITA Tomonori; +Cc: James.Bottomley, linux-kernel
On Wed, 17 Mar 2010, FUJITA Tomonori wrote:
> dma_sync_single_for_* can do a partial sync but dma_sync_sg_for_*
> doesn't support a partial sync.
>
>
> > So it isn't clear that dma_sync_sg_for_cpu(dev, sg, 1, dir) can be used
> > on a mapping created by dma_map_sg(dev, sg, n, dir),
>
> You should not do (though it might work).
>
> > and it isn't
> > clear that dma_sync_single_for_cpu() can be used on a mapping created
> > by dma_map_sg().
>
> You should not do (though it might work).
>
>
> > But if you guys say it will work, I'll go ahead and use
> > dma_sync_single_for_cpu().
> >
>
> Well, it's undocumented. It might work but might not.
It's a real problem. I need it to work correctly.
Here's the situation. The USB controller drivers don't all support
scatter-gather operation. So there's a library routine in the USB core
which calls dma_map_sg() and then creates a separate I/O request for
each scatterlist element. The driver can process these requests one at
a time, and when they are all finished the library routine calls
dma_unmap_sg().
However... For tracing purposes (usbmon -- like tcpdump but for USB),
we may need to copy the data from each I/O request's transfer buffer.
Unfortunately, this copying is done as each request is submitted (for
output) or as it completes (for input), at which times the buffers are
all mapped for DMA. That's the problem.
Would we be better off not using dma_map_sg() at all in this situation?
We could map each scatterlist buffer individually with dma_map_single()
as the request is submitted and then unmap the buffer when the request
completes, just like with non-sg transfers; then the problem wouldn't
arise.
Would there be any significant penalty for doing this? I realize it
would prevent adjacent buffers from getting coalesced, but that's
probably okay. Any other reason not to?
Alan Stern
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element
2010-03-17 15:22 ` Alan Stern
@ 2010-03-18 1:46 ` FUJITA Tomonori
2010-03-18 14:13 ` Alan Stern
0 siblings, 1 reply; 8+ messages in thread
From: FUJITA Tomonori @ 2010-03-18 1:46 UTC (permalink / raw)
To: stern; +Cc: fujita.tomonori, James.Bottomley, linux-kernel
On Wed, 17 Mar 2010 11:22:01 -0400 (EDT)
Alan Stern <stern@rowland.harvard.edu> wrote:
> Here's the situation. The USB controller drivers don't all support
> scatter-gather operation. So there's a library routine in the USB core
> which calls dma_map_sg() and then creates a separate I/O request for
> each scatterlist element. The driver can process these requests one at
> a time, and when they are all finished the library routine calls
> dma_unmap_sg().
>
> However... For tracing purposes (usbmon -- like tcpdump but for USB),
> we may need to copy the data from each I/O request's transfer buffer.
> Unfortunately, this copying is done as each request is submitted (for
> output) or as it completes (for input), at which times the buffers are
> all mapped for DMA. That's the problem.
>
> Would we be better off not using dma_map_sg() at all in this situation?
> We could map each scatterlist buffer individually with dma_map_single()
> as the request is submitted and then unmap the buffer when the request
> completes, just like with non-sg transfers; then the problem wouldn't
> arise.
>
> Would there be any significant penalty for doing this? I realize it
> would prevent adjacent buffers from getting coalesced, but that's
> probably okay. Any other reason not to?
No reason. About merging adjacent buffers, there are few IOMMU
implementations that do. The recent IOMMU implementations such as VT-d
and AMD IOMMU don't.
If drivers don't support scatter-gather operation, they had better use
dma_map_page() instead of forging scatter-gather lists and playing
with dma_map_sg().
^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element
2010-03-18 1:46 ` FUJITA Tomonori
@ 2010-03-18 14:13 ` Alan Stern
0 siblings, 0 replies; 8+ messages in thread
From: Alan Stern @ 2010-03-18 14:13 UTC (permalink / raw)
To: FUJITA Tomonori; +Cc: James.Bottomley, linux-kernel
On Thu, 18 Mar 2010, FUJITA Tomonori wrote:
> On Wed, 17 Mar 2010 11:22:01 -0400 (EDT)
> Alan Stern <stern@rowland.harvard.edu> wrote:
>
> > Here's the situation. The USB controller drivers don't all support
> > scatter-gather operation. So there's a library routine in the USB core
> > which calls dma_map_sg() and then creates a separate I/O request for
> > each scatterlist element. The driver can process these requests one at
> > a time, and when they are all finished the library routine calls
> > dma_unmap_sg().
> >
> > However... For tracing purposes (usbmon -- like tcpdump but for USB),
> > we may need to copy the data from each I/O request's transfer buffer.
> > Unfortunately, this copying is done as each request is submitted (for
> > output) or as it completes (for input), at which times the buffers are
> > all mapped for DMA. That's the problem.
> >
> > Would we be better off not using dma_map_sg() at all in this situation?
> > We could map each scatterlist buffer individually with dma_map_single()
> > as the request is submitted and then unmap the buffer when the request
> > completes, just like with non-sg transfers; then the problem wouldn't
> > arise.
> >
> > Would there be any significant penalty for doing this? I realize it
> > would prevent adjacent buffers from getting coalesced, but that's
> > probably okay. Any other reason not to?
>
> No reason. About merging adjacent buffers, there are few IOMMU
> implementations that do. The recent IOMMU implementations such as VT-d
> and AMD IOMMU don't.
>
> If drivers don't support scatter-gather operation, they had better use
> dma_map_page() instead of forging scatter-gather lists and playing
> with dma_map_sg().
Okay, I'll handle it that way. Thanks for the advice.
Alan Stern
^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2010-03-18 14:14 UTC | newest]
Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2010-03-15 21:30 dma_sync_sg_for_cpu applied to a single scatterlist element Alan Stern
2010-03-15 22:59 ` James Bottomley
2010-03-16 14:30 ` Alan Stern
2010-03-16 22:28 ` FUJITA Tomonori
2010-03-17 15:22 ` Alan Stern
2010-03-18 1:46 ` FUJITA Tomonori
2010-03-18 14:13 ` Alan Stern
2010-03-15 23:16 ` FUJITA Tomonori
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®