* dma_sync_sg_for_cpu applied to a single scatterlist element @ 2010-03-15 21:30 Alan Stern 2010-03-15 22:59 ` James Bottomley 2010-03-15 23:16 ` FUJITA Tomonori 0 siblings, 2 replies; 8+ messages in thread From: Alan Stern @ 2010-03-15 21:30 UTC (permalink / raw) To: James Bottomley; +Cc: Kernel development list This is addressed to James Bottomley as he is the author of Documentation/DMA-API.txt, but anyone else who can contribute is invited to do so. Suppose a scatter-gather transfer with multiple scatterlist elements has been mapped via dma_map_sg(). Is it then valid to call dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the mapped scatterlist elements (not necessarily the first one) and the "nelems" argument set to 1? Thanks, Alan Stern ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element 2010-03-15 21:30 dma_sync_sg_for_cpu applied to a single scatterlist element Alan Stern @ 2010-03-15 22:59 ` James Bottomley 2010-03-16 14:30 ` Alan Stern 2010-03-15 23:16 ` FUJITA Tomonori 1 sibling, 1 reply; 8+ messages in thread From: James Bottomley @ 2010-03-15 22:59 UTC (permalink / raw) To: Alan Stern; +Cc: Kernel development list On Mon, 2010-03-15 at 17:30 -0400, Alan Stern wrote: > This is addressed to James Bottomley as he is the author of > Documentation/DMA-API.txt, but anyone else who can contribute is > invited to do so. > > Suppose a scatter-gather transfer with multiple scatterlist elements > has been mapped via dma_map_sg(). Is it then valid to call > dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the > mapped scatterlist elements (not necessarily the first one) and the > "nelems" argument set to 1? It's not the design of the API, but I'm guessing, given the way the API works on most arch's that it will work. However, if you just want a single element sync'd, won't dma_sync_single_for_cpu do that transparently (as in just feed in the address and length from the sg list), without mucking with the sg API? James ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element 2010-03-15 22:59 ` James Bottomley @ 2010-03-16 14:30 ` Alan Stern 2010-03-16 22:28 ` FUJITA Tomonori 0 siblings, 1 reply; 8+ messages in thread From: Alan Stern @ 2010-03-16 14:30 UTC (permalink / raw) To: James Bottomley, FUJITA Tomonori; +Cc: Kernel development list On Mon, 15 Mar 2010, James Bottomley wrote: > On Mon, 2010-03-15 at 17:30 -0400, Alan Stern wrote: > > This is addressed to James Bottomley as he is the author of > > Documentation/DMA-API.txt, but anyone else who can contribute is > > invited to do so. > > > > Suppose a scatter-gather transfer with multiple scatterlist elements > > has been mapped via dma_map_sg(). Is it then valid to call > > dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the > > mapped scatterlist elements (not necessarily the first one) and the > > "nelems" argument set to 1? > > It's not the design of the API, but I'm guessing, given the way the API > works on most arch's that it will work. However, if you just want a > single element sync'd, won't dma_sync_single_for_cpu do that > transparently (as in just feed in the address and length from the sg > list), without mucking with the sg API? On Tue, 16 Mar 2010, FUJITA Tomonori wrote: > As James said, probably it works. As long as passed scatterlist > elements points to mapped regions, it works. > > However, I think that the latest DMA-API.txt makes it clear that only > dma_sync_single_for_{cpu|device} supports a partial sync. So it's not > recommended, I guess. What the documentation actually says about the dma_sync_* functions is: All the parameters must be the same as those passed into the single mapping API. So it isn't clear that dma_sync_sg_for_cpu(dev, sg, 1, dir) can be used on a mapping created by dma_map_sg(dev, sg, n, dir), and it isn't clear that dma_sync_single_for_cpu() can be used on a mapping created by dma_map_sg(). But if you guys say it will work, I'll go ahead and use dma_sync_single_for_cpu(). Thanks, Alan Stern ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element 2010-03-16 14:30 ` Alan Stern @ 2010-03-16 22:28 ` FUJITA Tomonori 2010-03-17 15:22 ` Alan Stern 0 siblings, 1 reply; 8+ messages in thread From: FUJITA Tomonori @ 2010-03-16 22:28 UTC (permalink / raw) To: stern; +Cc: James.Bottomley, fujita.tomonori, linux-kernel On Tue, 16 Mar 2010 10:30:47 -0400 (EDT) Alan Stern <stern@rowland.harvard.edu> wrote: > On Mon, 15 Mar 2010, James Bottomley wrote: > > > On Mon, 2010-03-15 at 17:30 -0400, Alan Stern wrote: > > > This is addressed to James Bottomley as he is the author of > > > Documentation/DMA-API.txt, but anyone else who can contribute is > > > invited to do so. > > > > > > Suppose a scatter-gather transfer with multiple scatterlist elements > > > has been mapped via dma_map_sg(). Is it then valid to call > > > dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the > > > mapped scatterlist elements (not necessarily the first one) and the > > > "nelems" argument set to 1? > > > > It's not the design of the API, but I'm guessing, given the way the API > > works on most arch's that it will work. However, if you just want a > > single element sync'd, won't dma_sync_single_for_cpu do that > > transparently (as in just feed in the address and length from the sg > > list), without mucking with the sg API? > > On Tue, 16 Mar 2010, FUJITA Tomonori wrote: > > > As James said, probably it works. As long as passed scatterlist > > elements points to mapped regions, it works. > > > > However, I think that the latest DMA-API.txt makes it clear that only > > dma_sync_single_for_{cpu|device} supports a partial sync. So it's not > > recommended, I guess. > > What the documentation actually says about the dma_sync_* functions is: > > All the parameters must be the same as those passed into the > single mapping API. It's true to dma_sync_sg_for_*. You can see there: With the sync_single API, you can use dma_handle and size parameters that aren't identical to those passed into the single mapping API to do a partial sync. dma_sync_single_for_* can do a partial sync but dma_sync_sg_for_* doesn't support a partial sync. > So it isn't clear that dma_sync_sg_for_cpu(dev, sg, 1, dir) can be used > on a mapping created by dma_map_sg(dev, sg, n, dir), You should not do (though it might work). > and it isn't > clear that dma_sync_single_for_cpu() can be used on a mapping created > by dma_map_sg(). You should not do (though it might work). > But if you guys say it will work, I'll go ahead and use > dma_sync_single_for_cpu(). > Well, it's undocumented. It might work but might not. ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element 2010-03-16 22:28 ` FUJITA Tomonori @ 2010-03-17 15:22 ` Alan Stern 2010-03-18 1:46 ` FUJITA Tomonori 0 siblings, 1 reply; 8+ messages in thread From: Alan Stern @ 2010-03-17 15:22 UTC (permalink / raw) To: FUJITA Tomonori; +Cc: James.Bottomley, linux-kernel On Wed, 17 Mar 2010, FUJITA Tomonori wrote: > dma_sync_single_for_* can do a partial sync but dma_sync_sg_for_* > doesn't support a partial sync. > > > > So it isn't clear that dma_sync_sg_for_cpu(dev, sg, 1, dir) can be used > > on a mapping created by dma_map_sg(dev, sg, n, dir), > > You should not do (though it might work). > > > and it isn't > > clear that dma_sync_single_for_cpu() can be used on a mapping created > > by dma_map_sg(). > > You should not do (though it might work). > > > > But if you guys say it will work, I'll go ahead and use > > dma_sync_single_for_cpu(). > > > > Well, it's undocumented. It might work but might not. It's a real problem. I need it to work correctly. Here's the situation. The USB controller drivers don't all support scatter-gather operation. So there's a library routine in the USB core which calls dma_map_sg() and then creates a separate I/O request for each scatterlist element. The driver can process these requests one at a time, and when they are all finished the library routine calls dma_unmap_sg(). However... For tracing purposes (usbmon -- like tcpdump but for USB), we may need to copy the data from each I/O request's transfer buffer. Unfortunately, this copying is done as each request is submitted (for output) or as it completes (for input), at which times the buffers are all mapped for DMA. That's the problem. Would we be better off not using dma_map_sg() at all in this situation? We could map each scatterlist buffer individually with dma_map_single() as the request is submitted and then unmap the buffer when the request completes, just like with non-sg transfers; then the problem wouldn't arise. Would there be any significant penalty for doing this? I realize it would prevent adjacent buffers from getting coalesced, but that's probably okay. Any other reason not to? Alan Stern ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element 2010-03-17 15:22 ` Alan Stern @ 2010-03-18 1:46 ` FUJITA Tomonori 2010-03-18 14:13 ` Alan Stern 0 siblings, 1 reply; 8+ messages in thread From: FUJITA Tomonori @ 2010-03-18 1:46 UTC (permalink / raw) To: stern; +Cc: fujita.tomonori, James.Bottomley, linux-kernel On Wed, 17 Mar 2010 11:22:01 -0400 (EDT) Alan Stern <stern@rowland.harvard.edu> wrote: > Here's the situation. The USB controller drivers don't all support > scatter-gather operation. So there's a library routine in the USB core > which calls dma_map_sg() and then creates a separate I/O request for > each scatterlist element. The driver can process these requests one at > a time, and when they are all finished the library routine calls > dma_unmap_sg(). > > However... For tracing purposes (usbmon -- like tcpdump but for USB), > we may need to copy the data from each I/O request's transfer buffer. > Unfortunately, this copying is done as each request is submitted (for > output) or as it completes (for input), at which times the buffers are > all mapped for DMA. That's the problem. > > Would we be better off not using dma_map_sg() at all in this situation? > We could map each scatterlist buffer individually with dma_map_single() > as the request is submitted and then unmap the buffer when the request > completes, just like with non-sg transfers; then the problem wouldn't > arise. > > Would there be any significant penalty for doing this? I realize it > would prevent adjacent buffers from getting coalesced, but that's > probably okay. Any other reason not to? No reason. About merging adjacent buffers, there are few IOMMU implementations that do. The recent IOMMU implementations such as VT-d and AMD IOMMU don't. If drivers don't support scatter-gather operation, they had better use dma_map_page() instead of forging scatter-gather lists and playing with dma_map_sg(). ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element 2010-03-18 1:46 ` FUJITA Tomonori @ 2010-03-18 14:13 ` Alan Stern 0 siblings, 0 replies; 8+ messages in thread From: Alan Stern @ 2010-03-18 14:13 UTC (permalink / raw) To: FUJITA Tomonori; +Cc: James.Bottomley, linux-kernel On Thu, 18 Mar 2010, FUJITA Tomonori wrote: > On Wed, 17 Mar 2010 11:22:01 -0400 (EDT) > Alan Stern <stern@rowland.harvard.edu> wrote: > > > Here's the situation. The USB controller drivers don't all support > > scatter-gather operation. So there's a library routine in the USB core > > which calls dma_map_sg() and then creates a separate I/O request for > > each scatterlist element. The driver can process these requests one at > > a time, and when they are all finished the library routine calls > > dma_unmap_sg(). > > > > However... For tracing purposes (usbmon -- like tcpdump but for USB), > > we may need to copy the data from each I/O request's transfer buffer. > > Unfortunately, this copying is done as each request is submitted (for > > output) or as it completes (for input), at which times the buffers are > > all mapped for DMA. That's the problem. > > > > Would we be better off not using dma_map_sg() at all in this situation? > > We could map each scatterlist buffer individually with dma_map_single() > > as the request is submitted and then unmap the buffer when the request > > completes, just like with non-sg transfers; then the problem wouldn't > > arise. > > > > Would there be any significant penalty for doing this? I realize it > > would prevent adjacent buffers from getting coalesced, but that's > > probably okay. Any other reason not to? > > No reason. About merging adjacent buffers, there are few IOMMU > implementations that do. The recent IOMMU implementations such as VT-d > and AMD IOMMU don't. > > If drivers don't support scatter-gather operation, they had better use > dma_map_page() instead of forging scatter-gather lists and playing > with dma_map_sg(). Okay, I'll handle it that way. Thanks for the advice. Alan Stern ^ permalink raw reply [flat|nested] 8+ messages in thread
* Re: dma_sync_sg_for_cpu applied to a single scatterlist element 2010-03-15 21:30 dma_sync_sg_for_cpu applied to a single scatterlist element Alan Stern 2010-03-15 22:59 ` James Bottomley @ 2010-03-15 23:16 ` FUJITA Tomonori 1 sibling, 0 replies; 8+ messages in thread From: FUJITA Tomonori @ 2010-03-15 23:16 UTC (permalink / raw) To: stern; +Cc: James.Bottomley, linux-kernel On Mon, 15 Mar 2010 17:30:19 -0400 (EDT) Alan Stern <stern@rowland.harvard.edu> wrote: > This is addressed to James Bottomley as he is the author of > Documentation/DMA-API.txt, but anyone else who can contribute is > invited to do so. > > Suppose a scatter-gather transfer with multiple scatterlist elements > has been mapped via dma_map_sg(). Is it then valid to call > dma_sync_sg_for_cpu() with the "sg" argument pointing to one of the > mapped scatterlist elements (not necessarily the first one) and the > "nelems" argument set to 1? As James said, probably it works. As long as passed scatterlist elements points to mapped regions, it works. However, I think that the latest DMA-API.txt makes it clear that only dma_sync_single_for_{cpu|device} supports a partial sync. So it's not recommended, I guess. ^ permalink raw reply [flat|nested] 8+ messages in thread
end of thread, other threads:[~2010-03-18 14:14 UTC | newest] Thread overview: 8+ messages (download: mbox.gz / follow: Atom feed) -- links below jump to the message on this page -- 2010-03-15 21:30 dma_sync_sg_for_cpu applied to a single scatterlist element Alan Stern 2010-03-15 22:59 ` James Bottomley 2010-03-16 14:30 ` Alan Stern 2010-03-16 22:28 ` FUJITA Tomonori 2010-03-17 15:22 ` Alan Stern 2010-03-18 1:46 ` FUJITA Tomonori 2010-03-18 14:13 ` Alan Stern 2010-03-15 23:16 ` FUJITA Tomonori
This is a public inbox, see mirroring instructions for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®