From mboxrd@z Thu Jan 1 00:00:00 1970 Received: from us-smtp-delivery-124.mimecast.com (us-smtp-delivery-124.mimecast.com [170.10.133.124]) (using TLSv1.2 with cipher ECDHE-RSA-AES256-GCM-SHA384 (256/256 bits)) (No client certificate requested) by smtp.subspace.kernel.org (Postfix) with ESMTPS id 1502417332C for ; Wed, 4 Jun 2025 15:34:34 +0000 (UTC) Authentication-Results: smtp.subspace.kernel.org; arc=none smtp.client-ip=170.10.133.124 ARC-Seal:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1749051276; cv=none; b=OMirVx7WIprLlCf0NpELcZDF10/mVeBZQTFzmrGE+sa1XHU32CQM9vzL0dntp9SiPDkLhqxv8TaAs3EhOM/ofeCt6c1ajB2uQL1ro/ZV4aYlzxg1DyiSm7hLXG/JWf5SuOZ7ZhKMOXWQFu6tC61Ssi022f0C+RAQ1zQ6v91Zf8E= ARC-Message-Signature:i=1; a=rsa-sha256; d=subspace.kernel.org; s=arc-20240116; t=1749051276; c=relaxed/simple; bh=qTXEgU4faK8V3+o1webIL89G6tLPcijOJ2B4kkypv04=; h=From:In-Reply-To:References:To:Cc:Subject:MIME-Version: Content-Type:Date:Message-ID; b=SXD4mIQ1Pp7OnM3ZpwjMXDib7/ei+ms022T10sXtkmHj3JZb0MsMdu2lBWjYhnOHRyI0LPF68w1uPDHttFVupS1nY1t554G0wtKR1DfEeBwZo5/B+Sh4vrq+Sr0jjw0SBA/TZ1K3coF4qDSQ7qxwQNmi2ZJUZSt0z8y6qcOxoKA= ARC-Authentication-Results:i=1; smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com; spf=pass smtp.mailfrom=redhat.com; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b=aHSdrBfd; arc=none smtp.client-ip=170.10.133.124 Authentication-Results: smtp.subspace.kernel.org; dmarc=pass (p=quarantine dis=none) header.from=redhat.com Authentication-Results: smtp.subspace.kernel.org; spf=pass smtp.mailfrom=redhat.com Authentication-Results: smtp.subspace.kernel.org; dkim=pass (1024-bit key) header.d=redhat.com header.i=@redhat.com header.b="aHSdrBfd" DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=redhat.com; s=mimecast20190719; t=1749051274; h=from:from:reply-to:subject:subject:date:date:message-id:message-id: to:to:cc:cc:mime-version:mime-version:content-type:content-type: in-reply-to:in-reply-to:references:references; bh=o2KiML0jpnv35I2ryANTISLMOhOT0NTQ+BIj5SUwYgU=; b=aHSdrBfdpWo+VlBMNNWLM7vQrYqmLAhHjLfHW3BIKRcNXwgCJ+8JB/VLGOI/IVOhV0Xs+6 p9/Z3AUkfvexKcKsu8azXafIuEEjYMHYLyBThi1g7u61UqMOFgFmiudpSUxqKvtK+9wpGL wapQT8R9p1LTdooCJyidUqYer8+IYk0= Received: from mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (ec2-54-186-198-63.us-west-2.compute.amazonaws.com [54.186.198.63]) by relay.mimecast.com with ESMTP with STARTTLS (version=TLSv1.3, cipher=TLS_AES_256_GCM_SHA384) id us-mta-16-NZydtSNtNuymBvMkhyhlXA-1; Wed, 04 Jun 2025 11:34:31 -0400 X-MC-Unique: NZydtSNtNuymBvMkhyhlXA-1 X-Mimecast-MFC-AGG-ID: NZydtSNtNuymBvMkhyhlXA_1749051269 Received: from mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com [10.30.177.17]) (using TLSv1.3 with cipher TLS_AES_256_GCM_SHA384 (256/256 bits) key-exchange X25519 server-signature RSA-PSS (2048 bits) server-digest SHA256) (No client certificate requested) by mx-prod-mc-01.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTPS id 41659195608C; Wed, 4 Jun 2025 15:34:29 +0000 (UTC) Received: from warthog.procyon.org.uk (unknown [10.42.28.2]) by mx-prod-int-05.mail-002.prod.us-west-2.aws.redhat.com (Postfix) with ESMTP id DCD6919560AE; Wed, 4 Jun 2025 15:34:26 +0000 (UTC) Organization: Red Hat UK Ltd. Registered Address: Red Hat UK Ltd, Amberley Place, 107-111 Peascod Street, Windsor, Berkshire, SI4 1TE, United Kingdom. Registered in England and Wales under Company Registration No. 3798903 From: David Howells In-Reply-To: References: <770012.1748618092@warthog.procyon.org.uk> To: Mina Almasry Cc: dhowells@redhat.com, willy@infradead.org, hch@infradead.org, Jakub Kicinski , Eric Dumazet , netdev@vger.kernel.org, linux-mm@kvack.org, linux-kernel@vger.kernel.org Subject: Re: Device mem changes vs pinning/zerocopy changes Precedence: bulk X-Mailing-List: linux-kernel@vger.kernel.org List-Id: List-Subscribe: List-Unsubscribe: MIME-Version: 1.0 Content-Type: text/plain; charset="us-ascii" Content-ID: <1098852.1749051265.1@warthog.procyon.org.uk> Date: Wed, 04 Jun 2025 16:34:25 +0100 Message-ID: <1098853.1749051265@warthog.procyon.org.uk> X-Scanned-By: MIMEDefang 3.0 on 10.30.177.17 Mina Almasry wrote: > Hi David! Yes, very happy to collaborate. :-) > FWIW, my initial gut feeling is that the work doesn't conflict that much. > The tcp devmem netmem/net_iov stuff is designed to follow the page stuff, > and as the usage of struct page changes we're happy moving net_iovs and > netmems to do the same thing. My read is that it will take a small amount of > extra work, but there are no in-principle design conflicts, at least AFAICT > so far. The problem is more the code you changed in the current merge window I'm also wanting to change, so merge conflicts will arise. However, I'm also looking to move the points at which refs are taken/dropped which will directly inpinge on the design of the code that's currently upstream. Would it help if I created some diagrams to show what I'm thinking of? > I believe the main challenge here is that there are many code paths in > the net stack that expect to be able to grab a ref on skb frags. See > all the callers of skb_frag_ref/skb_frag_unref: > > tcp_grow_skb, __skb_zcopy_downgrade_managed, __pskb_copy_fclone, > pskb_expand_head, skb_zerocopy, skb_split, pksb_carve_inside_header, > pskb_care_inside_nonlinear, tcp_clone_payload, skb_segment. Oh, yes, I've come to appreciate that well. A chunk of that can actually go away, I think. > I think to accomplish what you're describing we need to modify > skb_frag_ref to do something else other than taking a reference on the > page or net_iov. I think maybe taking a reference on the skb itself > may be acceptable, and the skb can 'guarantee' that the individual > frags underneath it don't disappear while these functions are > executing. Maybe. There is an issue with that, though it may not be insurmountable: If a userspace process does, say, a MSG_ZEROCOPY send of a page worth of data over TCP, under a typicalish MTU, say, 1500, this will be split across at least three skbuffs. This would involve making a call into GUP to get a pin - but we'd need a separate pin for each skbuff and we might (in fact we currently do) end up calling into GUP thrice to do the address translation and page pinning. What I want to do is to put this outside of the skbuff so that GUP pin can be shared - but if, instead, we attach a pin to each skbuff, we need to get that extra pin in some way. Now, it may be reasonable to add a "get me an extra pin for such-and-such a range" thing and store the {physaddr,len} in the skbuff fragment, but we also have to be careful not to overrun the pin count - if there's even a pin count per se. > But devmem TCP doesn't get much in the way here, AFAICT. It's really > the fact that so many of the networking code paths want to obtain page > refs via skb_frag_ref so there is potentially a lot of code to touch. Yep. > But, AFAICT, skb_frag_t needs a struct page inside of it, not just a > physical address. skb_frags can mmap'd into userspace for TCP > zerocopy, see tcp_zerocopy_vm_insert_batch (which is a very old > feature, it's not a recent change). There may be other call paths in > the net stack that require a full page and just a physical address > will do. (unless somehow we can mmap a physical address to the > userspace). Yeah - I think this needs very careful consideration and will need some adjustment. Some of the pages that may, in the future, get zerocopied or spliced into the socket *really* shouldn't be spliced out into some random process's address space - and, in fact, may not even have a refcount (say they come from the slab). Basically, TCP has implemented async vmsplice()... > Is struct net_txbuf intended to replace struct sk_buff in the tx path > only? If so, I'm not sure that works. No - the idea is that it runs a parallel track to it and holds "references" to the buffer memory. This is then divided up amongst a number of sk_buffs that hold refs on the first txbuf that it uses memory from. txbufs would allow us to take and hold a single ref or pin (or even nothing, just a destructor) on each piece of supplied buffer memory and for that to be shared between a sequence of skbufs. > Currently TX and RX memory share a single data structure (sk_buff), and I > believe that is critical. Yep. And we need to keep sk_buff because an sk_buff is more than just a memory pinning device - it also retains the metadata for a packet. > ... So I think, maybe, instead of introducing a new struct, you have to make > the modifications you envision to struct sk_buff itself? It may be possible. But see above. I want to be able to share pins between sk_buffs. > OK, you realize that TX packets can be forwarded to RX. The opposite > is true, RX can be forwarded to TX. And it's not just AF_UNIX and > loopback. Packets can be forwarded via ip forwarding, and tc, and > probably another half dozen features in the net stack I don't know > about. Yeah, I'd noticed that. > I think you need to modify the existing sk_buff. I think adding > a new struct and migrating the entire net stack to use that is a bit > too ambitious. But up to you. Just my 2 cents here.