* [PATCH net v4 0/3] net: don't strip zerocopy frag markers from a forwarded skb
@ 2026-08-22 9:10 Norbert Szetei
2026-08-22 9:12 ` [PATCH net v4 1/3] openvswitch: only skb_tx_error() a packet we are about to drop Norbert Szetei
` (2 more replies)
0 siblings, 3 replies; 6+ messages in thread
From: Norbert Szetei @ 2026-08-22 9:10 UTC (permalink / raw)
To: netdev
Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Aaron Conole, Eelco Chaudron, Ilya Maximets,
Steffen Klassert, Kuan-Ting Chen, Michael S. Tsirkin,
Willem de Bruijn, linux-kernel, dev, Jongmin Jang
queue_userspace_packet() calls skb_tx_error() on the packet skb in its
error path, but it only borrows that skb: on the OVS_ACTION_ATTR_USERSPACE
action path do_execute_actions() ignores output_userspace()'s return value
and keeps forwarding the same skb through the flow's remaining actions.
skb_tx_error() completes the zerocopy uarg and clears SKBFL_ALL_ZEROCOPY,
and with it SKBFL_SHARED_FRAG.
For a MSG_ZEROCOPY skb carrying page-cache frags, SKBFL_SHARED_FRAG is
what makes esp_input() skb_cow_data() instead of taking the in-place AEAD
path. Once it is stripped, a later local ESP delivery decrypts in place
over pages the sender still shares with the page cache.
Patch 1 moves the skb_tx_error() into the one path that does drop the
packet, the "default" arm of ovs_dp_process_packet()'s switch(error).
Patch 2 removes a second such strip, in skb_zerocopy(), which calls
skb_tx_error() on its source when skb_orphan_frags() fails. A copy helper
should not perform a destructive action on its source, and both callers
already report the error on their own drop path. MSG_ZEROCOPY skbs cannot
reach that one -- SKBFL_DONT_ORPHAN makes skb_orphan_frags() return early
-- but producers that do not set that flag, such as vhost-net, can.
Patch 3 is new in v2. It stops skb_tx_error() from touching skb_shinfo()
state that is shared with clones, so patch 1's new call site cannot reach
a live skb either. For a non-last OVS_ACTION_ATTR_RECIRC action
clone_execute() sends a skb_clone() into ovs_dp_process_packet() while
do_execute_actions() keeps forwarding the original, and skb_clone() does
not privatise the frags for these skbs -- skb_orphan_frags() returns early
on SKBFL_DONT_ORPHAN -- so a flow miss on the clone strips
SKBFL_SHARED_FRAG from the packet still in flight. Confirmed on a KASAN
build with a flow matching recirc_id 0 and actions RECIRC(1),OUTPUT(0):
with patches 1 and 2 applied it still reproduces the page-cache write,
with patch 3 on top it no longer does (5/5 runs). A kprobe on
skb_tx_error() shows the datapath drop path is still reached in both
cases, so the difference is the guard and not the reproducer.
As Ilya noted, that makes patch 3 the general fix -- an skb can enter any
skb_tx_error() caller already cloned elsewhere in the stack -- while
patches 1 and 2 keep the callers from acting on an skb they do not own.
Removing skb_tx_error() altogether looks like the right long-term cleanup
and is planned as a net-next follow-up.
v4:
- rebased on net after commit 68d8c6532659 ("net: core: propagate
unreadable flag in skb_zerocopy"); patch 2 now only drops the
skb_tx_error() call from the error block, the -EFAULT/put_page()
handling added there stays (Ilya Maximets, Jakub Kicinski)
- Reviewed-by from Ilya Maximets picked up on patch 3
- v3: https://lore.kernel.org/netdev/F3B9E5BA-0AC1-4AD1-A7D9-F38033304270@doyensec.com/
v3:
- patch 3: Fixes tag corrected to 25121173f7b1 ("skb: api to report
errors for zero copy skbs"), the commit that added skb_tx_error()
(Ilya Maximets)
- Tested-by from Jongmin Jang picked up on patches 1 and 3
- v2: https://lore.kernel.org/netdev/AD1B7BEE-C04C-4A1B-982C-8385F1908911@doyensec.com/
v2:
- new patch 3: skip the shared skb_shinfo() work in skb_tx_error() when
the skb is cloned, which also covers the OVS_ACTION_ATTR_RECIRC path
that patch 1 alone leaves open (suggested by Ilya Maximets)
- patches 1 and 2 unchanged, Reviewed-by from Ilya Maximets picked up
- v1: https://lore.kernel.org/netdev/8063260C-05C9-4997-B9B6-2135063C4858@doyensec.com/
Norbert Szetei (3):
openvswitch: only skb_tx_error() a packet we are about to drop
net: skbuff: don't skb_tx_error() the source skb in skb_zerocopy()
net: skbuff: don't touch shared zerocopy state in skb_tx_error()
net/core/skbuff.c | 6 ++++--
net/openvswitch/datapath.c | 3 +--
2 files changed, 5 insertions(+), 4 deletions(-)
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH net v4 1/3] openvswitch: only skb_tx_error() a packet we are about to drop
2026-08-22 9:10 [PATCH net v4 0/3] net: don't strip zerocopy frag markers from a forwarded skb Norbert Szetei
@ 2026-08-22 9:12 ` Norbert Szetei
2026-08-22 9:13 ` [PATCH net v4 2/3] net: skbuff: don't skb_tx_error() the source skb in skb_zerocopy() Norbert Szetei
2026-08-22 9:15 ` [PATCH net v4 3/3] net: skbuff: don't touch shared zerocopy state in skb_tx_error() Norbert Szetei
2 siblings, 0 replies; 6+ messages in thread
From: Norbert Szetei @ 2026-08-22 9:12 UTC (permalink / raw)
To: netdev
Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Aaron Conole, Eelco Chaudron, Ilya Maximets,
Steffen Klassert, Kuan-Ting Chen, Michael S. Tsirkin,
Willem de Bruijn, linux-kernel, dev, Jongmin Jang
queue_userspace_packet() borrows the packet skb -- it only copies it into
a private netlink message (user_skb) and does not own it; on return
do_execute_actions() keeps forwarding it through the flow's remaining
actions. Its error path nevertheless calls skb_tx_error(skb), which via
skb_zcopy_clear() does skb_shinfo(skb)->flags &= ~SKBFL_ALL_ZEROCOPY,
stripping SKBFL_SHARED_FRAG from that live skb (skb_tx_error()'s kerneldoc
says "skb must be freed afterwards").
For a MSG_ZEROCOPY skb carrying page-cache frags, SKBFL_SHARED_FRAG is
what makes esp_input() skb_cow_data() before in-place AEAD; once it is
stripped a later local ESP-in-UDP delivery decrypts in place over pages
the sender does not own -- an unprivileged page-cache write (the
"Fragnesia" primitive).
do_execute_actions() ignores output_userspace()'s return value, so any
action after a failed USERSPACE upcall inherits the stripped skb.
Move the skb_tx_error() to the flow-miss drop path - the "default"
branch of ovs_dp_process_packet()'s switch(error), before kfree_skb().
The call has been here since commit 36d5fe6a0007 ("core, nfqueue,
openvswitch: Orphan frags in skb_zerocopy and handle errors") but was
harmless until esp_input() began relying on SKBFL_SHARED_FRAG to gate
in-place decrypt; only then did stripping it on a still-forwarded skb
become a page-cache write primitive.
Fixes: 36d5fe6a0007 ("core, nfqueue, openvswitch: Orphan frags in skb_zerocopy and handle errors")
Fixes: f4c50a4034e6 ("xfrm: esp: avoid in-place decrypt on shared skb frags")
Cc: stable@vger.kernel.org
Assisted-by: Claude:claude-opus-5
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Tested-by: Jongmin Jang <payload.jang@gmail.com>
---
net/openvswitch/datapath.c | 3 +--
1 file changed, 1 insertion(+), 2 deletions(-)
diff --git a/net/openvswitch/datapath.c b/net/openvswitch/datapath.c
index f2d5b5ab38de..2fc9ef6c321f 100644
--- a/net/openvswitch/datapath.c
+++ b/net/openvswitch/datapath.c
@@ -285,6 +285,7 @@ void ovs_dp_process_packet(struct sk_buff *skb, struct sw_flow_key *key)
consume_skb(skb);
break;
default:
+ skb_tx_error(skb);
kfree_skb(skb);
break;
}
@@ -604,8 +605,6 @@ static int queue_userspace_packet(struct datapath *dp, struct sk_buff *skb,
err = genlmsg_unicast(ovs_dp_get_net(dp), user_skb, upcall_info->portid);
user_skb = NULL;
out:
- if (err)
- skb_tx_error(skb);
consume_skb(user_skb);
consume_skb(nskb);
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH net v4 2/3] net: skbuff: don't skb_tx_error() the source skb in skb_zerocopy()
2026-08-22 9:10 [PATCH net v4 0/3] net: don't strip zerocopy frag markers from a forwarded skb Norbert Szetei
2026-08-22 9:12 ` [PATCH net v4 1/3] openvswitch: only skb_tx_error() a packet we are about to drop Norbert Szetei
@ 2026-08-22 9:13 ` Norbert Szetei
2026-08-22 20:58 ` Willem de Bruijn
2026-08-22 9:15 ` [PATCH net v4 3/3] net: skbuff: don't touch shared zerocopy state in skb_tx_error() Norbert Szetei
2 siblings, 1 reply; 6+ messages in thread
From: Norbert Szetei @ 2026-08-22 9:13 UTC (permalink / raw)
To: netdev
Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Aaron Conole, Eelco Chaudron, Ilya Maximets,
Steffen Klassert, Kuan-Ting Chen, Michael S. Tsirkin,
Willem de Bruijn, linux-kernel, dev, Jongmin Jang
skb_zerocopy() copies frags from @from into @to. On an
skb_orphan_frags() failure it calls skb_tx_error(@from), a destructive
operation on the source skb the copy helper does not own. That completes
@from's zerocopy uarg and clears SKBFL_ALL_ZEROCOPY, including the
SKBFL_SHARED_FRAG page-ownership marker.
Both callers already report the failure on their own drop path.
nfnetlink_queue does it at nla_put_failure, and Open vSwitch does it in
the flow-miss drop arm of ovs_dp_process_packet(), so nothing is lost by
dropping it here.
On Open vSwitch's OVS_ACTION_ATTR_USERSPACE path the skb is not freed on
this error: do_execute_actions() ignores output_userspace()'s return
value and, unless the upcall was the last action, keeps forwarding the
same skb through the flow's remaining actions. The uarg is completed
while that skb is still in flight, telling the producer its buffers are
free, and SKBFL_SHARED_FRAG is cleared on an skb the rest of the stack
still handles. That flag is what makes esp_input() call skb_cow_data()
instead of decrypting in place, so a later local ESP delivery can
decrypt over frags the skb does not own privately.
Leave error reporting to the callers.
Fixes: 36d5fe6a0007 ("core, nfqueue, openvswitch: Orphan frags in skb_zerocopy and handle errors")
Cc: stable@vger.kernel.org
Suggested-by: Ilya Maximets <i.maximets@ovn.org>
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
---
net/core/skbuff.c | 1 -
1 file changed, 1 deletion(-)
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index d4382b68d56e..ab3d161247b9 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -3914,7 +3914,6 @@ skb_zerocopy(struct sk_buff *to, struct sk_buff *from, int len, int hlen)
skb_len_add(to, len + plen);
if (unlikely(skb_orphan_frags(from, GFP_ATOMIC))) {
- skb_tx_error(from);
if (j > 0)
put_page(virt_to_head_page(from->head));
return -ENOMEM;
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* [PATCH net v4 3/3] net: skbuff: don't touch shared zerocopy state in skb_tx_error()
2026-08-22 9:10 [PATCH net v4 0/3] net: don't strip zerocopy frag markers from a forwarded skb Norbert Szetei
2026-08-22 9:12 ` [PATCH net v4 1/3] openvswitch: only skb_tx_error() a packet we are about to drop Norbert Szetei
2026-08-22 9:13 ` [PATCH net v4 2/3] net: skbuff: don't skb_tx_error() the source skb in skb_zerocopy() Norbert Szetei
@ 2026-08-22 9:15 ` Norbert Szetei
2026-08-23 18:26 ` Willem de Bruijn
2 siblings, 1 reply; 6+ messages in thread
From: Norbert Szetei @ 2026-08-22 9:15 UTC (permalink / raw)
To: netdev
Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Aaron Conole, Eelco Chaudron, Ilya Maximets,
Steffen Klassert, Kuan-Ting Chen, Michael S. Tsirkin,
Willem de Bruijn, linux-kernel, dev, Jongmin Jang
skb_tx_error() completes the zerocopy uarg and clears
SKBFL_ALL_ZEROCOPY, and skb_zcopy_downgrade_managed() clears
SKBFL_MANAGED_FRAG_REFS. Both live in skb_shinfo(), which every clone
shares, while the caller only owns the reference it is about to drop.
Through a clone it tells the producer its pages are free and drops
SKBFL_SHARED_FRAG for an skb that is still in flight.
Open vSwitch reaches this with a non-last OVS_ACTION_ATTR_RECIRC:
clone_execute() sends a skb_clone() into ovs_dp_process_packet() while
do_execute_actions() keeps forwarding the original, and skb_clone()
does not privatise the frags here -- skb_orphan_frags() returns early
on SKBFL_DONT_ORPHAN. A flow miss on the clone then strips the marker
from the packet still being forwarded, and a later local ESP delivery
decrypts in place over frags it does not own privately.
Skip it for a cloned skb. Nothing is lost: skb_release_data() clears
the zerocopy state once the last reference to the shared data goes.
Fixes: 25121173f7b1 ("skb: api to report errors for zero copy skbs")
Cc: stable@vger.kernel.org
Suggested-by: Ilya Maximets <i.maximets@ovn.org>
Signed-off-by: Norbert Szetei <norbert@doyensec.com>
Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Tested-by: Jongmin Jang <payload.jang@gmail.com>
---
net/core/skbuff.c | 5 ++++-
1 file changed, 4 insertions(+), 1 deletion(-)
diff --git a/net/core/skbuff.c b/net/core/skbuff.c
index ab3d161247b9..b9541329f1a7 100644
--- a/net/core/skbuff.c
+++ b/net/core/skbuff.c
@@ -1417,10 +1417,13 @@ EXPORT_SYMBOL(skb_dump);
*
* Report xmit error if a device callback is tracking this skb.
* skb must be freed afterwards.
+ *
+ * Does nothing for a cloned skb: the zerocopy state lives in
+ * skb_shinfo(), which the clones share.
*/
void skb_tx_error(struct sk_buff *skb)
{
- if (skb) {
+ if (skb && !skb_cloned(skb)) {
skb_zcopy_downgrade_managed(skb);
skb_zcopy_clear(skb, true);
}
--
2.55.0
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net v4 2/3] net: skbuff: don't skb_tx_error() the source skb in skb_zerocopy()
2026-08-22 9:13 ` [PATCH net v4 2/3] net: skbuff: don't skb_tx_error() the source skb in skb_zerocopy() Norbert Szetei
@ 2026-08-22 20:58 ` Willem de Bruijn
0 siblings, 0 replies; 6+ messages in thread
From: Willem de Bruijn @ 2026-08-22 20:58 UTC (permalink / raw)
To: Norbert Szetei, netdev
Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Aaron Conole, Eelco Chaudron, Ilya Maximets,
Steffen Klassert, Kuan-Ting Chen, Michael S. Tsirkin,
Willem de Bruijn, linux-kernel, dev, Jongmin Jang
Norbert Szetei wrote:
> skb_zerocopy() copies frags from @from into @to. On an
> skb_orphan_frags() failure it calls skb_tx_error(@from), a destructive
> operation on the source skb the copy helper does not own. That completes
> @from's zerocopy uarg and clears SKBFL_ALL_ZEROCOPY, including the
> SKBFL_SHARED_FRAG page-ownership marker.
>
> Both callers already report the failure on their own drop path.
> nfnetlink_queue does it at nla_put_failure, and Open vSwitch does it in
> the flow-miss drop arm of ovs_dp_process_packet(), so nothing is lost by
> dropping it here.
>
> On Open vSwitch's OVS_ACTION_ATTR_USERSPACE path the skb is not freed on
> this error: do_execute_actions() ignores output_userspace()'s return
> value and, unless the upcall was the last action, keeps forwarding the
> same skb through the flow's remaining actions. The uarg is completed
> while that skb is still in flight, telling the producer its buffers are
> free, and SKBFL_SHARED_FRAG is cleared on an skb the rest of the stack
> still handles. That flag is what makes esp_input() call skb_cow_data()
> instead of decrypting in place, so a later local ESP delivery can
> decrypt over frags the skb does not own privately.
>
> Leave error reporting to the callers.
>
> Fixes: 36d5fe6a0007 ("core, nfqueue, openvswitch: Orphan frags in skb_zerocopy and handle errors")
> Cc: stable@vger.kernel.org
> Suggested-by: Ilya Maximets <i.maximets@ovn.org>
> Signed-off-by: Norbert Szetei <norbert@doyensec.com>
> Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
Reviewed-by: Willem de Bruijn <willemb@google.com>
^ permalink raw reply [flat|nested] 6+ messages in thread
* Re: [PATCH net v4 3/3] net: skbuff: don't touch shared zerocopy state in skb_tx_error()
2026-08-22 9:15 ` [PATCH net v4 3/3] net: skbuff: don't touch shared zerocopy state in skb_tx_error() Norbert Szetei
@ 2026-08-23 18:26 ` Willem de Bruijn
0 siblings, 0 replies; 6+ messages in thread
From: Willem de Bruijn @ 2026-08-23 18:26 UTC (permalink / raw)
To: Norbert Szetei, netdev
Cc: David S. Miller, Eric Dumazet, Jakub Kicinski, Paolo Abeni,
Simon Horman, Aaron Conole, Eelco Chaudron, Ilya Maximets,
Steffen Klassert, Kuan-Ting Chen, Michael S. Tsirkin,
Willem de Bruijn, linux-kernel, dev, Jongmin Jang
Norbert Szetei wrote:
> skb_tx_error() completes the zerocopy uarg and clears
> SKBFL_ALL_ZEROCOPY, and skb_zcopy_downgrade_managed() clears
> SKBFL_MANAGED_FRAG_REFS. Both live in skb_shinfo(), which every clone
> shares, while the caller only owns the reference it is about to drop.
> Through a clone it tells the producer its pages are free and drops
> SKBFL_SHARED_FRAG for an skb that is still in flight.
>
> Open vSwitch reaches this with a non-last OVS_ACTION_ATTR_RECIRC:
> clone_execute() sends a skb_clone() into ovs_dp_process_packet() while
> do_execute_actions() keeps forwarding the original, and skb_clone()
> does not privatise the frags here -- skb_orphan_frags() returns early
> on SKBFL_DONT_ORPHAN. A flow miss on the clone then strips the marker
> from the packet still being forwarded, and a later local ESP delivery
> decrypts in place over frags it does not own privately.
>
> Skip it for a cloned skb. Nothing is lost: skb_release_data() clears
> the zerocopy state once the last reference to the shared data goes.
>
> Fixes: 25121173f7b1 ("skb: api to report errors for zero copy skbs")
> Cc: stable@vger.kernel.org
> Suggested-by: Ilya Maximets <i.maximets@ovn.org>
> Signed-off-by: Norbert Szetei <norbert@doyensec.com>
> Reviewed-by: Ilya Maximets <i.maximets@ovn.org>
> Tested-by: Jongmin Jang <payload.jang@gmail.com>
Reviewed-by: Willem de Bruijn <willemb@google.com>
Took me some time to wrap my head around this one, because
There are two independent types of zerocopy in this context:
1. skb_zerocopy(), used by nfqueue and ovs to create a derived skb
2. skb_zcopy(), skbs with "zerocopy" page frags
And there second has two variants:
2A. original, such as vhost-net, that do not support refcounting and
thus must be downgraded on skb_clone() and such
2B. SKBFL_DONT_ORPHAN, that support clones through refcounting
The bug here is modifying shared shinfo fields of cloned skbs, so
affects type 2B skbs only.
skb_tx_error was introduced for type 2A skbs, predates refcounting.
For type 2B, the signal is indeed generated at skb_release_data.
So LGTM.
> ---
> net/core/skbuff.c | 5 ++++-
> 1 file changed, 4 insertions(+), 1 deletion(-)
>
> diff --git a/net/core/skbuff.c b/net/core/skbuff.c
> index ab3d161247b9..b9541329f1a7 100644
> --- a/net/core/skbuff.c
> +++ b/net/core/skbuff.c
> @@ -1417,10 +1417,13 @@ EXPORT_SYMBOL(skb_dump);
> *
> * Report xmit error if a device callback is tracking this skb.
> * skb must be freed afterwards.
> + *
> + * Does nothing for a cloned skb: the zerocopy state lives in
> + * skb_shinfo(), which the clones share.
> */
> void skb_tx_error(struct sk_buff *skb)
> {
> - if (skb) {
> + if (skb && !skb_cloned(skb)) {
> skb_zcopy_downgrade_managed(skb);
> skb_zcopy_clear(skb, true);
> }
> --
> 2.55.0
>
^ permalink raw reply [flat|nested] 6+ messages in thread
end of thread, other threads:[~2026-08-23 18:26 UTC | newest]
Thread overview: 6+ messages (download: mbox.gz / follow: Atom feed)
-- links below jump to the message on this page --
2026-08-22 9:10 [PATCH net v4 0/3] net: don't strip zerocopy frag markers from a forwarded skb Norbert Szetei
2026-08-22 9:12 ` [PATCH net v4 1/3] openvswitch: only skb_tx_error() a packet we are about to drop Norbert Szetei
2026-08-22 9:13 ` [PATCH net v4 2/3] net: skbuff: don't skb_tx_error() the source skb in skb_zerocopy() Norbert Szetei
2026-08-22 20:58 ` Willem de Bruijn
2026-08-22 9:15 ` [PATCH net v4 3/3] net: skbuff: don't touch shared zerocopy state in skb_tx_error() Norbert Szetei
2026-08-23 18:26 ` Willem de Bruijn
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox
all inboxes | Powered by JetHome®