mirror of https://lore.kernel.org/lkml/
 help / color / mirror / Atom feed
From: Xin Xie <xiexinet@gmail.com>
To: netdev@vger.kernel.org, linux-kselftest@vger.kernel.org,
	linux-kernel@vger.kernel.org
Cc: davem@davemloft.net, edumazet@google.com, kuba@kernel.org,
	pabeni@redhat.com, horms@kernel.org, andrew+netdev@lunn.ch,
	shuah@kernel.org, kees@kernel.org, petr.wozniak@gmail.com,
	qingfang.deng@linux.dev, fmaurer@redhat.com,
	luka.gejak@linux.dev, bigeasy@linutronix.de,
	xiaoliang.yang_1@nxp.com, skhawaja@google.com,
	liuhangbin@gmail.com, stable@vger.kernel.org,
	sdf.kernel@gmail.com, xiexinet@gmail.com
Subject: [PATCH net v7 0/4] net: hsr: fix super-packet forwarding and ordering
Date: Fri,  9 Oct 2026 22:13:20 +0200	[thread overview]
Message-ID: <20261009201324.17-1-xiexinet@gmail.com> (raw)

HSR/PRP requires per-wire-frame tags/RCTs and sequence numbers, and
duplicate discard is per frame. RX GRO and TX GSO can present multiple
frames as one skb and violate that assumption: a super-skb is either
rejected by a constrained lower device, or forwarded without valid
per-frame trailers and sequence numbers. An oversized PRP aggregate
can also truncate the RCT's 12-bit LSDU size field.

Patch 1 keeps software GRO off while devices are direct HSR/PRP
members, without rewriting wanted features. Fixed-on or driver-required
GRO_HW may remain enabled.

Patch 2 replaces the forwarding lock with one consumer to avoid
stacked-HSR deadlocks. It preserves per-lower submission order for
locally numbered frames, so old HSR peers that drop late sequence
numbers keep working. Only sequence allocation holds the short counter
lock; forwarding runs after it is released.

Patch 3 segments valid GSO before per-frame processing, so each wire
frame gets its own tag/RCT and sequence number. Patch 4 tests member
GRO policy, GSO forwarding and local submission order.

A typical GSO source is a container connected to a PRP RedBox
interlink through veth. An existing TCP/UDP GSO skb can reach the
interlink with GRO disabled:

  container/netns                  host PRP RedBox
  prp0-peer ---- veth ---- prp0-int (interlink RX)
                                  |
                           split GSO into frames
                                  |
                           sequence number + RCT
                              /          \
                           LAN A        LAN B

Compatibility with older PRP senders:

Valid PRP frames from older senders remain supported. Patch 2
preserves submission order for locally numbered frames without
changing the PRP frame format or requiring peers to upgrade.

Some old Linux PRP senders incorrectly append one RCT to an entire
GSO skb. This violates the per-frame PRP requirement. Patch 3 drops
such intact aggregates when trailing bytes extend past a known,
nonzero IP length; zero lengths and GSO_PARTIAL are not covered by
that check.

The mixed-version test uses two PRP VMs; the host runs no PRP.
On virtio-net/TAP paths, disabling GRO_HW in patch 1 can make the
host segment an old sender's malformed aggregate before it reaches
the receiving VM. The RCT then becomes transport payload, which the
receiver cannot reliably identify or remove. In our QEMU 8.2.2 UDP
test this produced an extra six-byte datagram; the base receiver
delivered the original data correctly. Upgrade the sender or disable
GSO/TSO on its HSR/PRP master; the latter avoided the failure in the
test.

Validation:

On the net test kernel (7.3.0-rc5-gb0fe53dd6370), the installed
ordered selftest passed 5/5 cases, including supervision, concurrent
producers and sequence wrap, with complete, loss-free captures.

GRO passed 11/11 with the preceding helper. The later cleanup-only
fix does not affect GRO cases. Failure-injection checks confirmed
error reporting and cleanup.

Earlier queue-stress and PREEMPT_RT results are retained. Paired W=1
builds had no additional diagnostics. These checks were not repeated
for the helper update.

Based on net 6dc989ea46b9 (2026-10-05). Integration preserves the
upstream interlink promiscuous-mode exception and setup-failure
cleanup. The series applies directly to this base.

Patch 3 depends on patch 2; no stable backport is requested here.
The remaining HSR dev->stats races are left to a separate series,
as Paolo suggested [1]. Pre-existing shared-skb mutations are also
handled separately.

[1] https://lore.kernel.org/netdev/4fc3b9f1-4bef-4b34-ae7a-e89037cce829@redhat.com/

Changes since v6, including the NIPA and Gemini review responses:

- Replace recursive, one-time GRO disabling with persistent member
  policy and restoration of the latest wanted state.
- Preserve submission order with one consumer; remove the bitmap
  prerequisite and use core per-CPU counters for new RX/TX drops.
- Handle bounded 802.1Q/802.1AD VLAN stacks and reject trailing data
  beyond a known nonzero IP length before segmentation.
- Replace child-process signalling and fixed readiness sleeps with
  namespace-bound sockets registered before traffic. Bound cleanup
  under signals and preserve failures.
- Check packet contents, retransmissions and capture completeness,
  rather than throughput or average sizes. Use a standalone Python
  helper, with no embedded Python in the shell entries.

Veth tests do not validate hardware GRO_HW. The generic GSO bit need
not be off, so tests now check GSO_MASK member types. The claim that
the old test always failed was not supported by its actual feature
state and recorded passing runs. Shared-skb mutation is a separate
existing issue; both reports about new drop accounting are addressed
by the per-CPU counters.

Previous postings (earlier design, newest first):
v6: https://lore.kernel.org/netdev/20260809121455.1745-1-xiexinet@gmail.com/
v5: https://lore.kernel.org/netdev/20260807140751.1351-1-xiexinet@gmail.com/
v4: https://lore.kernel.org/netdev/20260803222211.877-1-xiexinet@gmail.com/
v3: https://lore.kernel.org/netdev/20260731090224.18-1-xiexinet@gmail.com/
v2: https://lore.kernel.org/netdev/20260724161253.79-1-xiexinet@gmail.com/
v1: https://lore.kernel.org/netdev/20260722171836.196-1-xiexinet@gmail.com/

Xin Xie (4):
  net: hsr: keep GRO disabled on HSR/PRP ports
  net: hsr: preserve submission order without a forwarding lock
  net: hsr: segment GSO before per-frame forwarding
  selftests: net: hsr: verify GRO policy and ordered forwarding

 .../networking/net_cachelines/net_device.rst      |    1 +
 include/linux/netdevice.h                         |    6 +
 net/core/dev.c                                    |   10 +
 net/hsr/Makefile                                  |    3 +-
 net/hsr/hsr_device.c                              |   66 +-
 net/hsr/hsr_forward.c                             |  233 +++-
 net/hsr/hsr_forward.h                             |   37 +
 net/hsr/hsr_forward_queue.c                       |  294 +++++
 net/hsr/hsr_main.h                                |   18 +
 net/hsr/hsr_netlink.c                             |    4 +-
 net/hsr/hsr_slave.c                               |   60 +-
 tools/testing/selftests/net/hsr/Makefile          |    3 +
 tools/testing/selftests/net/hsr/config            |    1 +
 .../selftests/net/hsr/hsr_gro_superpacket.py      | 1088 +++++++++++++++++++
 .../selftests/net/hsr/hsr_gro_superpacket.sh      |    9 +
 .../selftests/net/hsr/hsr_ordered_forwarding.sh   |    9 +
 16 files changed, 1784 insertions(+), 58 deletions(-)
 create mode 100644 net/hsr/hsr_forward_queue.c
 create mode 100755 tools/testing/selftests/net/hsr/hsr_gro_superpacket.py
 create mode 100755 tools/testing/selftests/net/hsr/hsr_gro_superpacket.sh
 create mode 100755 tools/testing/selftests/net/hsr/hsr_ordered_forwarding.sh


base-commit: 6dc989ea46b96ce170840174b4a38c4a387fb005
-- 
2.43.0

             reply	other threads:[~2026-10-09 20:13 UTC|newest]

Thread overview: 5+ messages / expand[flat|nested]  mbox.gz  Atom feed  top
2026-10-09 20:13 Xin Xie [this message]
2026-10-09 20:13 ` [PATCH net v7 1/4] net: hsr: keep GRO disabled on HSR/PRP ports Xin Xie
2026-10-09 20:13 ` [PATCH net v7 2/4] net: hsr: preserve submission order without a forwarding lock Xin Xie
2026-10-09 20:13 ` [PATCH net v7 3/4] net: hsr: segment GSO before per-frame forwarding Xin Xie
2026-10-09 20:13 ` [PATCH net v7 4/4] selftests: net: hsr: verify GRO policy and ordered forwarding Xin Xie

Reply instructions:

You may reply publicly to this message via plain-text email
using any one of the following methods:

* Save the following mbox file, import it into your mail client,
  and reply-to-all from there: mbox

  Avoid top-posting and favor interleaved quoting:
  https://en.wikipedia.org/wiki/Posting_style#Interleaved_style

* Reply using the --to, --cc, and --in-reply-to
  switches of git-send-email(1):

  git send-email \
    --in-reply-to=20261009201324.17-1-xiexinet@gmail.com \
    --to=xiexinet@gmail.com \
    --cc=andrew+netdev@lunn.ch \
    --cc=bigeasy@linutronix.de \
    --cc=davem@davemloft.net \
    --cc=edumazet@google.com \
    --cc=fmaurer@redhat.com \
    --cc=horms@kernel.org \
    --cc=kees@kernel.org \
    --cc=kuba@kernel.org \
    --cc=linux-kernel@vger.kernel.org \
    --cc=linux-kselftest@vger.kernel.org \
    --cc=liuhangbin@gmail.com \
    --cc=luka.gejak@linux.dev \
    --cc=netdev@vger.kernel.org \
    --cc=pabeni@redhat.com \
    --cc=petr.wozniak@gmail.com \
    --cc=qingfang.deng@linux.dev \
    --cc=sdf.kernel@gmail.com \
    --cc=shuah@kernel.org \
    --cc=skhawaja@google.com \
    --cc=stable@vger.kernel.org \
    --cc=xiaoliang.yang_1@nxp.com \
    /path/to/YOUR_REPLY

  https://kernel.org/pub/software/scm/git/docs/git-send-email.html

* If your mail client supports setting the In-Reply-To header
  via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line before the message body.
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox

all inboxes | Powered by JetHome®