A GSO skb which exceeds an egress device limit loses its GSO feature mask
and is segmented into individual packets.  This is unnecessarily expensive
when the device can still offload smaller TCP GSO skbs, which is easy to
hit once one hop of a BIG TCP path raises gso_max_size and the next one
does not.

For an unencapsulated TCP GSO skb which exceeds gso_max_size or
gso_max_segs, work out how many MSS segments each output skb may carry and
re-segment the skb with that max_segs instead.  The bound is measured from
the transport header, so an skb whose transport header is unset or stale
keeps today's segmentation.

An encapsulated or frag-list skb, a GSO type the device cannot offload and
a max_segs which leaves room for a single MSS all keep today's segmentation
as well.  The GSO type test runs on the features without the limit checks,
because gso_features_check() has already cleared the GSO bits of an skb
which exceeds them.

This path emits a plain GSO skb, not a BIG TCP one: inet_gso_segment() and
ipv6_gso_segment() write the whole length of each output into the 16-bit L3
length field.  An egress limit above 64 KiB would give outputs whose length
truncates, so the size limit is also capped at what that field can express.

The helper runs on the skb which is handed to the driver, after
validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
netif_needs_gso() branch: an skb which the device takes as it is pays
nothing.  A skb which reaches that branch pays one device limit test, and
an over-limit one pays the TCP header read and the features recomputation
for max_segs, in exchange for keeping the output a GSO skb.

Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
enabled on the veth endpoints and left off in the guest, so the skbs which
the veth hop accepts have to be segmented before the TAP device.  A single
iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
affinity and port tuple).  The middle column is the same tree with the
re-segmentation disabled:

  protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
  TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
  TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps

Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
IPv4 and 0.44% and 0.44% for IPv6.  A BIG TCP hop which feeds a 64 KiB hop
loses 69% of the throughput of a path which never enables BIG TCP at all;
re-segmentation recovers it, 3.3x over the existing segmentation path and
within noise of the no BIG TCP baseline.

Assisted-by: LLM
Signed-off-by: Wang Zhan <[email protected]>

---
v4:
- drop the comment above the frag_list check
- drop the mac header test, it is always set on this path
- take the TCP header from the checksum start the GSO engine uses
- test the transport header first, skb_transport_header() warns when unset
- shorten the comment above the features lookup
v3: https://lore.kernel.org/[email protected]/
v2: https://lore.kernel.org/[email protected]/
v1: https://lore.kernel.org/[email protected]/
---
 net/core/dev.c | 53 +++++++++++++++++++++++++++++++++++++++++++++++++-
 1 file changed, 52 insertions(+), 1 deletion(-)

diff --git a/net/core/dev.c b/net/core/dev.c
index a6213c9ed5e721..ff9bce2506f670 100644
--- a/net/core/dev.c
+++ b/net/core/dev.c
@@ -3977,6 +3977,53 @@ __netif_skb_features(struct sk_buff *skb, bool 
check_gso_limits)
 }
 EXPORT_SYMBOL(__netif_skb_features);
 
+static unsigned int
+skb_gso_output_max_segs(struct sk_buff *skb, struct net_device *dev)
+{
+       unsigned int mss = skb_shinfo(skb)->gso_size;
+       unsigned int gso_max_size, hdr_len, max_segs;
+       netdev_features_t features;
+       struct tcphdr _tcph, *th;
+
+       if (!skb_is_gso(skb) || !skb_is_gso_tcp(skb) ||
+           skb->encapsulation || skb_has_frag_list(skb))
+               return 0;
+
+       /*
+        * The transport offset can be unset or stale, so hdr_len is only taken
+        * from a header at the checksum start.
+        */
+       if (!skb_transport_header_was_set(skb) ||
+           skb_checksum_start(skb) != skb_transport_header(skb))
+               return 0;
+
+       th = skb_header_pointer(skb, skb_transport_offset(skb), sizeof(_tcph),
+                               &_tcph);
+       if (!th || __tcp_hdrlen(th) < sizeof(*th))
+               return 0;
+
+       hdr_len = skb_transport_header(skb) - skb_mac_header(skb) +
+                 __tcp_hdrlen(th);
+
+       /* The output is still a GSO skb: recheck without the limit checks */
+       features = __netif_skb_features(skb, false) | NETIF_F_GSO_ROBUST;
+       if (!net_gso_ok(features, skb_shinfo(skb)->gso_type))
+               return 0;
+
+       gso_max_size = netif_get_gso_max_size(dev, skb);
+       gso_max_size = min(gso_max_size, GSO_LEGACY_MAX_SIZE);
+
+       /*
+        * gso_within_dev_limits() accepts gso_segs == gso_max_segs but
+        * rejects skb->len >= gso_max_size, so only the size limit needs the
+        * - 1; the inner min() keeps that subtraction from wrapping.
+        */
+       max_segs = (gso_max_size - min(gso_max_size, hdr_len + 1)) / mss;
+       max_segs = min(max_segs, READ_ONCE(dev->gso_max_segs));
+
+       return max_segs;
+}
+
 static int xmit_one(struct sk_buff *skb, struct net_device *dev,
                    struct netdev_queue *txq, bool more)
 {
@@ -4141,9 +4188,13 @@ static struct sk_buff *validate_xmit_skb(struct sk_buff 
*skb, struct net_device
                goto out_null;
 
        if (netif_needs_gso(skb, features)) {
+               unsigned int max_segs = 0;
                struct sk_buff *segs;
 
-               segs = skb_gso_segment(skb, features);
+               if (unlikely(!gso_within_dev_limits(skb, dev)))
+                       max_segs = skb_gso_output_max_segs(skb, dev);
+
+               segs = __skb_gso_segment(skb, features, true, max_segs);
                if (IS_ERR(segs)) {
                        goto out_kfree_skb;
                } else if (segs) {
-- 
2.47.3

_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev

Reply via email to