Wang Zhan wrote:
> A GSO skb which exceeds an egress device limit loses its GSO feature mask
> and is segmented into individual packets.  This is unnecessarily expensive
> when the device can still offload smaller TCP GSO skbs, which is easy to
> hit once one hop of a BIG TCP path raises gso_max_size and the next one
> does not.
> 
> For an unencapsulated TCP GSO skb which exceeds gso_max_size or
> gso_max_segs, work out how many MSS segments each output skb may carry and
> re-segment the skb with that max_segs instead.  The bound is measured from
> the transport header, so an skb whose transport header is unset or stale
> keeps today's segmentation.
> 
> An encapsulated or frag-list skb, a GSO type the device cannot offload and
> a max_segs which leaves room for a single MSS all keep today's segmentation
> as well.  The GSO type test runs on the features without the limit checks,
> because gso_features_check() has already cleared the GSO bits of an skb
> which exceeds them.
> 
> This path emits a plain GSO skb, not a BIG TCP one: inet_gso_segment() and
> ipv6_gso_segment() write the whole length of each output into the 16-bit L3
> length field.  An egress limit above 64 KiB would give outputs whose length
> truncates, so the size limit is also capped at what that field can express.
> 
> The helper runs on the skb which is handed to the driver, after
> validate_xmit_vlan() and sk_validate_xmit_skb(), and only from the
> netif_needs_gso() branch: an skb which the device takes as it is pays
> nothing.  A skb which reaches that branch pays one device limit test, and
> an over-limit one pays the TCP header read and the features recomputation
> for max_segs, in exchange for keeping the output a GSO skb.
> 
> Measured on a veth -> bridge -> TAP -> guest virtio-net path, with BIG TCP
> enabled on the veth endpoints and left off in the guest, so the skbs which
> the veth hop accepts have to be segmented before the TAP device.  A single
> iperf3 TCP flow, six alternating runs per state (`-t 15 -O 5`, fixed CPU
> affinity and port tuple).  The middle column is the same tree with the
> re-segmentation disabled:
> 
>   protocol  no BIG TCP   mixed, no reseg  mixed, resegmented
>   TCP/IPv4  51.550 Gbps  15.850 Gbps      52.617 Gbps
>   TCP/IPv6  52.050 Gbps  15.783 Gbps      51.933 Gbps
> 
> Coefficient of variation for the two mixed columns was 0.48% and 0.82% for
> IPv4 and 0.44% and 0.44% for IPv6.  A BIG TCP hop which feeds a 64 KiB hop
> loses 69% of the throughput of a path which never enables BIG TCP at all;
> re-segmentation recovers it, 3.3x over the existing segmentation path and
> within noise of the no BIG TCP baseline.
> 
> Assisted-by: LLM
> Signed-off-by: Wang Zhan <[email protected]>
> 
> ---
> v4:
> - drop the comment above the frag_list check
> - drop the mac header test, it is always set on this path
> - take the TCP header from the checksum start the GSO engine uses
> - test the transport header first, skb_transport_header() warns when unset
> - shorten the comment above the features lookup
> v3: https://lore.kernel.org/[email protected]/
> v2: https://lore.kernel.org/[email protected]/
> v1: https://lore.kernel.org/[email protected]/

Reviewed-by: Willem de Bruijn <[email protected]>
_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev

Reply via email to