Currently, check_iphdr() sets the transport header offset for all IPv4 packets, including later fragments that do not contain an L4 header.
Stop setting the transport header for non-first IPv4 fragments. This leaves skb->transport_header at whatever value it had when the packet entered OVS. That is only safe if no later code relies on it. update_ip_l4_checksum() read the transport offset before checking for a later fragment, so do the check first. set_tcp(), set_udp() and set_sctp() had no such check. An IPv4 later fragment keeps ip.proto set to the real protocol, and the flow may wildcard the frag field, so validate_set() accepts these actions for it. They then modified the start of the fragment payload as if it were an L4 header, corrupting the packet. Skip the actions for OVS_FRAG_TYPE_LATER, as there is no L4 header to modify. IPv6 later fragments are not affected, since their ip.proto is NEXTHDR_FRAGMENT and validate_set() rejects L4 set actions for them. Link: https://lore.kernel.org/netdev/[email protected]/ Suggested-by: Ilya Maximets <[email protected]> Fixes: ccb1352e76cf ("net: Add Open vSwitch kernel components.") Fixes: a175a723301a ("openvswitch: Add SCTP support") Signed-off-by: Fernando Fernandez Mancera <[email protected]> --- Note: since the data corruption issue was actually found by sashiko, I suggest we still take this on net-next instead of net, this is just a corner-case most users won't ever hit. Sashiko: If you notice a similar problem at IPv6 path, we have fixed it already in net tree, here is submission: https://lore.kernel.org/netdev/[email protected]/ --- net/openvswitch/actions.c | 19 ++++++++++++++++--- net/openvswitch/flow.c | 9 +++++++-- 2 files changed, 23 insertions(+), 5 deletions(-) diff --git a/net/openvswitch/actions.c b/net/openvswitch/actions.c index dc5ff859f114..8556a5a74ccd 100644 --- a/net/openvswitch/actions.c +++ b/net/openvswitch/actions.c @@ -322,11 +322,13 @@ static int pop_nsh(struct sk_buff *skb, struct sw_flow_key *key) static void update_ip_l4_checksum(struct sk_buff *skb, struct iphdr *nh, __be32 addr, __be32 new_addr) { - int transport_len = skb->len - skb_transport_offset(skb); + int transport_len; if (nh->frag_off & htons(IP_OFFSET)) return; + transport_len = skb->len - skb_transport_offset(skb); + if (nh->protocol == IPPROTO_TCP) { if (likely(transport_len >= sizeof(struct tcphdr))) inet_proto_csum_replace4(&tcp_hdr(skb)->check, skb, @@ -590,6 +592,9 @@ static int set_udp(struct sk_buff *skb, struct sw_flow_key *flow_key, __be16 src, dst; int err; + if (flow_key->ip.frag == OVS_FRAG_TYPE_LATER) + return 0; + err = skb_ensure_writable(skb, skb_transport_offset(skb) + sizeof(struct udphdr)); if (unlikely(err)) @@ -633,6 +638,9 @@ static int set_tcp(struct sk_buff *skb, struct sw_flow_key *flow_key, __be16 src, dst; int err; + if (flow_key->ip.frag == OVS_FRAG_TYPE_LATER) + return 0; + err = skb_ensure_writable(skb, skb_transport_offset(skb) + sizeof(struct tcphdr)); if (unlikely(err)) @@ -658,11 +666,16 @@ static int set_sctp(struct sk_buff *skb, struct sw_flow_key *flow_key, const struct ovs_key_sctp *key, const struct ovs_key_sctp *mask) { - unsigned int sctphoff = skb_transport_offset(skb); - struct sctphdr *sh; __le32 old_correct_csum, new_csum, old_csum; + unsigned int sctphoff; + struct sctphdr *sh; int err; + if (flow_key->ip.frag == OVS_FRAG_TYPE_LATER) + return 0; + + sctphoff = skb_transport_offset(skb); + err = skb_ensure_writable(skb, sctphoff + sizeof(struct sctphdr)); if (unlikely(err)) return err; diff --git a/net/openvswitch/flow.c b/net/openvswitch/flow.c index 1c4f3a079044..52cc63837445 100644 --- a/net/openvswitch/flow.c +++ b/net/openvswitch/flow.c @@ -189,6 +189,7 @@ static bool arphdr_ok(struct sk_buff *skb) static int check_iphdr(struct sk_buff *skb) { unsigned int nh_ofs = skb_network_offset(skb); + const struct iphdr *nh; unsigned int ip_len; int err; @@ -201,7 +202,10 @@ static int check_iphdr(struct sk_buff *skb) skb->len < nh_ofs + ip_len)) return -EINVAL; - skb_set_transport_header(skb, nh_ofs + ip_len); + nh = ip_hdr(skb); + if (!(nh->frag_off & htons(IP_OFFSET))) + skb_set_transport_header(skb, nh_ofs + ip_len); + return 0; } @@ -898,7 +902,8 @@ static int key_extract_l3l4(struct sk_buff *skb, struct sw_flow_key *key) * - skb->transport_header: If key->eth.type is ETH_P_IP or ETH_P_IPV6 * on output, then just past the IP header, if one is present and * of a correct length, otherwise the same as skb->network_header. - * For other key->eth.type values it is left untouched. + * For later IP fragments and for other key->eth.type values it is + * left untouched. * * - skb->protocol: the type of the data starting at skb->network_header. * Equals to key->eth.type. -- 2.55.0 _______________________________________________ dev mailing list [email protected] https://mail.openvswitch.org/mailman/listinfo/ovs-dev
