Multichassis ports force traffic through tunnels during live migration.
Deriving the packet size limit from the VIF MTU unnecessarily reduces the
PMTU on provider networks whose tunnel underlay can carry larger frames.

Read NB_Global options:tunnel_mtu through the existing SB_Global options
propagation and allow a local external_ids:ovn-tunnel-mtu override. Use
the configured value minus the largest overhead of the chassis' tunnels
while retaining oversized-packet checks and ICMP errors. Preserve the
existing VIF-based behavior when no usable value is set. Recompute
physical flows when either configuration source changes.

Values below 1280 + tunnel overhead + 18 bytes fall back to the VIF MTU,
so generated ICMP errors fit and the configured limit advertises an IPv6
PMTU of at least 1280. When the option is set, install the checks once
for each multichassis port in consider_port_binding(), on every chassis
with the switch as a local datapath, and own them by that port binding.
The checks match all traffic to and from the port, so use the largest
overhead of the chassis' tunnels. Unlike a maximum over the switch's
port bindings, it only depends on tunnels, whose changes already trigger
a full recompute, so port binding changes cannot leave stale or competing
limits behind. Without a tunnel MTU, keep the existing VIF-based checks,
which only chassis that host the port can compute.

Document that the default can only limit packet sizes on chassis that
host the port, that the option applies on chassis with a bridge mapping
for the switch's network, and that the minimum uses the same largest
tunnel overhead as the limit, below which chassis that do not host the
port do not check packet sizes. Also document that a chassis override
below the minimum falls back to the VIF MTU.

Extend multichassis tests for jumbo underlays, precedence, runtime
changes, invalid values, and the safety floor, including delivery of
1280-byte IPv6 packets. Add a three-chassis test in which the only
sender does not host the multichassis port. It covers ICMP errors
without delivery to either chassis, full-size delivery with a jumbo
tunnel MTU, removal and re-addition of the additional chassis, and a
chassis-local tunnel MTU. Also check 1500-byte IPv4 packets with DF set
and 1500-byte IPv6 packets, which are rejected without a tunnel MTU and
delivered with a 9000-byte one. Add a fourth chassis without a port on
the switch that hv3 reaches with the other encapsulation, and check that
a Geneve tunnel lowers hv3's limit, that a VXLAN tunnel does not, and
that the limit follows tunnel changes without differences after a
recompute. Also check that an override below the minimum and
always_tunnel remove the checks from the chassis that does not host the
port.

Submitted-at: https://github.com/ovn-org/ovn/pull/329
Assisted-by: Claude Opus 5.5 (claude-opus-5-5), Cursor Grok Bot / Ultimum 
harness assistants
Signed-off-by: Premysl Kouril <[email protected]>
---
 NEWS                            |   3 +
 controller/ovn-controller.8.xml |  15 ++
 controller/ovn-controller.c     |  33 +++
 controller/physical.c           |  67 ++++-
 controller/physical.h           |   1 +
 ovn-architecture.7.xml          |  10 +
 ovn-nb.xml                      |  34 +++
 ovn-sb.xml                      |   5 +
 tests/ovn.at                    | 442 +++++++++++++++++++++++++++++++-
 9 files changed, 593 insertions(+), 17 deletions(-)

diff --git a/NEWS b/NEWS
index f1c56dd14..b73624af2 100644
--- a/NEWS
+++ b/NEWS
@@ -2,6 +2,9 @@ Post v26.09.0
 --------------
    - Added a new "options:ttl" key on the NB DNS table to make the TTL of
      DNS replies from OVN's native DNS resolver configurable per row.
+   - Add NB_Global options:tunnel_mtu and a per-chassis
+     external_ids:ovn-tunnel-mtu override to base multichassis path MTU
+     discovery on the tunnel underlay MTU instead of the VIF MTU.
    - Removed implementations of the commit_ecmp_nh, chk_ecmp_nh, and
      chk_ecmp_nh_mac actions from the code.
    - Mark tunnel ports as transient (other_config:transient=true) when the
diff --git a/controller/ovn-controller.8.xml b/controller/ovn-controller.8.xml
index 3c33654ff..071bccf8a 100644
--- a/controller/ovn-controller.8.xml
+++ b/controller/ovn-controller.8.xml
@@ -197,6 +197,21 @@
         </p>
       </dd>
 
+      <dt><code>external_ids:ovn-tunnel-mtu</code></dt>
+      <dd>
+        <p>
+          Optional underlay MTU in bytes for path MTU discovery on multichassis
+          logical ports.  An integer from 1 through 65535 overrides
+          <code>NB_Global.options:tunnel_mtu</code>; an unset or invalid
+          override uses the global setting.  An override below the minimum
+          safe value falls back to the VIF MTU instead, so a chassis that does
+          not host a multichassis port then does not check its packet sizes.
+          See
+          <code>ovn-nb</code>(5) for the tunnel path requirements and minimum
+          safe value.
+        </p>
+      </dd>
+
       <dt><code>external_ids:ovn-encap-ip</code></dt>
       <dd>
         <p>
diff --git a/controller/ovn-controller.c b/controller/ovn-controller.c
index c601f89dc..3a104a3c1 100644
--- a/controller/ovn-controller.c
+++ b/controller/ovn-controller.c
@@ -3743,8 +3743,25 @@ struct ed_type_northd_options {
                          * be tunnelled or sent via the localnet
                          * port.  Default value is 'false'. */
     bool enable_ch_nb_cfg_update;
+    uint16_t tunnel_mtu;
 };
 
+/* Zero means that multichassis PMTU discovery uses the VIF MTU. */
+static uint16_t
+parse_tunnel_mtu(const char *value, uint16_t def)
+{
+    if (!value) {
+        return def;
+    }
+
+    int mtu;
+    if (!str_to_int(value, 10, &mtu) || mtu <= 0 || mtu > UINT16_MAX) {
+        static struct vlog_rate_limit rl = VLOG_RATE_LIMIT_INIT(1, 1);
+        VLOG_WARN_RL(&rl, "Invalid tunnel MTU: %s", value);
+        return def;
+    }
+    return mtu;
+}
 
 static void *
 en_northd_options_init(struct engine_node *node OVS_UNUSED,
@@ -3780,6 +3797,9 @@ en_northd_options_run(struct engine_node *node, void 
*data)
                         true)
         : true;
 
+    n_opts->tunnel_mtu = parse_tunnel_mtu(
+        sb_global ? smap_get(&sb_global->options, "tunnel_mtu") : NULL, 0);
+
     return EN_UPDATED;
 }
 
@@ -3815,6 +3835,13 @@ en_northd_options_sb_sb_global_handler(struct 
engine_node *node, void *data)
         result = EN_HANDLED_UPDATED;
     }
 
+    uint16_t tunnel_mtu = parse_tunnel_mtu(
+        sb_global ? smap_get(&sb_global->options, "tunnel_mtu") : NULL, 0);
+    if (tunnel_mtu != n_opts->tunnel_mtu) {
+        n_opts->tunnel_mtu = tunnel_mtu;
+        result = EN_HANDLED_UPDATED;
+    }
+
     return result;
 }
 
@@ -4806,6 +4833,12 @@ static void init_physical_ctx(struct engine_node *node,
     p_ctx->flow_tunnels = non_vif_data->flow_tunnels;
     p_ctx->use_flow_based_tunnels = non_vif_data->use_flow_based_tunnels;
     p_ctx->always_tunnel = n_opts->always_tunnel;
+    const struct ovsrec_open_vswitch *cfg =
+        ovsrec_open_vswitch_table_first(ovs_table);
+    p_ctx->tunnel_mtu = parse_tunnel_mtu(
+        cfg ? get_chassis_external_id_value(&cfg->external_ids, chassis_id,
+                                             "ovn-tunnel-mtu", NULL) : NULL,
+        n_opts->tunnel_mtu);
     p_ctx->evpn_bindings = &eb_data->bindings;
     p_ctx->evpn_multicast_groups = &eb_data->multicast_groups;
     p_ctx->evpn_fdbs = &efdb_data->fdbs;
diff --git a/controller/physical.c b/controller/physical.c
index 8c5d9b940..ac75b4cc4 100644
--- a/controller/physical.c
+++ b/controller/physical.c
@@ -2069,15 +2069,8 @@ get_tunnel_overhead(struct chassis_tunnel const *tun)
 static uint16_t
 get_effective_mtu(const struct sbrec_port_binding *mcp,
                   struct vector *remote_tunnels,
-                  const struct if_status_mgr *if_mgr)
+                  const struct physical_ctx *ctx)
 {
-    /* Use interface MTU as a base for calculation */
-    uint16_t iface_mtu = if_status_mgr_iface_get_mtu(if_mgr,
-                                                     mcp->logical_port);
-    if (!iface_mtu) {
-        return 0;
-    }
-
     /* Iterate over all peer tunnels and find the biggest tunnel overhead */
     uint16_t overhead = 0;
     const struct chassis_tunnel *tun;
@@ -2088,7 +2081,26 @@ get_effective_mtu(const struct sbrec_port_binding *mcp,
         return 0;
     }
 
-    return iface_mtu - overhead;
+    uint16_t mtu = ctx->tunnel_mtu;
+    /* Both IP versions use this limit.  Leave room for the minimum IPv6 MTU
+     * advertised by reply_icmp_error_if_pkt_too_big(), including Ethernet
+     * overhead.  Otherwise, PMTU discovery cannot converge and the generated
+     * ICMP errors can themselves exceed the limit and trigger more errors. */
+    uint16_t min_mtu = 1280 + overhead + ETHERNET_OVERHEAD;
+    if (mtu && mtu < min_mtu) {
+        static struct vlog_rate_limit rl = VLOG_RATE_LIMIT_INIT(1, 1);
+        VLOG_WARN_RL(&rl, "Tunnel MTU %"PRIu16" is too small; minimum is "
+                     "%"PRIu16" for overhead %"PRIu16, mtu, min_mtu,
+                     overhead + ETHERNET_OVERHEAD);
+        mtu = 0;
+    }
+    if (!mtu) {
+        /* Preserve the VIF-based calculation when no usable tunnel MTU
+         * is configured. */
+        mtu = if_status_mgr_iface_get_mtu(ctx->if_mgr, mcp->logical_port);
+    }
+
+    return mtu ? mtu - overhead : 0;
 }
 
 static void
@@ -2115,9 +2127,9 @@ handle_pkt_too_big(struct ovn_desired_flow_table 
*flow_table,
                    struct vector *remote_tunnels,
                    const struct sbrec_port_binding *binding,
                    const struct sbrec_port_binding *mcp,
-                   const struct if_status_mgr *if_mgr)
+                   const struct physical_ctx *ctx)
 {
-    uint16_t mtu = get_effective_mtu(mcp, remote_tunnels, if_mgr);
+    uint16_t mtu = get_effective_mtu(mcp, remote_tunnels, ctx);
     if (!mtu) {
         return;
     }
@@ -2125,6 +2137,25 @@ handle_pkt_too_big(struct ovn_desired_flow_table 
*flow_table,
     handle_pkt_too_big_for_ip_version(flow_table, binding, mcp, mtu, true);
 }
 
+/* Checks packet sizes for multichassis port 'mcp' based on the tunnel MTU,
+ * which unlike the VIF MTU is also known on chassis that do not host the
+ * port.  The checks match all of the port's traffic, whatever the peer, so
+ * use the largest overhead of all tunnels. */
+static void
+handle_multichassis_pkt_too_big(struct ovn_desired_flow_table *flow_table,
+                                const struct sbrec_port_binding *mcp,
+                                const struct physical_ctx *ctx)
+{
+    struct vector tuns =
+        VECTOR_EMPTY_INITIALIZER(const struct chassis_tunnel *);
+    const struct chassis_tunnel *tun;
+    HMAP_FOR_EACH (tun, hmap_node, ctx->chassis_tunnels) {
+        vector_push(&tuns, &tun);
+    }
+    handle_pkt_too_big(flow_table, &tuns, mcp, mcp, ctx);
+    vector_destroy(&tuns);
+}
+
 /* XXX: Need to support flow-based tunnel for this function. */
 static void
 enforce_tunneling_for_multichassis_ports(
@@ -2176,7 +2207,11 @@ enforce_tunneling_for_multichassis_ports(
                         &binding->header_.uuid);
         ofpbuf_uninit(&ofpacts);
 
-        handle_pkt_too_big(flow_table, &tuns, binding, mcp, ctx->if_mgr);
+        /* With a tunnel MTU, consider_port_binding() checks packet sizes
+         * for each multichassis port. */
+        if (!ctx->tunnel_mtu) {
+            handle_pkt_too_big(flow_table, &tuns, binding, mcp, ctx);
+        }
     }
     vector_destroy(&tuns);
 }
@@ -2212,6 +2247,14 @@ consider_port_binding(const struct physical_ctx *ctx,
         return;
     }
 
+    /* Switches with a localnet port tunnel the traffic of multichassis
+     * ports, also from chassis that do not host them.  With a tunnel MTU,
+     * all of these chassis check packet sizes. */
+    if (ctx->tunnel_mtu && binding->n_additional_chassis
+        && ld->localnet_port && !ctx->always_tunnel) {
+        handle_multichassis_pkt_too_big(flow_table, binding, ctx);
+    }
+
     if (type == LP_VIF) {
         /* Table 104, priority 100.
          * ========================
diff --git a/controller/physical.h b/controller/physical.h
index 8218096d5..696f5f83f 100644
--- a/controller/physical.h
+++ b/controller/physical.h
@@ -73,6 +73,7 @@ struct physical_ctx {
     const char **encap_ips;
     struct physical_debug debug;
     bool always_tunnel;
+    uint16_t tunnel_mtu;
     const struct hmap *evpn_bindings;
     const struct hmap *evpn_multicast_groups;
     const struct hmap *evpn_fdbs;
diff --git a/ovn-architecture.7.xml b/ovn-architecture.7.xml
index 00728ee9b..90efe17f7 100644
--- a/ovn-architecture.7.xml
+++ b/ovn-architecture.7.xml
@@ -864,6 +864,16 @@
     Set the <code>redirect-type</code> option on a distributed gateway port.
   </p>
 
+  <p>
+    During live migration, multichassis logical ports use tunnels even on
+    networks bridged to a physical VLAN.  By default, OVN can only limit
+    packet sizes on chassis that host such a port, using the VIF MTU minus
+    tunnel overhead, and returns ICMP errors for oversized IP packets.  Set
+    <code>NB_Global.options:tunnel_mtu</code> to use the tunnel path's MTU
+    instead, also on chassis that send to the port without hosting it; see
+    <code>ovn-nb</code>(5) for details.
+  </p>
+
   <h4>Using Distributed Gateway Ports For Scalability</h4>
 
   <p>
diff --git a/ovn-nb.xml b/ovn-nb.xml
index 078bd6068..05affb1c4 100644
--- a/ovn-nb.xml
+++ b/ovn-nb.xml
@@ -427,6 +427,40 @@
         to true.
       </column>
 
+      <column name="options" key="tunnel_mtu">
+        <p>
+          Underlay MTU, in bytes, used for path MTU discovery on multichassis
+          logical ports.  This must be the smallest MTU supported along the
+          entire tunnel path between chassis.  It only affects the packet
+          size limit for multichassis ports on switches with a
+          <code>localnet</code> port and <ref column="options"
+          key="always_tunnel"/> disabled, on chassis with a bridge mapping
+          for that port's network.
+          A chassis can override this value with
+          <code>external_ids:ovn-tunnel-mtu</code>;
+          see <code>ovn-controller</code>(8).
+        </p>
+        <p>
+          Each chassis subtracts the largest overhead of its tunnels from this
+          value to determine the packet size limit.  It checks the limit for
+          traffic to and from multichassis ports, including ports that it does
+          not host.  Larger IP packets are dropped and receive ICMP
+          Fragmentation Needed or Packet Too Big errors.  For example, 9000
+          with Geneve over IPv4 gives a packet size limit of 8942 and an
+          advertised IP MTU of 8924.
+        </p>
+        <p>
+          When unset, invalid, or below 1280 plus the same tunnel overhead and
+          18 bytes of Ethernet overhead (1356 for Geneve over IPv4, 1376 if a
+          chassis also has tunnels with Geneve over IPv6), OVN falls back to
+          the VIF MTU.  Only chassis that host a multichassis port know its VIF
+          MTU, so other chassis then do not check packet sizes, and it can be
+          smaller than the tunnel path supports.  Upgrading OVN
+          therefore does not raise the packet size limit until this option is
+          set.
+        </p>
+      </column>
+
       <column name="options" key="always_tunnel"
            type='{"type": "boolean"}'>
         <p>
diff --git a/ovn-sb.xml b/ovn-sb.xml
index 2096fc3e3..05659552f 100644
--- a/ovn-sb.xml
+++ b/ovn-sb.xml
@@ -193,6 +193,11 @@
         </p>
       </column>
 
+      <column name="options" key="tunnel_mtu">
+        Copied from <ref db="OVN_Northbound" table="NB_Global"
+        column="options" key="tunnel_mtu"/> by <code>ovn-northd</code>.
+      </column>
+
       <group title="Options for configuring BFD">
         <p>
           These options apply when <code>ovn-controller</code> configures
diff --git a/tests/ovn.at b/tests/ovn.at
index 13e95f9db..943abc25b 100644
--- a/tests/ovn.at
+++ b/tests/ovn.at
@@ -16929,7 +16929,7 @@ AT_CLEANUP
 m4_define([MULTICHASSIS_PATH_MTU_DISCOVERY_TEST],
   [OVN_FOR_EACH_NORTHD([
    AT_SETUP([localnet connectivity with multiple requested-chassis, path mtu 
discovery (ip=$1, tunnel=$2, mtu=$3)])
-   AT_KEYWORDS([multi-chassis])
+   AT_KEYWORDS([multi-chassis tunnel-mtu])
    CHECK_SCAPY
 
    ovn_start
@@ -17000,16 +17000,16 @@ m4_define([MULTICHASSIS_PATH_MTU_DISCOVERY_TEST],
    done
 
    send_ip_packet() {
-       local inport=${1} hv=${2} eth_src=${3} eth_dst=${4} ipv4_src=${5} 
ipv4_dst=${6} data=${7} fail=${8} mtu=${9:-$3}
+       local inport=${1} hv=${2} eth_src=${3} eth_dst=${4} ipv4_src=${5} 
ipv4_dst=${6} data=${7} fail=${8} mtu=${9:-$3} flags=${10:-0}
        packet=$(fmt_pkt "
            Ether(dst='${eth_dst}', src='${eth_src}') /
-           IP(src='${ipv4_src}', dst='${ipv4_dst}') /
+           IP(src='${ipv4_src}', dst='${ipv4_dst}', flags=${flags}) /
            ICMP(type=8) / bytes.fromhex('${data}')
        ")
        as hv${hv} ovs-appctl netdev-dummy/receive ${inport} ${packet}
        if [[ x"${fail}" != x0 ]]; then
          original_ip_frame=$(fmt_pkt "
-           IP(src='${ipv4_src}', dst='${ipv4_dst}') /
+           IP(src='${ipv4_src}', dst='${ipv4_dst}', flags=${flags}) /
            ICMP(type=8) / bytes.fromhex('${data}')
          ")
          # IP(flags=2) means DF (Don't Fragment) = 1
@@ -17247,8 +17247,197 @@ m4_define([MULTICHASSIS_PATH_MTU_DISCOVERY_TEST],
    echo $packet >> hv1/multi1.expected
 
    check_pkts
+   reset_env
+
+   # Check the programmed limit explicitly so packet tests never race flow
+   # updates, including those triggered only by local OVS configuration.
+   wait_for_mtu() {
+       local hv=${1} limit=$((${2} + 18))
+       OVS_WAIT_UNTIL([test "$(as ${hv} ovs-ofctl dump-flows br-int |
+           sed -n 's/.*check_pkt_larger(\([[0-9]]*\)).*/\1/p' |
+           sort -u)" = "${limit}"])
+   }
+
+   # Exercise both IP versions, both directions, and local and remote peers.
+   # Payload lengths are in hexadecimal digits (6000 means 3000 bytes).
+   # "full" sends 1500-byte IPv4 packets with DF set and 1500-byte IPv6
+   # packets.
+   check_tunnel_mtu_packets() {
+       local hv=${1} len=${2} fail=${3} mtu=${4}
+       local port peer peer_hv mac peer_mac ip ip6 peer_ip peer_ip6
+       if test "${hv}" = 1; then
+           port=first peer=second peer_hv=2
+           mac=$first_mac peer_mac=$second_mac
+           ip=$first_ip ip6=$first_ip6
+           peer_ip=$second_ip peer_ip6=$second_ip6
+       else
+           port=second peer=first peer_hv=1
+           mac=$second_mac peer_mac=$first_mac
+           ip=$second_ip ip6=$second_ip6
+           peer_ip=$first_ip peer_ip6=$first_ip6
+       fi
+
+       for version in 4 6; do
+           local send src dst remote data_len=${len} flags=0
+           if test "${version}" = 4; then
+               send=send_ip_packet
+               src=$multi1_ip dst=$ip remote=$peer_ip
+               if test "${len}" = full; then
+                   data_len=2944 flags=2
+               fi
+           else
+               send=send_ip6_packet
+               src=$multi1_ip6 dst=$ip6 remote=$peer_ip6
+               if test "${len}" = full; then
+                   data_len=2904
+               fi
+           fi
+           packet=$($send multi1 $hv $multi1_mac $mac $src $dst $(payload 
$data_len) $fail $mtu $flags)
+           if test "${fail}" = 1; then
+               echo $packet >> hv${hv}/multi1.expected
+           else
+               echo $packet >> hv${hv}/${port}.expected
+           fi
+
+           packet=$($send multi1 $hv $multi1_mac $peer_mac $src $remote 
$(payload $data_len) $fail $mtu $flags)
+           if test "${fail}" = 1; then
+               echo $packet >> hv${hv}/multi1.expected
+           else
+               echo $packet >> hv${peer_hv}/${peer}.expected
+           fi
+
+           packet=$($send $port $hv $mac $multi1_mac $dst $src $(payload 
$data_len) $fail $mtu $flags)
+           if test "${fail}" = 1; then
+               echo $packet >> hv${hv}/${port}.expected
+           else
+               echo $packet >> hv1/multi1.expected
+               echo $packet >> hv2/multi1.expected
+           fi
+       done
+       check_pkts
+       reset_env
+   }
+
+   AS_BOX([Jumbo tunnel MTU permits large packets with 1500-byte VIF MTUs])
+   set_mtu_for_all_ports 1500
+   for hv in hv1 hv2; do
+       # Model a jumbo underlay independently of the VIF MTUs.
+       as $hv check ovs-vsctl set Interface br-phys mtu_request=10000 \
+           -- set Interface br-phys_n1 mtu_request=10000
+       as main check ovs-vsctl set Interface ${hv}_br-phys mtu_request=10000
+       wait_for_mtu $hv $3
+       AT_CHECK([as $hv ovs-vsctl get Interface multi1 mtu], [0], [1500
+])
+   done
+   # Without a tunnel MTU, the VIF MTU still limits full-size packets.
+   for hv in 1 2; do
+       check_tunnel_mtu_packets $hv full 1 $3
+   done
+   check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=9000
+   AT_CHECK([ovn-sbctl get SB_Global . options:tunnel_mtu], [0], ["9000"
+])
+   jumbo_mtu=$(($3 + 7500))
+   for hv in hv1 hv2; do
+       wait_for_mtu $hv $jumbo_mtu
+   done
+   for hv in 1 2; do
+       check_tunnel_mtu_packets $hv 2880 0 $jumbo_mtu
+       check_tunnel_mtu_packets $hv full 0 $jumbo_mtu
+   done
+
+   # Enlarge the dummy netdevs for jumbo packet delivery.
+   set_mtu_for_all_ports 10000
+   for hv in hv1 hv2; do
+       wait_for_mtu $hv $jumbo_mtu
+   done
+   for hv in 1 2; do
+       check_tunnel_mtu_packets $hv 6000 0 $jumbo_mtu
+       # The cap remains in place and advertises the underlay-based MTU.
+       check_tunnel_mtu_packets $hv 18000 1 $jumbo_mtu
+   done
+
+   AS_BOX([Tunnel MTU precedence and runtime changes])
+   as hv1 check ovs-vsctl set Open_vSwitch . external_ids:ovn-tunnel-mtu=1500
+   wait_for_mtu hv1 $3
+   wait_for_mtu hv2 $jumbo_mtu
+   check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=1600
+   wait_for_mtu hv1 $3
+   wait_for_mtu hv2 $(($3 + 100))
+   as hv1 check ovs-vsctl remove Open_vSwitch . external_ids ovn-tunnel-mtu
+   wait_for_mtu hv1 $(($3 + 100))
+   check ovn-nbctl --wait=hv remove NB_Global . options tunnel_mtu
+   AT_CHECK([ovn-sbctl get SB_Global . options:tunnel_mtu], [1], [], [ignore])
+   for hv in hv1 hv2; do
+       wait_for_mtu $hv $(($3 + 8500))
+   done
+
+   AS_BOX([A 1500-byte tunnel MTU retains the original ICMP MTU])
+   check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=1500
+   for hv in hv1 hv2; do
+       wait_for_mtu $hv $3
+   done
+   for hv in 1 2; do
+       check_tunnel_mtu_packets $hv 6000 1 $3
+       check_tunnel_mtu_packets $hv full 1 $3
+   done
+
+   AS_BOX([Invalid overrides use the global MTU without integer wrapping])
+   check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=9000
+   wait_for_mtu hv1 $jumbo_mtu
+   as hv1 check ovs-vsctl set Open_vSwitch . external_ids:ovn-tunnel-mtu=1500
+   wait_for_mtu hv1 $3
+   as hv1 check ovs-vsctl set Open_vSwitch . external_ids:ovn-tunnel-mtu=65536
+   wait_for_mtu hv1 $jumbo_mtu
+   as hv1 check ovs-vsctl remove Open_vSwitch . external_ids ovn-tunnel-mtu
+
+   AS_BOX([Invalid global settings restore VIF-based behavior])
+   set_mtu_for_all_ports 1500
+   for value in invalid 65536; do
+       check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=$value
+       for hv in hv1 hv2; do
+           wait_for_mtu $hv $3
+       done
+   done
+
+   AS_BOX([Low positive tunnel MTUs fall back without ICMP error loops])
+   # Both IP versions share a limit.  Leave room for an advertised IPv6 MTU
+   # of at least 1280, in addition to tunnel and Ethernet overhead.
+   min_tunnel_mtu=$((1500 - $3 + 1280))
+   check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=9000
+   for hv in hv1 hv2; do
+       wait_for_mtu $hv $jumbo_mtu
+   done
+   for value in 1280 $((min_tunnel_mtu - 1)); do
+       check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=$value
+       for hv in hv1 hv2; do
+           wait_for_mtu $hv $3
+       done
+       # Each request must produce exactly one error, including the
+       # 1280-byte IPv6 error frame.
+       packet=$(send_ip_packet first 1 $first_mac $multi1_mac $first_ip 
$multi1_ip $(payload 2880) 1 $3)
+       echo $packet >> hv1/first.expected
+       packet=$(send_ip6_packet first 1 $first_mac $multi1_mac $first_ip6 
$multi1_ip6 $(payload 2880) 1 $3)
+       echo $packet >> hv1/first.expected
+       check_pkts
+       reset_env
+   done
+
+   AS_BOX([The minimum safe tunnel MTU gives usable IPv6 PMTU feedback])
+   check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=$min_tunnel_mtu
+   for hv in hv1 hv2; do
+       wait_for_mtu $hv 1280
+   done
+   # Oversized packets get feedback; IPv6 packets with 1280 L3 bytes fit.
+   check_tunnel_mtu_packets 1 2880 1 1280
+   check_tunnel_mtu_packets 1 2464 0 1280
 
-   OVN_CLEANUP([hv1],[hv2])
+   OVN_CLEANUP([hv1
+/Invalid tunnel MTU:/d
+/Tunnel MTU .* is too small; minimum is/d
+], [hv2
+/Invalid tunnel MTU:/d
+/Tunnel MTU .* is too small; minimum is/d
+])
 
    AT_CLEANUP
    ])])
@@ -17258,6 +17447,249 @@ MULTICHASSIS_PATH_MTU_DISCOVERY_TEST([ipv6], 
[geneve], [1404])
 MULTICHASSIS_PATH_MTU_DISCOVERY_TEST([ipv4], [vxlan], [1432])
 MULTICHASSIS_PATH_MTU_DISCOVERY_TEST([ipv6], [vxlan], [1412])
 
+m4_define([MULTICHASSIS_TUNNEL_MTU_REMOTE_SENDER_TEST],
+  [OVN_FOR_EACH_NORTHD([
+   AT_SETUP([localnet connectivity with multiple requested-chassis, tunnel mtu 
on non-hosting chassis (ip=$1, tunnel=$2, mtu=$3)])
+   AT_KEYWORDS([multi-chassis tunnel-mtu])
+   CHECK_SCAPY
+
+   ovn_start
+
+   net_add n1
+   for i in 1 2 3; do
+       sim_add hv$i
+       as hv$i
+       check ovs-vsctl add-br br-phys
+       if test "x$1" = "xipv6"; then
+           ovn_attach n1 br-phys fd00::$i 64 $2
+       else
+           ovn_attach n1 br-phys 192.168.0.$i 24 $2
+       fi
+       check ovs-vsctl set open . external-ids:ovn-bridge-mappings=phys:br-phys
+       # Model a jumbo underlay, so that only OVN limits packet sizes.
+       check ovs-vsctl set Interface br-phys mtu_request=10000 \
+           -- set Interface br-phys_n1 mtu_request=10000
+       as main check ovs-vsctl set Interface hv${i}_br-phys mtu_request=10000
+   done
+
+   multi_mac=00:00:00:00:00:f0
+   third_mac=00:00:00:00:00:03
+   multi_ip=10.0.0.10
+   third_ip=10.0.0.3
+   multi_ip6=abcd::f0
+   third_ip6=abcd::3
+
+   # Only hv1 and hv2 host the multichassis port.  hv3 has no other remote
+   # port on the switch, so it only tunnels to the multichassis port.
+   check ovn-nbctl ls-add ls0
+   check ovn-nbctl lsp-add-localnet-port ls0 public phys
+   check ovn-nbctl lsp-add ls0 multi
+   check ovn-nbctl lsp-set-addresses multi "${multi_mac} ${multi_ip} 
${multi_ip6}"
+   check ovn-nbctl lsp-set-options multi requested-chassis=hv1,hv2
+   check ovn-nbctl lsp-add ls0 third
+   check ovn-nbctl lsp-set-addresses third "${third_mac} ${third_ip} 
${third_ip6}"
+   check ovn-nbctl lsp-set-options third requested-chassis=hv3
+
+   for hv in hv1 hv2; do
+       as $hv check ovs-vsctl -- add-port br-int multi -- \
+           set Interface multi external-ids:iface-id=multi \
+           options:tx_pcap=$hv/multi-tx.pcap \
+           options:rxq_pcap=$hv/multi-rx.pcap
+   done
+   as hv3 check ovs-vsctl -- add-port br-int third -- \
+       set Interface third external-ids:iface-id=third \
+       options:tx_pcap=hv3/third-tx.pcap \
+       options:rxq_pcap=hv3/third-rx.pcap
+
+   wait_for_ports_up
+   OVN_POPULATE_ARP
+   hv2_uuid=$(fetch_column Chassis _uuid name=hv2)
+   wait_column "$hv2_uuid" Port_Binding additional_chassis logical_port=multi
+   check ovn-nbctl --wait=hv sync
+   multi_key=$(printf "0x%x" $(fetch_column Port_Binding tunnel_key 
logical_port=multi))
+
+   # Sends an IPv4 packet with DF set and an IPv6 packet, each 'size' bytes
+   # long without the Ethernet header, from hv3 to the multichassis port.
+   send_packets() {
+       local size=${1} fail=${2} mtu=${3}
+       local ip ip_error packet original
+
+       ip="IP(src='${third_ip}', dst='${multi_ip}', flags=2) /
+           ICMP(type=8) / Raw(b'x' * $((size - 28)))"
+       ip_error="IP(src='${multi_ip}', dst='${third_ip}', ttl=255, flags=2, 
id=0) /
+           ICMP(type=3, code=4, nexthopmtu=${mtu})"
+       packet=$(fmt_pkt "Ether(dst='${multi_mac}', src='${third_mac}') / 
${ip}")
+       as hv3 ovs-appctl netdev-dummy/receive third ${packet}
+       if test "${fail}" = 1; then
+           original=$(fmt_pkt "${ip}")
+           fmt_pkt "Ether(dst='${third_mac}', src='${multi_mac}') /
+               ${ip_error} / bytes.fromhex('${original:0:$((534 * 2))}')" \
+               >> hv3/third.expected
+       else
+           echo ${packet} >> hv1/multi.expected
+           echo ${packet} >> hv2/multi.expected
+       fi
+
+       ip="IPv6(src='${third_ip6}', dst='${multi_ip6}') /
+           ICMPv6EchoRequest() / Raw(b'y' * $((size - 48)))"
+       ip_error="IPv6(src='${multi_ip6}', dst='${third_ip6}', hlim=255) /
+           ICMPv6PacketTooBig(mtu=${mtu})"
+       packet=$(fmt_pkt "Ether(dst='${multi_mac}', src='${third_mac}') / 
${ip}")
+       as hv3 ovs-appctl netdev-dummy/receive third ${packet}
+       if test "${fail}" = 1; then
+           original=$(fmt_pkt "${ip}")
+           fmt_pkt "Ether(dst='${third_mac}', src='${multi_mac}') /
+               ${ip_error} / bytes.fromhex('${original:0:$((1218 * 2))}')" \
+               >> hv3/third.expected
+       else
+           echo ${packet} >> hv1/multi.expected
+           echo ${packet} >> hv2/multi.expected
+       fi
+   }
+
+   reset_env() {
+       for hv in hv1 hv2; do
+           as $hv reset_pcap_file multi $hv/multi
+       done
+       as hv3 reset_pcap_file third hv3/third
+       for port in hv1/multi hv2/multi hv3/third; do
+           : > $port.expected
+       done
+   }
+
+   # Oversized packets are sent first, followed by packets that fit, so
+   # receiving the latter shows that the former were not delivered.
+   check_pkts() {
+       for port in hv1/multi hv2/multi hv3/third; do
+           OVN_CHECK_PACKETS_REMOVE_BROADCAST([${port}-tx.pcap], 
[${port}.expected])
+       done
+       reset_env
+   }
+
+   # Waits until all packet size checks on 'hv' use the limit for the IP
+   # MTU 'mtu', or until there are no checks if 'mtu' is empty.
+   wait_for_mtu() {
+       local hv=${1} limit=${2:+$((${2} + 18))}
+       OVS_WAIT_UNTIL([test "$(as ${hv} ovs-ofctl dump-flows br-int |
+           sed -n 's/.*check_pkt_larger(\([[0-9]]*\)).*/\1/p' |
+           sort -u)" = "${limit}"])
+   }
+
+   reset_env
+
+   AS_BOX([Without a tunnel MTU, only hosting chassis check packet sizes])
+   wait_for_mtu hv1 $3
+   wait_for_mtu hv2 $3
+   wait_for_mtu hv3
+
+   AS_BOX([Tunnel MTU checks packets from a non-hosting chassis])
+   check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=1500
+   for hv in hv1 hv2 hv3; do
+       wait_for_mtu $hv $3
+   done
+   # Four detection and four ICMP flows, all for the multichassis port.
+   AT_CHECK([as hv3 ovs-ofctl dump-flows br-int 
table=OFTABLE_OUTPUT_LARGE_PKT_DETECT |
+       grep check_pkt_larger | grep -c "reg1[[45]]=${multi_key},"], [0], [4
+])
+   AT_CHECK([as hv3 ovs-ofctl dump-flows br-int 
table=OFTABLE_OUTPUT_LARGE_PKT_PROCESS |
+       grep controller | grep -c "reg1[[45]]=${multi_key},"], [0], [4
+])
+   send_packets 1500 1 $3
+   send_packets 1000 0
+   check_pkts
+
+   AS_BOX([Jumbo tunnel MTU delivers full-size packets to both chassis])
+   check ovn-nbctl --wait=hv set NB_Global . options:tunnel_mtu=9000
+   jumbo_mtu=$(($3 + 7500))
+   for hv in hv1 hv2 hv3; do
+       wait_for_mtu $hv $jumbo_mtu
+   done
+   send_packets 1500 0
+   check_pkts
+
+   AS_BOX([Checks follow the additional chassis])
+   check ovn-nbctl lsp-set-options multi requested-chassis=hv2
+   wait_column "" Port_Binding additional_chassis logical_port=multi
+   for hv in hv1 hv2 hv3; do
+       wait_for_mtu $hv
+   done
+   check ovn-nbctl lsp-set-options multi requested-chassis=hv2,hv1
+   hv1_uuid=$(fetch_column Chassis _uuid name=hv1)
+   wait_column "$hv1_uuid" Port_Binding additional_chassis logical_port=multi
+   for hv in hv1 hv2 hv3; do
+       wait_for_mtu $hv $jumbo_mtu
+       CHECK_FLOWS_AFTER_RECOMPUTE([$hv], [$hv])
+   done
+
+   AS_BOX([Checks use the largest overhead of all tunnels])
+   # hv4 has no port on the switch.  hv3 reaches it with the other
+   # encapsulation, which lowers the limit only if it has more overhead.
+   if test "$2" = geneve; then
+       other=vxlan delta=0
+   else
+       other=geneve delta=8
+   fi
+   as hv3 check ovs-vsctl set Open_vSwitch . \
+       external-ids:ovn-encap-type=geneve,vxlan
+   sim_add hv4
+   as hv4
+   check ovs-vsctl add-br br-phys
+   if test "x$1" = "xipv6"; then
+       ovn_attach n1 br-phys fd00::4 64 $other
+   else
+       ovn_attach n1 br-phys 192.168.0.4 24 $other
+   fi
+   OVN_POPULATE_ARP
+   OVS_WAIT_UNTIL([test -n "$(as hv3 ovs-vsctl --bare --columns=name \
+       find Interface type=$other)"])
+   wait_for_mtu hv3 $(($jumbo_mtu - $delta))
+   CHECK_FLOWS_AFTER_RECOMPUTE([hv3], [hv3])
+   for hv in hv1 hv2; do
+       wait_for_mtu $hv $jumbo_mtu
+   done
+   as hv4 check ovs-vsctl set Open_vSwitch . external-ids:ovn-encap-type=$2
+   OVS_WAIT_UNTIL([test -z "$(as hv3 ovs-vsctl --bare --columns=name \
+       find Interface type=$other)"])
+   wait_for_mtu hv3 $jumbo_mtu
+   CHECK_FLOWS_AFTER_RECOMPUTE([hv3], [hv3])
+
+   AS_BOX([A local tunnel MTU suffices on a non-hosting chassis])
+   check ovn-nbctl --wait=hv remove NB_Global . options tunnel_mtu
+   wait_for_mtu hv3
+   as hv3 check ovs-vsctl set Open_vSwitch . external_ids:ovn-tunnel-mtu=1500
+   wait_for_mtu hv3 $3
+   send_packets 1500 1 $3
+   send_packets 1000 0
+   check_pkts
+
+   AS_BOX([Non-hosting chassis have no VIF MTU to fall back to])
+   as hv3 check ovs-vsctl set Open_vSwitch . external_ids:ovn-tunnel-mtu=1300
+   wait_for_mtu hv3
+   as hv3 check ovs-vsctl set Open_vSwitch . external_ids:ovn-tunnel-mtu=1500
+   wait_for_mtu hv3 $3
+
+   AS_BOX([Tunnel MTU does not apply with always_tunnel])
+   check ovn-nbctl --wait=hv set NB_Global . options:always_tunnel=true
+   for hv in hv1 hv2 hv3; do
+       wait_for_mtu $hv
+   done
+   check ovn-nbctl --wait=hv remove NB_Global . options always_tunnel
+   for hv in hv1 hv2 hv3; do
+       wait_for_mtu $hv $3
+   done
+
+   OVN_CLEANUP([hv1],[hv2],[hv3
+/Tunnel MTU .* is too small; minimum is/d
+],[hv4])
+
+   AT_CLEANUP
+   ])])
+
+MULTICHASSIS_TUNNEL_MTU_REMOTE_SENDER_TEST([ipv4], [geneve], [1424])
+MULTICHASSIS_TUNNEL_MTU_REMOTE_SENDER_TEST([ipv6], [geneve], [1404])
+MULTICHASSIS_TUNNEL_MTU_REMOTE_SENDER_TEST([ipv4], [vxlan], [1432])
+MULTICHASSIS_TUNNEL_MTU_REMOTE_SENDER_TEST([ipv6], [vxlan], [1412])
+
 m4_define([ACTIVATION_STRATEGY_TEST],
   [OVN_FOR_EACH_NORTHD([
     AT_SETUP([options:activation-strategy=$1 for logical port])
-- 
2.48.1

_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev

Reply via email to