On Tue, Oct 6, 2026 at 1:41 PM Mairtin O'Loingsigh <[email protected]>
wrote:

> By default only the chassis currently active for a distributed gateway
> port advertises the routes that track it, so a failover lasts as long
> as the newly active chassis needs to originate the routes and the
> fabric to reconverge on them.
>
> Make every member of a distributed gateway port's HA chassis group
> install and advertise the port's routes, not only the active one.
> ovn-controller derives the membership from the ha_chassis_group on
> the chassisredirect port binding; no explicit opt-in is required.
>
> Standby members advertise with a priority offset of PRIORITY_DEFAULT
> above the base, which keeps every standby priority strictly above
> every active one.  The routing daemon translates that into a higher
> metric (BGP MED), so the fabric holds a precomputed backup path
> towards each standby without ever preferring it over the active
> chassis (BGP PIC Edge).
>
> The VRF maintenance (dynamic-routing-maintain-vrf) is similarly
> extended to all group members so the routing session is established
> before a failover happens rather than while traffic is blackholed.
>
> Assisted-by: Claude Opus 5, Claude Code
> Reported-at: https://redhat.atlassian.net/browse/FDP-3745
> Signed-off-by: Mairtin O'Loingsigh <[email protected]>
> ---
>

Hi Mairtin,

thank you for the v3, there are still some things down below.


> v2:
> - Remove option to enable advertisement of standby routes.
>
> v3:
> - Only check logical_port for locality.
> - Remove unnecessary helper function.
> - Remove unrelated change.
> - Add update to BGP documentation.
> - Remove unneeded test.
>
>  .../topics/dynamic-routing/architecture.rst   |  36 ++
>  NEWS                                          |   9 +
>  controller/route.c                            | 107 ++++-
>  ovn-nb.xml                                    |  26 +
>  ovn-sb.xml                                    |  10 +
>  tests/multinode-bgp-macros.at                 | 136 ++++++
>  tests/multinode-macros.at                     |  81 ++++
>  tests/multinode.at                            | 445 +++++++++++++-----
>  8 files changed, 734 insertions(+), 116 deletions(-)
>
> diff --git a/Documentation/topics/dynamic-routing/architecture.rst
> b/Documentation/topics/dynamic-routing/architecture.rst
> index cf7de2b79..eda4adcbe 100644
> --- a/Documentation/topics/dynamic-routing/architecture.rst
> +++ b/Documentation/topics/dynamic-routing/architecture.rst
> @@ -337,6 +337,36 @@ to ensure host routes are only announced from the
> chassis that owns the
>  workload, providing optimal traffic forwarding and avoiding unnecessary
>  traffic tromboning.
>
> +Advertisement from Standby Gateway Chassis
> +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~
> +
> +When the advertising port of an ``Advertised_Route`` is a distributed
> +gateway port, the route is installed and advertised by **every** chassis
> +in the port's HA chassis group, not only by the one the
> ``chassisredirect``
> +port is currently bound to.  ``ovn-controller`` derives the membership
> from
> +the ``ha_chassis_group`` column of the ``chassisredirect``
> ``Port_Binding``;
> +there is no option to opt in or out.
> +
> +The chassis the ``chassisredirect`` port is bound to is the active one.
> +Every other member is a standby and installs the route one full priority
> +band (1000) above the priority it would otherwise use, which the routing
> +daemon translates into a correspondingly higher metric and, for BGP, a
> +higher MED.  Since the band exceeds the spread of the ordinary priorities,
> +every standby route sorts strictly below every active one from the
> fabric's
> +point of view.
> +
> +The fabric therefore keeps preferring the active chassis while already
> +holding a resolved backup path towards each standby, so a failover is a
> +local repair at the peer (BGP PIC Edge) instead of a reconvergence that
> +waits for the newly active chassis to originate the prefix.  The cost is
> +that the number of paths the fabric holds per advertised prefix grows with
> +the size of the HA chassis group.
> +
> +Note that this is distinct from the ``tracked_port`` priority described
> +above: that one differentiates chassis by where the *tracked* workload or
> +gateway port lives, and applies even when the advertising port itself is
> an
> +ordinary, chassis-local logical router port.
> +
>  IP Route Learning
>  -----------------
>
> @@ -430,6 +460,12 @@ interface on the chassis where the port is bound.
> This includes:
>  - Deleting the VRF interface when dynamic routing is disabled or the
>    port is unbound.
>
> +If the logical router port is a distributed gateway port, the VRF is
> +created on every chassis of its HA chassis group, not only on the one the
> +``chassisredirect`` port is currently bound to.  The routing daemon on a
> +standby chassis can therefore establish its session in the VRF ahead of a
> +failover, rather than while traffic is already blackholed.
> +
>  If ``dynamic-routing-maintain-vrf`` is ``false`` (the default), the VRF
>  is expected to already exist on the chassis, managed by external tooling
>  or configuration management.
> diff --git a/NEWS b/NEWS
> index 7e320d361..1323c8d3e 100644
> --- a/NEWS
> +++ b/NEWS
> @@ -21,6 +21,15 @@ Post v26.09.0
>         now written to the SB MAC_Binding table and consumed at the
>         same priority as dynamic entries, making the preference option
>         obsolete.
> +     * The routes of a distributed gateway port are now installed and
> +       advertised by every member of its HA chassis group instead of only
> +       by the active one.  The standby members use a higher priority band
> +       (and hence a higher metric/BGP MED) so the fabric prefers the
> active
> +       chassis while holding a precomputed backup path (BGP PIC Edge).
> +     * "options:dynamic-routing-maintain-vrf" on a distributed gateway
> port
> +       now creates the VRF on every member of the port's HA chassis group
> +       instead of only on the active one, so the routing daemon can bring
> +       up its session in the VRF before a failover.
>     - Removed OVN's ovs-bugtool plugin and helper scripts.
>     - Removed ovn-sim utility scripts.
>     - Removed ovn-docker utility scripts.
> diff --git a/controller/route.c b/controller/route.c
> index c7df5b6ce..7b0e2827f 100644
> --- a/controller/route.c
> +++ b/controller/route.c
> @@ -42,6 +42,10 @@ VLOG_DEFINE_THIS_MODULE(exchange);
>  #define PRIORITY_DEFAULT 1000
>  #define PRIORITY_LOCAL_BOUND 100
>
> +/* Name of the Logical_Router_Port option asking ovn-controller to create
> and
> + * remove the VRF the routes of the Logical_Router are exchanged in. */
> +#define OPT_MAINTAIN_VRF "dynamic-routing-maintain-vrf"
> +
>  /* Discover the veth peer interface name of 'iface' using the
>   * status:peer_ifindex value that OVS populates for veth devices.
>   *
> @@ -111,14 +115,59 @@ route_is_distributed_lb(const struct
> sbrec_advertised_route *route)
>                           OVN_AR_DISTRIBUTED_LB_ID, false);
>  }
>
> +/* Returns true if 'cr_pb' is a chassisredirect port whose routes are
> + * arbitrated by an HA chassis group.
> + *
> + * Every member of that group installs and advertises the port's routes;
> the
> + * standby members do so in a strictly higher priority band, so the fabric
> + * pre-computes a backup path (BGP PIC Edge) without ever preferring it
> over
> + * the active chassis. */
> +static bool
> +route_cr_port_is_ha(const struct sbrec_port_binding *cr_pb)
> +{
> +    return cr_pb && cr_pb->ha_chassis_group;
> +}
> +
> +/* Returns true if this chassis advertises 'route'.
> + *
> + * '*is_standby', if non-NULL, receives whether this chassis advertises
> the
> + * route as a standby member of the advertising port's HA chassis group.
> The
> + * locality of the advertising port is the only arbiter: the chassis the
> + * chassisredirect port is bound to is the active one, every other member
> of
> + * the group is a standby.  Both are derived from Port_Binding.chassis,
> so all
> + * members agree on the outcome even while their local BFD state differs.
> */
>  static bool
>  route_advertising_port_is_local(
>      const struct sbrec_advertised_route *route,
>      struct ovsdb_idl_index *sbrec_port_binding_by_name,
> -    const struct sbrec_chassis *chassis)
> +    const struct sbrec_chassis *chassis,
> +    bool *is_standby)
>  {
> -    return lport_is_local(sbrec_port_binding_by_name, chassis,
> -                          route->logical_port->logical_port);
> +    if (is_standby) {
> +        *is_standby = false;
> +    }
> +
> +    if (lport_is_local(sbrec_port_binding_by_name, chassis,
> +                       route->logical_port->logical_port)) {
> +        return true;
> +    }
> +
> +    /* The advertising port is resident elsewhere.  It is still advertised
> +     * from here if it is a distributed gateway port this chassis is a
> standby
> +     * member of, so that the fabric pre-computes a backup path (BGP PIC
> +     * Edge). */
> +    const struct sbrec_port_binding *cr_pb =
> +        lport_get_cr_port(sbrec_port_binding_by_name, route->logical_port,
> +                          NULL);
> +    if (!route_cr_port_is_ha(cr_pb) ||
> +        !ha_chassis_group_contains(cr_pb->ha_chassis_group, chassis)) {
> +        return false;
> +    }
> +
> +    if (is_standby) {
> +        *is_standby = true;
> +    }
> +    return true;
>

The update of this function doesn't really fit IMO.
First, the function's name is largely inconsistent with its logic.
Second, it also changes the logic for build_lb_route_gates().
We should do the HA check in the route_run itself.
That will make it actually more visible.

 }
>
>  static bool
> @@ -162,8 +211,8 @@ build_lb_route_gates(struct hmap *gates,
>          if (!route->tracked_port || !route_has_health_checks(route)) {
>              continue;
>          }
> -        if (!route_advertising_port_is_local(
> -                route, sbrec_port_binding_by_name, chassis)) {
> +        if (!route_advertising_port_is_local(route,
> sbrec_port_binding_by_name,
> +                                             chassis, NULL)) {
>              continue;
>          }
>
> @@ -412,6 +461,18 @@ route_exchange_find_port(struct ovsdb_idl_index
> *sbrec_port_binding_by_name,
>              smap_get(&cr_pb->options, "dynamic-routing-port-name");
>      }
>
> +    /* Every member of the HA chassis group processes the port's routes,
> not
> +     * just the resident one.  The active/standby distinction is
> expressed as
> +     * a route priority in route_run(), so the standby's routes are
> advertised
> +     * with a higher metric and the fabric can pre-compute a backup path.
> +     *
> +     * This is also what puts the VRF of a standby member in place before
> a
> +     * failover rather than while traffic is already blackholed. */
> +    if (route_cr_port_is_ha(cr_pb) &&
> +        ha_chassis_group_contains(cr_pb->ha_chassis_group, chassis)) {
> +        return route_exchange_relevant_port(cr_pb) ? cr_pb : NULL;
> +    }
> +
>      if (!lport_pb_is_chassis_resident(chassis, cr_pb)) {
>          return NULL;
>      }
> @@ -676,9 +737,7 @@ route_run(struct route_ctx_in *r_ctx_in,
>              }
>
>              ad->maintain_vrf |=
> -                smap_get_bool(&repb->options,
> -                              "dynamic-routing-maintain-vrf",
> -                              false);
> +                smap_get_bool(&repb->options, OPT_MAINTAIN_VRF, false);
>
>              const char *vrf_name = smap_get(&repb->options,
>                                              "dynamic-routing-vrf-name");
> @@ -757,9 +816,13 @@ route_run(struct route_ctx_in *r_ctx_in,
>              continue;
>          }
>
> +        /* Set when the advertising port is a distributed gateway port
> that
> +         * is resident on another chassis of an HA chassis group this
> chassis
> +         * is a member of. */
> +        bool is_standby;
>          if (!route_advertising_port_is_local(
> -                route, r_ctx_in->sbrec_port_binding_by_name,
> -                r_ctx_in->chassis)) {
> +                route, r_ctx_in->sbrec_port_binding_by_name,
> r_ctx_in->chassis,
> +                &is_standby)) {
>              sset_add(r_ctx_out->tracked_ports_remote,
>                       route->logical_port->logical_port);
>              continue;
> @@ -793,6 +856,30 @@ route_run(struct route_ctx_in *r_ctx_in,
>              }
>          }
>
> +        /* Single site where the standby band is applied, once the base
> +         * priority is final. Routes installed by a standby member of the
> HA
> +         * chassis group land in a strictly higher priority band, which
> the
> +         * routing daemon translates into a higher metric (and hence a
> higher
> +         * BGP MED), so the fabric pre-computes the backup path without
> ever
> +         * preferring it over the active chassis.
> +         *
> +         * PRIORITY_DEFAULT is the smallest band that preserves the
> invariant
> +         * "every standby priority exceeds every active priority": the
> lowest
> +         * standby priority is PRIORITY_LOCAL_BOUND + PRIORITY_DEFAULT,
> which
> +         * is still above the highest active one, PRIORITY_DEFAULT. */
> +        if (is_standby) {
> +            unsigned int base = priority;
> +
> +            priority += PRIORITY_DEFAULT;
> +
> +            static struct vlog_rate_limit rl = VLOG_RATE_LIMIT_INIT(5,
> 20);
> +            VLOG_DBG_RL(&rl,
> +                        "Advertising route %s from standby chassis %s "
> +                        "with priority %u (base priority %u)",
> +                        route->ip_prefix, r_ctx_in->chassis->name,
> priority,
> +                        base);
> +        }
> +
>          if (distributed_lb &&
>              !smap_get_bool(&route->external_ids, "enabled", true)) {
>              continue;
> diff --git a/ovn-nb.xml b/ovn-nb.xml
> index 003bbf414..dd885a5ae 100644
> --- a/ovn-nb.xml
> +++ b/ovn-nb.xml
> @@ -3545,6 +3545,20 @@ or
>          <li><ref column="options" key="dynamic-routing-redistribute"
>                   table="Logical_Router_Port"/> on Logical_Router_Port</li>
>          </ul>
> +
> +        <p>
> +          The routes of a distributed gateway port are installed and
> +          advertised by every member of the port's HA chassis group, not
> only
> +          by the chassis that is currently active for it. The standby
> members
> +          do so in a strictly higher priority band, which the routing
> daemon
> +          translates into a higher metric and hence, for BGP, a higher
> MED.
> +          The fabric consequently prefers the active chassis while already
> +          holding a precomputed backup path towards each standby, which
> allows
> +          for a fast local repair (BGP PIC Edge) when the active chassis
> +          fails. Note that this increases the number of paths the fabric
> has
> +          to hold for each advertised prefix by the size of the HA chassis
> +          group.
> +        </p>
>        </column>
>
>        <column name="options" key="dynamic-routing-redistribute"
> @@ -4916,6 +4930,18 @@ or
>            appended to it.
>          </p>
>
> +        <p>
> +          If this LRP is a distributed gateway port, the VRF is created on
> +          every chassis of its HA chassis group, not only on the one that
> is
> +          currently active for the port.  Every member of the group
> advertises
> +          the port's routes from that vrf, so the routing daemon has its
> +          session established and the fabric has a backup path in place
> before
> +          a failover happens; see <ref column="options"
> key="dynamic-routing"
> +          table="Logical_Router"/>.  A standby chassis removes the vrf
> again
> +          once this option is cleared; the chassis the port is resident on
> +          keeps it, as described below.
> +        </p>
> +
>          <p>
>            If the setting is not set or false the ovn-controller will
> expect
>            this VRF to already exist. Some tooling outside of OVN needs to
> diff --git a/ovn-sb.xml b/ovn-sb.xml
> index 2096fc3e3..a05c043f6 100644
> --- a/ovn-sb.xml
> +++ b/ovn-sb.xml
> @@ -4144,6 +4144,16 @@ tcp.flags = RST;
>            <ref table="Logical_Router_Port" db="OVN_Northbound"/>
>            <code>options:dynamic-routing-maintain-vrf</code> option.
>          </p>
> +
> +        <p>
> +          On a <code>chassisredirect</code> port the VRF is created by
> every
> +          <code>ovn-controller</code> whose chassis is a member of the
> port's
> +          <ref column="ha_chassis_group"/>, so that it is in place before
> the
> +          port fails over. Every one of those chassis also advertises the
> +          port's routes in it; the ones the port is not resident on do so
> in a
> +          strictly higher priority band, so that the fabric holds a backup
> +          path without ever preferring it over the resident chassis.
> +        </p>
>        </column>
>
>        <column name="options" key="dynamic-routing-vrf-name">
> diff --git a/tests/multinode-bgp-macros.at b/tests/multinode-bgp-macros.at
> index dba811f50..452dec4b0 100644
> --- a/tests/multinode-bgp-macros.at
> +++ b/tests/multinode-bgp-macros.at
> @@ -620,4 +620,140 @@ m_config_host_frr_router_l3() {
>      m_setup_host_frr_vrf $node $vni $vxlan_ip $bgp_mac $bgp_ip
>  }
>
> +# m_setup_bgp_unnumbered_gws
> +#
> +# Sets up the pair of BGP unnumbered gateways used by the multinode BGP
> +# tests: on ovn-gw-1 and ovn-gw-2 an external FRR router (the simulated
> +# fabric peer) and an OVN FRR router, peering over an unnumbered (IPv6
> +# link-local) session in ovnvrf10 and ovnvrf20 respectively.
> +#
> +# Callers should wait for the sessions to come up themselves, e.g.
> +#   OVS_WAIT_UNTIL([m_as ovn-gw-1 vtysh -c 'show bgp vrf ovnvrf10
> neighbors' \
> +#                   | grep -qE 'Connections established 1'])
> +# so that a failure is reported against the test's own line number.
> +m_setup_bgp_unnumbered_gws() {
> +    m_setup_external_frr_router  ovn-gw-1
> 41.41.41.41/32
> +    m_config_external_frr_router ovn-gw-1 4200000100 41.41.41.41
> 41.41.41.41/32 41::41/64 12:fb:d6:66:99:0c
> +    m_setup_ovn_frr_router       ovn-gw-1
>                  12:fb:d6:66:99:1c 10
> +    m_config_ovn_frr_router      ovn-gw-1 4210000000 14.14.14.14
>                                   10
> +
> +    m_setup_external_frr_router  ovn-gw-2
> 42.42.42.42/32
> +    m_config_external_frr_router ovn-gw-2 4200000200 42.42.42.42
> 42.42.42.42/32 42::42/64 22:fb:d6:66:99:0c
> +    m_setup_ovn_frr_router       ovn-gw-2
>                  22:fb:d6:66:99:2c 20
> +    m_config_ovn_frr_router      ovn-gw-2 4210000000 24.24.24.24
>                                   20
> +}
> +
> +# m_add_guest_vm_and_connections NODE VRF_ID DEFAULT_ROUTE_MAC
> DEFAULT_ROUTE
> +#                                DEFAULT_ROUTE_GW GUEST_GW_IP GUEST_IP
> +#
> +# Connects the OVN FRR router of NODE to the join switch, hangs a guest
> +# logical switch off the shared guest router and creates a fake VM on it.
> +# Adds the default route on the gateway router towards the external FRR
> +# speaker (learned from the unnumbered session in ovnvrf-VRF_ID) and the
> +# route back out of the guest router via the DGP.
> +#
> +# The gateway router port is tagged "dynamic-routing-redistribute=nat", so
> +# NATs on the guest router are advertised from every node.
> +#
> +# Expects these shell variables to be set by m_setup_bgp_guest_topology:
> +#   join_ls, lr_guest, lrp_guest_join, guest_vm_ns
> +m_add_guest_vm_and_connections() {
> +    local node=$1 vrf_id=$2 default_route_mac=$3 default_route=$4
> +    local default_route_gw=$5 guest_gw_ip=$6 guest_ip=$7
> +
> +    local gw_router=$(m_ovn_frr_router_name $node)
> +    local gw_router_lrp=$(m_ovn_frr_router_port_name $node)
> +
> +    local gw_lr=$(m_ovn_frr_router_name $node)
> +    local lrp_to_join=lrp-$node-to-join
> +    local lsp_join_to_lrp=join-to-lrp-$node
> +
> +    local ls_g=ls-guest-$node
> +    local lsp_g_lrg=lsp-guest-$node-lr-guest
> +    local lsp_g_iface=lsp-guest-$node-guest-vm
> +    local lrp_g_lsg=lrp-guest-ls-guest-$node
> +
> +    local guest_gw_cidr="$guest_gw_ip/24"
> +    local guest_cidr="$guest_ip/24"
> +
> +    # Set up connections to connect the new vm.
> +    check multinode_nbctl lrp-add $gw_lr $lrp_to_join $default_route_mac
> +    check multinode_nbctl lrp-set-options $lrp_to_join \
> +        dynamic-routing-redistribute=nat
> +    check multinode_nbctl lsp-add-router-port $join_ls $lsp_join_to_lrp
> $lrp_to_join
> +
> +    check multinode_nbctl ls-add $ls_g
> +    check multinode_nbctl lrp-add $lr_guest $lrp_g_lsg \
> +        00:16:03:01:03:03 $guest_gw_cidr
> +    check multinode_nbctl lsp-add-router-port $ls_g $lsp_g_lrg $lrp_g_lsg
> +    check multinode_nbctl lsp-add $ls_g $lsp_g_iface
> +    check multinode_nbctl lsp-set-addresses $lsp_g_iface \
> +        '00:16:01:00:02:02 '$guest_cidr''
> +
> +    # Create the new vm.
> +    m_as $node /data/create_fake_vm.sh $lsp_g_iface $guest_vm_ns \
> +        00:16:01:00:02:02 1342 $guest_ip 24 $guest_gw_ip 1000::13/64
> 1000::a
> +    local neighbor_lla=$(m_as $node vtysh -c "show bgp vrf
> ovnvrf${vrf_id} neighbor ext0-bgp" | grep "^Foreign host:" | awk '{print
> $3}' | tr -d ',')
> +
> +    check multinode_nbctl lr-route-add $gw_router "0.0.0.0/0" \
> +        $neighbor_lla $gw_router_lrp
> +    check multinode_nbctl lr-route-add $lr_guest \
> +        $default_route $default_route_gw $lrp_guest_join
> +}
> +
> +# m_setup_bgp_guest_topology GW1_PRIO GW2_PRIO
> +#
> +# Builds the guest topology shared by the multinode BGP unnumbered tests
> on
> +# top of the gateways created by m_setup_bgp_unnumbered_gws:
> +#
> +#                guest-1          guest-2
> +#                       \        /
> +#                        lr-guest
> +#                          DGP (priority: gw-1=GW1_PRIO, gw-2=GW2_PRIO)
> +#                           |
> +#                        ls-join
> +#                       /       \
> +# tor <-> lr-ovn-gw-1-ext0*    lr-ovn-gw-2-ext0 <-> tor
> +#               |                     |
> +#         ls-ovn-gw-1-ext0     ls-ovn-gw-2-ext0
> +#
> +# The DGP is hosted by an HA chassis group holding both gateways.  Equal
> +# priorities leave the choice of active chassis to OVN; unequal ones pin
> it,
> +# which is what a test asserting per-chassis route metrics needs.
> +#
> +# A dnat_and_snat for 172.16.10.2 on lr-guest is redistributed onto both
> +# gateway routers, so both nodes advertise it into the fabric.
> +#
> +# Exports the topology names for the caller: join_ls, lsp_join_guest,
> +# lr_guest, lrp_guest_join, guest_vm_iface, guest_vm_ns.
> +m_setup_bgp_guest_topology() {
> +    local gw1_prio=$1 gw2_prio=$2
> +
> +    join_ls="ls-join"
> +    lsp_join_guest="lsp-join-guest"
> +
> +    lr_guest="lr-guest"
> +    lrp_guest_join="lrp-guest-join-dgp"
> +
> +    guest_vm_iface="guest-vm"
> +    guest_vm_ns="ns-guest"
> +
> +    check multinode_nbctl ls-add $join_ls
> +
> +    check multinode_nbctl lr-add $lr_guest
> +    check multinode_nbctl lrp-add $lr_guest $lrp_guest_join
> 00:16:06:12:f0:0d
> +    check multinode_nbctl lsp-add-router-port $join_ls $lsp_join_guest
> $lrp_guest_join
> +    check multinode_nbctl lrp-set-gateway-chassis $lrp_guest_join
> ovn-gw-1 $gw1_prio
> +    check multinode_nbctl lrp-set-gateway-chassis $lrp_guest_join
> ovn-gw-2 $gw2_prio
> +
> +    m_add_guest_vm_and_connections ovn-gw-1 10 00:00:ff:00:00:01
> 41.0.0.0/8 \
> +        fe80::200:ffff:fe00:1 192.168.10.1 192.168.10.10
> +    m_add_guest_vm_and_connections ovn-gw-2 20 00:00:ff:00:00:02
> 42.0.0.0/8 \
> +        fe80::200:ffff:fe00:2 192.168.20.1 192.168.20.10
> +
> +    # NAT that gets advertised via BGP from both gateways.
> +    check multinode_nbctl --gateway-port $lrp_guest_join --add-route \
> +        lr-nat-add $lr_guest dnat_and_snat 172.16.10.2 192.168.10.10
> +}
> +
>  OVS_END_SHELL_HELPERS
> diff --git a/tests/multinode-macros.at b/tests/multinode-macros.at
> index 74f2d7894..be177ed43 100644
> --- a/tests/multinode-macros.at
> +++ b/tests/multinode-macros.at
> @@ -463,6 +463,87 @@ m_check_column() {
>      fi
>  }
>
> +# m_route_metric NODE TABLE PREFIX
> +#
> +# Prints the metric of every route for PREFIX in routing table TABLE on
> NODE.
> +# The prefix is not always the first field: a route OVN redistributes
> without
> +# a locally resolvable nexthop, which is what a NAT or a static route
> with an
> +# off-link nexthop installs, renders as
> +# "blackhole PREFIX proto ovn metric N".
> +m_route_metric() {
> +    m_as $1 ip route show table $2 | awk -v p="$3" '
> +        { match_found = 0
> +          for (i = 1; i <= NF; i++) if ($i == p) match_found = 1
> +          if (match_found) {
> +              for (i = 1; i <= NF; i++) if ($i == "metric") print $(i + 1)
> +          } }'
> +}
> +
> +# m_wait_route_metric NODE TABLE PREFIX EXPECTED
> +#
> +# Waits until the metric of PREFIX in routing table TABLE on NODE is
> EXPECTED.
> +m_wait_route_metric() {
> +    # OVS_WAIT_UNTIL runs its body in a context where the positional
> +    # parameters are its own, so copy them out first.
> +    local rm_node=$1 rm_table=$2 rm_prefix=$3 rm_expected=$4
> +
> +    echo "Waiting for $rm_prefix in table $rm_table on $rm_node to have" \
> +         "metric $rm_expected..."
> +    OVS_WAIT_UNTIL([test "$rm_expected" = \
> +                        "$(m_route_metric $rm_node $rm_table
> $rm_prefix)"], [
> +        echo "Routes in table $rm_table on $rm_node:"
> +        m_as $rm_node ip route show table $rm_table])
> +}
> +
> +# m_route_count NODE TABLE PREFIX
> +#
> +# Prints the number of routes for PREFIX in routing table TABLE on NODE.
> See
> +# m_route_metric() for why the prefix is looked for in every field.
> +m_route_count() {
> +    m_as $1 ip route show table $2 | awk -v p="$3" '
> +        { for (i = 1; i <= NF; i++) if ($i == p) { found++; break } }
> +        END { print found + 0 }'
> +}
> +
> +# m_wait_route_count NODE TABLE PREFIX EXPECTED
> +#
> +# Waits until routing table TABLE on NODE holds EXPECTED routes for
> PREFIX.
> +m_wait_route_count() {
> +    local rc_node=$1 rc_table=$2 rc_prefix=$3 rc_expected=$4
> +
> +    echo "Waiting until table $rc_table on $rc_node has $rc_expected" \
> +         "route(s) for $rc_prefix..."
> +    OVS_WAIT_UNTIL([test "$rc_expected" = \
> +                        "$(m_route_count $rc_node $rc_table
> $rc_prefix)"], [
> +        echo "Routes in table $rc_table on $rc_node:"
> +        m_as $rc_node ip route show table $rc_table])
> +}
> +
> +# m_vrf_exists NODE DEV
> +#
> +# Prints "yes" if the vrf device DEV exists on NODE, "no" otherwise.
> +m_vrf_exists() {
> +    if m_as $1 ip link show dev $2 type vrf > /dev/null 2>&1; then
> +        echo yes
> +    else
> +        echo no
> +    fi
> +}
> +
> +# m_wait_vrf NODE DEV EXPECTED
> +#
> +# Waits until the presence of the vrf device DEV on NODE is EXPECTED,
> which is
> +# either "yes" or "no".
> +m_wait_vrf() {
> +    local wv_node=$1 wv_dev=$2 wv_expected=$3
> +
> +    echo "Waiting until vrf $wv_dev on $wv_node is
> present=$wv_expected..."
> +    OVS_WAIT_UNTIL([test "$wv_expected" = \
> +                        "$(m_vrf_exists $wv_node $wv_dev)"], [
> +        echo "vrf devices on $wv_node:"
> +        m_as $wv_node ip -o link show type vrf])
> +}
> +
>  # m_add_internal_port NODE NETNS OVS_BRIDGE PORT IP [GW]
>  #
>  # Adds an OVS internal PORT to OVS_BRIDGE on NODE and moves the resulting
> diff --git a/tests/multinode.at b/tests/multinode.at
> index bb9e47749..47c6b1316 100644
> --- a/tests/multinode.at
> +++ b/tests/multinode.at
> @@ -2909,132 +2909,365 @@ cleanup_multinode_resources
>
>  CHECK_VRF()
>
> -add_guest_vm_and_connections() {
> -    node=$1
> -    vrf_id=$2
> -    default_route_mac=$3
> -    default_route=$4
> -    default_route_gw=$5
> -    guest_gw_ip=$6
> -    guest_ip=$7
> -    gw_router=$(m_ovn_frr_router_name $node)
> -    gw_router_lrp=$(m_ovn_frr_router_port_name $node)
> -
> -    gw_lr=$(m_ovn_frr_router_name $node)
> -    lrp_to_join=lrp-$node-to-join
> -    lsp_join_to_lrp=join-to-lrp-$node
> -    lrp_guest=lrp-guest-$node
> -
> -    ls_g=ls-guest-$node
> -    lsp_g_lrg=lsp-guest-$node-lr-guest
> -    lsp_g_iface=lsp-guest-$node-guest-vm
> -    lrp_g_lsg=lrp-guest-ls-guest-$node
> -
> -    guest_gw_cidr="$guest_gw_ip/24"
> -    guest_cidr="$guest_ip/24"
> -
> -    # set up connections to connect the new vm
> -    check multinode_nbctl lrp-add $gw_lr $lrp_to_join $default_route_mac
> -    check multinode_nbctl lrp-set-options $lrp_to_join \
> -        dynamic-routing-redistribute=nat
> -    check multinode_nbctl lsp-add-router-port $join_ls $lsp_join_to_lrp
> $lrp_to_join
> -
> -    check multinode_nbctl ls-add $ls_g
> -    check multinode_nbctl lrp-add $lr_guest $lrp_g_lsg \
> -        00:16:03:01:03:03 $guest_gw_cidr
> -    check multinode_nbctl lsp-add-router-port $ls_g $lsp_g_lrg $lrp_g_lsg
> -    check multinode_nbctl lsp-add $ls_g $lsp_g_iface
> -    check multinode_nbctl lsp-set-addresses $lsp_g_iface \
> -        '00:16:01:00:02:02 '$guest_cidr''
> -
> -    # create the new vm
> -    m_as $node /data/create_fake_vm.sh $lsp_g_iface $guest_vm_ns \
> -        00:16:01:00:02:02 1342 $guest_ip 24 $guest_gw_ip 1000::13/64
> 1000::a
> -    neighbor_lla=$(m_as $node vtysh -c "show bgp vrf ovnvrf${vrf_id}
> neighbor ext0-bgp" | grep "^Foreign host:" | awk '{print $3}' | tr -d ',')
> -
> -    check multinode_nbctl lr-route-add $gw_router "0.0.0.0/0" \
> -        $neighbor_lla $gw_router_lrp
> -    check multinode_nbctl lr-route-add $lr_guest \
> -        $default_route $default_route_gw $lrp_guest_join
> -}
> -
> -m_setup_external_frr_router  ovn-gw-1
> 41.41.41.41/32
> -m_config_external_frr_router
> <http://41.41.41.41/32-m_config_external_frr_router> ovn-gw-1 4200000100
> 41.41.41.41 41.41.41.41/32 41::41/64 12:fb:d6:66:99:0c
> -m_setup_ovn_frr_router       ovn-gw-1
>              12:fb:d6:66:99:1c 10
> -m_config_ovn_frr_router      ovn-gw-1 4210000000 14.14.14.14
>                               10
> -
> -m_setup_external_frr_router  ovn-gw-2
> 42.42.42.42/32
> -m_config_external_frr_router
> <http://42.42.42.42/32-m_config_external_frr_router> ovn-gw-2 4200000200
> 42.42.42.42 42.42.42.42/32 42::42/64 22:fb:d6:66:99:0c
> -m_setup_ovn_frr_router       ovn-gw-2
>              22:fb:d6:66:99:2c 20
> -m_config_ovn_frr_router      ovn-gw-2 4210000000 24.24.24.24
>                               20
> +m_setup_bgp_unnumbered_gws
>
>  OVS_WAIT_UNTIL([m_as ovn-gw-2 vtysh -c 'show bgp vrf ovnvrf20 neighbors'
> | grep -qE 'Connections established 1'])
>  OVS_WAIT_UNTIL([m_as ovn-gw-1 vtysh -c 'show bgp vrf ovnvrf10 neighbors'
> | grep -qE 'Connections established 1'])
>
> -# Tor <-> ovn-gw via bgp
> -# lr-guest with distributed gateway port
> -# bgp on lr-ovn-gw-2-ext0
> +# Tor <-> ovn-gw via bgp, lr-guest with a distributed gateway port.  See
> +# m_setup_bgp_guest_topology for the topology diagram.  Both gateways get
> +# the same HA chassis group priority here.
> +m_setup_bgp_guest_topology 20 20
> +
> +OVS_WAIT_UNTIL([m_central_as ovn-sbctl list Advertised_Route | grep -q
> 172.16.10.2])
> +OVS_WAIT_UNTIL([m_as ovn-gw-1 ip netns exec frr-ns ip route | grep -q
> 'ext1'])
> +OVS_WAIT_UNTIL([m_as ovn-gw-1 ip netns exec frr-ns ping -W 1 -c 1
> 172.16.10.2])
> +OVS_WAIT_UNTIL([m_as ovn-gw-2 ip netns exec frr-ns ip route | grep -q
> 'ext1'])
> +OVS_WAIT_UNTIL([m_as ovn-gw-2 ip netns exec frr-ns ping -W 1 -c 1
> 172.16.10.2])
> +
> +AT_CLEANUP
> +
> +AT_SETUP([ovn multinode bgp ha-chassis standby routes])
> +
> +# Check that ovn-fake-multinode setup is up and running.
> +check_fake_multinode_setup
> +
> +CHECK_VRF()
> +
> +# Delete the multinode NB and OVS resources before starting the test.
> +cleanup_multinode_resources
> +
> +CHECK_VRF()
> +
> +# Verifies that the route priority of a NAT tracked to a distributed
> gateway
> +# port reaches the fabric as a BGP MED, and that it follows the DGP
> across a
> +# failover: the chassis the DGP is resident on advertises the lower
> metric,
> +# the other one keeps advertising a backup path at a higher metric so the
> +# fabric can pre-compute it (BGP PIC Edge).
>  #
> +# Topology:
>  #                guest-1          guest-2
>  #                       \        /
>  #                        lr-guest
> -#                          DGP
> +#                          DGP (priority: gw-1=30, gw-2=10)
>  #                           |
>  #                        ls-join
>  #                       /       \
> -# tor <-> lr-ovn-gw-2-ext0*    lr-ovn-gw-1-ext0* <-> tor
> -#               |                     |
> -#         ls-ovn-gw-2-ext0     ls-ovn-gw-1-ext0
> +# tor <-> lr-ovn-gw-1-ext0*    lr-ovn-gw-2-ext0 <-> tor
> +#     (AS 4200000100)   |        | (AS 4200000200)
> +#    (ACTIVE pri=30)    |        | (STANDBY pri=10)
> +#                 ls-ovn-gw-1  ls-ovn-gw-2
> +#
> +# The NAT on lr-guest is redistributed onto both lrp-ovn-gw-N-to-join
> ports,
> +# so each chassis advertises the prefix out of its own gateway router.
> Those
> +# LRPs are ordinary patch ports, local to their chassis; what differs
> between
> +# the two is the locality of the route's tracked_port, the DGP.
>  #
> +# Expected kernel route metrics for the NAT IP:
>  #
> +#   DGP resident here       PRIORITY_LOCAL_BOUND =  100
> +#   DGP resident elsewhere  PRIORITY_DEFAULT     = 1000
>  #
>
> -join_ls="ls-join"
> -lsp_join_guest="lsp-join-guest"
> -
> -lr_guest="lr-guest"
> -lrp_guest_join="lrp-guest-join-dgp"
> -
> -guest_vm_iface="guest-vm"
> -guest_vm_ns="ns-guest"
> -
> -check multinode_nbctl ls-add $join_ls
> -
> -check multinode_nbctl lr-add $lr_guest
> -check multinode_nbctl lrp-add $lr_guest $lrp_guest_join 00:16:06:12:f0:0d
> -check multinode_nbctl lsp-add-router-port $join_ls $lsp_join_guest
> $lrp_guest_join
> -check multinode_nbctl lrp-set-gateway-chassis $lrp_guest_join ovn-gw-1 20
> -check multinode_nbctl lrp-set-gateway-chassis $lrp_guest_join ovn-gw-2 20
> -
> -vrf_id=10
> -default_route_mac=00:00:ff:00:00:01
> -default_route=41.0.0.0/8
> -default_route_gw=fe80::200:ffff:fe00:1
> -guest_gw_ip=192.168.10.1
> -guest_ip=192.168.10.10
> -add_guest_vm_and_connections ovn-gw-1 $vrf_id           \
> -    $default_route_mac $default_route $default_route_gw \
> -    $guest_gw_ip $guest_ip
> -
> -vrf_id=20
> -default_route_mac=00:00:ff:00:00:02
> -default_route=42.0.0.0/8
> -default_route_gw=fe80::200:ffff:fe00:2
> -guest_gw_ip=192.168.20.1
> -guest_ip=192.168.20.10
> -add_guest_vm_and_connections ovn-gw-2 $vrf_id           \
> -    $default_route_mac $default_route $default_route_gw \
> -    $guest_gw_ip $guest_ip
> -
> -check multinode_nbctl --gateway-port $lrp_guest_join --add-route
> lr-nat-add \
> -    $lr_guest dnat_and_snat 172.16.10.2 192.168.10.10
> +# Setup external FRR routers and OVN FRR routers with BGP unnumbered.
> +m_setup_bgp_unnumbered_gws
>
> +echo "Waiting for BGP sessions to establish..."
> +OVS_WAIT_UNTIL([m_as ovn-gw-2 vtysh -c 'show bgp vrf ovnvrf20 neighbors'
> | grep -qE 'Connections established 1'])
> +echo "BGP session established on gw-2"
> +OVS_WAIT_UNTIL([m_as ovn-gw-1 vtysh -c 'show bgp vrf ovnvrf10 neighbors'
> | grep -qE 'Connections established 1'])
> +echo "BGP session established on gw-1"
> +
> +# Create topology with unequal priorities: ovn-gw-1 = 30 (active),
> ovn-gw-2 = 10 (standby).
> +m_setup_bgp_guest_topology 30 10
> +
> +# Wait for routes to be advertised.
>  OVS_WAIT_UNTIL([m_central_as ovn-sbctl list Advertised_Route | grep -q
> 172.16.10.2])
> -OVS_WAIT_UNTIL([m_as ovn-gw-1 ip netns exec frr-ns ip route | grep -q
> 'ext1'])
> +
> +# Get chassis UUIDs and tunnel interface names for BFD testing
> +gw1_chassis=$(m_fetch_column Chassis _uuid name=ovn-gw-1)
> +gw2_chassis=$(m_fetch_column Chassis _uuid name=ovn-gw-2)
> +
> +ip_gw1=$(m_as ovn-gw-1 ip a show dev eth1 | grep "inet " | awk '{print
> $2}'| cut -d '/' -f1)
> +ip_gw2=$(m_as ovn-gw-2 ip a show dev eth1 | grep "inet " | awk '{print
> $2}'| cut -d '/' -f1)
> +
> +tunnel_gw1_to_gw2=$(m_as ovn-gw-1 ovs-vsctl --bare --columns=name find
> interface \
> +    options:remote_ip=$ip_gw2 type=geneve)
> +tunnel_gw2_to_gw1=$(m_as ovn-gw-2 ovs-vsctl --bare --columns=name find
> interface \
> +    options:remote_ip=$ip_gw1 type=geneve)
> +
> +# Wait for BFD to come up between gateways
> +OVS_WAIT_UNTIL([
> +    state=$(m_as ovn-gw-1 ovs-vsctl get interface $tunnel_gw1_to_gw2
> bfd_status:state 2>/dev/null || echo "down")
> +    echo "BFD gw-1 -> gw-2: $state"
> +    test "$state" = "up"
> +])
> +
> +OVS_WAIT_UNTIL([
> +    state=$(m_as ovn-gw-2 ovs-vsctl get interface $tunnel_gw2_to_gw1
> bfd_status:state 2>/dev/null || echo "down")
> +    echo "BFD gw-2 -> gw-1: $state"
> +    test "$state" = "up"
> +])
> +
> +# Verify CR port is bound to ovn-gw-1 (highest priority)
> +m_wait_row_count Port_Binding 1 logical_port=cr-$lrp_guest_join
> chassis=$gw1_chassis
> +
> +# bgp_med NODE PREFIX.
> +bgp_med() {
> +    m_as $1 vtysh $(m_frr_ns_flags frr-ns) \
> +        -c "show bgp ipv4 unicast $2" |
> +        sed -n 's/.*metric \([[0-9]][[0-9]]*\).*/\1/p' | head -1
> +}
> +
> +AS_BOX([Verifying the metrics are differentiated without any option set.])
> +
> +# Both chassis should have routes learned via BGP in frr-ns namespace
> +OVS_WAIT_UNTIL([m_as ovn-gw-1 ip netns exec frr-ns ip route | grep -q
> '172.16.10.2'])
> +OVS_WAIT_UNTIL([m_as ovn-gw-2 ip netns exec frr-ns ip route | grep -q
> '172.16.10.2'])
> +
> +# Both chassis can ping the NAT IP
>  OVS_WAIT_UNTIL([m_as ovn-gw-1 ip netns exec frr-ns ping -W 1 -c 1
> 172.16.10.2])
> -OVS_WAIT_UNTIL([m_as ovn-gw-2 ip netns exec frr-ns ip route | grep -q
> 'ext1'])
>  OVS_WAIT_UNTIL([m_as ovn-gw-2 ip netns exec frr-ns ping -W 1 -c 1
> 172.16.10.2])
>
> +# Both chassis advertise exactly one route each for the prefix.
> +m_wait_row_count Advertised_Route 2 ip_prefix=172.16.10.2
> +
> +# gw-1 hosts the tracked DGP, so it uses PRIORITY_LOCAL_BOUND.  gw-2 does
> +# not, so its copy of the route stays at PRIORITY_DEFAULT.
> +m_wait_route_metric ovn-gw-1 10 172.16.10.2 100
> +m_wait_route_metric ovn-gw-2 20 172.16.10.2 1000
> +
> +# The kernel metric must reach the fabric as a BGP MED, otherwise the
> backup
> +# path would be indistinguishable from the primary one to the external
> peer.
> +echo "Verifying BGP MED seen by the external speakers"
> +OVS_WAIT_UNTIL([test "100" = "$(bgp_med ovn-gw-1 172.16.10.2/32)"], [
> +    m_as ovn-gw-1 vtysh $(m_frr_ns_flags frr-ns) \
> +        -c "show bgp ipv4 unicast 172.16.10.2/32"])
> +OVS_WAIT_UNTIL([test "1000" = "$(bgp_med ovn-gw-2 172.16.10.2/32)"], [
> +    m_as ovn-gw-2 vtysh $(m_frr_ns_flags frr-ns) \
> +        -c "show bgp ipv4 unicast 172.16.10.2/32"])
> +
> +AS_BOX([Fail the DGP over to gw-2 and verify the metrics swap.])
> +check multinode_nbctl lrp-set-gateway-chassis $lrp_guest_join ovn-gw-2 40
> +
> +m_wait_row_count Port_Binding 1 logical_port=cr-$lrp_guest_join
> chassis=$gw2_chassis
> +
> +# Verify connectivity is maintained.
> +OVS_WAIT_UNTIL([m_as ovn-gw-2 ip netns exec frr-ns ping -W 1 -c 1
> 172.16.10.2])
> +
> +# Roles are now reversed: gw-2 hosts the tracked DGP, gw-1 falls back to
> +# advertising the backup path.
> +m_wait_route_metric ovn-gw-2 20 172.16.10.2 100
> +m_wait_route_metric ovn-gw-1 10 172.16.10.2 1000
> +
> +m_wait_row_count Advertised_Route 2 ip_prefix=172.16.10.2
> +
> +OVS_WAIT_UNTIL([test "100" = "$(bgp_med ovn-gw-2 172.16.10.2/32)"], [
> +    m_as ovn-gw-2 vtysh $(m_frr_ns_flags frr-ns) \
> +        -c "show bgp ipv4 unicast 172.16.10.2/32"])
> +OVS_WAIT_UNTIL([test "1000" = "$(bgp_med ovn-gw-1 172.16.10.2/32)"], [
> +    m_as ovn-gw-1 vtysh $(m_frr_ns_flags frr-ns) \
> +        -c "show bgp ipv4 unicast 172.16.10.2/32"])
> +
> +AS_BOX([Fail back to gw-1.])
> +
> +check multinode_nbctl lrp-set-gateway-chassis $lrp_guest_join ovn-gw-2 10
> +
> +m_wait_row_count Port_Binding 1 logical_port=cr-$lrp_guest_join
> chassis=$gw1_chassis
> +echo "CR port migrated back to ovn-gw-1"
> +
> +OVS_WAIT_UNTIL([m_as ovn-gw-1 ip netns exec frr-ns ping -W 1 -c 1
> 172.16.10.2])
> +OVS_WAIT_UNTIL([m_as ovn-gw-2 ip netns exec frr-ns ping -W 1 -c 1
> 172.16.10.2])
> +
> +m_wait_route_metric ovn-gw-1 10 172.16.10.2 100
> +m_wait_route_metric ovn-gw-2 20 172.16.10.2 1000
> +
> +m_wait_row_count Advertised_Route 2 ip_prefix=172.16.10.2
> +
> +OVS_WAIT_UNTIL([test "100" = "$(bgp_med ovn-gw-1 172.16.10.2/32)"], [
> +    m_as ovn-gw-1 vtysh $(m_frr_ns_flags frr-ns) \
> +        -c "show bgp ipv4 unicast 172.16.10.2/32"])
> +OVS_WAIT_UNTIL([test "1000" = "$(bgp_med ovn-gw-2 172.16.10.2/32)"], [
> +    m_as ovn-gw-2 vtysh $(m_frr_ns_flags frr-ns) \
> +        -c "show bgp ipv4 unicast 172.16.10.2/32"])
> +
> +AT_CLEANUP
> +
> +AT_SETUP([ovn multinode dynamic-routing - VRF maintained on standby
> gateway chassis])
> +
> +# Check that ovn-fake-multinode setup is up and running.
> +check_fake_multinode_setup
> +
> +CHECK_VRF()
> +
> +# Delete the multinode NB and OVS resources before starting the test.
> +cleanup_multinode_resources
> +
> +CHECK_VRF()
> +
> +# Verifies "options:dynamic-routing-maintain-vrf" on a distributed gateway
> +# port: the VRF is created by every member of the DGP's HA chassis group,
> not
> +# only by the chassis the port is currently resident on, so the routing
> daemon
> +# can bring its session up in the VRF before a failover happens instead of
> +# while traffic is already blackholed.
> +#
> +#                 ls-vrf-guest (192.168.30.0/24)
> +#                       |
> +#                  lrp-vrf-guest
> +#                     lr-vrf  (dynamic-routing, vrf-id 1030,
> +#                              redistribute static)
> +#                  lrp-vrf-public  (DGP, HA chassis group:
> +#                       |           ovn-gw-1 prio 30 -> active
> +#                  ls-vrf-public    ovn-gw-2 prio 10 -> standby)
> +#
> +# No BGP speaker is involved: what is under test is purely which chassis
> +# ovn-controller creates the VRF netdev on, and which one syncs routes
> into
> +# the corresponding routing table.  Both members of the HA chassis group
> +# advertise the routes unconditionally; the standby one does so a full
> +# priority band above the resident one.
> +
> +vrf_id=1030
> +vrf_dev=ovnvrf$vrf_id
> +prefix=10.30.0.0/24
> +
> +# Clearing maintain-vrf makes ovn-controller disown the device rather than
> +# delete it, on both nodes, so clean them up unconditionally.
> +for gw in ovn-gw-1 ovn-gw-2; do
> +    on_exit "m_as $gw ip link del $vrf_dev > /dev/null 2>&1"
> +done
> +
> +# The router doing the route exchange, in routing table $vrf_id.
> +check multinode_nbctl lr-add lr-vrf
> +check multinode_nbctl set Logical_Router lr-vrf \
> +    options:dynamic-routing=true                \
> +    options:dynamic-routing-vrf-id=$vrf_id      \
> +    options:dynamic-routing-redistribute=static
> +
> +# The DGP the routes are exchanged from. ovn-gw-1 has the higher gateway
> +# chassis priority, so it is the active member and ovn-gw-2 the standby
> one.
> +# "dynamic-routing-maintain-vrf" is intentionally left unset for now.
> +check multinode_nbctl lrp-add lr-vrf lrp-vrf-public 00:00:00:00:30:01 \
> +    20.30.0.1/24
> +check multinode_nbctl lrp-set-gateway-chassis lrp-vrf-public ovn-gw-1 30
> +check multinode_nbctl lrp-set-gateway-chassis lrp-vrf-public ovn-gw-2 10
> +
> +check multinode_nbctl ls-add ls-vrf-public
> +check multinode_nbctl lsp-add-router-port ls-vrf-public public-lr-vrf \
> +    lrp-vrf-public
> +
> +# A tenant subnet behind the router, so the DGP is not the only LRP.
> +check multinode_nbctl lrp-add lr-vrf lrp-vrf-guest 00:00:00:00:30:02 \
> +    192.168.30.1/24
> +check multinode_nbctl ls-add ls-vrf-guest
> +check multinode_nbctl lsp-add-router-port ls-vrf-guest guest-lr-vrf \
> +    lrp-vrf-guest
> +
> +# The static route that gets redistributed out of the DGP.
> Advertised_Route
> +# carries no nexthop column, so ovn-controller installs it as a blackhole
> +# route ("blackhole 10.30.0.0/24 proto ovn metric N"); all this test
> cares
> +# about is which chassis installs it, in which table.
> +check multinode_nbctl --wait=hv lr-route-add lr-vrf $prefix 20.30.0.25 \
> +    lrp-vrf-public
> +
> +gw1_chassis=$(m_fetch_column Chassis _uuid name=ovn-gw-1)
> +gw2_chassis=$(m_fetch_column Chassis _uuid name=ovn-gw-2)
> +
> +# The chassisredirect port is the record ovn-controller consults, both
> for the
> +# active/standby decision and for the VRF options.
> +m_wait_row_count Port_Binding 1 logical_port=cr-lrp-vrf-public \
> +    chassis=$gw1_chassis
> +m_wait_row_count Port_Binding 1 logical_port=cr-lrp-vrf-public \
> +    'options:dynamic-routing=true'
> +
> +AS_BOX([Without maintain-vrf no chassis creates the VRF.])
> +
> +# A routing table exists independently of a VRF netdev, so both chassis
> +# already sync the route into table $vrf_id; ovn-controller just expects
> the
> +# VRF device itself to be provided externally.
> +m_wait_route_count ovn-gw-1 $vrf_id $prefix 1
> +m_wait_route_count ovn-gw-2 $vrf_id $prefix 1
> +
> +m_wait_vrf ovn-gw-1 $vrf_dev no
> +m_wait_vrf ovn-gw-2 $vrf_dev no
> +
> +# The standby lands one full band (PRIORITY_DEFAULT) above the resident
> +# chassis, so the fabric holds the backup path without ever preferring it.
> +# A static route carries no tracked_port, so both chassis start from the
> same
> +# base priority and the delta is exactly one band; read the resident
> metric
> +# rather than hardcoding the band's absolute value.
> +active_metric=$(m_route_metric ovn-gw-1 $vrf_id $prefix)
> +AT_CHECK([test -n "$active_metric"])
> +m_wait_route_metric ovn-gw-2 $vrf_id $prefix $(($active_metric + 1000))
> +
> +AS_BOX([Enable maintain-vrf: both HA chassis group members create the
> VRF.])
> +
> +check multinode_nbctl lrp-set-options lrp-vrf-public \
> +    dynamic-routing-maintain-vrf=true
> +
> +# northd must propagate the option onto the chassisredirect port binding.
> +m_wait_row_count Port_Binding 1 logical_port=cr-lrp-vrf-public \
> +    'options:dynamic-routing-maintain-vrf=true'
> +
> +# This is the point of the feature: the standby holds the VRF too, so the
> +# routing daemon can bring its session up in it ahead of a failover.
> +m_wait_vrf ovn-gw-1 $vrf_dev yes
> +m_wait_vrf ovn-gw-2 $vrf_dev yes
> +
> +m_wait_route_count ovn-gw-1 $vrf_id $prefix 1
> +m_wait_route_count ovn-gw-2 $vrf_id $prefix 1
> +
> +AS_BOX([Fail the DGP over to ovn-gw-2.])
> +
> +check multinode_nbctl lrp-set-gateway-chassis lrp-vrf-public ovn-gw-2 40
> +
> +m_wait_row_count Port_Binding 1 logical_port=cr-lrp-vrf-public \
> +    chassis=$gw2_chassis
> +
> +# Both VRFs survive the failover: ovn-gw-2 already had the one it now
> +# exchanges routes in, and ovn-gw-1 keeps its as the new standby.  This is
> +# what the feature buys: no netdev is created or torn down on the data
> path
> +# of a failover.
> +m_wait_vrf ovn-gw-1 $vrf_dev yes
> +m_wait_vrf ovn-gw-2 $vrf_dev yes
> +
> +# Both chassis keep a route; only the band they sit in swaps over.
> +m_wait_route_count ovn-gw-2 $vrf_id $prefix 1
> +m_wait_route_count ovn-gw-1 $vrf_id $prefix 1
> +
> +m_wait_route_metric ovn-gw-2 $vrf_id $prefix $active_metric
> +m_wait_route_metric ovn-gw-1 $vrf_id $prefix $(($active_metric + 1000))
> +
> +AS_BOX([Fail back to ovn-gw-1.])
> +
> +check multinode_nbctl lrp-set-gateway-chassis lrp-vrf-public ovn-gw-2 10
> +
> +m_wait_row_count Port_Binding 1 logical_port=cr-lrp-vrf-public \
> +    chassis=$gw1_chassis
> +
> +m_wait_vrf ovn-gw-1 $vrf_dev yes
> +m_wait_vrf ovn-gw-2 $vrf_dev yes
> +
> +m_wait_route_metric ovn-gw-1 $vrf_id $prefix $active_metric
> +m_wait_route_metric ovn-gw-2 $vrf_id $prefix $(($active_metric + 1000))
> +
> +AS_BOX([Dropping maintain-vrf makes OVN disown the VRF on both chassis.])
> +
> +check multinode_nbctl remove Logical_Router_Port lrp-vrf-public options \
> +    dynamic-routing-maintain-vrf
> +
> +m_wait_row_count Port_Binding 0 logical_port=cr-lrp-vrf-public \
> +    'options:dynamic-routing-maintain-vrf=true'
> +
> +# Clearing the option means "OVN no longer manages this VRF", not "delete
> it
> +# from under the routing daemon", so neither member removes the netdev.
> +m_wait_vrf ovn-gw-1 $vrf_dev yes
> +m_wait_vrf ovn-gw-2 $vrf_dev yes
> +
> +# Route exchange itself is unaffected: the routing table is independent of
> +# the netdev, exactly as in the first phase of this test.
> +m_wait_route_metric ovn-gw-1 $vrf_id $prefix $active_metric
> +m_wait_route_metric ovn-gw-2 $vrf_id $prefix $(($active_metric + 1000))
> +
>  AT_CLEANUP
>
>  AT_SETUP([ovn multinode dynamic-routing - BGP learned routes with router
> filter name and multiple DGPs])
> --
> 2.55.0
>
>
I think we have an issue when the HA group changes we are
missing a recompute of the routes, could you take a look into that.

Regards,
Ales
_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev

Reply via email to