Only the ovn-controller that runs a BFD session writes its status in the
Southbound BFD table, and only when the state of the session changes.
So a wrong status stays until the next change of state:

- A session that starts, after ovn-controller restarted or after the
  gateway port moved to this chassis, goes from ADMIN_DOWN to DOWN
  without writing anything.  If the peer never answers, a stale "up"
  from the earlier session stays forever, and ovn-northd keeps the
  route's nexthop in the ECMP set.
- A status that another client writes stays even when it is wrong.
- A state change is written only once.  If its transaction fails, the
  status is not written again.
- The pinctrl thread checks the detection time of a session only while
  it is connected to ovs-vswitchd, and ovn-controller does not call
  pinctrl_run() at all while it cannot reach ovs-vswitchd over OpenFlow.
  So while ovs-vswitchd is down, a session that was up stays "up".

bfd_monitor_run() now compares the Southbound status of each session
that this chassis runs with the state of the session and corrects it
when they differ:

- "admin_down", in Southbound or in the session, is never corrected.
- The update has no "verify", because a failed verify would abort the
  whole transaction of the pass.
- "down" is not written over "up" or "init" while the session starts.
  The correction waits until the session has been sending for the
  longer of 5 seconds and the detection time configured in the row, and
  while the peer keeps answering without the session coming up, until
  the detection time after its last packet, but at most 30 seconds.
- bfd_monitor_run() also checks the detection time, so a correction
  never publishes "up" for a session without packets.  When
  ovn-controller cannot call pinctrl_run(), for example because
  ovs-vswitchd is down, it now runs the BFD part alone, through the new
  pinctrl_bfd_run().
- Each correction is logged at INFO and counted in the coverage counter
  pinctrl_bfd_status_reconcile.  Corrections of a row that keeps its
  status are spaced 5, 10, 20 and 40 seconds apart, and the fifth is the
  last one: a WARN then names the row, and the counter
  pinctrl_bfd_status_reconcile_stuck counts it.  Corrections resume when
  the status or the chassis_name of the row changes, when the session
  is created again, or after 10 minutes.  When chassis_name becomes this
  chassis after such corrections, the next one comes at once.

With SB RBAC, the SB database accepts the update of a row only if its
chassis_name is this chassis, and a rejected update fails the whole
transaction.  So corrections of other rows are never written in the
same pass as an update of a row whose chassis_name is this chassis.

ovn-controller advertises the new behavior with
other_config:ovn-bfd-status-reconcile=true in its Chassis record.  A
CMS that writes "down" itself for a gateway that it believes dead can
check it: such a chassis sets the status back if its session is up.

Fixes: 02839c4d8934 ("controller: bfd: introduce BFD state machine.")
Reported-at: https://github.com/ovn-org/ovn/issues/322
Submitted-at: https://github.com/ovn-org/ovn/pull/333
Assisted-by: Claude Opus 5.5 (claude-opus-5-5), Cursor Grok Bot / Ultimum 
harness assistants
Signed-off-by: Premysl Kouril <[email protected]>
---
 NEWS                        |  11 +
 controller/chassis.c        |   8 +
 controller/ovn-controller.c |  11 +
 controller/pinctrl.c        | 452 ++++++++++++++++++++++++++++++++++-
 controller/pinctrl.h        |   4 +
 include/ovn/features.h      |   1 +
 ovn-sb.xml                  |  43 ++++
 tests/ovn.at                | 461 ++++++++++++++++++++++++++++++++++++
 8 files changed, 980 insertions(+), 11 deletions(-)

diff --git a/NEWS b/NEWS
index 7f94d0b14..9c1cc54a7 100644
--- a/NEWS
+++ b/NEWS
@@ -19,6 +19,17 @@ Post v26.09.0
    - Removed OVN's ovs-bugtool plugin and helper scripts.
    - Removed ovn-sim utility scripts.
    - Removed ovn-docker utility scripts.
+   - ovn-controller now corrects the Southbound status of the BFD sessions
+     that it runs when the status differs from the state of the session,
+     for example a stale "up" left by an earlier session that cannot reach
+     its peer, or a status written by another client.  A starting session
+     gets a few seconds to come up before "down" is written over "up" or
+     "init", so a restart does not remove its routes.  ovn-controller also
+     times out BFD sessions while ovs-vswitchd is down.  Chassis that do
+     this have other_config:ovn-bfd-status-reconcile=true.  With SB RBAC,
+     the Southbound database rejects these corrections for rows whose
+     chassis_name is not the chassis, as it already rejects the status
+     updates on state changes; they stop after five tries with a warning.
 
 OVN v26.09.0 - xxx xx xxxx
 --------------------------
diff --git a/controller/chassis.c b/controller/chassis.c
index b82683064..6b2eaf5d6 100644
--- a/controller/chassis.c
+++ b/controller/chassis.c
@@ -424,6 +424,7 @@ chassis_build_other_config(const struct ovs_chassis_cfg 
*ovs_cfg,
     smap_replace(config, OVN_FEATURE_CT_LABEL_FLUSH,
                  ovs_cfg->ct_label_flush ? "true" :"false");
     smap_replace(config, OVN_FEATURE_CT_STATE_SAVE, "true");
+    smap_replace(config, OVN_FEATURE_BFD_STATUS_RECONCILE, "true");
 }
 
 /*
@@ -597,6 +598,12 @@ chassis_other_config_changed(const struct ovs_chassis_cfg 
*ovs_cfg,
         return true;
     }
 
+    if (!smap_get_bool(&chassis_rec->other_config,
+                       OVN_FEATURE_BFD_STATUS_RECONCILE,
+                       false)) {
+        return true;
+    }
+
     return false;
 }
 
@@ -779,6 +786,7 @@ update_supported_sset(struct sset *supported)
     sset_add(supported, OVN_FEATURE_CT_NEXT_ZONE);
     sset_add(supported, OVN_FEATURE_CT_LABEL_FLUSH);
     sset_add(supported, OVN_FEATURE_CT_STATE_SAVE);
+    sset_add(supported, OVN_FEATURE_BFD_STATUS_RECONCILE);
 }
 
 static void
diff --git a/controller/ovn-controller.c b/controller/ovn-controller.c
index d57ff316d..de8bc8fb5 100644
--- a/controller/ovn-controller.c
+++ b/controller/ovn-controller.c
@@ -8366,6 +8366,7 @@ main(int argc, char *argv[])
                                         ovs_feature_max_meters_get());
             }
 
+            bool pinctrl_ran = false;
             if (br_int) {
                 ct_zones_data = engine_get_data(&en_ct_zones);
                 if (ofctrl_run(br_int_remote.target,
@@ -8554,6 +8555,7 @@ main(int argc, char *argv[])
                                     ovsrec_open_vswitch_table_get(
                                             ovs_idl_loop.idl),
                                     ovnsb_idl_loop.cur_cfg);
+                        pinctrl_ran = true;
                         stopwatch_stop(PINCTRL_RUN_STOPWATCH_NAME,
                                        time_msec());
                         mirror_run(ovs_idl_txn,
@@ -8686,6 +8688,15 @@ main(int argc, char *argv[])
                 }
             }
 
+            if (chassis && !pinctrl_ran) {
+                /* pinctrl_run() could not run, for example because
+                 * ovs-vswitchd is down.  Still time out the BFD sessions
+                 * that get no packets and publish their state. */
+                pinctrl_bfd_run(ovnsb_idl_txn,
+                                sbrec_bfd_table_get(ovnsb_idl_loop.idl),
+                                sbrec_port_binding_by_name, chassis);
+            }
+
             if (!engine_has_run()) {
                 if (engine_need_run()) {
                     VLOG_DBG("engine did not run, force recompute next time: "
diff --git a/controller/pinctrl.c b/controller/pinctrl.c
index 5c9634b24..eb0838244 100644
--- a/controller/pinctrl.c
+++ b/controller/pinctrl.c
@@ -58,6 +58,7 @@
 #include "openvswitch/poll-loop.h"
 #include "openvswitch/rconn.h"
 #include "socket-util.h"
+#include "sat-math.h"
 #include "seq.h"
 #include "timeval.h"
 #include "vswitch-idl.h"
@@ -357,6 +358,10 @@ static void notify_pinctrl_handler(void);
 
 static bool bfd_monitor_should_inject(void);
 static void bfd_monitor_wait(long long int timeout);
+static void bfd_monitor_main_wait(void) OVS_REQUIRES(pinctrl_mutex);
+/* Whether the pinctrl thread is connected to ovs-vswitchd, so that it
+ * checks the detection times of the BFD sessions. */
+static atomic_bool bfd_thread_connected = false;
 static void bfd_monitor_init(void);
 static void bfd_monitor_destroy(void);
 static void bfd_monitor_send_msg(struct rconn *swconn, long long int *bfd_time)
@@ -368,7 +373,8 @@ pinctrl_handle_bfd_msg(struct rconn *swconn, const struct 
flow *ip_flow,
 static void bfd_monitor_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
                             const struct sbrec_bfd_table *bfd_table,
                             struct ovsdb_idl_index *sbrec_port_binding_by_name,
-                            const struct sbrec_chassis *chassis)
+                            const struct sbrec_chassis *chassis,
+                            bool can_send)
                             OVS_REQUIRES(pinctrl_mutex);
 static void init_fdb_entries(void);
 static void destroy_fdb_entries(void);
@@ -401,6 +407,8 @@ COVERAGE_DEFINE(pinctrl_drop_put_vport_binding);
 COVERAGE_DEFINE(pinctrl_notify_main_thread);
 COVERAGE_DEFINE(pinctrl_notify_handler_thread);
 COVERAGE_DEFINE(pinctrl_total_pin_pkts);
+COVERAGE_DEFINE(pinctrl_bfd_status_reconcile);
+COVERAGE_DEFINE(pinctrl_bfd_status_reconcile_stuck);
 
 /* DNS query statistics - thread-safe coverage counters */
 COVERAGE_DEFINE(dns_query_total);
@@ -4072,6 +4080,8 @@ pinctrl_handler(void *arg_)
 
         rconn_run(swconn);
         new_seq = seq_read(pinctrl_handler_seq);
+        atomic_store_relaxed(&bfd_thread_connected,
+                             rconn_is_connected(swconn));
         if (rconn_is_connected(swconn)) {
             if (conn_seq_no != rconn_get_connection_seqno(swconn)) {
                 pinctrl_setup(swconn);
@@ -4243,7 +4253,7 @@ pinctrl_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
     sync_svc_monitors(ovnsb_idl_txn, svc_mon_table, sbrec_port_binding_by_name,
                       chassis);
     bfd_monitor_run(ovnsb_idl_txn, bfd_table, sbrec_port_binding_by_name,
-                    chassis);
+                    chassis, true);
     run_activated_ports(ovnsb_idl_txn, sbrec_datapath_binding_by_key,
                         sbrec_port_binding_by_key, chassis);
     ovs_mutex_unlock(&pinctrl_mutex);
@@ -4768,6 +4778,7 @@ pinctrl_wait(struct ovsdb_idl_txn *ovnsb_idl_txn)
         seq_wait(pinctrl_main_seq, main_seq);
     }
     wait_activated_ports();
+    bfd_monitor_main_wait();
     ovs_mutex_unlock(&pinctrl_mutex);
 
     if (ovnsb_idl_txn) {
@@ -7674,8 +7685,43 @@ struct bfd_entry {
     uint32_t detection_timeout;
     long long int last_rx;
     long long int next_tx;
+
+    /* Correction of the SB status, see bfd_reconcile_due(). */
+    long long int active_since;    /* When the session started sending. */
+    long long int next_reconcile;  /* Next correction or end of a hold,
+                                    * 0 if none is pending. */
+    unsigned int reconcile_attempts; /* Corrections since the SB status last
+                                      * matched the session. */
+    bool corrected_other_name;     /* The last correction was made while the
+                                    * SB chassis_name was not this chassis. */
+    bool start_pending;            /* Started while pinctrl could not
+                                    * send, see pinctrl_bfd_run(). */
+    bool parked;                   /* Corrections stopped, see
+                                    * bfd_reconcile_park(). */
+    char *parked_status;           /* SB status when parked. */
+    char *parked_chassis_name;     /* SB chassis_name when parked. */
+    long long int parked_until;    /* Corrections resume at this time. */
 };
 
+/* Corrections without the SB status matching the session after which
+ * corrections stop ("park") for the row. */
+#define BFD_RECONCILE_PARK_THRESHOLD 5
+/* Longest wait between two corrections of the same row.  Parking comes
+ * first with the current threshold; this bounds the wait if it is raised. */
+#define BFD_RECONCILE_MAX_BACKOFF    60000LL
+/* How long a parked row stays parked unless its SB row changes. */
+#define BFD_RECONCILE_PARK_TIME      (10 * 60 * 1000LL)
+/* Longest a handshake in progress may hold a downward correction, counted
+ * from when the session started sending. */
+#define BFD_RECONCILE_HANDSHAKE_MAX  30000LL
+
+/* Time of the last bfd_monitor_run() pass, and whether one ran since the
+ * last bfd_monitor_main_wait(). */
+static long long int bfd_monitor_last_run;
+static bool bfd_monitor_ran;
+/* Number of entries whose corrections are parked. */
+static unsigned int bfd_monitor_n_parked;
+
 static void
 bfd_monitor_init(void)
 {
@@ -7683,12 +7729,23 @@ bfd_monitor_init(void)
     bfd_last_update = time_msec();
 }
 
+static void
+bfd_entry_destroy(struct bfd_entry *entry)
+{
+    if (entry->parked) {
+        bfd_monitor_n_parked--;
+    }
+    free(entry->parked_status);
+    free(entry->parked_chassis_name);
+    free(entry);
+}
+
 static void
 bfd_monitor_destroy(void)
 {
     struct bfd_entry *entry;
     HMAP_FOR_EACH_POP (entry, node, &bfd_monitor_map) {
-        free(entry);
+        bfd_entry_destroy(entry);
     }
     hmap_destroy(&bfd_monitor_map);
 }
@@ -7744,6 +7801,38 @@ bfd_monitor_wait(long long int timeout)
     }
 }
 
+/* Called in the main thread.  Wakes it up when a correction or the end of a
+ * hold is due, so that neither waits for an unrelated wake-up.  A deadline
+ * that the last bfd_monitor_run() already saw expire, but could not act on
+ * (no SB transaction, or a write postponed), is left to the completion of
+ * the SB transaction, so that nothing spins.  While the pinctrl thread is
+ * not connected to ovs-vswitchd and so does not check the detection times
+ * of the sessions, also wakes the main thread up when one of them expires:
+ * bfd_monitor_run() checks them too. */
+static void
+bfd_monitor_main_wait(void)
+    OVS_REQUIRES(pinctrl_mutex)
+{
+    if (!bfd_monitor_ran) {
+        return;
+    }
+
+    bool thread_connected;
+    atomic_read_relaxed(&bfd_thread_connected, &thread_connected);
+
+    struct bfd_entry *entry;
+    HMAP_FOR_EACH (entry, node, &bfd_monitor_map) {
+        if (entry->next_reconcile > bfd_monitor_last_run) {
+            poll_timer_wait_until(entry->next_reconcile);
+        }
+        if (!thread_connected && entry->detection_timeout &&
+            (entry->state == BFD_STATE_UP || entry->state == BFD_STATE_INIT)) {
+            poll_timer_wait_until(entry->last_rx + entry->detection_timeout);
+        }
+    }
+    bfd_monitor_ran = false;
+}
+
 static void
 bfd_monitor_put_bfd_msg(struct bfd_entry *entry, struct dp_packet *packet,
                         bool final)
@@ -7851,28 +7940,30 @@ update:
     return true;
 }
 
-static void
+/* Moves 'entry' from UP or INIT to DOWN if no packet came from the peer
+ * within the detection time.  Returns true if it did. */
+static bool
 bfd_check_detection_timeout(struct bfd_entry *entry)
 {
     if (entry->state == BFD_STATE_ADMIN_DOWN ||
         entry->state == BFD_STATE_DOWN) {
-        return;
+        return false;
     }
 
     if (!entry->detection_timeout) {
-        return;
+        return false;
     }
 
     long long int cur_time = time_msec();
     if (cur_time < entry->last_rx + entry->detection_timeout) {
-        return;
+        return false;
     }
 
     entry->state = BFD_STATE_DOWN;
     entry->change_state = true;
     bfd_last_update = cur_time;
     bfd_pending_update = 0;
-    notify_pinctrl_main();
+    return true;
 }
 
 static void
@@ -7889,7 +7980,9 @@ bfd_monitor_send_msg(struct rconn *swconn, long long int 
*bfd_time)
     HMAP_FOR_EACH (entry, node, &bfd_monitor_map) {
         unsigned long tx_timeout;
 
-        bfd_check_detection_timeout(entry);
+        if (bfd_check_detection_timeout(entry)) {
+            notify_pinctrl_main();
+        }
 
         if (cur_time < entry->next_tx) {
             goto next;
@@ -8143,13 +8236,278 @@ bfd_monitor_check_sb_conf(const struct sbrec_bfd 
*sb_bt,
     }
 }
 
+/* Forgets the corrections of 'entry' and lifts its parking.  'why' says why
+ * corrections may resume, for the log if the row was parked. */
+static void
+bfd_reconcile_reset(struct bfd_entry *entry, const struct sbrec_bfd *bt,
+                    const char *why)
+{
+    if (entry->parked) {
+        static struct vlog_rate_limit rl = VLOG_RATE_LIMIT_INIT(20, 100);
+        VLOG_INFO_RL(&rl, "BFD %s dst %s: %s, SB status corrections resume",
+                     bt->logical_port, bt->dst_ip, why);
+        bfd_monitor_n_parked--;
+        free(entry->parked_status);
+        free(entry->parked_chassis_name);
+        entry->parked_status = NULL;
+        entry->parked_chassis_name = NULL;
+        entry->parked = false;
+        entry->parked_until = 0;
+    }
+    entry->reconcile_attempts = 0;
+    entry->corrected_other_name = false;
+    entry->next_reconcile = 0;
+}
+
+/* Returns the time to wait after the last correction before the next one:
+ * BFD_UPDATE_TIMEOUT, doubling with each correction, at most
+ * BFD_RECONCILE_MAX_BACKOFF. */
+static long long int
+bfd_reconcile_backoff(const struct bfd_entry *entry)
+{
+    unsigned int shift = MIN(entry->reconcile_attempts - 1, 4);
+    return MIN(BFD_UPDATE_TIMEOUT << shift, BFD_RECONCILE_MAX_BACKOFF);
+}
+
+/* Returns the time until which the correction of SB status 'bt->status' to
+ * the state of 'entry' must wait, or 0 if it need not wait.
+ *
+ * Only a correction to "down" over "up" or "init" waits.  A session that was
+ * just created, for example after ovn-controller restarted or the gateway
+ * moved here, is DOWN until the handshake with the peer completes, while SB
+ * still has the status of the previous session, which is usually right.
+ * Writing "down" right away would remove the route for no reason.  So wait
+ * until the session has been sending for the longer of BFD_UPDATE_TIMEOUT and
+ * the detection time configured in the row, and while a handshake is in
+ * progress (a packet came from the peer since the session started sending,
+ * within the detection time), until the detection time after that packet,
+ * but at most BFD_RECONCILE_HANDSHAKE_MAX after the session started sending.
+ * The wait counts from when the session started sending, not from when it
+ * was created, because the two can be far apart. */
+static long long int
+bfd_reconcile_hold_until(const struct bfd_entry *entry,
+                         const struct sbrec_bfd *bt, long long int now)
+{
+    if (entry->state != BFD_STATE_DOWN ||
+        (strcmp(bt->status, "up") && strcmp(bt->status, "init"))) {
+        return 0;
+    }
+
+    long long int detect_time = llsat_mul(bt->detect_mult,
+                                          MAX(bt->min_rx, bt->min_tx));
+    long long int hold_until = llsat_add(entry->active_since,
+                                         MAX(BFD_UPDATE_TIMEOUT, detect_time));
+
+    if (entry->last_rx > entry->active_since &&
+        now < entry->last_rx + entry->detection_timeout) {
+        long long int handshake_until =
+            MIN(entry->last_rx + entry->detection_timeout,
+                entry->active_since + BFD_RECONCILE_HANDSHAKE_MAX);
+        hold_until = MAX(hold_until, handshake_until);
+    }
+    return hold_until;
+}
+
+/* Called before the last of BFD_RECONCILE_PARK_THRESHOLD corrections of
+ * 'entry' is written, while 'bt' still has the status that the earlier
+ * ones did not change.  Stops the corrections after that one ("parks" the
+ * row).  A rejected update leaves the SB row as it was, while another client
+ * that keeps writing a different status changes it, so corrections resume
+ * when the row's status or chassis_name changes, when the entry is recreated
+ * (for example because the gateway moved), or after
+ * BFD_RECONCILE_PARK_TIME. */
+static void
+bfd_reconcile_park(struct bfd_entry *entry, const struct sbrec_bfd *bt,
+                   const struct sbrec_chassis *chassis, long long int now)
+{
+    static struct vlog_rate_limit rl = VLOG_RATE_LIMIT_INIT(20, 100);
+
+    entry->parked = true;
+    entry->parked_status = xstrdup(bt->status);
+    entry->parked_chassis_name = xstrdup(bt->chassis_name);
+    entry->parked_until = now + BFD_RECONCILE_PARK_TIME;
+    entry->next_reconcile = entry->parked_until;
+    bfd_monitor_n_parked++;
+    COVERAGE_INC(pinctrl_bfd_status_reconcile_stuck);
+
+    if (!VLOG_DROP_WARN(&rl)) {
+        struct ds why = DS_EMPTY_INITIALIZER;
+        if (strcmp(bt->chassis_name, chassis->name)) {
+            ds_put_format(&why, "The SB server may be rejecting them: with "
+                          "SB RBAC, the row's chassis_name (\"%s\") must be "
+                          "this chassis (\"%s\")", bt->chassis_name,
+                          chassis->name);
+        } else {
+            ds_put_cstr(&why, "The SB transactions that carried them may "
+                        "have failed; see the earlier log messages");
+        }
+        VLOG_WARN("BFD %s dst %s: SB status \"%s\" did not follow %u "
+                  "corrections to the session state \"%s\"; making one last "
+                  "correction, then none until the row changes or for %lld "
+                  "minutes (%u row(s) of this chassis stopped).  %s",
+                  bt->logical_port, bt->dst_ip, bt->status,
+                  entry->reconcile_attempts - 1, bfd_get_status(entry->state),
+                  BFD_RECONCILE_PARK_TIME / (60 * 1000), bfd_monitor_n_parked,
+                  ds_cstr(&why));
+        ds_destroy(&why);
+    }
+}
+
+/* Returns true if corrections of 'entry' are parked.  Lifts the parking if
+ * the row changed or the parking time is over. */
+static bool
+bfd_reconcile_parked(struct bfd_entry *entry, const struct sbrec_bfd *bt,
+                     long long int now)
+{
+    if (!entry->parked) {
+        return false;
+    }
+    if (strcmp(bt->status, entry->parked_status)) {
+        bfd_reconcile_reset(entry, bt, "SB status changed");
+    } else if (strcmp(bt->chassis_name, entry->parked_chassis_name)) {
+        bfd_reconcile_reset(entry, bt, "SB chassis_name changed");
+    } else if (now >= entry->parked_until) {
+        bfd_reconcile_reset(entry, bt, "parking time is over");
+    }
+    return entry->parked;
+}
+
+/* Correcting the SB status of the sessions that this chassis runs.
+ *
+ * The session's state machine writes the status only when the state
+ * changes.  So without this, a wrong status would stay until the next
+ * change: a status left over from an earlier session (after ovn-controller
+ * restarted, or after the gateway moved here) when the new session cannot
+ * reach its peer, a status that another client wrote while the session
+ * stays in the same state, or a state change whose write failed.
+ *
+ * An "admin_down" in SB or in the session is never corrected: ovn-northd
+ * sets it for rows that no route uses, and the session follows it.  The
+ * update has no "verify", because a failed one would abort the whole SB
+ * transaction of this pass; ovn-northd changes the status only to and from
+ * "admin_down", and sets "admin_down" again by itself if needed.
+ * Corrections are spaced out (see bfd_reconcile_backoff()) because a
+ * rejected update, for example by SB RBAC, fails the whole SB transaction
+ * of its pass, and they stop after a few failures (see
+ * bfd_reconcile_park()).  Only a matching status, an "admin_down", a new
+ * entry, a lifted parking, or a chassis_name that became this chassis after
+ * corrections made under another one resets them; a pass without an SB
+ * transaction or with a state change still to be written does not. */
+
+/* A correction that is due in this pass of bfd_monitor_run(). */
+struct bfd_correction {
+    struct bfd_entry *entry;
+    const struct sbrec_bfd *bt;
+};
+
+/* Returns true if the SB status of 'bt' must be corrected to the state of
+ * 'entry' now.  Otherwise updates the corrections of 'entry' as needed. */
+static bool
+bfd_reconcile_due(struct ovsdb_idl_txn *ovnsb_idl_txn,
+                  const struct sbrec_bfd *bt, struct bfd_entry *entry,
+                  const struct sbrec_chassis *chassis, long long int now)
+{
+    if (entry->state == BFD_STATE_ADMIN_DOWN ||
+        !strcmp(bt->status, "admin_down")) {
+        bfd_reconcile_reset(entry, bt, "SB status is admin_down");
+        return false;
+    }
+    if (!strcmp(bt->status, bfd_get_status(entry->state))) {
+        bfd_reconcile_reset(entry, bt, "SB status matches the session");
+        return false;
+    }
+
+    if (!ovnsb_idl_txn || entry->change_state ||
+        bfd_reconcile_parked(entry, bt, now)) {
+        return false;
+    }
+    if (entry->corrected_other_name &&
+        !strcmp(bt->chassis_name, chassis->name)) {
+        /* The earlier corrections went to a row whose chassis_name was
+         * empty or named another chassis, which SB RBAC rejects.  Now the
+         * row names this chassis, so start over instead of waiting out the
+         * back-off.  This happens at most once for each change of
+         * chassis_name to this chassis. */
+        static struct vlog_rate_limit rl = VLOG_RATE_LIMIT_INIT(10, 20);
+        VLOG_INFO_RL(&rl, "BFD %s dst %s: SB chassis_name is now this "
+                     "chassis, SB status corrections start over",
+                     bt->logical_port, bt->dst_ip);
+        bfd_reconcile_reset(entry, bt, "SB chassis_name is this chassis");
+    }
+    if (now < entry->next_reconcile) {
+        return false;
+    }
+
+    long long int hold_until = bfd_reconcile_hold_until(entry, bt, now);
+    if (hold_until > now) {
+        entry->next_reconcile = hold_until;
+        return false;
+    }
+    return true;
+}
+
+static void
+bfd_reconcile_write(const struct sbrec_bfd *bt, struct bfd_entry *entry,
+                    const struct sbrec_chassis *chassis, long long int now)
+{
+    static struct vlog_rate_limit rl = VLOG_RATE_LIMIT_INIT(10, 20);
+    const char *state = bfd_get_status(entry->state);
+
+    entry->reconcile_attempts++;
+    entry->corrected_other_name = strcmp(bt->chassis_name, chassis->name) != 0;
+    VLOG_INFO_RL(&rl, "BFD %s dst %s: correcting SB status \"%s\" to the "
+                 "session state \"%s\" (correction %u)", bt->logical_port,
+                 bt->dst_ip, bt->status, state, entry->reconcile_attempts);
+    entry->next_reconcile = now + bfd_reconcile_backoff(entry);
+    if (entry->reconcile_attempts >= BFD_RECONCILE_PARK_THRESHOLD) {
+        bfd_reconcile_park(entry, bt, chassis, now);
+    }
+    sbrec_bfd_set_status(bt, state);
+    COVERAGE_INC(pinctrl_bfd_status_reconcile);
+}
+
+/* Writes the 'n' corrections due in this pass.  'own_written' says whether
+ * the pass already wrote the status of a row whose chassis_name is this
+ * chassis.
+ *
+ * With SB RBAC, the SB server rejects the update of a row whose chassis_name
+ * is not this chassis, and one rejected update fails the whole transaction.
+ * So such corrections are not written in a pass that writes a row whose
+ * chassis_name is this chassis: they wait for the next pass, which the
+ * completion of this pass's transaction brings.  Otherwise a row that the
+ * server always rejects would keep the corrections of the other rows from
+ * ever taking effect. */
+static void
+bfd_reconcile_write_all(const struct bfd_correction *corrections, size_t n,
+                        bool own_written, const struct sbrec_chassis *chassis,
+                        long long int now)
+{
+    for (size_t i = 0; i < n; i++) {
+        if (!strcmp(corrections[i].bt->chassis_name, chassis->name)) {
+            own_written = true;
+        }
+    }
+    for (size_t i = 0; i < n; i++) {
+        const struct bfd_correction *c = &corrections[i];
+        if (!own_written || !strcmp(c->bt->chassis_name, chassis->name)) {
+            bfd_reconcile_write(c->bt, c->entry, chassis, now);
+        }
+    }
+}
+
+/* Runs the BFD sessions of the rows in 'bfd_table' that this chassis owns.
+ * 'can_send' is false if the pinctrl thread cannot send their packets yet
+ * (see pinctrl_bfd_run()). */
 static void
 bfd_monitor_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
                 const struct sbrec_bfd_table *bfd_table,
                 struct ovsdb_idl_index *sbrec_port_binding_by_name,
-                const struct sbrec_chassis *chassis)
+                const struct sbrec_chassis *chassis, bool can_send)
     OVS_REQUIRES(pinctrl_mutex)
 {
+    struct bfd_correction *corrections = NULL;
+    size_t n_corrections = 0, allocated_corrections = 0;
+    bool own_written = false;
     struct bfd_entry *entry;
     long long int cur_time = time_msec();
     bool changed = false;
@@ -8189,6 +8547,26 @@ bfd_monitor_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
 
         entry = pinctrl_find_bfd_monitor_entry_by_port(
                 bt->dst_ip, bt->src_port);
+        if (entry) {
+            /* The pinctrl thread checks the detection time only while it
+             * is connected to ovs-vswitchd.  Check it here too, so that an
+             * UP or INIT session always heard from its peer within the
+             * detection time.  A timeout found here is written below. */
+            bfd_check_detection_timeout(entry);
+
+            if (can_send && entry->start_pending) {
+                /* The session started while it could not send: its hold
+                 * starts now. */
+                entry->start_pending = false;
+                if (entry->state == BFD_STATE_DOWN) {
+                    entry->active_since = cur_time;
+                    VLOG_DBG("BFD %s dst %s: session can send now",
+                             bt->logical_port, bt->dst_ip);
+                }
+            }
+        }
+
+        bool status_written = false;
         if (!entry) {
             struct eth_addr ea = eth_addr_zero;
             struct lport_addresses dst_addr;
@@ -8245,6 +8623,8 @@ bfd_monitor_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
             entry->local_min_rx = bt->min_rx;
             entry->remote_min_rx = 1; /* RFC5880 page 29 */
             entry->local_mult = bt->detect_mult;
+            entry->active_since = cur_time;
+            entry->start_pending = !can_send;
 
             uint32_t hash = hash_string(bt->dst_ip, 0);
             hmap_insert(&bfd_monitor_map, &entry->node, hash);
@@ -8258,6 +8638,11 @@ bfd_monitor_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
             entry->state = BFD_STATE_DOWN;
             entry->change_state = false;
             entry->remote_disc = 0;
+            /* The session starts sending now. */
+            entry->active_since = cur_time;
+            entry->start_pending = !can_send;
+            VLOG_DBG("BFD %s dst %s: session starts%s", bt->logical_port,
+                     bt->dst_ip, can_send ? "" : " (cannot send yet)");
             changed = true;
         } else if (entry->change_state && ovnsb_idl_txn) {
             if (entry->state == BFD_STATE_DOWN) {
@@ -8265,23 +8650,68 @@ bfd_monitor_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
             }
             sbrec_bfd_set_status(bt, bfd_get_status(entry->state));
             entry->change_state = false;
+            status_written = true;
+            if (!strcmp(bt->chassis_name, chassis->name)) {
+                own_written = true;
+            }
+        }
+        if (!status_written &&
+            bfd_reconcile_due(ovnsb_idl_txn, bt, entry, chassis,
+                              cur_time)) {
+            if (n_corrections >= allocated_corrections) {
+                corrections = x2nrealloc(corrections, &allocated_corrections,
+                                         sizeof *corrections);
+            }
+            corrections[n_corrections++] = (struct bfd_correction) {
+                .entry = entry,
+                .bt = bt,
+            };
         }
         bfd_monitor_check_sb_conf(bt, entry);
         entry->erase = false;
     }
 
+    bfd_reconcile_write_all(corrections, n_corrections, own_written, chassis,
+                            cur_time);
+    free(corrections);
+
     HMAP_FOR_EACH_SAFE (entry, node, &bfd_monitor_map) {
         if (entry->erase) {
             hmap_remove(&bfd_monitor_map, &entry->node);
-            free(entry);
+            bfd_entry_destroy(entry);
         }
     }
 
+    bfd_monitor_last_run = cur_time;
+    bfd_monitor_ran = true;
+
     if (changed) {
         notify_pinctrl_handler();
     }
 }
 
+/* Called by ovn-controller instead of pinctrl_run() in the iterations in
+ * which it cannot call that, for example because ovs-vswitchd is down.  The
+ * BFD sessions get no packets then, and the pinctrl thread, which needs the
+ * OpenFlow connection, does not check their detection times.  Check them
+ * here, and keep the SB status of the sessions in line with them, so that
+ * a status stays "up" only while packets arrive.  A session started here
+ * cannot send yet, so its start-up hold starts again when pinctrl_run()
+ * first runs it.  Also registers the wake-ups, because ovn-controller may
+ * not call pinctrl_wait() in such an iteration. */
+void
+pinctrl_bfd_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
+                const struct sbrec_bfd_table *bfd_table,
+                struct ovsdb_idl_index *sbrec_port_binding_by_name,
+                const struct sbrec_chassis *chassis)
+{
+    ovs_mutex_lock(&pinctrl_mutex);
+    bfd_monitor_run(ovnsb_idl_txn, bfd_table, sbrec_port_binding_by_name,
+                    chassis, false);
+    bfd_monitor_main_wait();
+    ovs_mutex_unlock(&pinctrl_mutex);
+}
+
 static uint16_t
 get_random_src_port(void)
 {
diff --git a/controller/pinctrl.h b/controller/pinctrl.h
index a638fe29f..1b5827ef0 100644
--- a/controller/pinctrl.h
+++ b/controller/pinctrl.h
@@ -61,6 +61,10 @@ void pinctrl_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
                  const struct shash *local_active_ports_ras,
                  const struct ovsrec_open_vswitch_table *ovs_table,
                  int64_t cur_cfg);
+void pinctrl_bfd_run(struct ovsdb_idl_txn *ovnsb_idl_txn,
+                     const struct sbrec_bfd_table *,
+                     struct ovsdb_idl_index *sbrec_port_binding_by_name,
+                     const struct sbrec_chassis *chassis);
 void pinctrl_wait(struct ovsdb_idl_txn *ovnsb_idl_txn);
 void pinctrl_destroy(void);
 void pinctrl_seqno_run(void);
diff --git a/include/ovn/features.h b/include/ovn/features.h
index 7207da67f..6b456c8ed 100644
--- a/include/ovn/features.h
+++ b/include/ovn/features.h
@@ -27,6 +27,7 @@
 #define OVN_FEATURE_CT_NEXT_ZONE "ct-next-zone"
 #define OVN_FEATURE_CT_LABEL_FLUSH "ct-label-flush"
 #define OVN_FEATURE_CT_STATE_SAVE "ct-state-save"
+#define OVN_FEATURE_BFD_STATUS_RECONCILE "ovn-bfd-status-reconcile"
 
 /* DEPRECATED: The following features can be removed
  * after the next LTS version release. */
diff --git a/ovn-sb.xml b/ovn-sb.xml
index 2096fc3e3..13262c216 100644
--- a/ovn-sb.xml
+++ b/ovn-sb.xml
@@ -399,6 +399,18 @@
       table. See <code>ovn-controller</code>(8) for more information.
     </column>
 
+    <column name="other_config" key="ovn-bfd-status-reconcile">
+      <code>ovn-controller</code> sets this key to <code>true</code> if it
+      corrects the <ref table="BFD" column="status"/> of the BFD sessions
+      that it runs when that differs from the state of the session.  A CMS
+      that sets a BFD status to <code>down</code> itself can then expect
+      this chassis to set it back if the session is in fact up, as long as
+      the chassis is connected to the Southbound database and allowed to
+      update the row: with role-based access control, the row's <ref
+      table="BFD" column="chassis_name"/> must be the chassis's name.  Other
+      applications should treat this key as read-only.
+    </column>
+
     <group title="Common Columns">
       The overall purpose of these columns is described under <code>Common
       Columns</code> at the beginning of this document.
@@ -5590,6 +5602,37 @@ tcp.flags = RST;
             </li>
           </ul>
         </p>
+
+        <p>
+          The <code>ovn-controller</code> that runs the session updates this
+          column when the state of the session changes.  It also corrects
+          the column when it differs from the state of the session, for
+          example after another client wrote it, unless one of them is
+          <code>admin_down</code>.  A session is <code>up</code> or
+          <code>init</code> only while packets from the peer arrive within
+          the detection time, including while <code>ovs-vswitchd</code> is
+          down.  When a session starts, for example after
+          <code>ovn-controller</code> restarted or the gateway port moved, it
+          is <code>down</code> until the peer answers, so
+          <code>ovn-controller</code> waits until the session has been
+          sending for the longer of 5 seconds and the detection time
+          (<ref column="detect_mult"/> times the larger of <ref
+          column="min_rx"/> and <ref column="min_tx"/>), and while the peer
+          keeps answering without the session coming up, up to 30 seconds,
+          before it corrects an <code>up</code> or <code>init</code> to
+          <code>down</code>.  Corrections that do not take effect, for
+          example because the SB database rejects them, are retried 5, 10,
+          20 and 40 seconds apart, and then stop with a warning in the log
+          until the row's status or <ref column="chassis_name"/> changes or
+          for 10 minutes.  If <code>ovn-controller</code> made the
+          corrections while <ref column="chassis_name"/> was empty or named
+          another chassis, and <ref column="chassis_name"/> then changes to
+          its own chassis, it makes the next correction at once.  The
+          coverage counter
+          <code>pinctrl_bfd_status_reconcile</code> counts the corrections,
+          and <code>pinctrl_bfd_status_reconcile_stuck</code> counts how
+          many times they stopped for a row.
+        </p>
       </column>
     </group>
   </table>
diff --git a/tests/ovn.at b/tests/ovn.at
index 136b825f9..3b1fe4d92 100644
--- a/tests/ovn.at
+++ b/tests/ovn.at
@@ -12953,6 +12953,467 @@ ignored_tables=OFTABLE_PHY_TO_LOG
 AT_CLEANUP
 ])
 
+dnl BFD_RECONCILE_SETUP([rbac])
+dnl
+dnl Sandbox gw1 runs the BFD session of distributed gateway port r0-ext
+dnl (10.0.0.1/24) towards 10.0.0.254, the nexthop of a default route.  No
+dnl peer answers unless a test plays it with bfd_rc_peer_start.  Unless
+dnl "rbac" is given, gw1's ovn-controller uses an SB connection without an
+dnl RBAC role.  (With SSL, the default connection has role ovn-controller,
+dnl see ovn_start.)
+m4_define([BFD_RECONCILE_SETUP], [
+CHECK_SCAPY
+ovn_start
+net_add n1
+net_add provider
+sim_add gw1
+as gw1
+check ovs-vsctl add-br br-phys
+ovn_attach n1 br-phys 192.168.0.1
+check ovs-vsctl add-br br-ex
+net_attach provider br-ex
+check ovs-vsctl set open . external-ids:ovn-bridge-mappings=phys:br-ex
+if test "$1" != rbac && test X$HAVE_OPENSSL = Xyes; then
+    check ovs-vsctl set open . \
+        external-ids:ovn-remote=unix:$ovs_base/ovn-sb/ovn-sb.sock
+    OVS_WAIT_UNTIL([grep -q 'ovn-sb.sock: connected' gw1/ovn-controller.log])
+fi
+
+check ovn-nbctl lr-add r0
+check ovn-nbctl ls-add ext
+check ovn-nbctl lsp-add-localnet-port ext ln-ext phys
+check ovn-nbctl lrp-add r0 r0-ext 00:00:00:00:00:01 10.0.0.1/24
+check ovn-nbctl lsp-add-router-port ext ext-r0 r0-ext
+check ovn-nbctl lrp-set-gateway-chassis r0-ext gw1
+check ovn-nbctl static-mac-binding-add r0-ext 10.0.0.254 00:00:00:00:00:fe
+check ovn-nbctl --bfd lr-route-add r0 0.0.0.0/0 10.0.0.254 r0-ext
+wait_column "$(fetch_column Chassis _uuid name=gw1)" Port_Binding chassis \
+    logical_port=cr-r0-ext
+wait_row_count BFD 1
+bfd=$(fetch_column BFD _uuid dst_ip=10.0.0.254)
+wait_column down BFD status dst_ip=10.0.0.254
+OVN_WAIT_PATCH_PORT_FLOWS([ln-ext], [gw1])
+check ovn-nbctl --wait=hv sync
+
+# bfd_rc_peer_pkt STATE MULT
+#
+# Prints a BFD control packet from the peer 10.0.0.254 to gw1's session, in
+# state STATE (0 admin_down, 1 down, 2 init, 3 up), with 1 s intervals and
+# detection multiplier MULT.  The session's detection time is then MULT s.
+bfd_rc_peer_pkt() {
+    local disc flags mult
+    disc=$(printf %08x $(fetch_column BFD disc dst_ip=10.0.0.254))
+    flags=$(printf %02x $((${1} * 64)))
+    mult=$(printf %02x ${2})
+    fmt_pkt "Ether(dst='00:00:00:00:00:01', src='00:00:00:00:00:fe')/ \
+             IP(src='10.0.0.254', dst='10.0.0.1', ttl=255)/ \
+             UDP(sport=49152, dport=3784)/ \
+             Raw(load=bytes.fromhex('20${flags}${mult}18000000fe${disc}' \
+                                    '000f4240000f424000000000'))"
+}
+
+# bfd_rc_peer_start [STATE [MULT]]
+#
+# Plays the peer: sends a packet in STATE (default 2, init) with detection
+# multiplier MULT (default 3) every 0.3 s, until bfd_rc_peer_stop.  "init"
+# brings the session up and keeps it up.
+bfd_rc_peer_start() {
+    local pkt
+    pkt=$(bfd_rc_peer_pkt ${1:-2} ${2:-3})
+    (while :; do
+         as gw1 ovs-appctl netdev-dummy/receive br-ex_provider $pkt \
+             >/dev/null 2>&1
+         sleep 0.3
+     done) &
+    echo $! > bfd-peer.pid
+    on_exit 'test -e bfd-peer.pid && kill $(cat bfd-peer.pid)'
+}
+
+bfd_rc_peer_stop() {
+    kill $(cat bfd-peer.pid)
+    wait $(cat bfd-peer.pid) 2>/dev/null
+    rm -f bfd-peer.pid
+}
+
+# bfd_rc_stop
+#
+# Stops gw1's ovn-controller without cleanup: its gateway stays bound.
+bfd_rc_stop() {
+    OVN_CONTROLLER_EXIT([gw1], [--restart])
+}
+
+# bfd_rc_start
+#
+# Starts gw1's ovn-controller with debug logs from pinctrl.  Sets
+# bfd_rc_first to the first line of its log.
+bfd_rc_start() {
+    bfd_rc_first=$(($(wc -l < gw1/ovn-controller.log) + 1))
+    as gw1 start_daemon ovn-controller --enable-dummy-vif-plug \
+        -vpinctrl:file:dbg
+}
+
+# bfd_rc_writes STATUS
+#
+# Prints how many transactions from ovn-controller set a BFD status to
+# STATUS, as received by the SB database.
+bfd_rc_writes() {
+    grep 'received request, method="transact"' ovn-sb/ovsdb-server.log \
+        | grep '"comment":"ovn-controller' \
+        | grep -c "\"row\":{\"status\":\"${1}\"},\"table\":\"BFD\""
+}
+
+# bfd_rc_counter COUNTER
+#
+# Prints the value of coverage counter COUNTER in gw1's ovn-controller.
+bfd_rc_counter() {
+    as gw1 ovn-appctl -t ovn-controller coverage/read-counter ${1}
+}
+
+# bfd_rc_ms TIMESTAMP...
+#
+# Prints each log TIMESTAMP in ms since the epoch.
+bfd_rc_ms() {
+    local ts
+    for ts in "${@}"; do
+        date -d "$ts" +%s%3N
+    done
+}
+
+# bfd_rc_log_ms PATTERN [FIRST_LINE]
+#
+# Prints the times, in ms since the epoch, of the lines of gw1's
+# ovn-controller log that match PATTERN, from line FIRST_LINE on.
+bfd_rc_log_ms() {
+    bfd_rc_ms $(tail -n +${2:-1} gw1/ovn-controller.log | grep "${1}" \
+                | cut -d'|' -f1)
+}
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - reconcile SB status with the session state])
+AT_KEYWORDS([bfd bfd-reconcile])
+BFD_RECONCILE_SETUP
+
+AS_BOX([The session comes up])
+bfd_rc_peer_start
+wait_column up BFD status dst_ip=10.0.0.254
+
+AS_BOX([A status written by another client is corrected])
+# The session stays up, so its state does not change: only the correction
+# writes the status back.
+for status in down init; do
+    check ovn-sbctl set BFD $bfd status=$status
+    wait_column up BFD status dst_ip=10.0.0.254
+    OVS_WAIT_UNTIL([grep -q "BFD r0-ext dst 10.0.0.254: correcting SB status 
\"$status\" to the session state \"up\" (correction 1)" \
+                        gw1/ovn-controller.log])
+done
+OVS_WAIT_UNTIL([test "$(bfd_rc_counter pinctrl_bfd_status_reconcile)" = 2])
+
+AS_BOX([No writes while the status is right])
+writes=$(bfd_rc_writes up)
+sleep 6
+check_column up BFD status dst_ip=10.0.0.254
+check test "$(bfd_rc_counter pinctrl_bfd_status_reconcile)" = 2
+check test "$(bfd_rc_writes up)" = "$writes"
+check test "$(bfd_rc_writes down)" = 0
+# The corrections do not "verify" the status: a failed verify would abort
+# the whole SB transaction of ovn-controller.
+AT_CHECK([grep 'received request, method="transact"' ovn-sb/ovsdb-server.log \
+              | grep '"comment":"ovn-controller' | grep '"table":"BFD"' \
+              | grep -c '"columns":."status".,"op":"wait"'], [1], [0
+])
+
+AS_BOX([A restart while the peer answers does not change the status])
+# The new session is down until the peer answers it.  Correcting SB "up"
+# to "down" in the meantime would remove the route for no reason.
+bfd_rc_stop
+bfd_rc_start
+OVS_WAIT_UNTIL([tail -n +$bfd_rc_first gw1/ovn-controller.log | \
+                    grep -q 'rx BFD packets from 10.0.0.254, remote state 
init, local state up'])
+check_column up BFD status dst_ip=10.0.0.254
+check test "$(bfd_rc_writes down)" = 0
+AT_CHECK([tail -n +$bfd_rc_first gw1/ovn-controller.log | grep -c 'correcting 
SB status'],
+         [1], [0
+])
+
+AS_BOX([The hold counts from when the session starts sending])
+# The session was created long ago.  Stop it with "admin_down" and start it
+# again while SB says "init": the new handshake must not see a "down"
+# written over "init", which a hold counted from the session's creation
+# would allow.  Keep ovn-northd from setting the status itself meanwhile.
+check as northd ovn-appctl -t ovn-northd pause
+check ovn-sbctl set BFD $bfd status=admin_down
+first=$(($(wc -l < gw1/ovn-controller.log) + 1))
+OVS_WAIT_UNTIL([tail -n +$first gw1/ovn-controller.log | \
+                    grep -q 'local state admin_down'])
+# A stopped session ignores the peer.  Let its last packet from the peer
+# become older than the detection time (3 s), as for a new session.
+sleep 4
+check ovn-sbctl set BFD $bfd status=init
+# The session goes down, starts sending and comes up: "up" is written by the
+# state change, nothing is written before.
+wait_column up BFD status dst_ip=10.0.0.254
+check test "$(bfd_rc_writes down)" = 0
+AT_CHECK([tail -n +$first gw1/ovn-controller.log | grep -c 'correcting SB 
status'],
+         [1], [0
+])
+check as northd ovn-appctl -t ovn-northd resume
+
+AS_BOX([ovn-controller advertises that it corrects the status])
+AT_CHECK([ovn-sbctl get Chassis gw1 other_config:ovn-bfd-status-reconcile],
+         [0], [dnl
+"true"
+])
+
+# Let the session time out before the final flow checks.
+bfd_rc_peer_stop
+wait_column down BFD status dst_ip=10.0.0.254
+check ovn-nbctl --wait=hv sync
+OVN_CLEANUP([gw1
+ignored_tables=OFTABLE_PHY_TO_LOG
+])
+AT_CLEANUP
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - reconcile SB status when the peer does not come up])
+AT_KEYWORDS([bfd bfd-reconcile])
+BFD_RECONCILE_SETUP
+# No SB probes, which would wake ovn-controller up regularly: nothing else
+# happens in the databases, so only the correction's own timer can wake it
+# up in time.
+check as gw1 ovs-vsctl set open . external-ids:ovn-remote-probe-interval=0
+
+AS_BOX([A stale "up" from an earlier session is corrected after the hold])
+# The session starts sending when pinctrl_run() first runs it, which can be
+# a little after it was created: measure from the later of the two.  Log
+# times can be a few ms later than the times that the hold counts from.
+bfd_rc_stop
+check ovn-sbctl set BFD $bfd status=up
+bfd_rc_start
+wait_column down BFD status dst_ip=10.0.0.254
+started=$(bfd_rc_log_ms 'BFD r0-ext dst 10.0.0.254: session \(starts\|can send 
now\)' \
+              $bfd_rc_first | tail -1)
+corrected=$(bfd_rc_log_ms 'correcting SB status "up" to the session state 
"down" (correction 1)' $bfd_rc_first)
+echo "corrected $((corrected - started)) ms after the session started"
+check test $((corrected - started)) -ge 4900
+check test $((corrected - started)) -lt 7000
+check test "$(bfd_rc_writes down)" = 1
+OVS_WAIT_UNTIL([test "$(bfd_rc_counter pinctrl_bfd_status_reconcile)" = 1])
+
+AS_BOX([An "admin_down" is never overwritten])
+check as northd ovn-appctl -t ovn-northd pause
+check ovn-sbctl set BFD $bfd status=admin_down
+sleep 6
+check_column admin_down BFD status dst_ip=10.0.0.254
+check test "$(bfd_rc_counter pinctrl_bfd_status_reconcile)" = 1
+check test "$(bfd_rc_writes down)" = 1
+check as northd ovn-appctl -t ovn-northd resume
+# The route uses the row, so ovn-northd sets "down" again.
+wait_column down BFD status dst_ip=10.0.0.254
+
+AS_BOX([A peer that answers but does not come up holds it for 30 s at most])
+# The peer stays admin_down: the session hears from it but stays down.
+bfd_rc_peer_start 0
+bfd_rc_stop
+check ovn-sbctl set BFD $bfd status=up
+bfd_rc_start
+OVS_WAIT_UNTIL([tail -n +$bfd_rc_first gw1/ovn-controller.log | \
+                    grep -q 'rx BFD packets from 10.0.0.254, remote state 
admin_down, local state down'])
+OVS_CTL_TIMEOUT_SAVED=$OVS_CTL_TIMEOUT
+OVS_CTL_TIMEOUT=60
+wait_column down BFD status dst_ip=10.0.0.254
+OVS_CTL_TIMEOUT=$OVS_CTL_TIMEOUT_SAVED
+started=$(bfd_rc_log_ms 'BFD r0-ext dst 10.0.0.254: session \(starts\|can send 
now\)' \
+              $bfd_rc_first | tail -1)
+corrected=$(bfd_rc_log_ms 'correcting SB status "up" to the session state 
"down" (correction 1)' $bfd_rc_first)
+echo "corrected $((corrected - started)) ms after the session started"
+check test $((corrected - started)) -ge 29900
+check test $((corrected - started)) -lt 32000
+bfd_rc_peer_stop
+
+check ovn-nbctl --wait=hv sync
+OVN_CLEANUP([gw1
+ignored_tables=OFTABLE_PHY_TO_LOG
+])
+AT_CLEANUP
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - session times out while ovs-vswitchd is down])
+AT_KEYWORDS([bfd bfd-reconcile])
+BFD_RECONCILE_SETUP
+
+# Detection time 10 s.  ovn-controller's reconnection attempts to
+# ovs-vswitchd wake it up 1, 3, 7 and 15 s after the disconnection, so
+# only the detection timer can wake it up in time.
+bfd_rc_peer_start 2 10
+wait_column up BFD status dst_ip=10.0.0.254
+
+AS_BOX([ovs-vswitchd stops: the session times out and "down" is written])
+as gw1
+OVS_APP_EXIT_AND_WAIT([ovs-vswitchd])
+bfd_rc_peer_stop
+stopped=$(date +%s%3N)
+wait_column down BFD status dst_ip=10.0.0.254
+written=$(bfd_rc_ms $(grep 'received request, method="transact"' \
+                          ovn-sb/ovsdb-server.log \
+                      | grep '"comment":"ovn-controller' \
+                      | grep '"row":{"status":"down"},"table":"BFD"' \
+                      | tail -1 | cut -d'|' -f1))
+echo "down written $((written - stopped)) ms after ovs-vswitchd stopped"
+check test $((written - stopped)) -lt 12000
+
+AS_BOX([No "up" without packets from the peer])
+sleep 6
+check_column down BFD status dst_ip=10.0.0.254
+check test "$(bfd_rc_writes up)" = 1
+check test "$(bfd_rc_counter pinctrl_bfd_status_reconcile)" = 0
+
+AS_BOX([ovn-controller starts while ovs-vswitchd is down])
+# Its new session cannot send; a stale "up" is corrected after the hold.
+bfd_rc_stop
+check ovn-sbctl set BFD $bfd status=up
+bfd_rc_start
+wait_column down BFD status dst_ip=10.0.0.254
+started=$(bfd_rc_log_ms 'BFD r0-ext dst 10.0.0.254: session starts (cannot 
send yet)' $bfd_rc_first)
+corrected=$(bfd_rc_log_ms 'correcting SB status "up" to the session state 
"down" (correction 1)' $bfd_rc_first)
+echo "corrected $((corrected - started)) ms after the session started"
+check test $((corrected - started)) -ge 4900
+check test $((corrected - started)) -lt 7000
+
+AS_BOX([ovs-vswitchd and the peer come back])
+as gw1
+start_daemon ovs-vswitchd --enable-dummy=system -vvconn -vofproto_dpif 
-vunixctl
+# ovn-controller acknowledges this only after it installed its flows again.
+check ovn-nbctl --wait=hv sync
+OVS_WAIT_UNTIL([tail -n +$bfd_rc_first gw1/ovn-controller.log | \
+                    grep -q 'BFD r0-ext dst 10.0.0.254: session can send now'])
+check_column down BFD status dst_ip=10.0.0.254
+bfd_rc_peer_start
+wait_column up BFD status dst_ip=10.0.0.254
+
+# Let the session time out before the final flow checks.
+bfd_rc_peer_stop
+wait_column down BFD status dst_ip=10.0.0.254
+check ovn-nbctl --wait=hv sync
+OVN_CLEANUP([gw1
+ignored_tables=OFTABLE_PHY_TO_LOG
+])
+AT_CLEANUP
+])
+
+OVN_FOR_EACH_NORTHD([
+AT_SETUP([BFD - reconcile SB status when SB RBAC rejects the update])
+AT_KEYWORDS([bfd bfd-reconcile])
+AT_SKIP_IF([test "$HAVE_OPENSSL" = no])
+BFD_RECONCILE_SETUP([rbac])
+# Two more sessions, towards 10.0.0.253 and 10.0.0.252.  SB RBAC allows gw1
+# to update only the row whose chassis_name is gw1, here 10.0.0.253.  Keep
+# ovn-northd from setting chassis_name meanwhile.
+check ovn-nbctl --bfd lr-route-add r0 10.1.0.0/16 10.0.0.253 r0-ext
+check ovn-nbctl --bfd lr-route-add r0 10.2.0.0/16 10.0.0.252 r0-ext
+wait_row_count BFD 3
+for ip in 10.0.0.254 10.0.0.253 10.0.0.252; do
+    wait_column down BFD status dst_ip=$ip
+done
+own=$(fetch_column BFD _uuid dst_ip=10.0.0.253)
+other=$(fetch_column BFD _uuid dst_ip=10.0.0.252)
+check as northd ovn-appctl -t ovn-northd pause
+check ovn-sbctl set BFD $bfd chassis_name=\"\" \
+    -- set BFD $own chassis_name=gw1 -- set BFD $other chassis_name=\"\"
+# Let the start-up holds of the new sessions end, so that all three rows are
+# corrected in the same pass below.
+sleep 6
+
+AS_BOX([Corrections back off, then stop, and do not block other rows])
+# One transaction sets all three rows to a stale "up".  The row that gw1 may
+# update must not fail with the others.
+check ovn-sbctl set BFD $bfd status=up -- set BFD $own status=up \
+    -- set BFD $other status=up
+wait_column down BFD status dst_ip=10.0.0.253
+OVS_CTL_TIMEOUT_SAVED=$OVS_CTL_TIMEOUT
+OVS_CTL_TIMEOUT=120
+OVS_WAIT_UNTIL([test $(grep -c 'did not follow 4 corrections' 
gw1/ovn-controller.log) = 2])
+# The last corrections fail too: 5 failed transactions, as the two rows that
+# gw1 may not update are corrected together.
+OVS_WAIT_UNTIL([test $(grep -c 'OVNSB commit failed' gw1/ovn-controller.log) 
-ge 5])
+OVS_CTL_TIMEOUT=$OVS_CTL_TIMEOUT_SAVED
+check_column up BFD status dst_ip=10.0.0.254
+check_column up BFD status dst_ip=10.0.0.252
+check_column down BFD status dst_ip=10.0.0.253
+
+bfd_rc_log_ms 'BFD r0-ext dst 10.0.0.254: correcting SB status "up" to the 
session state "down"' > times
+AT_CHECK([wc -l < times], [0], [5
+])
+# 5, 10, 20 and 40 s apart, so at most 3 corrections in the first 30 s.
+prev=
+for gap in 0 5000 10000 20000 40000; do
+    read t
+    if test -n "$prev"; then
+        echo "correction $((t - prev)) ms after the previous one"
+        check test $((t - prev)) -ge $((gap - 50))
+        check test $((t - prev)) -lt $((gap + 3000))
+    fi
+    prev=$t
+done < times
+AT_CHECK([grep -c 'BFD r0-ext dst 10.0.0.253: correcting SB status' \
+              gw1/ovn-controller.log], [0], [1
+])
+AT_CHECK([grep 'BFD r0-ext dst 10.0.0.254: SB status "up" did not follow' \
+              gw1/ovn-controller.log | sed 's/.*|WARN|//; s/@{:@. row/@{:@N 
row/'],
+         [0], [dnl
+BFD r0-ext dst 10.0.0.254: SB status "up" did not follow 4 corrections to the 
session state "down"; making one last correction, then none until the row 
changes or for 10 minutes (N row(s) of this chassis stopped).  The SB server 
may be rejecting them: with SB RBAC, the row's chassis_name ("") must be this 
chassis ("gw1")
+])
+OVS_WAIT_UNTIL([test "$(bfd_rc_counter pinctrl_bfd_status_reconcile_stuck)" = 
2])
+OVS_WAIT_UNTIL([test "$(bfd_rc_counter pinctrl_bfd_status_reconcile)" = 11])
+
+AS_BOX([Stopped rows cost no failed SB transactions])
+failed=$(grep -c 'OVNSB commit failed' gw1/ovn-controller.log)
+sleep 10
+check test "$(grep -c 'OVNSB commit failed' gw1/ovn-controller.log)" = 
"$failed"
+check test "$(bfd_rc_counter pinctrl_bfd_status_reconcile)" = 11
+
+AS_BOX([Another status in a stopped row resumes its corrections])
+check ovn-sbctl set BFD $other status=init
+OVS_WAIT_UNTIL([grep -q 'BFD r0-ext dst 10.0.0.252: SB status changed, SB 
status corrections resume' \
+                    gw1/ovn-controller.log])
+OVS_WAIT_UNTIL([grep -q 'BFD r0-ext dst 10.0.0.252: correcting SB status 
"init" to the session state "down" (correction 1)' \
+                    gw1/ovn-controller.log])
+
+AS_BOX([chassis_name becoming gw1 restarts the corrections at once])
+OVS_WAIT_UNTIL([grep -q 'BFD r0-ext dst 10.0.0.252: correcting SB status 
"init" to the session state "down" (correction 2)' \
+                    gw1/ovn-controller.log])
+# The next correction would come 10 s after the second one.
+t0=$(date +%s%3N)
+check ovn-sbctl set BFD $other chassis_name=gw1
+wait_column down BFD status dst_ip=10.0.0.252
+t1=$(date +%s%3N)
+echo "corrected $((t1 - t0)) ms after chassis_name changed"
+check test $((t1 - t0)) -lt 5000
+AT_CHECK([grep -c 'BFD r0-ext dst 10.0.0.252: correcting SB status "init" to 
the session state "down" (correction 3)' \
+              gw1/ovn-controller.log], [1], [0
+])
+
+AS_BOX([Setting chassis_name lets the correction through])
+check ovn-sbctl set BFD $bfd chassis_name=gw1
+wait_column down BFD status dst_ip=10.0.0.254
+check grep -q 'BFD r0-ext dst 10.0.0.254: SB chassis_name changed, SB status 
corrections resume' \
+    gw1/ovn-controller.log
+AT_CHECK([grep -c 'BFD r0-ext dst 10.0.0.254: correcting SB status' \
+              gw1/ovn-controller.log], [0], [6
+])
+
+check as northd ovn-appctl -t ovn-northd resume
+check ovn-nbctl --wait=hv sync
+OVN_CLEANUP([gw1
+/did not follow 4 corrections/d
+/transaction error/d
+ignored_tables=OFTABLE_PHY_TO_LOG
+])
+AT_CLEANUP
+])
+
 OVN_FOR_EACH_NORTHD([
 AT_SETUP([4 HV, 1 LS, 1 LR, packet test with HA distributed router gateway 
port])
 ovn_start
-- 
2.48.1

_______________________________________________
dev mailing list
[email protected]
https://mail.openvswitch.org/mailman/listinfo/ovs-dev

Reply via email to