=== EXECUTIVE SUMMARY ===

Operational Status: Ingress service outages during Keepalived failovers
are currently prevented by deploying a periodic Gratuitous ARP (GARP)
workaround (vrrp_garp_master_refresh > 0).

> ARP (Address Resolution Protocol): Network communication protocol connecting 
> a logical IP address to a physical MAC address on a local area network (LAN).
> Gratuitous ARP (GARP): Unprompted ARP broadcast used to announce a device's 
> IP-to-MAC mapping to the entire network without a prior request.

* Immediate Failover Recovery (Sub-Second): Failover recovery is sub-second. 
Upon winning the master election, Keepalived immediately fires an initial burst 
of GARPs, restoring traffic routing instantly.
* The 60-Second Refresh as a Safety Net: The 60-second periodic timer is an 
automated self-healing mechanism built into the Keepalived daemon. If the 
initial GARP packet is dropped or missed by OVN, the periodic refresh caps the 
worst-case recovery time at 60 seconds.
* The Underlying Issue: OVN contains a core networking bug where virtual port 
ownership moves to the new chassis correctly during failover, but the OVN 
Southbound Mac_Binding table fails to update automatically. Without GARPs 
forcing an update, traffic routes to the old host's MAC address.
* Strategic Mitigation & Next Steps:
  - Periodic GARP (Workaround): Currently being deployed via GitOps to 100 
ingress deployments (300 total units). Delivers immediate failover recovery 
while providing a 60s backstop against packet loss, subject to multi-day 
control-plane monitoring.
  - Shared Virtual MAC (VMAC): Evaluated as a structural fix, but triggers an 
OVN routing failure when OpenStack Port Security is enabled (misrouting traffic 
to inactive backup hosts).
  - Permanent Resolution: Upstream OVN fixes to ensure correct traffic 
forwarding after Keepalived failover. Once fixed, GARP can be retained as a 
resilience mechanism rather than a primary workaround.

--------------------------------------------------------------------------------

=== IMPACT SUMMARY MATRIX ===

+--------------------------+-----------------------------+-----------------------------------------------------------------------+
| Area                     | Current Status & Assessment | Operational & 
Technical Impact                                        |
+--------------------------+-----------------------------+-----------------------------------------------------------------------+
| Failover Latency         | Sub-Second Recovery         | Immediate traffic 
restoration via initial GARP burst upon VRRP state  |
|                          |                             | change.              
                                                 |
| 60s Timer Function       | Automated Safety Net        | Probabilistic Safety 
Net: Native daemon timer provides repeated       |
|                          |                             | update attempts. 
Recovery is probabilistic rather than guaranteed     |
|                          |                             | if OVS continues 
dropping GARP packets.                               |
| Service Availability     | Mitigated                   | Expected to remain 
available in most failovers; prolonged outages     |
|                          |                             | should be rare with 
periodic GARP enabled.                            |
| Physical Fabric          | Negligible Overhead         | ~3–5 packets/60s per 
unit (~70B each); trivial load on 400 Gbps       |
|                          |                             | physical switches.   
                                                 |
| Hypervisor / OVN Plane   | Under Active Monitoring     | Evaluating whether 
extra broadcast traffic impacts OVN daemons at     |
|                          |                             | scale.               
                                                 |
| Split-Brain State        | Symptom-Mitigated           | GARP prevents 
stale-MAC blackholing, but does not solve the           |
|                          |                             | split-brain state 
itself.                                             |
| Shared VMAC (PortSec ON) | Failed                      | OVN misroutes 
traffic to inactive backup nodes (Host-M3).             |
| Shared VMAC (PortSec OFF)| Functional (Unsafe)         | Functions reliably, 
but removing Port Security introduces cloud       |
|                          |                             | security risks 
(discarded for production).                            |
| Permanent Fix            | Pending                     | Requires upstream 
OVN kernel/control-plane bug fixes and validation.  |
+--------------------------+-----------------------------+-----------------------------------------------------------------------+

--------------------------------------------------------------------------------

=== DETAILED TECHNICAL BRIEFING ===

1. Root Cause Analysis: What Breaks Without the Workaround

During a standard Keepalived VRRP failover on OpenStack (Caracal) with
OVN (24.03.2 / 24.03.6):

1. Port Binding Succeeds: OVN correctly moves the virtual port binding and 
chassis claim to the new active Keepalived node (ovn-controller logs confirm 
Host-B claimed the lport).
2. MAC Binding Stays Stale: The OVN Southbound Mac_Binding database table fails 
to refresh. It continues to point the VIP's IP (10.152.44.43) to the previous 
owner's MAC (Host-C).
3. Consequence: Inbound traffic is directed to the wrong physical chassis, 
creating a full service outage until Keepalived is manually restarted on the 
affected node.

    [Failover Triggered] ---> OVN Lport Claims Host-B (Correct)
                         ---> OVN Mac_Binding Retains Host-C (Bug)
                         ---> Ingress Traffic Sent to Host-C (Outage)


2. Approach 1: Periodic GARP (Current Functional Workaround)

Recovery Timing & The 60-Second Timer:
There is a critical operational distinction between immediate failover 
execution and the 60-second periodic timer:

* Immediate Failover Execution (t = 0s): When a backup node takes over as 
master, Keepalived instantly broadcasts an initial burst of GARP packets. OVS 
processes this immediately, updating its MAC table in sub-second time. Under 
normal operation, there is no 60-second wait.
* The 60-Second Periodic Safety Net (t = 60s): Configured via 
vrrp_garp_master_refresh > 0, Keepalived continuously re-broadcasts GARPs every 
60 seconds.
  - Without Periodic Refresh: If the initial t = 0s GARP packet is lost in the 
fabric, OVN remains stuck, causing an indefinite outage until manual 
intervention.
  - With Periodic Refresh: If the initial packet is lost, the daemon self-heals 
at the next tick, capping the absolute worst-case recovery window at 60 seconds.
* Daemon-Native: This timer is executed directly inside the Keepalived C-binary 
event loop (zero OS scheduling latency, script overhead, or cron lag).

Fabric Traversal & Split-Brain Handling:
* Fabric Traversal: Physical leaf switches and Geneve tunnels propagate GARPs 
cleanly; underlay switch arp-nd-suppress settings do not interfere.
* Split-Brain Behavior: In tests where inter-node communication was blocked via 
iptables, both nodes claimed ownership and sent GARPs. OVS continuously updated 
to whichever node sent the latest GARP. While periodic GARP does not resolve 
the underlying network partition, it prevents OVS from routing traffic to a 
dead MAC address.

Control-Plane Overhead & Risk Profile:
* Bandwidth Overhead: Across 100 ingress services (300 units), the workaround 
generates approximately 300–500 GARP packets per minute (~70 bytes each). This 
represents less than 0.001% of physical switch capacity (400 Gbps).
* Hypervisor Risk: The primary unknown is whether continuous broadcast 
processing adds friction to hypervisor CPU cycles or OVN daemons, which are 
already experiencing intermittent stability issues.


3. Alternative Evaluated: External Connectivity Checks

Adding health checks in Keepalived to ping external targets (e.g., core
network VIPs) was evaluated to automatically drop VIP ownership if a
node loses connectivity.

* Outcome: Inadequate. In network partition (split-brain) tests,
isolated Keepalived nodes could still reach external core VIPs while
being unable to talk to each other. Both nodes continued claiming the
VIP, proving that external ping checks cannot reliably prevent state
desynchronization.


4. Approach 2: Shared Virtual MAC (VMAC / RFC 9568)

The Concept:
Using a shared Virtual MAC (VMAC) across all Keepalived nodes keeps the MAC 
address static while the IP transitions between chassis. In theory, this 
eliminates the need to update OVN Mac_Binding tables during failover.

Findings & Test Results:
* With Port Security Disabled: VIP migrations completed repeatedly without any 
packet loss or service disruption.
* With Port Security Enabled: Triggered a severe OVN control-plane bug. Upon 
failover:
  1. Unit M2 became VRRP master and sent GARPs with the VMAC.
  2. OVN incorrectly assigned the egress chassis to Unit M3 (the 
lowest-priority backup node, which never won master election or sent GARPs).
  3. Traffic was directed to M3 (where the VIP was inactive), causing complete 
service failure.

    [VMAC Failover with Port Security ON]
    Master M2 Wins Election ---> Sends VMAC GARP ---> OVN Assigns Egress to 
Inactive Host-M3 ---> Outage


5. Deployment Plan & Operational Next Steps

1. GitOps Rollout: The periodic GARP workaround (vrrp_garp_master_refresh > 0) 
is currently being deployed to 100 ingress services (300 units).
2. 5–7 Day Burn-in Observation: The rollout will remain active in production 
for several days to measure:
   - System stability across repeated automated failovers.
   - Quantifiable CPU/memory impact of broadcast traffic on hypervisors and OVN 
controllers.
3. Upstream Patch Tracking: Tracking OVN upstream fixes (specifically patches 
related to commits 60b58842ed and 8e825d936c) to resolve core Mac_Binding 
staleness and VMAC port-security behavior, paving the way for a permanent 
solution.

-- 
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2166984

Title:
  OVN Mac_Binding table not refreshed after keepalived VIP failover
  between virtual port parents

To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/ovn/+bug/2166984/+subscriptions


-- 
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs

Reply via email to