=== EXECUTIVE SUMMARY ===
Operational Status: Ingress service outages during Keepalived failovers
are currently prevented by deploying a periodic Gratuitous ARP (GARP)
workaround (vrrp_garp_master_refresh > 0).
> ARP (Address Resolution Protocol): Network communication protocol connecting
> a logical IP address to a physical MAC address on a local area network (LAN).
> Gratuitous ARP (GARP): Unprompted ARP broadcast used to announce a device's
> IP-to-MAC mapping to the entire network without a prior request.
* Immediate Failover Recovery (Sub-Second): Failover recovery is sub-second.
Upon winning the master election, Keepalived immediately fires an initial burst
of GARPs, restoring traffic routing instantly.
* The 60-Second Refresh as a Safety Net: The 60-second periodic timer is an
automated self-healing mechanism built into the Keepalived daemon. If the
initial GARP packet is dropped or missed by OVN, the periodic refresh caps the
worst-case recovery time at 60 seconds.
* The Underlying Issue: OVN contains a core networking bug where virtual port
ownership moves to the new chassis correctly during failover, but the OVN
Southbound Mac_Binding table fails to update automatically. Without GARPs
forcing an update, traffic routes to the old host's MAC address.
* Strategic Mitigation & Next Steps:
- Periodic GARP (Workaround): Currently being deployed via GitOps to 100
ingress deployments (300 total units). Delivers immediate failover recovery
while providing a 60s backstop against packet loss, subject to multi-day
control-plane monitoring.
- Shared Virtual MAC (VMAC): Evaluated as a structural fix, but triggers an
OVN routing failure when OpenStack Port Security is enabled (misrouting traffic
to inactive backup hosts).
- Permanent Resolution: Upstream OVN fixes to ensure correct traffic
forwarding after Keepalived failover. Once fixed, GARP can be retained as a
resilience mechanism rather than a primary workaround.
--------------------------------------------------------------------------------
=== IMPACT SUMMARY MATRIX ===
+--------------------------+-----------------------------+-----------------------------------------------------------------------+
| Area | Current Status & Assessment | Operational &
Technical Impact |
+--------------------------+-----------------------------+-----------------------------------------------------------------------+
| Failover Latency | Sub-Second Recovery | Immediate traffic
restoration via initial GARP burst upon VRRP state |
| | | change.
|
| 60s Timer Function | Automated Safety Net | Probabilistic Safety
Net: Native daemon timer provides repeated |
| | | update attempts.
Recovery is probabilistic rather than guaranteed |
| | | if OVS continues
dropping GARP packets. |
| Service Availability | Mitigated | Expected to remain
available in most failovers; prolonged outages |
| | | should be rare with
periodic GARP enabled. |
| Physical Fabric | Negligible Overhead | ~3–5 packets/60s per
unit (~70B each); trivial load on 400 Gbps |
| | | physical switches.
|
| Hypervisor / OVN Plane | Under Active Monitoring | Evaluating whether
extra broadcast traffic impacts OVN daemons at |
| | | scale.
|
| Split-Brain State | Symptom-Mitigated | GARP prevents
stale-MAC blackholing, but does not solve the |
| | | split-brain state
itself. |
| Shared VMAC (PortSec ON) | Failed | OVN misroutes
traffic to inactive backup nodes (Host-M3). |
| Shared VMAC (PortSec OFF)| Functional (Unsafe) | Functions reliably,
but removing Port Security introduces cloud |
| | | security risks
(discarded for production). |
| Permanent Fix | Pending | Requires upstream
OVN kernel/control-plane bug fixes and validation. |
+--------------------------+-----------------------------+-----------------------------------------------------------------------+
--------------------------------------------------------------------------------
=== DETAILED TECHNICAL BRIEFING ===
1. Root Cause Analysis: What Breaks Without the Workaround
During a standard Keepalived VRRP failover on OpenStack (Caracal) with
OVN (24.03.2 / 24.03.6):
1. Port Binding Succeeds: OVN correctly moves the virtual port binding and
chassis claim to the new active Keepalived node (ovn-controller logs confirm
Host-B claimed the lport).
2. MAC Binding Stays Stale: The OVN Southbound Mac_Binding database table fails
to refresh. It continues to point the VIP's IP (10.152.44.43) to the previous
owner's MAC (Host-C).
3. Consequence: Inbound traffic is directed to the wrong physical chassis,
creating a full service outage until Keepalived is manually restarted on the
affected node.
[Failover Triggered] ---> OVN Lport Claims Host-B (Correct)
---> OVN Mac_Binding Retains Host-C (Bug)
---> Ingress Traffic Sent to Host-C (Outage)
2. Approach 1: Periodic GARP (Current Functional Workaround)
Recovery Timing & The 60-Second Timer:
There is a critical operational distinction between immediate failover
execution and the 60-second periodic timer:
* Immediate Failover Execution (t = 0s): When a backup node takes over as
master, Keepalived instantly broadcasts an initial burst of GARP packets. OVS
processes this immediately, updating its MAC table in sub-second time. Under
normal operation, there is no 60-second wait.
* The 60-Second Periodic Safety Net (t = 60s): Configured via
vrrp_garp_master_refresh > 0, Keepalived continuously re-broadcasts GARPs every
60 seconds.
- Without Periodic Refresh: If the initial t = 0s GARP packet is lost in the
fabric, OVN remains stuck, causing an indefinite outage until manual
intervention.
- With Periodic Refresh: If the initial packet is lost, the daemon self-heals
at the next tick, capping the absolute worst-case recovery window at 60 seconds.
* Daemon-Native: This timer is executed directly inside the Keepalived C-binary
event loop (zero OS scheduling latency, script overhead, or cron lag).
Fabric Traversal & Split-Brain Handling:
* Fabric Traversal: Physical leaf switches and Geneve tunnels propagate GARPs
cleanly; underlay switch arp-nd-suppress settings do not interfere.
* Split-Brain Behavior: In tests where inter-node communication was blocked via
iptables, both nodes claimed ownership and sent GARPs. OVS continuously updated
to whichever node sent the latest GARP. While periodic GARP does not resolve
the underlying network partition, it prevents OVS from routing traffic to a
dead MAC address.
Control-Plane Overhead & Risk Profile:
* Bandwidth Overhead: Across 100 ingress services (300 units), the workaround
generates approximately 300–500 GARP packets per minute (~70 bytes each). This
represents less than 0.001% of physical switch capacity (400 Gbps).
* Hypervisor Risk: The primary unknown is whether continuous broadcast
processing adds friction to hypervisor CPU cycles or OVN daemons, which are
already experiencing intermittent stability issues.
3. Alternative Evaluated: External Connectivity Checks
Adding health checks in Keepalived to ping external targets (e.g., core
network VIPs) was evaluated to automatically drop VIP ownership if a
node loses connectivity.
* Outcome: Inadequate. In network partition (split-brain) tests,
isolated Keepalived nodes could still reach external core VIPs while
being unable to talk to each other. Both nodes continued claiming the
VIP, proving that external ping checks cannot reliably prevent state
desynchronization.
4. Approach 2: Shared Virtual MAC (VMAC / RFC 9568)
The Concept:
Using a shared Virtual MAC (VMAC) across all Keepalived nodes keeps the MAC
address static while the IP transitions between chassis. In theory, this
eliminates the need to update OVN Mac_Binding tables during failover.
Findings & Test Results:
* With Port Security Disabled: VIP migrations completed repeatedly without any
packet loss or service disruption.
* With Port Security Enabled: Triggered a severe OVN control-plane bug. Upon
failover:
1. Unit M2 became VRRP master and sent GARPs with the VMAC.
2. OVN incorrectly assigned the egress chassis to Unit M3 (the
lowest-priority backup node, which never won master election or sent GARPs).
3. Traffic was directed to M3 (where the VIP was inactive), causing complete
service failure.
[VMAC Failover with Port Security ON]
Master M2 Wins Election ---> Sends VMAC GARP ---> OVN Assigns Egress to
Inactive Host-M3 ---> Outage
5. Deployment Plan & Operational Next Steps
1. GitOps Rollout: The periodic GARP workaround (vrrp_garp_master_refresh > 0)
is currently being deployed to 100 ingress services (300 units).
2. 5–7 Day Burn-in Observation: The rollout will remain active in production
for several days to measure:
- System stability across repeated automated failovers.
- Quantifiable CPU/memory impact of broadcast traffic on hypervisors and OVN
controllers.
3. Upstream Patch Tracking: Tracking OVN upstream fixes (specifically patches
related to commits 60b58842ed and 8e825d936c) to resolve core Mac_Binding
staleness and VMAC port-security behavior, paving the way for a permanent
solution.
--
You received this bug notification because you are a member of Ubuntu
Bugs, which is subscribed to Ubuntu.
https://bugs.launchpad.net/bugs/2166984
Title:
OVN Mac_Binding table not refreshed after keepalived VIP failover
between virtual port parents
To manage notifications about this bug go to:
https://bugs.launchpad.net/ubuntu/+source/ovn/+bug/2166984/+subscriptions
--
ubuntu-bugs mailing list
[email protected]
https://lists.ubuntu.com/mailman/listinfo/ubuntu-bugs